A collection of training corpus and models for "Multilingual Language Model Pretraining using Machine-translated Data".
BritLLM
community
AI & ML interests
contact@llm.org.uk
datasets 18
britllm/TransWebEdu
Updated • 53 • 3
britllm/TransWeb-Edu-English
Viewer • Updated • 36M • 258
britllm/TransWeb-Edu-Spanish
Viewer • Updated • 35.2M • 483 • 3
britllm/TransWeb-Edu-French
Viewer • Updated • 36M • 338
britllm/TransWeb-Edu-German
Viewer • Updated • 36M • 474 • 1
britllm/xnli_brit
Viewer • Updated • 9.69k • 13
britllm/piqa_scottish_gaelic
Updated • 1
britllm/piqa_welsh
Updated • 1
britllm/piqa_irish
Updated • 2
britllm/arc_scottish_gaelic
Viewer • Updated • 7.56k • 30