MaLA corpus
MaLA Corpus for Massive Language Adaptation of Large Language Models https://mala-lm.github.io
Viewer • Updated • 2.14B • 4.93k • 2Note The MaLA monolingual corpus's noisy version that integrates texts from different sources without cleaning.
MaLA-LM/mala-monolingual-filter
Updated • 1.48k • 3Note The MaLA monolingual corpus's filtered version that performs further data filtering
MaLA-LM/mala-monolingual-dedup
Updated • 1.06k • 2Note The MaLA monolingual corpus's deduplicated version that removes repeated data points
MaLA-LM/mala-monolingual-split
Viewer • Updated • 825M • 4.05k • 4Note The MaLA monolingual corpus's final version is processed by splitting the filtered and deduplicated version into training and test sets
MaLA-LM/mala-bilingual-translation-corpus
Viewer • Updated • 16.5B • 835 • 8Note The MaLA bilingual translation corpus contains parallel data in more than 2,500 language pairs (500+ languages).
MaLA-LM/mala-code-reasoning
Viewer • Updated • 44.9M • 27 • 5Note The first version of the MaLA code and reasoning dataset used for training https://huggingface.co/MaLA-LM/emma-500-llama2-7b
MaLA-LM/mala-code-reasoning-v2
Viewer • Updated • 89.7M • 63 • 9Note The 2nd version of the MaLA code and reasoning dataset used for training EMMA-500 Llama 3(.1) Mono/Bi model series.
MaLA-LM/mala-opus-dedup-2410-sample
Viewer • Updated • 9.5B • 281Note A sampled set of MaLA-LM/mala-opus-dedup-2410
-
MaLA-LM/mala-code-reasoning-v3
Viewer • Updated • 168M • 51 • 3