Small Multilingual Pretraining Copora used in the ICML 2025 Paper: Banyan: Improved Representation Learning with Explicit Structure
It contains the following languages:
Afrikaans: af
Amharic: am
Arabic: ar
English: en
Spanish: es
Hausa: ha
Hindi: hi
Indonesian: id
Marathi: mr
Telugu: te
Each language contains roughly between 10-100 million tokens of pre-training data.
English is a sourced from a subsample of human written English Wikpedia.