Logo Lanfrica

Small-Multilingual-Corpora

Domain:

natural language processing

Record type:

dataset
Creator:
m0p
Host:
Small Multilingual Pretraining Copora used in the ICML 2025 Paper: Banyan: Improved Representation Learning with Explicit Structure It contains the following languages: Afrikaans: af Amharic: am Arabic: ar English: en Spanish: es Hausa: ha Hindi: hi Indonesian: id Marathi: mr Telugu: te Each language contains roughly between 10-100 million tokens of pre-training data. English is a sourced from a subsample of human written English Wikpedia.