This dataset is part of the BabyLM multilingual collection.
Language: nso
Script: Latin
Number of Documents: 26772
Total Tokens: 1067761
child-books: 122083 tokens
child-news: 130 tokens
educational: 92589 tokens
padding-mt: 206703 tokens
padding-news: 150960 tokens
padding-wikipedia: 495296 tokens
text: The document text