This dataset is part of the BabyLM multilingual collection.
Language: zul
Script: Latin
Number of Documents: 16108
Total Tokens: 742449
child-books: 96383 tokens
educational: 56641 tokens
padding-wikipedia: 584023 tokens
qed: 5402 tokens
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)