This dataset is part of the BabyLM multilingual collection.
Language: afr
Script: Latin
Number of Documents: 96599
Total Tokens: 9314216
child-books: 153914 tokens
child-directed-speech: 240864 tokens
educational: 116380 tokens
padding-wikipedia: 7487317 tokens
qed: 464009 tokens
subtitles: 851732 tokens
text: The document text