Filtered and sorted by date, 56 articles were identified out of 23,000 english articles in the globa
This consist of a monolingual news corpus for 19 languages from various sources like VOA, B
Language modeling, topic classification, AI training for Swahili NLP, digital literacy tools Notes
To build distributional semantic models, a large amount of text is required. These days, an enormous amount of texts are being generated continuously from different sources. As we want to build general-purpose semantic models, we collected datasets from different c