A large-scale Swahili text corpus for language model pretraining and NLP research.
The Swahili Corpus Dataset is a large-scale collection of Swahili (Kiswahili) text designed to support Natural Language Processing (NLP) research and the development of large language models (LLMs) for a low-resource African language.