The Swahili Large Corpus (v1) is one of the largest and most diverse open pretraining datasets
for the Swahili language. It was engineered for training compute-optimal large language models (LLMs),
combining multiple high-quality Swahili and multilingual sources into a single, rigorously
deduplicated Parquet dataset.
Following Chinchilla scaling laws (D = 20N), this corpus is ideally suited for training