The Swahili Large Corpus (v1) is one of the largest and most diverse open pretraining datasets for t
This dataset consists of clean, structured, and filtered Somali language text compiled from various