A large-scale, open-source speech dataset for the Tigre language (ISO 639-3: tig), developed to supp
A Swahili corpus made from Swahili Wikipedia articles # wikipedia-swahili-corpus A Swahili corpus m
Cleaned Somali Wikipedia corpus (~9,500 articles) for NLP, LLM training, and linguistic research #
This repository introduces the Monolingual Text component of the Tigre language resource collection.