A cleaned and curated corpus of Hausa-language news articles from 2020-2025. This dataset is designed for training and fine-tuning language models for Hausa text generation and understanding tasks.
The dataset is split into training and validation sets:
Training set: ~57K articles
Validation set: ~3K articles
text: Article content in the format <|title|>... <|section|>... body