Curated index of Somali NLP: datasets, models, benchmarks, papers, and the gaps still waiting to be filled.
# Awesome Somali NLP
A curated index of every dataset, model, benchmark, and paper for **Somali natural
language processing** — and an honest map of what's still missing.
Somali is spoken by more than 20 million people and carries one of the world's great
oral traditions, yet it remains severely under-resourced in NLP. This list exists so
that nobody researching Somali starts from zero again.
Maintained by Unkad Labs, a non-profit AI research laboratory in
Mogadishu. Contributions welcome — see Contributing.
## Contents
- Text corpora
- Task datasets
- Parallel data & machine translation
- Models
- Benchmarks & evaluation
- Papers
- The gaps
- Community & infrastructure
## Text corpora
- SomaliWeb v1 — deduplicated, quality-filtered Somali web corpus: 819K documents, ~303M tokens (HPLT v2 + CC100-so + Wikipedia), CC BY-SA 4.0, with a matched 16K-vocab Somali tokenizer. Paper.
- HPLT v2 — large multilingual web corpus with a substantial Somali portion; the largest single source of raw Somali text.
- CC100 (so) — the Somali split of the CC100 CommonCrawl corpus.
- Somali Wikipedia — small (thousands of articles) but clean, openly licensed encyclopedic Somali.
## Task datasets
- MasakhaNEWS — news topic classification for 16 African languages; 2,915 labeled Somali articles.
- SomBERTa fake-news & toxicity sets — human-annotated Somali Facebook data: ~1.9K fake-news instances and ~3K toxicity-labeled comments, introduced with the SomBERTa model. Paper.
- SIB-200 — topic classification across 205 languages and dialects, Somali included.
- Belebele — multiple-choice reading comprehension in 122 language variants, Somali included.
- Aya Dataset — human-curated multilingual instruction data (65 languages) with Somali coverage.
## Parallel data & machine translation
- FLORES-200 — the standard many-to-many MT evaluation benchmark; Somali (`som_Latn`) is one of its 200+ languages.
- OPUS — aggregated parallel corpora; multiple collections include English–Somali pai …