Logo Lanfrica

Unkadlabs/awesome-somali-nlp

Domaine:

natural language processing

Type de record:

paper
Créateur:
Unk
Hôte:
Curated index of Somali NLP: datasets, models, benchmarks, papers, and the gaps still waiting to be filled. # Awesome Somali NLP A curated index of every dataset, model, benchmark, and paper for **Somali natural language processing** — and an honest map of what's still missing. Somali is spoken by more than 20 million people and carries one of the world's great oral traditions, yet it remains severely under-resourced in NLP. This list exists so that nobody researching Somali starts from zero again. Maintained by Unkad Labs, a non-profit AI research laboratory in Mogadishu. Contributions welcome — see Contributing. ## Contents - Text corpora - Task datasets - Parallel data & machine translation - Models - Benchmarks & evaluation - Papers - The gaps - Community & infrastructure ## Text corpora - SomaliWeb v1 — deduplicated, quality-filtered Somali web corpus: 819K documents, ~303M tokens (HPLT v2 + CC100-so + Wikipedia), CC BY-SA 4.0, with a matched 16K-vocab Somali tokenizer. Paper. - HPLT v2 — large multilingual web corpus with a substantial Somali portion; the largest single source of raw Somali text. - CC100 (so) — the Somali split of the CC100 CommonCrawl corpus. - Somali Wikipedia — small (thousands of articles) but clean, openly licensed encyclopedic Somali. ## Task datasets - MasakhaNEWS — news topic classification for 16 African languages; 2,915 labeled Somali articles. - SomBERTa fake-news & toxicity sets — human-annotated Somali Facebook data: ~1.9K fake-news instances and ~3K toxicity-labeled comments, introduced with the SomBERTa model. Paper. - SIB-200 — topic classification across 205 languages and dialects, Somali included. - Belebele — multiple-choice reading comprehension in 122 language variants, Somali included. - Aya Dataset — human-curated multilingual instruction data (65 languages) with Somali coverage. ## Parallel data & machine translation - FLORES-200 — the standard many-to-many MT evaluation benchmark; Somali (`som_Latn`) is one of its 200+ languages. - OPUS — aggregated parallel corpora; multiple collections include English–Somali pai …