Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Somali Web Corpus V1

Domain:

natural language processing

Record type:

dataset
Creator:
maa
Host:
This dataset consists of clean, structured, and filtered Somali language text compiled from various online sources. It is designed for training and fine-tuning Somali language models (LLMs) and supporting natural language processing (NLP) research for the Somali language. Language: Somali (so) Format: JSON lines (.jsonl) Data Structure: Each record has a single text field containing a cleaned paragraph.

Visit

huggingface.co

Tasks

language modeling

Languages

Somali

Tags

somalicorpustext-generationllm

Licenses

mit

Similar

SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification BenchmarkMorphologically-informed Somali Lemmatization Corpus built with a Web-based Crowdsourcing PlatformSwahili Large Corpus (v1)nolashii143/somali-tts-webOdhuso/ddd-kenya-somali-asr-v1datawise-africa/sheria-corpus-v1

SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark

Somali is a Cushitic language of the Horn of Africa with ~25 million speakers, yet no documented ded

Morphologically-informed Somali Lemmatization Corpus built with a Web-based Crowdsourcing Platform

Swahili Large Corpus (v1)

The Swahili Large Corpus (v1) is one of the largest and most diverse open pretraining datasets for t

nolashii143/somali-tts-web

Somali TTS web app with Supabase auth and Hugging Face Spaces # Somali TTS Web Next.js dashboard f

Odhuso/ddd-kenya-somali-asr-v1

datawise-africa/sheria-corpus-v1

The Sheria Corpus v1 is a curated collection of Kenyan legal case summaries from both the High Court