Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark

Domaine:

natural language processing

Type de record:

paperdatasetsoftware
Créateur:
Dah
Hôte:avatar
Somali is a Cushitic language of the Horn of Africa with ~25 million speakers, yet no documented dedicated Somali pretraining corpus with a companion tokenizer and language-identification benchmark has been publicly released. Existing Somali text appears either inside multilingual distributions (HPLT v2, CC100, MADLAD-400, OSCAR, mC4) or in small, undocumented Somali-only uploads on Hugging Face. We introduce SomaliWeb v1, a quality-filtered Somali corpus of 819,322 documents (~303M tokens) built from three upstream sources (HPLT v2, CC100, Somali Wikipedia) through a six-stage reproducible pipeline. We release (i) the corpus, (ii) a matched BPE-16K tokenizer, and (iii) the first public side-by-side Somali benchmark of three production language identifiers. Our measurements reveal concrete quality defects in existing distributions: HPLT v2's "cleaned" Somali release retains 17.3% byte-exact duplicates, 56.1% of its documents contain fixable mojibake, and 10.7% of its byte-unique documents are near-duplicates at Jaccard tau=0.80. Our BPE-16K tokenizer emits 40.2% fewer tokens than GPT-4's cl100k_base on FLORES-200 Somali devtest as a tokenizer-level measurement; downstream language-model perplexity comparisons are deferred to a follow-up release. 16 pages, 6 figures, 6 tables. Code: khaledyusuf44/somali-corpus Dataset: SomaliWeb v1

Visit

arxiv.org

Tasks

language identificationlanguage modeling

Languages

Somali

Tags

Computation and LanguageArtificial IntelligenceInformation RetrievalI.2.7

Similaires

Somali Web Corpus V1SomaliWeb v1Morphologically-informed Somali Lemmatization Corpus built with a Web-based Crowdsourcing PlatformSomali Dialect Identification: A Low-Resource Benchmark for MAXAA TIRI and MAAY Using Machine and Deep LearningAnanseLabs-Org/ghana-english-speech-filtered-v1ZigZeug/wolof-tokenizer-v1

Somali Web Corpus V1

This dataset consists of clean, structured, and filtered Somali language text compiled from various

SomaliWeb v1

📄 Paper: arXiv:2605.18232 — SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokeni

Morphologically-informed Somali Lemmatization Corpus built with a Web-based Crowdsourcing Platform

Somali Dialect Identification: A Low-Resource Benchmark for MAXAA TIRI and MAAY Using Machine and Deep Learning

Abstract This study addresses the task of automatic dialect identification within the Soma

AnanseLabs-Org/ghana-english-speech-filtered-v1

Samples: 87,000 Total duration: 210.20 hours Format: WAV audio embedded via Hugging Face Audio featu

ZigZeug/wolof-tokenizer-v1