Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

Domaine:

natural language processing

Type de record:

datasetpaper

Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data often used to train multilingual language models. In this paper, we introduce CommonLID, a community-driven, human-annotated LID benchmark for the web domain, covering 109 languages. Many of the included languages have been previously under-served, making CommonLID a key resource for developing more representative high-quality text corpora. We show CommonLID's value by using it, alongside five other common evaluation sets, to test eight popular LID models. We analyse our results to situate our contribution and to provide an overview of the state of the art. In particular, we highlight that existing evaluations overestimate LID accuracy for many languages in the web domain. We make CommonLID and the code used to create it available under an open, permissive license.

Visit

arxiv.org

Tasks

language identification

Languages

AfrikaansAmazighAmharicArabic, Egyptian SpokenArabic, Moroccan SpokenArabic, Sudanese SpokenFulfulde, NigerianGandaGikuyuGun+19

Similaires

CommonLIDA State of the Art Review on Natural Language Processing applied to the Malagasy LanguageEVALUATING THE EFFECTIVENESS OF GREEN PROCUREMENT STRATEGIES ON STATE CORPORATION PERFORMANCE IN RWANDAIdentification of Cultural Markers on Tunisian Web SitesContinued Pretraining for Low-Resource Swahili ASR: Achieving State-of-the-Art Performance with Minimal Labeled DataVegetative Identification of Tropical Woody Plants: State of the Art and Annotated Bibliography<sup>1</sup>

CommonLID

CommonLID is a community-created language identification (LID) benchmark. CommonLID consists of web

A State of the Art Review on Natural Language Processing applied to the Malagasy Language

Based on the growing mass of information of all kinds to be processed and to facilitate human/machin

EVALUATING THE EFFECTIVENESS OF GREEN PROCUREMENT STRATEGIES ON STATE CORPORATION PERFORMANCE IN RWANDA

Green procurement has become a critical strategy for organizations that aim to achieve environmental

Identification of Cultural Markers on Tunisian Web Sites

Continued Pretraining for Low-Resource Swahili ASR: Achieving State-of-the-Art Performance with Minimal Labeled Data

We investigate continued pretraining (CPT) for adapting wav2vec2-bert-2.0 to Swahili automatic speec

Vegetative Identification of Tropical Woody Plants: State of the Art and Annotated Bibliography<sup>1</sup>

ABSTRACT This annotated bibliography is provided in order to assess the achievements and gaps in th