Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

CommonLID

Domain:

natural language processing

Record type:

dataset
Creator:
com
Host:
CommonLID is a community-created language identification (LID) benchmark. CommonLID consists of web text manually annotated for the language that it is written in. CommonLID contains annotations for 109 languages, where 78 of those languages have at least 100 lines of data. The number of lines available for each language is provided in Appendix A of the preprint. Dataset construction details

Visit

huggingface.co

Tasks

language identification

Languages

AfrikaansAmazighAmharicArabic, Egyptian SpokenArabic, Moroccan SpokenArabic, Sudanese SpokenArabic, Tunisian SpokenFulfulde, NigerianGandaGikuyu+18

Tags

text

Licenses

other

Similar

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data often used to train multilingual language models. In this paper,