Logo Lanfrica

CommonLID

Domaine:

natural language processing

Type de record:

dataset
Créateur:
com
Hôte:
CommonLID is a community-created language identification (LID) benchmark. CommonLID consists of web text manually annotated for the language that it is written in. CommonLID contains annotations for 109 languages, where 78 of those languages have at least 100 lines of data. The number of lines available for each language is provided in Appendix A of the preprint. Dataset construction details