Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Short Text Language Identification for Under Resourced Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
Duv
Éditeur:
arXiv
Hôte:avatar
The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text language identification (LID) useful for under resourced languages. The algorithm is evaluated on short pieces of text for the 11 official South African languages some of which are similar languages. The algorithm is compared to recent approaches using test sets from previous works on South African languages as well as the Discriminating between Similar Languages (DSL) shared tasks' datasets. Remaining research opportunities and pressing concerns in evaluating and comparing LID approaches are also discussed. Presented at NeurIPS 2019 Workshop on Machine Learning for the Developing World

Visit

doi.orgarxiv.org

Tasks

language identification

Tags

Computation and Language (cs.CL)FOS: Computer and information sciencesFOS: Computer and information sciences68T50

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/