Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

LIMIT: Language Identification, Misidentification, and Translation using Hierarchical Models in 350+ Languages

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
AgaAlam, Md Mahfuz IbnAnastasopoulos, Antonios
Hôte:avatar
Knowing the language of an input text/audio is a necessary first step for using almost every NLP tool such as taggers, parsers, or translation systems. Language identification is a well-studied problem, sometimes even considered solved; in reality, due to lack of data and computational challenges, current systems cannot accurately identify most of the world's 7000 languages. To tackle this bottleneck, we first compile a corpus, MCS-350, of 50K multilingual and parallel children's stories in 350+ languages. MCS-350 can serve as a benchmark for language identification of short texts and for 1400+ new translation directions in low-resource Indian and African languages. Second, we propose a novel misprediction-resolution hierarchical model, LIMIt, for language identification that reduces error by 55% (from 0.71 to 0.32) on our compiled children's stories dataset and by 40% (from 0.23 to 0.14) on the FLORES-200 benchmark. Our method can expand language identification coverage into low-resource languages by relying solely on systemic misprediction patterns, bypassing the need to retrain large models from scratch. To appear at EMNLP 2023. 24 pages, 2 figures, 12 tables

Visit

arxiv.org

Tasks

language identification

Tags

Computation and LanguageArtificial Intelligence

Similaires

Goldfish: Monolingual Language Models for 350 LanguagesMachine Translation Hallucination Detection for Low and High Resource Languages using Large Language ModelsA Multilingual Sentiment Lexicon for Low-Resource Language Translation using Large Languages Models and Explainable AIInteractive Machine Translation with Large Language Models for Low-resource LanguagesfastText Language Identification Modelsautomatic language identification for berber and arabic languages using prosodic features

Goldfish: Monolingual Language Models for 350 Languages

For many low-resource languages, the only available language models are large multilingual models tr

Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models

Recent advancements in massively multilingual machine translation systems have significantly enhance

A Multilingual Sentiment Lexicon for Low-Resource Language Translation using Large Languages Models and Explainable AI

South Africa and the Democratic Republic of Congo (DRC) present a complex linguistic landscape with

Interactive Machine Translation with Large Language Models for Low-resource Languages

Large language models (LLM) have been applied to machine translation with notable success. However,

fastText Language Identification Models

We distribute two models for language identification, which can recognize 176 languages (see the list of ISO codes below). These models were trained on data from Wikipedia, Tatoeba and SETimes, used under CC-BY-SA.

automatic language identification for berber and arabic languages using prosodic features