Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

LexC-Gen: Generating Data for Extremely Low-Resource Languages with Large Language Models and Bilingual Lexicons

Domaine:

natural language processing

Type de record:

paper
Créateur:
YonMenBac
Hôte:avatar
Data scarcity in low-resource languages can be addressed with word-to-word translations from labeled task data in high-resource languages using bilingual lexicons. However, bilingual lexicons often have limited lexical overlap with task data, which results in poor translation coverage and lexicon utilization. We propose lexicon-conditioned data generation LexC-Gen, a method that generates low-resource-language classification task data at scale. Specifically, LexC-Gen first uses high-resource-language words from bilingual lexicons to generate lexicon-compatible task data, and then it translates them into low-resource languages with bilingual lexicons via word translation. Across 17 extremely low-resource languages, LexC-Gen generated data is competitive with expert-translated gold data, and yields on average 5.6 and 8.9 points improvement over existing lexicon-based word translation methods on sentiment analysis and topic classification tasks respectively. Through ablation study, we show that conditioning on bilingual lexicons is the key component of LexC-Gen. LexC-Gen serves as a potential solution to close the performance gap between open-source multilingual models, such as BLOOMZ and Aya-101, and state-of-the-art commercial models like GPT-4o on low-resource-language tasks. EMNLP Findings 2024

Visit

arxiv.org

Tags

Computation and LanguageArtificial IntelligenceMachine Learning

Similaires

Dict-NMT: Bilingual Dictionary based NMT for Extremely Low Resource LanguagesInteractive Machine Translation with Large Language Models for Low-resource LanguagesEffective vocabulary expansion of multilingual language models for extremely low-resource languagesBilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource LanguagesScaling Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource LanguagesZero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource Languages

Dict-NMT: Bilingual Dictionary based NMT for Extremely Low Resource Languages

Neural Machine Translation (NMT) models have been effective on large bilingual datasets. However, th

Interactive Machine Translation with Large Language Models for Low-resource Languages

Large language models (LLM) have been applied to machine translation with notable success. However,

Effective vocabulary expansion of multilingual language models for extremely low-resource languages

Multilingual pre-trained language models(mPLMs) offer significant benefits for many low-resource lan

Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Zero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource Languages

Large language models (LLMs) have shown impressive zero-shot capabilities in various document rerank