Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

TharuChat: Bootstrapping Large Language Models for a Low-Resource Language via Synthetic Data and Human Validation

Domaine:

natural language processing

Type de record:

papermodeldataset
Créateur:
PanMai
Hôte:avatar
The rapid proliferation of Large Language Models (LLMs) has created a profound digital divide, effectively excluding indigenous languages of the Global South from the AI revolution. The Tharu language, an Indo-Aryan vernacular spoken by approximately 1.7 million people across the Terai belt of Nepal and India, exemplifies this crisis. Despite a rich oral tradition, Tharu suffers from severe data scarcity and linguistic fragmentation, causing state-of-the-art multilingual models to routinely "hallucinate" or default to dominant high-resource neighbors like Hindi and Nepali due to contamination in pre-training corpora. This paper presents Tharu-LLaMA (3B), a specialized instruction-following model designed to address this exclusion. We introduce TharuChat, a novel dataset constructed via a LLM-to-Human bootstrapping pipeline. We utilized prompt-engineered Gemini models, fed with Rana Tharu grammar and folklore, to synthesize training data. Unlike curated gold-standard corpora, TharuChat reflects the noisy, heterogeneous linguistic reality of the region: it is predominantly anchored in Rana Tharu (~70%) while integrating elements of Dangaura and Kochila dialects. We provide a transparent analysis of the dataset's limitations, including dialectal code-mixing and residual Awadhi/Hindi influence. Through a rigorous empirical ablation study, we demonstrate that despite these imperfections, small-scale synthetic data is highly effective, increasing the dataset volume from 25% to 100% results in a linear reduction in perplexity from 6.42 to 2.88. The resulting model serves as a proof-of-concept for the preservation of under-resourced Himalayan languages via generative AI, achievable on consumer-grade hardware. 6 pages, 1 figure, 2 tables. Preprint. Code and dataset available on Hugging Face

Visit

arxiv.org

Tasks

language modeling

Tags

Computation and LanguageArtificial IntelligenceMachine Learning

Similaires

NagaNLP: Bootstrapping NLP for Low-Resource Nagamese Creole with Human-in-the-Loop Synthetic DataHuman review for post-training improvement of low-resource language performance in large language modelsEmploying large language models in Swahili, a low-resource languageMining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbeshruthi0429/BLOOMZ-and-mT5-Fine-Tuning-Optimizing-Large-Language-Models-for-a-Low-Resource-LanguageAdaptive and Efficient Large Language Models for Low-Resource African Languages

NagaNLP: Bootstrapping NLP for Low-Resource Nagamese Creole with Human-in-the-Loop Synthetic Data

The vast majority of the world's languages, particularly creoles like Nagamese, remain severely unde

Human review for post-training improvement of low-resource language performance in large language models

Large language models (LLMs) have significantly improved natural language processing, holding the p

Employing large language models in Swahili, a low-resource language

Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe

Large language models (LLMs) are trained on data contributed by low-resource language communities, y

shruthi0429/BLOOMZ-and-mT5-Fine-Tuning-Optimizing-Large-Language-Models-for-a-Low-Resource-Language

# BLOOMZ and mT5 Fine-Tuning: Optimizing Large Language Models for a Low-Resource Language This rep

Adaptive and Efficient Large Language Models for Low-Resource African Languages

PAIDeF SuperAI 2025 Conference

Adaptive and Efficient Large Language Mod