Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe

Domain:

natural language processing

Record type:

paperdataset
Creator:
AdjEiselen, RoaldMit
Host:avatar
Large language models (LLMs) are trained on data contributed by low-resource language communities, yet the linguistic knowledge encoded in these models remains accessible only through commercial APIs. This paper investigates whether strategic prompting can extract usable text data from LLMs for two West African languages: Hausa (Afroasiatic, approximately 80 million speakers) and Fongbe (Niger-Congo, approximately 2 million speakers). We systematically compare six elicitation task types across two commercial LLMs (GPT-4o Mini and Gemini 2.5 Flash). GPT-4o Mini extracts 6-41 times more usable target-language words per API call than Gemini. Optimal strategies differ by language: Hausa benefits from functional text and dialogue, while Fongbe requires constrained generation prompts. We release all generated corpora and code. 11 pages, 5 figures, 6 tables; to appear in LREC-COLING 2026

Visit

arxiv.org

Languages

FonHausa

Tags

Computation and LanguageArtificial IntelligenceI.2.7; H.3.1; I.2.0

Similar

Post-ASR Correction in Hindi: Comparing Language Models and Large Language Models in Low-Resource ScenariosEvaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric ReliabilityTharuChat: Bootstrapping Large Language Models for a Low-Resource Language via Synthetic Data and Human ValidationAdaptive and Efficient Large Language Models for Low-Resource African LanguagesEvaluating Quantized Large Language Models for Code Generation on Low-Resource Language BenchmarksContinual-learning for Modelling Low-Resource Languages from Large Language Models

Post-ASR Correction in Hindi: Comparing Language Models and Large Language Models in Low-Resource Scenarios

Automatic Speech Recognition (ASR) systems for low-resource languages like Hindi often produce erron

Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability

We investigate the translation quality of current large language models (LLMs) for English-to-Hausa

TharuChat: Bootstrapping Large Language Models for a Low-Resource Language via Synthetic Data and Human Validation

The rapid proliferation of Large Language Models (LLMs) has created a profound digital divide, effec

Adaptive and Efficient Large Language Models for Low-Resource African Languages

PAIDeF SuperAI 2025 Conference

Adaptive and Efficient Large Language Mod

Evaluating Quantized Large Language Models for Code Generation on Low-Resource Language Benchmarks

Democratization of AI is an important topic within the broader topic of the digital divide. This iss

Continual-learning for Modelling Low-Resource Languages from Large Language Models

Modelling a language model for a multi-lingual scenario includes several potential challenges, among