Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study

Domain:

natural language processing

Record type:

papermodel
Creator:
KhaSaiVanKho
Host:avatar
Despite the widespread adoption of Large Language Models (LLMs), their strongest capabilities remain largely confined to a small number of high-resource languages for which there is abundant training data. Recently, continual pre-training (CPT) has emerged as a means to fine-tune these models to low-resource regional dialects. In this paper, we study the use of CPT for dialect learning under tight data and compute budgets. Using low-rank adaptation (LoRA) and compute-efficient continual pre-training, we adapt three LLMs to the Québec French dialect using a very small dataset and benchmark them on the COLE suite. Our experiments demonstrate an improvement on the minority dialect benchmarks with minimal regression on the prestige language benchmarks with around 1% of model parameters updated. Analysis of the results demonstrate that gains are highly contingent on corpus composition. These findings indicate that CPT with parameter-efficient fine-tuning (PEFT) can narrow the dialect gap by providing cost-effective and sustainable language resource creation, expanding high-quality LLM access to minority linguistic communities. To support reproducibility and broaden access, we release the first Québec French LLMs on Hugging Face. Accepted at LREC 2026

Visit

arxiv.org

Tasks

language modelingtransfer learning

Tags

Computation and LanguageArtificial Intelligence

Similar

Atlas-Chat: Adapting Large Language Models for Low-Resource Moroccan Arabic DialectLarge Language Models Adaptation for Low-resource Languages: The Case for African LanguagesImproving Low-Resource Dialect Identification through Audio Data Augmentation: A Case Study on Algerian LanguageEvaluation of Arabic Large Language Models on Moroccan DialectEmploying large language models in Swahili, a low-resource languageNeural Machine Translation for French–Mooré: Adapting Large Language Models to Low-Resource Languages

Atlas-Chat: Adapting Large Language Models for Low-Resource Moroccan Arabic Dialect

We introduce Atlas-Chat, the first-ever collection of LLMs specifically developed for dialectal Arab

Large Language Models Adaptation for Low-resource Languages: The Case for African Languages

David Ifeoluwa Adelani (Supervisor) Despite remarkable advances in Large Language Models (LLMs), Afr

Improving Low-Resource Dialect Identification through Audio Data Augmentation: A Case Study on Algerian Language

Evaluation of Arabic Large Language Models on Moroccan Dialect

Large Language Models (LLMs) have shown outstanding performance in many Natural Language Processing

Employing large language models in Swahili, a low-resource language

Neural Machine Translation for French–Mooré: Adapting Large Language Models to Low-Resource Languages