Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities

Domaine:

natural language processing

Type de record:

papermodeldataset
Créateur:
MekAtoNacShe
Hôte:avatar
Enhancing the linguistic capabilities of Large Language Models (LLMs) to include low-resource languages is a critical research area. Current research directions predominantly rely on synthetic data generated by translating English corpora, which, while demonstrating promising linguistic understanding and translation abilities, often results in models aligned with source language culture. These models frequently fail to represent the cultural heritage and values of local communities. This work proposes a methodology to create both synthetic and retrieval-based pre-training data tailored to a specific community, considering its (i) language, (ii) cultural heritage, and (iii) cultural values. We demonstrate our methodology using Egyptian and Moroccan dialects as testbeds, chosen for their linguistic and cultural richness and current underrepresentation in LLMs. As a proof-of-concept, we develop NileChat, a 3B parameter Egyptian and Moroccan Arabic LLM adapted for Egyptian and Moroccan communities, incorporating their language, cultural heritage, and values. Our results on various understanding, translation, and cultural and values alignment benchmarks show that NileChat outperforms existing Arabic-aware LLMs of similar size and performs on par with larger models. This work addresses Arabic dialect in LLMs with a focus on cultural and values alignment via controlled synthetic data generation and retrieval-augmented pre-training for Moroccan Darija and Egyptian Arabic, including Arabizi variants, advancing Arabic NLP for low-resource communities. We share our methods, data, and models with the community to promote the inclusion and coverage of more diverse communities in cultural LLM development: github.com . Accepted to EMNLP 2025 (Main Conference). Camera-ready version. Data & models: github.com

Visit

arxiv.org

Tasks

language modeling

Languages

Arabic, Algerian SpokenArabic, Moroccan Spoken

Tags

Computation and Language

Similaires

Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMsWorking towards Culturally and Linguistically Diverse Speech Assessments for South African Children: A Xhosa Case StudyCreating identity texts with young children across culturally and linguistically diverse contextsDemocratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse PromptsLocal Languages AI: Toward Linguistically Inclusive, Cognitively Diverse, and Regenerative Artificial IntelligenceA cultural neuropsychological approach to harmonization of cognitive data across culturally and linguistically diverse older adult populations.

Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs

As large language models (LLMs) become increasingly integrated into daily life, ensuring their cultu

Working towards Culturally and Linguistically Diverse Speech Assessments for South African Children: A Xhosa Case Study

Creating identity texts with young children across culturally and linguistically diverse contexts

This article addresses the ways young children in culturally and linguistically diverse settings wer

Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse Prompts

Large language models (LLMs) are known to effectively perform tasks by simply observing few exemplar

Local Languages AI: Toward Linguistically Inclusive, Cognitively Diverse, and Regenerative Artificial Intelligence

Artificial Intelligence (AI) has become one of the most influential cognitive infrastructures of the

A cultural neuropsychological approach to harmonization of cognitive data across culturally and linguistically diverse older adult populations.