Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

LKValues: Aligning Large Language Models with Sri Lankan Societal Values

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
MutSupRande
Éditeur:
arXiv
Hôte:avatar
Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societies such as Sri Lanka that have their unique cultural dynamics. Existing benchmarks overlook Sri Lankan-contextualized values in its official language Sinhala, hindering culturally sensitive evaluation and fine-tuning. To bridge this gap, we propose LKValues, the first survey-grounded resource suite for Sri Lankan value alignment. From a trilingual survey of 205 respondents, blending adapted global frameworks and LLM-elicited local constructs, we derive 40 majority-endorsed societal values. Using these values, we construct LKvaluesIT, a Sinhala-English news-derived instruction corpus containing 150k scenario-based instances, and LKvaluesBench, a value-sensitive evaluation benchmark of 1,000 instances. We evaluate a set of proprietary and open-weight LLMs with LKvaluesBench. We fine-tune three open-weight base models (Qwen3.5-4B-Base, Qwen3.5-9B-Base, and Aya-Expanse-8B-Base). Our experiments show that newer and larger LLMs still exhibit low-resource and cultural value-alignment gaps. LKValues fine-tuning improves Qwen-family models in English and Sinhala, reducing invalid outputs and cross-lingual disparities, though gains remain model-family dependent. These highlight LKValues efficacy in embedding Sri Lankan values, offering a replicable pipeline for low-resource, country-specific pluralist value alignment. The dataset is publicly available at github.com. 37 pages, 10 figures, and 15 tables. Includes appendices. Datasets are available at the project repository

Visit

doi.org

Tags

Computation and Language (cs.CL)FOS: Computer and information sciences

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/

Similaires

CeyNews: A Sri Lankan Trilingual News CorpusAdaptive Machine Translation with Large Language ModelsFinancial Product Ontology Population with Large Language ModelsStory Generation with Large Language Models for African LanguagesTech Meets Tradition: A Nigerian Perspective on Aligning AI Development with African Religious ValuesEthiopicEmotion: Multi-label Emotion Dataset with Large Language Models Evaluation

CeyNews: A Sri Lankan Trilingual News Corpus

CeyNews: A Trilingual Sri Lankan News Corpus This Zenodo release contains the CeyNews dataset, a la

Adaptive Machine Translation with Large Language Models

Consistency is a key requirement of high-quality translation. It is especially important to adhere t

Financial Product Ontology Population with Large Language Models

Ontology population, which aims to extract structured data to enrich domain-specific ontologies from

Story Generation with Large Language Models for African Languages

Tech Meets Tradition: A Nigerian Perspective on Aligning AI Development with African Religious Values

EthiopicEmotion: Multi-label Emotion Dataset with Large Language Models Evaluation

Large Language Models (LLMs) show promising learning and reasoning abilities. Compared to other NLP