Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Pak3H: Evaluating the Cost of Cultural Mismatch in LLM Alignment with a Human-Contextualized Urdu Benchmark

Domain:

natural language processing

Record type:

datasetpaper
Creator:
HasNasRaz
Publisher:
arXiv
Host:avatar
Large language models (LLMs) demonstrate strong Helpfulness, Harmlessness, and Honesty (3H) alignment in English-centric settings, but these gains transfer poorly to low-resource languages due to cultural mismatches. Existing multilingual 3H benchmarks rely predominantly on automated translation or LLM based synthesis, propagating source-language biases while sacrificing local relevance. To address this gap, we introduce Pak3H1, the first human-validated, culturally contextualized Urdu benchmark suite for 3H alignment, comprising PakAlpaca (helpfulness), PakBeaverTails (harmlessness), and PakTruthfulQA (honesty). Our multi-stage pipeline integrates manual cultural adaptation and dictionary-guided post editing to prioritize native speaker judgment, ensuring both semantic fidelity and contextual authenticity. Zero-shot evaluations across multiple open and proprietary LLM architectures reveal systematic cross-lingual alignment gaps: helpfulness win rates decline under localized contexts, harmlessness guardrails break down against regional safety risks, and composite honesty metrics degrade substantially due to localized factual constraints. These findings expose structural limitations in current alignment approaches, underscoring the necessity of human-guided localization for equitable multilingual evaluation. We introduce Pak3H, a human-validated Urdu benchmark for helpfulness, harmlessness, and honesty. Zero-shot evaluations show LLM performance degrades across all three dimensions in low-resourced contextualized settings

Visit

doi.org

Tags

Computation and Language (cs.CL)Artificial Intelligence (cs.AI)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Kambaata LLM Cultural BenchmarkUrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in UrduBehailuBerhanu/kambaata-llm-cultural-benchmarkI Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMsKhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Recordsrifaasa/hassaniya-llm-benchmark

Kambaata LLM Cultural Benchmark

Maintenance release for Zenodo archival of the Kambaata LLM Cultural Benchmark. This release contain

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu

Multilingual Large Language Models (LLMs) have shown remarkable performance across various languages

BehailuBerhanu/kambaata-llm-cultural-benchmark

A 77-item benchmark for evaluating cultural knowledge, hallucination, and epistemic behavior in larg

I Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs

We introduce MENAValues, a novel benchmark designed to evaluate the cultural alignment and multiling

KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records

Land ownership in Bangladesh is recorded in Ana-Ganda-Kora-Kranti-Til, a base-16 positional fraction

rifaasa/hassaniya-llm-benchmark

benchmark of modern LLMs on Hassaniya Arabic dialect » # hassaniya-llm-benchmark Code and evaluati