Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Cross-Lingual Stability and Bias in Instruction-Tuned Language Models for Humanitarian NLP

Domain:

natural language processingpeace and security

Record type:

paper
Creator:
NemAdhPeaSad
Host:avatar
Humanitarian organizations face a critical choice: invest in costly commercial APIs or rely on free open-weight models for multilingual human rights monitoring. While commercial systems offer reliability, open-weight alternatives lack empirical validation -- especially for low-resource languages common in conflict zones. This paper presents the first systematic comparison of commercial and open-weight large language models (LLMs) for human-rights-violation detection across seven languages, quantifying the cost-reliability trade-off facing resource-constrained organizations. Across 78,000 multilingual inferences, we evaluate six models -- four instruction-aligned (Claude-Sonnet-4, DeepSeek-V3, Gemini-Flash-2.0, GPT-4.1-mini) and two open-weight (LLaMA-3-8B, Mistral-7B) -- using both standard classification metrics and new measures of cross-lingual reliability: Calibration Deviation (CD), Decision Bias (B), Language Robustness Score (LRS), and Language Stability Score (LSS). Results show that alignment, not scale, determines stability: aligned models maintain near-invariant accuracy and balanced calibration across typologically distant and low-resource languages (e.g., Lingala, Burmese), while open-weight models exhibit significant prompt-language sensitivity and calibration drift. These findings demonstrate that multilingual alignment enables language-agnostic reasoning and provide practical guidance for humanitarian organizations balancing budget constraints with reliability in multilingual deployment.

Visit

arxiv.org

Tasks

text classification

Languages

Lingala

Tags

Computation and LanguageArtificial Intelligence

Similar

Cendol: Open Instruction-tuned Generative Large Language Models for Indonesian LanguagesCross-Lingual Bias in Large Language Models: A Comparative Analysis of English and SwahiliCross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective StabilityTednn493/Cross-Lingual-Bias-in-Large-Language-Models-A-Comparative-Analysis-of-English-and-SwahiliCross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on RomanianContrastive Cross-Lingual Calibration for Large Language Models

Cendol: Open Instruction-tuned Generative Large Language Models for Indonesian Languages

Large language models (LLMs) show remarkable human-like capability in various domains and languages.

Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili

Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability

Recent advances in multilingual representation learning aim to bridge the performance gap between hi

Tednn493/Cross-Lingual-Bias-in-Large-Language-Models-A-Comparative-Analysis-of-English-and-Swahili

# Cross Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili ## Des

Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian

Relation extraction (RE) for low-resource languages is typically constrained by the lack of annotate

Contrastive Cross-Lingual Calibration for Large Language Models

Large language models (LLMs) are increasingly deployed in multilingual settings, yet their probabili