Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Contrastive Cross-Lingual Calibration for Large Language Models

Domain:

natural language processing

Record type:

paper
Creator:
Mar
Publisher:
Qei
Host:
Large language models (LLMs) are increasingly deployed in multilingual settings, yet their probability estimates are often miscalibrated, particularly for low-resource languages and code-switched inputs. We present C³, a post-hoc calibration framework that reduces cross-lingual miscalibration by optimizing language-aware temperature and bias parameters using contrastive counterfactuals generated via translation/back-translation and meaning-preserving perturbations. C³ aligns confidence across languages without retraining the base model. On classification (XNLI) and extractive/generative QA (XQuAD, MLQA, TyDi QA GoldP), C³ lowers Expected Calibration Error by 35–57% and Brier score by 9–18%, with modest accuracy gains (0.7–2.1 pp). For generative QA, hallucination rate decreases by 21% while maintaining answer quality. Benefits are largest for Swahili and Arabic and persist under code-switch and spelling noise. Ablations show that contrastive counterfactuals and language-specific scaling both contribute, and isotonic fusion improves tails of the confidence distribution. We release calibration recipes and evaluation scripts to support responsible multilingual deployment.

Visit

doi.org

Languages

Swahili

Licenses

https://creativecommons.org/licenses/by/4.0

Similar

BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual TransferContrastive Learning for Cross-Lingual Alignment and Robustness in Multimodal ModelsZero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource LanguagesCode-Switching In-Context Learning for Cross-Lingual Transfer of Large Language ModelsMulti-lingual Functional Evaluation for Large Language ModelsFew-Shot Cross-Lingual Transfer for Prompting Large Language Models in Low-Resource Languages

BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer

Despite remarkable advancements in few-shot generalization in natural language processing, most mode

Contrastive Learning for Cross-Lingual Alignment and Robustness in Multimodal Models

Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potentia

Zero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource Languages

Large language models (LLMs) have shown impressive zero-shot capabilities in various document rerank

Code-Switching In-Context Learning for Cross-Lingual Transfer of Large Language Models

While large language models (LLMs) exhibit strong multilingual abilities, their reliance on English

Multi-lingual Functional Evaluation for Large Language Models

Multi-lingual competence in large language models is often evaluated via static data benchmarks such

Few-Shot Cross-Lingual Transfer for Prompting Large Language Models in Low-Resource Languages

Large pre-trained language models (PLMs) are at the forefront of advances in Natural Language Proces