Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment

Domain:

natural language processing

Record type:

paperdataset
Creator:
Ye,SunZhaZha
Host:avatar
Idioms are difficult to transfer across languages due to their non-compositionality and weak surface-form grounding, making literal mappings unreliable. We present G-IdiomAlign, a gloss-pivoted benchmark where each idiom is anchored by an English gloss from Wiktionary. We further construct a high-confidence reference alignment set for reproducible evaluation. G-IdiomAlign supports two protocols: (1) a controlled Multiple-Choice Idiom Equivalence with typed distractors for error attribution; and (2) a Gloss-Contrastive Generation contrasting No-gloss and With-gloss inputs to isolate the effect of an explicit semantic pivot. Across diverse LLMs, a bias to literal translation is a dominant failure mode, especially when the target is a low-resource language. Glosses consistently improve Gloss-Contrastive Generation under an embedding-based semantic proxy, but performance remains modest, indicating substantial headroom in the open output space. Subsequent analysis on Qwen3-8B further suggests that cross-condition differences are concentrated more in attention heads than in layers, while better With-gloss generations coincide with stronger gloss anchoring. Accepted to ACL 2026

Visit

arxiv.org

Tags

Computation and LanguageArtificial Intelligence

Similar

POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text TranslationXTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralizationOptimal Transport Alignment for Reducing Cross-Lingual Retrieval Performance GapsOptimal Transport Distillation for Cross-Lingual Performance Alignment in XQuADOptimal Transport Distillation for Cross-Lingual Alignment Robustness in MIRACLDomain Adaptation for Cross-Lingual NER Alignment with Synthetic Noise

POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation

Speech Large Language Models have achieved breakthroughs in multilingual speech-to-text translation.

XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization

The Cross-lingual Natural Language Inference (XNLI) corpus is a crowd-sourced collection of 5,000 test and 2,500 dev pairs for the MultiNLI corpus. The pairs are annotated with textual entailment and translated into 14 languages: French, Spanish, German, Greek, Bu

Optimal Transport Alignment for Reducing Cross-Lingual Retrieval Performance Gaps

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Optimal Transport Distillation for Cross-Lingual Performance Alignment in XQuAD

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Optimal Transport Distillation for Cross-Lingual Alignment Robustness in MIRACL

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Domain Adaptation for Cross-Lingual NER Alignment with Synthetic Noise

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident