Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages

Domain:

natural language processing

Record type:

paperdataset
Creator:
AlmElbPraNik
Host:avatar
Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpretation. Prior work has focused on high-resource languages typically evaluates isolated idiom-meaning questions, overlooking realistic discourse. We introduce MIDI, a multilingual idiom dataset spanning 3 high-, 3 medium-, and 12 low-resource languages, curated by native speakers. Unlike previous datasets, MIDI provides idioms embedded in both sentence-level and conversational contexts, capturing both literal and figurative readings. Benchmarking state-of-the-art models shows that idiom comprehension degrades in low-resource languages and that, in all resource tiers, literal interpretations are substantially harder than figurative ones. Conversational context improves performance but does not eliminate these disparities. Through controlled tests and interventions on hidden representations, we further separate memorization from reasoning, exposing core limitations of current models.

Visit

arxiv.org

Tags

Computation and LanguageArtificial Intelligence

Similar

Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource LanguagesSenWiCh: Sense-Annotated Sentences for WSD and WiC in Low-Resource LanguagesLanguage Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource LanguagesMulti-Method Validation of Large Language Model Medical Translation Across High- and Low-Resource LanguagesMultilingual Knowledge Graphs and Low-Resource Languages: A ReviewNatural Language Processing (NLP) tools - multilingual and low-resource languages

Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages

Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, ye

SenWiCh: Sense-Annotated Sentences for WSD and WiC in Low-Resource Languages

SenWiCh is a multilingual dataset of sense-annotated sentences designed to support research in Word

Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages

The development of Large Language Models (LLMs) relies on extensive text corpora, which are often un

Multi-Method Validation of Large Language Model Medical Translation Across High- and Low-Resource Languages

Language barriers affect 27.3 million U.S. residents with non-English language preference, yet profe

Multilingual Knowledge Graphs and Low-Resource Languages: A Review

There is a lack of multilingual data to support applications in a large number of languages, especia

Natural Language Processing (NLP) tools - multilingual and low-resource languages