Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

M-IFEval: On Multilingual Instruction-Following Capability of Large Language Models

Domain:

natural language processing

Record type:

datasetpaper
Creator:
AssGenLi,Li,
Publisher:
Und
Host:avatar
Instruction-following capability has become a major ability to be evaluated for Large Language Models. However, existing datasets, such as IFEval, are either predominantly monolingual and centered on English or simply machine translated to other languages, limiting their applicability in multilingual contexts. In this paper, we present an carefully-curated extension of IFEval to a localized multilingual version named Marco-Bench-MIF, covering 30 languages with varying levels of localization. Our benchmark addresses linguistic constraints (e.g., modifying capitalization requirements for Chinese) and cultural references (e.g., substituting region-specific company names in prompts) via a hybrid pipeline combining translation with verification. Through comprehensive evaluation of 20+ LLMs on our Marco-Bench-MIF, we found that: (1) 25-35% accuracy gap between high/low-resource languages, (2) model scales largely impact performance by 45-60% yet persists script-specific challenges, and (3) machine-translated data underestimates accuracy by 7-22% versus localized data. Our analysis identifies challenges in multilingual instruction following, including keyword consistency preservation and compositional constraint adherence across languages. Our Marco-Bench-MIF will be made publicly available to the community.

Visit

doi.orgunderline.io

Tags

Computational LinguisticsNatural Language ProcessingArtificial Intelligence

Similar

Evaluating the capability of base and large-scale language models for multilingual sarcasm detectionM-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language ModelsQuantifying Language Disparities in Multilingual Large Language ModelsBridging language gaps in multilingual large language modelsRomanization-based Large-scale Adaptation of Multilingual Language ModelsHealMed: Multilingual Evaluation of Large Language Models in Medicine

Evaluating the capability of base and large-scale language models for multilingual sarcasm detection

Even though natural language understanding has made significant progress, language models still stru

M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models

Multilingual language models are deployed across a hundred or more languages, yet most benchmarks te

Quantifying Language Disparities in Multilingual Large Language Models

Results reported in large-scale multilingual evaluations are often fragmented and confounded by fact

Bridging language gaps in multilingual large language models

Large language models (LLMs) have revolutionized natural language processing, yet significant perfor

Romanization-based Large-scale Adaptation of Multilingual Language Models

Large multilingual pretrained language models (mPLMs) have become the de facto state of the art for

HealMed: Multilingual Evaluation of Large Language Models in Medicine

We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language model