Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

HealMed: Multilingual Evaluation of Large Language Models in Medicine

Domain:

natural language processinghealthcare

Record type:

datasetpaper
Creator:
CheGaoTonZha
Publisher:
arXiv
Host:avatar
We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions. Each translation was evaluated and revised by two experts fluent in English and the corresponding target language. On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models. The strongest proprietary models were the most stable across languages, whereas many open-source and medically specialized models showed larger and less consistent gaps. Medical specialization alone did not ensure multilingual robustness. Furthermore, expert revision could either raise or lower measured performance, indicating that translation quality materially affects cross-language evaluation results.

Visit

doi.org

Tasks

question answering

Tags

Computation and Language (cs.CL)FOS: Computer and information sciences

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/

Similar

MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language ModelsToward Global Large Language Models in MedicineGlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language ModelsEthioLLM: Multilingual Large Language Models for Ethiopian Languages with Task EvaluationQuantifying Language Disparities in Multilingual Large Language ModelsBridging language gaps in multilingual large language models

MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models

Parameter Efficient Finetuning (PEFT) has emerged as a viable solution for improving the performance

Toward Global Large Language Models in Medicine

Despite continuous advances in medical technology, the global distribution of health care resources

GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models

Large language models (LLMs) are advancing at an unprecedented pace globally, with regions increasin

EthioLLM: Multilingual Large Language Models for Ethiopian Languages with Task Evaluation

Large language models (LLMs) have gained popularity recently due to their outstanding performance in

Quantifying Language Disparities in Multilingual Large Language Models

Results reported in large-scale multilingual evaluations are often fragmented and confounded by fact

Bridging language gaps in multilingual large language models

Large language models (LLMs) have revolutionized natural language processing, yet significant perfor