Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Mind the Gap: Evaluating the Representativeness of Quantitative Medical Language Reasoning LLM Benchmarks for African Disease Burdens

Domaine:

natural language processinghealthcare

Type de record:

paperdataset
Créateur:
MutGitSyoOig
Hôte:avatar
Introduction: Existing medical LLM benchmarks largely reflect examination syllabi and disease profiles from high income settings, raising questions about their validity for African deployment where malaria, HIV, TB, sickle cell disease and other neglected tropical diseases (NTDs) dominate burden and national guidelines drive care. Methodology: We systematically reviewed 31 quantitative LLM evaluation papers (Jan 2019 May 2025) identifying 19 English medical QA benchmarks. Alama Health QA was developed using a retrieval augmented generation framework anchored on the Kenyan Clinical Practice Guidelines. Six widely used sets (AfriMedQA, MMLUMedical, PubMedQA, MedMCQA, MedQAUSMLE, and guideline grounded Alama Health QA) underwent harmonized semantic profiling (NTD proportion, recency, readability, lexical diversity metrics) and blinded expert rating across five dimensions: clinical relevance, guideline alignment, clarity, distractor plausibility, and language/cultural fit. Results: Alama Health QA captured >40% of all NTD mentions across corpora and the highest within set frequencies for malaria (7.7%), HIV (4.1%), and TB (5.2%); AfriMedQA ranked second but lacked formal guideline linkage. Global benchmarks showed minimal representation (e.g., sickle cell disease absent in three sets) despite large scale. Qualitatively, Alama scored highest for relevance and guideline alignment; PubMedQA lowest for clinical utility. Discussion: Quantitative medical LLM benchmarks widely used in the literature underrepresent African disease burdens and regulatory contexts, risking misleading performance claims. Guideline anchored, regionally curated resources such as Alama Health QA and expanded disease specific derivatives are essential for safe, equitable model evaluation and deployment across African health systems. Preprint. 26 pages, includes appendix and tables

Visit

arxiv.org

Tasks

question answering

Tags

Artificial Intelligence

Similaires

Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural AdjustmentsMind the adoption gapMind the gap: translation in a fractured African societyEvaluating the representativeness of the Setswana corpus using behavioral datalucianompilo3-cyber/mind-the-gap-sa-employmentWho Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA

Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments

Large Language Models (LLMs) have shown remarkable performance across various tasks, yet significant

Mind the adoption gap

A field experiment designed to scale up the availability of fodder shrub seedlings in Malawi

Mind the gap: translation in a fractured African society

The spaces and tensions between races, ethnic groups, and communities in late apartheid and post-199

Evaluating the representativeness of the Setswana corpus using behavioral data

This paper presents efforts to evaluate the representativeness of the Setswana corpus data with meas

lucianompilo3-cyber/mind-the-gap-sa-employment

Analysis of South African provincial unemployment trends using Stats SA QLFS Q1 2026 data. Built wit

Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA

International audience

Automatic evaluation of medical open-ended question an