Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

How Good Are Large Language Models at Supporting Frontline Healthcare Workers in Low-Resource Settings – A Benchmarking Study & Dataset

Domain:

healthcarenatural language processing

Record type:

datasetpaper
Creator:
SamGwyKleFra
Publisher:
ope
Host:
Abstract Large language models (LLMs) have demonstrated strong performance in medical contexts; however, existing benchmarks often fail to reflect the real-world complexity of low-resource health systems accurately. This study developed a dataset of 5,609 clinical questions contributed by 101 community health workers (CHWs) across four Rwandan districts and compared responses generated by five large language models (LLMs) (Gemini-2, GPT-4o, o3 mini, Deepseek R1, and Meditron-70B) with those from local clinicians. A subset of 524 question-answer pairs was evaluated using a rubric of 11 expert-rated metrics, scored on a five-point Likert scale. Gemini-2 and GPT-4o were the best performers (achieving mean scores of 4.49 and 4.48 out of 5, respectively, across all 11 metrics). All LLMs significantly outperformed local clinicians (ps < 0.001) across all metrics, with Gemini-2, for example, surpassing local GPs by an average of 0.83 points on every metric (range: 0.38 – 1.10). While performance degraded slightly when LLMs communicated in Kinyarwanda, the LLMs remained superior to clinicians and were over 500 times cheaper per response. These findings support the potential of LLMs to strengthen frontline care quality in low-resource, multilingual health systems.

Visit

doi.org

Tasks

language modelingquestion answering

Languages

Kinyarwanda

Licenses

http://creativecommons.org/licenses/by/4.0/

Similar

Large language models for frontline healthcare support in low-resource settingsHow Good Are Large Language Models at Arithmetic Reasoning in Low-Resource Language Settings?—A Study on Yorùbá Numerical Probes with Minimal ContaminationHow Good are Commercial Large Language Models on African Languages?AfroBench: How Good are Large Language Models on African Languages?A ‘Silent Trial’ Assessing the Accuracy of Large Language Models for Assisting Community Health Workers in Low-Resource SettingsPsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language

Large language models for frontline healthcare support in low-resource settings

Abstract Large language models (LLMs) have demonstrated str

How Good Are Large Language Models at Arithmetic Reasoning in Low-Resource Language Settings?—A Study on Yorùbá Numerical Probes with Minimal Contamination

We study the performance of large language models (LLMs) in natural language understanding and natur

How Good are Commercial Large Language Models on African Languages?

Recent advancements in Natural Language Processing (NLP) has led to the proliferation of large pretr

AfroBench: How Good are Large Language Models on African Languages?

Large-scale multilingual evaluations, such as MEGA, often include only a handful of African language

A ‘Silent Trial’ Assessing the Accuracy of Large Language Models for Assisting Community Health Workers in Low-Resource Settings

Abstract Community health workers (CHWs) in low-resource settings deliver variable

PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language

This paper evaluates the performance of Large Multimodal Models (LMMs) on Optical Character Recognit