Abstract
Community health workers (CHWs) in low-resource settings deliver variable-quality care. This study used OpenAI’s o3 and Google’s Gemini Flash 2.5 to evaluate whether large language models (LLMs) ‘listening’ to CHW–patient interactions could generate accurate referral decisions. Across 150 participating Rwandan CHWs, 429 encounters were recorded (in Kinyarwanda) and then processed by LLMs. CHWs demonstrated high referral accuracy (97.9% [95% CI: 96.1%-98.9%]), and OpenAI’s o3 performed similarly to CHWs while Gemini 2.5-Flash showed low accuracy (47.3% [95% CI: 42.6%-52.1%]). Assessment of LLM-generated differential diagnoses and management plan quality showed superior performance from o3 compared with Gemini, though both models missed important conditions. In conclusion, the choice of LLM appears to be a critical design decision. Moreover, the high baseline performance of Rwandan CHWs suggests that LLMs are likely to have a limited impact in the current context but could be useful in less well-established CHW programmes. Trial Registration: PACTR202504601308784.