Logo Lanfrica

Evaluating Off-the-Shelf Machine Translation Systems for Kinyarwanda-to-English Clinical Dialogue: A Benchmark and Cautionary Analysis

Domaine:

natural language processinghealthcare

Type de record:

paper
Créateur:
HasRiyAndMat
Éditeur:
Elsevier BV
Hôte:
Background: Linguistic barriers between patients and clinicians cause diagnostic error, poor adherence, and adverse events, and disproportionately affect speakers of low-resource languages. Kinyarwanda is Rwanda's national language, spoken by more than 20 million people, yet the safety of off-the-shelf machine translation (MT) for Kinyarwanda clinical dialogue is uncharacterised. 

Methods: We benchmarked four widely deployed MT systems (Google Gemini 2.5 Flash, Microsoft Azure Translator, Google Translate, and Meta NLLB-200 distilled-600M) on 600 patient–provider dialogues collected in Rwanda during the COVID-19 pandemic through the WelTel SMS platform, scoring each against professional human reference translations using semantic similarity, BLEU, chrF, and translation edit rate (TER). A bilingual reviewer separately rated outputs from six systems on a 0–100 scale, and a second independent rater scored a stratified subsample for inter-rater reliability. Analyses used de-identified secondary data approved by the University of British Columbia Research Ethics Board (H20-01031) and the Rwanda National Ethics Committee (100/RNEC/2020); no patients were involved in study design. 

Findings: Gemini 2.5 Flash achieved the highest semantic similarity (0.814) and BLEU (29.10) and the lowest TER (78.28); Azure Translator was effectively tied on TER (78.69). Google Translate achieved the highest chrF (53.25) and the fastest throughput; NLLB-200 scored lowest on every automatic metric. All four produced TER near 80 (about four in five tokens needing correction). Among 3372 valid human ratings, two systems formed an upper cluster (means 72.5 and 70.5), separated from the remaining four (36.0–49.9) by about 28 points, sharper than automatic metrics indicated. The second rater reproduced the ranking exactly (Spearman ρ = 1.00) but absolute item-level agreement was limited (intraclass correlation 0.23; 26 points more lenient on average).

Interpretation: No off-the-shelf MT system evaluated is ready for unsupervised use in Kinyarwanda clinical communication. Automatic metrics rank systems but understate the quality gap human readers perceive, and qualified reviewers disagree on the absolute quality of individual translations. The strongest systems are candidates only for human-in-the-loop use, with a bilingual clinician verifying every translation before it reaches a patient, especially for dosage, negation, and symptoms. Future priorities are clinician-led safety adjudication, domain fine-tuning of open-source models, and prospective evaluation in live consultations.

Languages

Similaires