Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Human review for post-training improvement of low-resource language performance in large language models

Domain:

natural language processing
Creator:
LewDeRMisNyi
Publisher:
Dry
Host:avatar
Large language models (LLMs) have significantly improved natural language processing, holding the potential to support health workers and their clients directly. Unfortunately, there is a substantial and variable drop in performance for low-resource languages. Here we present results from an exploratory case study in Malawi, aiming to enhance the performance of LLMs in Chichewa through innovative prompt engineering techniques. By focusing on practical evaluations over traditional metrics, we assess the subjective utility of LLM outputs, prioritizing end-user satisfaction. Our findings suggest that tailored prompt engineering may improve LLM utility in underserved linguistic contexts, offering a promising avenue to bridge the language inclusivity gap in digital health interventions. We compared the reported performance of five variations of an LLM-based chatbot prototype through a two-step process. First, a cohort of 24 target end users, community health volunteers (CHVs) in Malawi, was recruited to generate transcripts by interacting with the prototypes. Second, an additional cohort of 22 CHVs was recruited to evaluate the transcripts and provide subjective feedback on the quality and utility of the language in the model’s responses.  CHVs in the first cohort were preassigned to one of the five chatbot variations and were instructed to generate a single transcript by interacting with their assined chatbot variation. A second set of CHVs was then recruited to conduct the transcript review. This second group of CHVs was also recruited through convenience sampling to form an intended cohort of 25 participants; however, three were unable to attend. A total of 22 CHVs participated in the transcript review, where each CHV was instructed to review and rate four of the transcripts generated by the previous cohort. Prior to the transcript review, duplicate transcripts and those with insufficient length were excluded from the evaluation pool. We used a stratified allocation method to assign transcripts to participants, ensuring that a single participant would neither receive the same transcript nor more than one transcript from the same bot variation. The order in which each CHV would review their assigned transcripts was then randomized, and transcripts were printed and labeled with a unique 4-digit identification code for CHVs to reference when providing their ratings. Participants were blinded to the authors of the original transcript, as well as variation that was used for the chatbot.  CHVs were asked to review each transcript in their assigned order and complete both a brief demographic survey and a language survey. The language survey dataset is shared here. All survey data were de-identified prior to analysis. Ratings were removed if the 4-digit transcript identification code entered into the survey by CHVs did not match predetermined assignments. # Human review for post-training improvement of low-resource language performance in large language models [doi.org](doi.org) This dataset comprises a single Excel file with transcript review survey results as reported by a cohort of community health volunteers (CHVs) in Malawi. CHVs were each asked to review and rate four pre-assigned transcripts, each generated from 1 of 5 Chichewa-speaking chatbot variations differing in temperature model parameter and/or changes to the system prompt. This survey was designed to collect CHV feedback on the quality of Chichewa spoken by the chatbot, given the performance gap between higher- and lower-resource languages. Results suggest that the use of specific prompt engineering techniques may improve foundational model utility when conversing using low-resource languages. 

Visit

doi.orgdatadryad.org

Tasks

language modeling

Languages

Chichewa

Tags

FOS: Electrical engineering, electronic engineering, information engineeringFOS: Electrical engineering, electronic engineering, information engineeringLarge Language ModelsPrompt engineeringLow-resource languages

Licenses

Creative Commons Zero v1.0 Universalhttps://creativecommons.org/publicdomain/zero/1.0/legalcode

Similar

Post-ASR Correction in Hindi: Comparing Language Models and Large Language Models in Low-Resource ScenariosTharuChat: Bootstrapping Large Language Models for a Low-Resource Language via Synthetic Data and Human ValidationEmploying large language models in Swahili, a low-resource languagePerformance Diminishment of Intermediate-Task Training in Large Multilingual Models for Low-Resource LanguagesLarge language models for frontline healthcare support in low-resource settingsTransliteration for Low-Resource Translation in the Age of Large Language Models

Post-ASR Correction in Hindi: Comparing Language Models and Large Language Models in Low-Resource Scenarios

Automatic Speech Recognition (ASR) systems for low-resource languages like Hindi often produce erron

TharuChat: Bootstrapping Large Language Models for a Low-Resource Language via Synthetic Data and Human Validation

The rapid proliferation of Large Language Models (LLMs) has created a profound digital divide, effec

Employing large language models in Swahili, a low-resource language

Performance Diminishment of Intermediate-Task Training in Large Multilingual Models for Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Large language models for frontline healthcare support in low-resource settings

Abstract Large language models (LLMs) have demonstrated str

Transliteration for Low-Resource Translation in the Age of Large Language Models

Neural machine translation (NMT) systems are widely used, but their performance remains strongly dep