Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Using large language models and speech-to-text models to facilitate the assessment of basic literacy in Ghana

Domain:

natural language processingeducation

Record type:

datasetpaper
Creator:
Hen
Editor:
HalMcG
Publisher:
University of Oxford
Host:avatar
This dissertation examines how recent advances in artificial intelligence, particularly in Natural Language Processing (NLP) and Automatic Speech Recognition (ASR), could enhance the assessment of reading ability in English in West Africa. Although regular diagnostic and formative assessment has proven essential for effective reading instruction, the substantial time and expertise needed to conduct and evaluate reading assessments often restricts their routine implementation, especially in resource-constrained education systems. Recent innovations in Natural Language Processing (NLP) and Automatic Speech Recognition (ASR) offer promising new approaches for making reading assessment more accessible and scalable in low-resource environments. To investigate this potential, this research develops and analyzes two novel datasets gathered in partnership with Rising Academies, a network of schools in Ghana. The primary dataset encompasses results from various reading assessment tasks administered to 162 students between the ages of 9 and 18, including audio recordings of oral reading, written and spoken responses to comprehension questions, and story retells. The secondary dataset contains high-quality human transcriptions of these same students’ performance on the oral reading fluency tasks, enabling comprehensive analysis of reading errors and fluency patterns. All assessment materials were selected from released items from the 2016 PrePIRLS assessment, which was specifically designed for use in low and middle-income countries. Using these datasets, I then conducted three interconnected empirical studies. The first study evaluates the viability of using state-of-the-art ASR models to assess oral reading fluency. Using the Ghanaian student dataset, the research determined that the Whisper v2 model could transcribe student reading with high accuracy (Word Error Rate of 10.3%) without any fine-tuning or adaptation. Significantly, the model's performance remained consistent across different student populations, demonstrating similar accuracy when transcribing recordings from both Ghanaian and American students. When these transcriptions were used to calculate Words Correct Per Minute scores, they demonstrated strong agreement with expert human raters (correlation of 0.98). The second study explores the capacity of Large Language Models to evaluate short answer reading comprehension questions. Drawing from a dataset of over 1,000 student responses, the research established that GPT-4 could achieve accuracy levels matching those of expert human raters. With minimal prompt engineering, the model attained a Quadratic Weighted kappa of 0.91 in the three-class condition and a Linear Weighted kappa of 0.87 in the two-class condition, surpassing existing benchmarks in automated short answer grading. The third study assesses the effectiveness of LLMs in grading story retell tasks while simultaneously examining the reliability and validity of various retell measures. The research documented strong inter-rater reliability among human raters using both three-class and five-class rubrics (Kendall's W of 0.81 and 0.85 respectively). It also revealed moderate correlations between different measures of retell and other indicators of reading ability, suggesting that retell tasks capture unique aspects of reading comprehension. In evaluating LLM performance, the study found that GPT-4 could effectively replicate human scoring, achieving agreement levels (QWK of 0.82) that nearly matched human-to-human agreement (0.85). Several themes emerged across these studies. First, the latest generation of AI models exhibited remarkably robust performance across different assessment types without requiring task-specific fine-tuning or extensive prompt engineering. This marks a substantial improvement over previous approaches that demanded considerable technical expertise and task-specific training data. Second, throughout all three studies, the automated approaches reached levels of agreement with human raters that paralleled inter-rater agreement among experts, indicating these technologies could serve as reliable tools for diagnostic and formative assessment purposes. Third, the studies consistently showed minimal evidence of systematic bias in model performance based on student demographics, though this finding warrants careful interpretation given the relatively limited sample sizes. The research underscores the crucial role of assessment design in successful automation. Specifically, more detailed rubrics that identify meaningful distinctions in student performance appear particularly well-suited to automated scoring. Furthermore, the studies illuminate how different assessment formats may capture distinct aspects of reading ability, reinforcing the importance of employing multiple assessment types to construct a comprehensive picture of student reading proficiency. These findings contribute to an understanding of how AI can support literacy assessment across diverse educational settings. While previous studies have shown AI's potential for educational assessment in well-resourced environments, this dissertation provides pioneering empirical evidence for the feasibility of applying these technologies to support reading assessment. The consistently strong performance across various assessment types and student populations suggests that recent AI advances could facilitate more regular and comprehensive evaluation of reading ability, particularly in contexts where traditional assessment methods face resource constraints. By demonstrating the viability of automated scoring across multiple assessment types while identifying areas requiring further study, this research enhances our understanding of how artificial intelligence can strengthen literacy assessment in diverse educational contexts.

Visit

doi.orgora.ox.ac.uk

Tasks

automatic speech recognitionspeech processing

Licenses

Creative Commons Attribution Non Commercial 4.0 Internationalhttps://creativecommons.org/licenses/by-nc/4.0/legalcode

Similar

FuaadBashi/speech-to-text-AI-models-for-the-Somali-languageSupporting Literacy Assessment in West Africa: Using State-of-the-Art Speech Models to Assess Oral Reading FluencyAM-DETOX: Analyzing Amharic Text Detoxification Using Large Language ModelsEnhancing Crowdsourced Audio for Text-to-Speech ModelsUsing State-of-the-Art Speech Models to Evaluate Oral Reading Fluency in GhanaArabic Large Language Models for Medical Text Generation

FuaadBashi/speech-to-text-AI-models-for-the-Somali-language

Train or adapt a speech-to-text model capable of accurately transcribing Somali audio. # Somali Spe

Supporting Literacy Assessment in West Africa: Using State-of-the-Art Speech Models to Assess Oral Reading Fluency

This paper reports on a set of three recent experiments utilizing large-scale speech models to asses

AM-DETOX: Analyzing Amharic Text Detoxification Using Large Language Models

Enhancing Crowdsourced Audio for Text-to-Speech Models

High-quality audio data is a critical prerequisite for training robust text-to-speech models, which

Using State-of-the-Art Speech Models to Evaluate Oral Reading Fluency in Ghana

This paper reports on a set of three recent experiments utilizing large-scale speech models to evalu

Arabic Large Language Models for Medical Text Generation

Efficient hospital management systems (HMS) are critical worldwide to address challenges such as ove