Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

An End-to-End Approach for Child Reading Assessment in the Xhosa Language

Domaine:

educationnatural language processing

Type de record:

paperdatasetmodel
Créateur:
CheNavValUba
Hôte:avatar
Child literacy is a strong predictor of life outcomes at the subsequent stages of an individual's life. This points to a need for targeted interventions in vulnerable low and middle income populations to help bridge the gap between literacy levels in these regions and high income ones. In this effort, reading assessments provide an important tool to measure the effectiveness of these programs and AI can be a reliable and economical tool to support educators with this task. Developing accurate automatic reading assessment systems for child speech in low-resource languages poses significant challenges due to limited data and the unique acoustic properties of children's voices. This study focuses on Xhosa, a language spoken in South Africa, to advance child speech recognition capabilities. We present a novel dataset composed of child speech samples in Xhosa. The dataset is available upon request and contains ten words and letters, which are part of the Early Grade Reading Assessment (EGRA) system. Each recording is labeled with an online and cost-effective approach by multiple markers and a subsample is validated by an independent EGRA reviewer. This dataset is evaluated with three fine-tuned state-of-the-art end-to-end models: wav2vec 2.0, HuBERT, and Whisper. The results indicate that the performance of these models can be significantly influenced by the amount and balancing of the available training data, which is fundamental for cost-effective large dataset collection. Furthermore, our experiments indicate that the wav2vec 2.0 performance is improved by training on multiple classes at a time, even when the number of available samples is constrained. Paper accepted on AIED 2025 containing 14 pages, 6 figures and 4 tables

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Languages

Xhosa

Tags

Machine LearningComputation and Language

Similaires

Large Scale Speech Recognition for Low Resource Language Amharic, an End-to-End ApproachAMorph: An End-to-End Morpheme Level Natural Language Processing Pipeline for AmharicSub-word Based End-to-End Speech Recognition for an Under-Resourced Language: AmharicAn end-to-end framework for translation of American sign language to low-resource languages in NigeriaNatiQ: An End-to-end Text-to-Speech System for ArabicAmharic OCR: An End-to-End Learning

Large Scale Speech Recognition for Low Resource Language Amharic, an End-to-End Approach

Speech recognition, or automatic speech recognition (ASR), is a technology designed to convert spoke

AMorph: An End-to-End Morpheme Level Natural Language Processing Pipeline for Amharic

Sub-word Based End-to-End Speech Recognition for an Under-Resourced Language: Amharic

In this work, we focused on end-to-end speech recognition for less-resourced language, Amharic. The result can be integrated with other tasks such as spoken content retrieval. We explored three models, which consist of Convolutional Neural Networks, Recurrent Neura

An end-to-end framework for translation of American sign language to low-resource languages in Nigeria

NatiQ: An End-to-end Text-to-Speech System for Arabic

NatiQ is end-to-end text-to-speech system for Arabic. Our speech synthesizer uses an encoder-decoder

Amharic OCR: An End-to-End Learning

In this paper, we introduce an end-to-end Amharic text-line image recognition approach based on recu