Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

EGRA-Xhosa-14.9k: Annotated Child Reading Audio Dataset

Domain:

natural language processingeducation

Record type:

dataset
Publisher:
Wes
Host:avatar
The project involves collecting the child reading dataset for the language is Xhosa, a South African Bantu language. The collected dataset is then processed with the help of native speakers and utilized to train state-of-the-art machine learning models focussed on assessing whether the child has spoken the word correctly or not. The dataset contains 14,972 recordings with an average of 4 seconds each. Each recording is annotated by three independent markers and consists of children speaking a particular word or letter from the Xhosa language in a classroom setting.

Visit

doi.orgresearch-data.westernsydney.edu.au

Tasks

automatic speech recognitionspeech processing

Languages

Xhosa

Tags

EGRA-AIEGRAChildrenEarly GradeAssessmentisiXhosaClassroomAnnotated

Licenses

http://creativecommons.org/licenses/by-nc-sa/3.0/au