Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

EGRA-Xhosa-14.9k: Annotated Child Reading Audio Dataset

Domaine:

natural language processingeducation

Type de record:

dataset
Éditeur:
Wes
Hôte:avatar
The project involves collecting the child reading dataset for the language is Xhosa, a South African Bantu language. The collected dataset is then processed with the help of native speakers and utilized to train state-of-the-art machine learning models focussed on assessing whether the child has spoken the word correctly or not. The dataset contains 14,972 recordings with an average of 4 seconds each. Each recording is annotated by three independent markers and consists of children speaking a particular word or letter from the Xhosa language in a classroom setting.

Visit

doi.orgresearch-data.westernsydney.edu.au

Tasks

automatic speech recognitionspeech processing

Languages

Xhosa

Tags

EGRA-AIEGRAChildrenEarly GradeAssessmentisiXhosaClassroomAnnotated

Licenses

http://creativecommons.org/licenses/by-nc-sa/3.0/au

Similaires

An End-to-End Approach for Child Reading Assessment in the Xhosa LanguageSilva3012/xhosa-nlp-datasetXhosa-English Translation DatasetDesign and Validation of Structural Causal Model: A Focus on EGRA DatasetShona Speech Dataset (SNA) - AnnotatedSesame Plant Segmentation Dataset: A YOLO Formatted Annotated Dataset

An End-to-End Approach for Child Reading Assessment in the Xhosa Language

Child literacy is a strong predictor of life outcomes at the subsequent stages of an individual's li

Silva3012/xhosa-nlp-dataset

My attempt at building a xhosa NLP dataset --- language: - xh - en license: other multilinguality:

Xhosa-English Translation Dataset

A High-Quality Parallel Corpus for Low-Resource Machine Translation

Design and Validation of Structural Causal Model: A Focus on EGRA Dataset

Designing and validating structural causal model (SCM) correctness from a dataset whose background k

Shona Speech Dataset (SNA) - Annotated

An annotated, speaker-relabelled, and loudness-normalised Shona (sna) speech dataset prepared throug

Sesame Plant Segmentation Dataset: A YOLO Formatted Annotated Dataset

This paper presents the Sesame Plant Segmentation Dataset, an open source annotated image dataset de