Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Arabic Little STT: Arabic Children Speech Recognition Dataset

Domain:

natural language processing

Record type:

paperdataset
Creator:
AlkDesJal
Host:avatar
The performance of Artificial Intelligence (AI) systems fundamentally depends on high-quality training data. However, low-resource languages like Arabic suffer from severe data scarcity. Moreover, the absence of child-specific speech corpora is an essential gap that poses significant challenges. To address this gap, we present our created dataset, Arabic Little STT, a dataset of Levantine Arabic child speech recorded in classrooms, containing 355 utterances from 288 children (ages 6 - 13). We further conduct a systematic assessment of Whisper, a state-of-the-art automatic speech recognition (ASR) model, on this dataset and compare its performance with adult Arabic benchmarks. Our evaluation across eight Whisper variants reveals that even the best-performing model (Large_v3) struggles significantly, achieving a 0.66 word error rate (WER) on child speech, starkly contrasting with its sub 0.20 WER on adult datasets. These results align with other research on English speech. Results highlight the critical need for dedicated child speech benchmarks and inclusive training data in ASR development. Emphasizing that such data must be governed by strict ethical and privacy frameworks to protect sensitive child information. We hope that this study provides an initial step for future work on equitable speech technologies for Arabic-speaking children. We hope that our publicly available dataset enrich the children's demographic representation in ASR datasets.

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageArtificial IntelligenceHuman-Computer InteractionMachine LearningSound

Similar

Tamazight-Arabic Speech Recognition DatasetTamazight-Arabic Speech Recognition DatasetAbubakrHassan/Arabic-Speech-Recognition-DatasetAhmedMnsour/Arabic-Speech-Recognition-DatasetA Novel Dataset for Arabic Speech Recognition Recorded by Tamazight SpeakersTamazight-Arabic Speech Translation Dataset

Tamazight-Arabic Speech Recognition Dataset

This dataset contains speech segments in Tamazight (specifically focusing on the Tachelhit dialect)

Tamazight-Arabic Speech Recognition Dataset

This is the EMINES organization-hosted version of the Tamazight-Arabic Speech Recognition Dataset, s

AbubakrHassan/Arabic-Speech-Recognition-Dataset

This data was collected as part of the Speech Recognition course at AIMS Ghana, taught by Emmanuel D

AhmedMnsour/Arabic-Speech-Recognition-Dataset

This data was collected as part of the Speech Recognition course during the African Master's in Mach

A Novel Dataset for Arabic Speech Recognition Recorded by Tamazight Speakers

Automatic Speech Recognition (ASR) is an area of research that's constantly evolving, thanks to impo

Tamazight-Arabic Speech Translation Dataset

This is the Tamazight-NLP organization-hosted version of the Tamazight-Arabic Speech Recognition Dat