Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Employing self-supervised learning models for cross-linguistic child speech maturity classification

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
ZhaSurWarHit
Éditeur:
LabUni
Éditeur:
CCSDarXiv
Hôte:avatar
International audience Speech technology systems struggle with many downstream tasks for child speech due to small training corpora and the difficulties that child speech pose. We apply a novel dataset, SpeechMaturity, to state-of-the-art transformer models to address a fundamental classification task: identifying child vocalizations. Unlike previous corpora, our dataset captures maximally ecologically-valid child vocalizations across an unprecedented sample, comprising children acquiring 25+ languages in the U.S., Bolivia, Vanuatu, Papua New Guinea, Solomon Islands, and France. The dataset contains 242,004 labeled vocalizations, magnitudes larger than previous work. Models were trained to distinguish between cry, laughter, mature (consonant+vowel), and immature speech (just consonant or vowel). Models trained on the dataset outperform state-of-the-art models trained on previous datasets, achieved classification accuracy comparable to humans, and were robust across rural and urban settings.

Visit

hal.science

Tasks

speech processing

Tags

FOS: Computer and information sciencesArtificial Intelligence (cs.AI)Computation and Language (cs.CL)[SCCO.PSYC]Cognitive science/Psychology[SCCO.LING]Cognitive science/Linguistics

Licenses

https://creativecommons.org/licenses/by-nc-nd/4.0/info:eu-repo/semantics/OpenAccess