Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Investigating data partitioning strategies for crosslinguistic low-resource ASR evaluation

Domaine:

natural language processing

Type de record:

paper
Créateur:
LiuSpePru
Hôte:avatar
Many automatic speech recognition (ASR) data sets include a single pre-defined test set consisting of one or more speakers whose speech never appears in the training set. This "hold-speaker(s)-out" data partitioning strategy, however, may not be ideal for data sets in which the number of speakers is very small. This study investigates ten different data split methods for five languages with minimal ASR training resources. We find that (1) model performance varies greatly depending on which speaker is selected for testing; (2) the average word error rate (WER) across all held-out speakers is comparable not only to the average WER over multiple random splits but also to any given individual random split; (3) WER is also generally comparable when the data is split heuristically or adversarially; (4) utterance duration and intensity are comparatively more predictive factors of variability regardless of the data split. These results suggest that the widely used hold-speakers-out approach to ASR data partitioning can yield results that do not reflect model performance on unseen data or speakers. Random splits can yield more reliable and generalizable estimates when facing data sparsity.

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageSoundAudio and Speech Processing

Similaires

Data-driven Model Generalizability in Crosslinguistic Low-resource Morphological Segmentationkibaraki/data-augmentation-for-low-resource-asrFrustratingly Easy Data Augmentation for Low-Resource ASRTowards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languageselerdg/ASR-for-low-resource-languagesEfficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data

Data-driven Model Generalizability in Crosslinguistic Low-resource Morphological Segmentation

Common designs of model evaluation typically focus on monolingual settings, where different models a

kibaraki/data-augmentation-for-low-resource-asr

Self-contained data augmentation for low-resource ASR # data-augmentation-for-low-resource-asr ##

Frustratingly Easy Data Augmentation for Low-Resource ASR

This paper introduces three self-contained data augmentation methods for low-resource Automatic Spee

Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages

Automatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, p

elerdg/ASR-for-low-resource-languages

Fine-tune wav2vec2-xls-r on data from low-resource-languages # ASR for Low-resource languages ## O

Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data

Automatic speech recognition for low-resource languages remains fundamentally constrained by the sca