Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

QuaSR: Quality-Aware Sample Reweighting for Pacific Indigenous Speech Recognition

Domain:

natural language processing

Record type:

paper
Creator:
Li,Xiao YangHuaHol
Publisher:
arXiv
Host:avatar
Training automatic speech recognition (ASR) models for low-resource languages is challenging due to limited data and highly variable supervision quality. In particular, Pacific Indigenous speech corpora often exhibit heterogeneous acoustic conditions, transcript inconsistencies, and varying degrees of acoustic-text alignment reliability, making standard fine-tuning approaches sensitive to noisy or misleading supervision signals. In this work, we propose QuaSR, a simple yet effective weighting framework that combines data-side reliability with model-side learnability to improve ASR adaptation. Specifically, we estimate data reliability from acoustic, transcription, and alignment, while measuring learnability using training loss from the model. These two complementary signals are integrated into a unified sample utility score to produce training weights for the samples. We also evaluated across four Pacific Indigenous languages, which shows that the proposed utility scores reliably correlate with adaptation performance. Furthermore, QuaSR consistently improves ASR adaptation over standard fine-tuning and alternative data selection strategies, highlighting a new way to leverage difficulty scores for low-resource speech learning. 6 pages, under peer review

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Tags

Audio and Speech Processing (eess.AS)FOS: Electrical engineering, electronic engineering, information engineering

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Context-Aware Dynamic Chunking for Streaming Tibetan Speech RecognitionVoxArabica: A Robust Dialect-Aware Arabic Speech Recognition SystemDevelopment of a diacritic-aware large vocabulary automatic speech recognition for Hausa languageAdapting Where It Matters: Depth-Aware Adaptation for Efficient Multilingual Speech Recognition in Low-Resource LanguagesSpeech Recognition for Hausa languageFramework for Hausa Speech Recognition

Context-Aware Dynamic Chunking for Streaming Tibetan Speech Recognition

In this work, we propose a streaming speech recognition framework for Amdo Tibetan, built upon a hyb

VoxArabica: A Robust Dialect-Aware Arabic Speech Recognition System

Arabic is a complex language with many varieties and dialects spoken by over 450 millions all around

Development of a diacritic-aware large vocabulary automatic speech recognition for Hausa language

Adapting Where It Matters: Depth-Aware Adaptation for Efficient Multilingual Speech Recognition in Low-Resource Languages

Recent speech foundation models excel at multilingual automatic speech recognition (ASR) for high-resource languages, but adapting them to low-resource languages remains challenging due to data scarcity and efficiency constraints. Full-model fine-tuning is computat

Speech Recognition for Hausa language

Speech Recognition for Hausa language

Poster presented at the Deep Learning Indaba 2022 by Umar Adam Ibrahim

Framework for Hausa Speech Recognition