Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Goodness-of-pronunciation without phoneme time alignment

Domain:

natural language processing

Record type:

papermodel
Creator:
WonChe
Host:avatar
In speech evaluation, an Automatic Speech Recognition (ASR) model often computes time boundaries and phoneme posteriors for input features. However, limited data for ASR training hinders expansion of speech evaluation to low-resource languages. Open-source weakly-supervised models are capable of ASR over many languages, but they are frame-asynchronous and not phonemic, hindering feature extraction for speech evaluation. This paper proposes to overcome incompatibilities for feature extraction with weakly-supervised models, easing expansion of speech evaluation to low-resource languages. Phoneme posteriors are computed by mapping ASR hypotheses to a phoneme confusion network. Word instead of phoneme-level speaking rate and duration are used. Phoneme and frame-level features are combined using a cross-attention architecture, obviating phoneme time alignment. This performs comparably with standard frame-synchronous features on English speechocean762 and low-resource Tamil datasets.

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageArtificial IntelligenceHuman-Computer InteractionMachine Learning

Similar

Arabic Language Phoneme Pronunciation Difficulties Among Upper Basic Hausa-Speaking Students in Kano State, NigeriaReal-time low-resource phoneme recognition on edge devicesPhoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment EvaluationA Review on Grapheme-to-Phoneme Modelling Techniques to Transcribe Pronunciation Variants for Under-Resourced Language Goodness-of-fit statistics for LCA models.Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax

Arabic Language Phoneme Pronunciation Difficulties Among Upper Basic Hausa-Speaking Students in Kano State, Nigeria

In the process of learning a foreign language, there are some indispensable learning problems, espec

Real-time low-resource phoneme recognition on edge devices

While speech recognition has seen a surge in interest and research over the last decade, most machin

Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation

Forced alignment evaluation typically requires manually annotated timestamps, limiting large-scale a

A Review on Grapheme-to-Phoneme Modelling Techniques to Transcribe Pronunciation Variants for Under-Resourced Language

A pronunciation dictionary (PD) is one of the components in an Automatic Speech Recognition (ASR) sy

Goodness-of-fit statistics for LCA models.

Maternal vaccination, or vaccination in pregnancy, offers a critical opportunity to provide

Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax

Extending large language models (LLMs) to low-resource languages often incurs an "alignment tax": im