Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Hard to Be Heard: Phoneme-Level ASR Analysis of Phonologically Complex, Low-Resource Endangered Languages

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
AkaDanJäg
Hôte:avatar
We present a phoneme-level analysis of automatic speech recognition (ASR) for two low-resourced and phonologically complex East Caucasian languages, Archi and Rutul, based on curated and standardized speech-transcript resources totaling approximately 50 minutes and 1 hour 20 minutes of audio, respectively. Existing recordings and transcriptions are consolidated and processed into a form suitable for ASR training and evaluation. We evaluate several state-of-the-art audio and audio-language models, including wav2vec2, Whisper, and Qwen2-Audio. For wav2vec2, we introduce a language-specific phoneme vocabulary with heuristic output-layer initialization, which yields consistent improvements and achieves performance comparable to or exceeding Whisper in these extremely low-resource settings. Beyond standard word and character error rates, we conduct a detailed phoneme-level error analysis. We find that phoneme recognition accuracy strongly correlates with training frequency, exhibiting a characteristic sigmoid-shaped learning curve. For Archi, this relationship partially breaks for Whisper, pointing to model-specific generalization effects beyond what is predicted by training frequency. Overall, our results indicate that many errors attributed to phonological complexity are better explained by data scarcity. These findings demonstrate the value of phoneme-level evaluation for understanding ASR behavior in low-resource, typologically complex languages. Accepted to ACL 2026 (Findings)

Visit

arxiv.org

Tags

Computation and LanguageI.2.7

Similaires

elerdg/ASR-for-low-resource-languagesRafat-decodis/Robust-ASR-for-Low-Resource-LanguagesAdvances in Low-Resource and Endangered LanguagesMultimodal In-context Learning for ASR of Low-resource LanguagesKrishnateja244/Fine-tuning-of-ASR-models-on-low-resource-languagesTask Arithmetic with Support Languages for Low-Resource ASR

elerdg/ASR-for-low-resource-languages

Fine-tune wav2vec2-xls-r on data from low-resource-languages # ASR for Low-resource languages ## O

Rafat-decodis/Robust-ASR-for-Low-Resource-Languages

Exploring Benchmark Gaps and Real-World Speech Generalization for Language in Low Resource # 🧠 A Ro

Advances in Low-Resource and Endangered Languages

This paper reports on the approaches and results for the collection, analysis, and processing of low

Multimodal In-context Learning for ASR of Low-resource Languages

Automatic speech recognition (ASR) still covers only a small fraction of the world's languages, main

Krishnateja244/Fine-tuning-of-ASR-models-on-low-resource-languages

Fine tuning ASR models such as Wave2Vec2.0, Whisper, Nemo and MMS models on low-resource languages

Task Arithmetic with Support Languages for Low-Resource ASR

The development of resource-constrained approaches to automatic speech recognition (ASR) is of great