Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses

Domain:

natural language processing

Record type:

papermodelsoftware
Creator:
Li,Vu,
Host:avatar
Training a semi-supervised end-to-end speech recognition system using noisy student training has significantly improved performance. However, this approach requires a substantial amount of paired speech-text and unlabeled speech, which is costly for low-resource languages. Therefore, this paper considers a more extreme case of semi-supervised end-to-end automatic speech recognition where there are limited paired speech-text, unlabeled speech (less than five hours), and abundant external text. Firstly, we observe improved performance by training the model using our previous work on semi-supervised learning "CycleGAN and inter-domain losses" solely with external text. Secondly, we enhance "CycleGAN and inter-domain losses" by incorporating automatic hyperparameter tuning, calling it "enhanced CycleGAN inter-domain losses." Thirdly, we integrate it into the noisy student training approach pipeline for low-resource scenarios. Our experimental results, conducted on six non-English languages from Voxforge and Common Voice, show a 20% word error rate reduction compared to the baseline teacher model and a 10% word error rate reduction compared to the baseline best student model, highlighting the significant improvements achieved through our proposed method. 10 pages (2 for references), 4 figures, published in SIGUL2024@LREC-COLING 2024

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageSoundAudio and Speech Processing

Similar

Improving End-to-End Speech Translation for the Low Resource Language Fongbe to FrenchAn end-to-end framework for translation of American sign language to low-resource languages in NigeriaLow-Resource End-to-end Sanskrit TTS using Tacotron2, WaveGlow and Transfer LearningScaling End-to-End Models for Large-Scale Multilingual ASREnd-to-End Amharic speech to text with Noisy data(Only sample data)Large Scale Speech Recognition for Low Resource Language Amharic, an End-to-End Approach

Improving End-to-End Speech Translation for the Low Resource Language Fongbe to French

This study addresses the challenges of end-to-end (E2E) Speech-to-Text Translation (STT) for the low

An end-to-end framework for translation of American sign language to low-resource languages in Nigeria

Low-Resource End-to-end Sanskrit TTS using Tacotron2, WaveGlow and Transfer Learning

End-to-end text-to-speech (TTS) systems have been developed for European languages like English and

Scaling End-to-End Models for Large-Scale Multilingual ASR

Building ASR models across many languages is a challenging multi-task learning problem due to large

End-to-End Amharic speech to text with Noisy data(Only sample data)

This dataset is a nuanced compilation of diverse and noisy data meticulously curated to cater to the

Large Scale Speech Recognition for Low Resource Language Amharic, an End-to-End Approach

Speech recognition, or automatic speech recognition (ASR), is a technology designed to convert spoke