Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Self-Supervised Learning for Low-Resource Voice Recognition in Regional Television Channels

Domaine:

natural language processing

Type de record:

paper
Créateur:
ArvHadNut
Éditeur:
Jou
Hôte:avatar
The development of an automatic speech recognition (ASR) system for regional television channels is still a difficult task because of insufficient labeled data, various dialects, and substantial accent differences between speakers. As opposed to high-resource languages, regional broadcasts usually do not have enough labeled subtitles for supervised training and, therefore, need more expensive procedures that may require extensive effort and money. This problem becomes more complicated when there is ambient studio noise, spontaneous speech, code-mixed lexicon, and diverse pronunciation of the anchors and other participants. In order to solve these issues, the current paper suggests developing an ASR model based on self-supervised learning (SSL) principles. This method involves pretraining with large amounts of unlabeled audio clips from regional broadcasts. Afterward, the obtained knowledge can be used to extract acoustic and context-dependent features, which are further refined by applying a small number of annotated data points. In this work, the authors use Wav2Vec 2.0 encoder to train an SSL-based architecture on raw speech data. Then, the pretrained encoder can be fine- tuned by providing a limited amount of manually annotated data. Such a transfer-learning approach allows achieving better results in ASR tasks with limited-label scenarios. The experiments show that the proposed architecture has higher accuracy in terms of Word Error Rate (WER) and Character Error Rate (CER) than classical CNN, LSTM, and hybrid ASR models. Besides, the developed system is adaptive to dialect diversity and individual speech peculiarities of anchors. The suggested approach can be applied to generate subtitles for regional television broadcasts.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode