Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Self-Supervised Learning for Low-Resource Voice Recognition in Regional Television Channels

Domain:

natural language processing

Record type:

paper
Creator:
ArvHadNut
Publisher:
Jou
Host:avatar
The development of an automatic speech recognition (ASR) system for regional television channels is still a difficult task because of insufficient labeled data, various dialects, and substantial accent differences between speakers. As opposed to high-resource languages, regional broadcasts usually do not have enough labeled subtitles for supervised training and, therefore, need more expensive procedures that may require extensive effort and money. This problem becomes more complicated when there is ambient studio noise, spontaneous speech, code-mixed lexicon, and diverse pronunciation of the anchors and other participants. In order to solve these issues, the current paper suggests developing an ASR model based on self-supervised learning (SSL) principles. This method involves pretraining with large amounts of unlabeled audio clips from regional broadcasts. Afterward, the obtained knowledge can be used to extract acoustic and context-dependent features, which are further refined by applying a small number of annotated data points. In this work, the authors use Wav2Vec 2.0 encoder to train an SSL-based architecture on raw speech data. Then, the pretrained encoder can be fine- tuned by providing a limited amount of manually annotated data. Such a transfer-learning approach allows achieving better results in ASR tasks with limited-label scenarios. The experiments show that the proposed architecture has higher accuracy in terms of Word Error Rate (WER) and Character Error Rate (CER) than classical CNN, LSTM, and hybrid ASR models. Besides, the developed system is adaptive to dialect diversity and individual speech peculiarities of anchors. The suggested approach can be applied to generate subtitles for regional television broadcasts.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Ensemble of learning in Self-supervised speech recognitionMultilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitchingConLID: Supervised Contrastive Learning for Low-Resource Language IdentificationImproving Low-Resource Morphological Inflection via Self-Supervised ObjectivesLacuna Reconstruction: Self-supervised Pre-training for Low-Resource Historical Document TranscriptionX-MAD: A Public Multi-Anatomy X-ray Dataset for Self-Supervised Learning in Low-Resource Radiographic Imaging

Ensemble of learning in Self-supervised speech recognition

Ensemble of learning in Self-supervised speech recognition

Poster presented at the Deep Learning Indaba 2023 by Ussen Kimanuka

Multilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching

While many speakers of low-resource languages regularly code-switch between their languages and othe

ConLID: Supervised Contrastive Learning for Low-Resource Language Identification

Language identification (LID) is a critical step in curating multilingual LLM pretraining corpora fr

Improving Low-Resource Morphological Inflection via Self-Supervised Objectives

Self-supervised objectives have driven major advances in NLP by leveraging large-scale unlabeled dat

Lacuna Reconstruction: Self-supervised Pre-training for Low-Resource Historical Document Transcription

We present a self-supervised pre-training approach for learning rich visual language representations

X-MAD: A Public Multi-Anatomy X-ray Dataset for Self-Supervised Learning in Low-Resource Radiographic Imaging

X-MAD is a public multi-anatomy X-ray dataset collected from Ibn Hayyan Specialist Hospital in Sana'