Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Effectiveness of text to speech pseudo labels for forced alignment and cross lingual pretrained models for low resource speech recognition

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
GupGauDhuCha
Hôte:avatar
In the recent years end to end (E2E) automatic speech recognition (ASR) systems have achieved promising results given sufficient resources. Even for languages where not a lot of labelled data is available, state of the art E2E ASR systems can be developed by pretraining on huge amounts of high resource languages and finetune on low resource languages. For a lot of low resource languages the current approaches are still challenging, since in many cases labelled data is not available in open domain. In this paper we present an approach to create labelled data for Maithili, Bhojpuri and Dogri by utilising pseudo labels from text to speech for forced alignment. The created data was inspected for quality and then further used to train a transformer based wav2vec 2.0 ASR model. All data and models are available in open domain. Submitted to InterSpeech 2022

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageSoundAudio and Speech Processing

Similaires

POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text TranslationExploiting Adapters for Cross-lingual Low-resource Speech RecognitionText-To-Speech Data Augmentation for Low Resource Speech RecognitionDonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech RecognitionXLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech RecognitionUnsupervised Cross-Lingual Speech Emotion Recognition Using Pseudo Multilabel

POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation

Speech Large Language Models have achieved breakthroughs in multilingual speech-to-text translation.

Exploiting Adapters for Cross-lingual Low-resource Speech Recognition

Cross-lingual speech adaptation aims to solve the problem of leveraging multiple rich-resource langu

Text-To-Speech Data Augmentation for Low Resource Speech Recognition

Nowadays, the main problem of deep learning techniques used in the development of automatic speech r

DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition

Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where mod

XLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech Recognition

In this paper, we propose a weakly supervised multilingual representation learning framework, called

Unsupervised Cross-Lingual Speech Emotion Recognition Using Pseudo Multilabel

Speech Emotion Recognition (SER) in a single language has achieved remarkable results through deep l