Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

ASR2K: Speech Recognition for Around 2000 Languages without Audio

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
Li,MetMorBla
Hôte:avatar
Most recent speech recognition models rely on large supervised datasets, which are unavailable for many low-resource languages. In this work, we present a speech recognition pipeline that does not require any audio for the target language. The only assumption is that we have access to raw text datasets or a set of n-gram statistics. Our speech pipeline consists of three components: acoustic, pronunciation, and language models. Unlike the standard pipeline, our acoustic and pronunciation models use multilingual models without any supervision. The language model is built using n-gram statistics or the raw text dataset. We build speech recognition for 1909 languages by combining it with Crubadan: a large endangered languages n-gram database. Furthermore, we test our approach on 129 languages across two datasets: Common Voice and CMU Wilderness dataset. We achieve 50% CER and 74% WER on the Wilderness dataset with Crubadan statistics only and improve them to 45% CER and 69% WER when using 10000 raw text utterances. INTERSPEECH 2022

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and Language

Similaires

Improving Speech Recognition for Under-resourced Languages Utilizing Audio-codecs for Data AugmentationSpeechless: Speech Instruction Training Without Speech for Low Resource LanguagesDiscrete Audio Tokens Enhance Cross-Lingual Speech Recognition in Low-Resource LanguagesSpeech Recognition Datasets for Congolese LanguagesDnn-Based Speech Recognition For Globalphone LanguagesRobust speech recognition for low-resource languages

Improving Speech Recognition for Under-resourced Languages Utilizing Audio-codecs for Data Augmentation

Presenter: Nirayo Hailu Gebreegziabher, Ingo Siegert, Andreas Nürnberger, MMSP 2020, Virtual Event,

Speechless: Speech Instruction Training Without Speech for Low Resource Languages

The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need f

Discrete Audio Tokens Enhance Cross-Lingual Speech Recognition in Low-Resource Languages

This report synthesises findings from 13 peer-reviewed papers addressing the following research ques

Speech Recognition Datasets for Congolese Languages

This dataset contains two new benchmark corpora designed for low-resource languages spoken in the De

Dnn-Based Speech Recognition For Globalphone Languages

Presenter: Martha Yifiru Tachbelie, ICASSP 2020, Virtual Event, May 4-8, 2020 This paper describes n

Robust speech recognition for low-resource languages

Process of human-machine interaction is an integral part of everyday human life in a modern world. T