Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Fast transcription of speech in low-resource languages

Domain:

natural language processing

Record type:

papersoftwaremodel
Creator:
HasGouLev
Host:avatar
We present software that, in only a few hours, transcribes forty hours of recorded speech in a surprise language, using only a few tens of megabytes of noisy text in that language, and a zero-resource grapheme to phoneme (G2P) table. A pretrained acoustic model maps acoustic features to phonemes; a reversed G2P maps these to graphemes; then a language model maps these to a most-likely grapheme sequence, i.e., a transcription. This software has worked successfully with corpora in Arabic, Assam, Kinyarwanda, Russian, Sinhalese, Swahili, Tagalog, and Tamil. 8 pages

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Languages

KinyarwandaSwahili

Tags

Computation and Language68T10

Similar

Robust speech recognition for low-resource languagesUser-friendly automatic transcription of low-resource languages: Plugging ESPnet into ElpisEnhancing Automatic Speech Recognition for Child Speech in Low-Resource LanguagesSpeechless: Speech Instruction Training Without Speech for Low Resource LanguagesSANTLR: Speech Annotation Toolkit for Low Resource LanguagesAdversarial Text-to-Speech for low-resource languages

Robust speech recognition for low-resource languages

Process of human-machine interaction is an integral part of everyday human life in a modern world. T

User-friendly automatic transcription of low-resource languages: Plugging ESPnet into Elpis

International audience This paper reports on progress integrating the speech recognit

Enhancing Automatic Speech Recognition for Child Speech in Low-Resource Languages

Automatic speech recognition (ASR) for children is demanding because their speech differs c

Speechless: Speech Instruction Training Without Speech for Low Resource Languages

The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need f

SANTLR: Speech Annotation Toolkit for Low Resource Languages

While low resource speech recognition has attracted a lot of attention from the speech community, th

Adversarial Text-to-Speech for low-resource languages

Improving the adversarial TTS models for low-resource languages by utilizing the high-frequency similarities between the different languages.