Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages

Domain:

natural language processing

Record type:

papermodelsoftware
Creator:
OmnKerKozMen
Host:avatar
Automatic speech recognition (ASR) has advanced in high-resource languages, but most of the world's 7,000+ languages remain unsupported, leaving thousands of long-tail languages behind. Expanding ASR coverage has been costly and limited by architectures that restrict language support, making extension inaccessible to most--all while entangled with ethical concerns when pursued without community collaboration. To transcend these limitations, we introduce Omnilingual ASR, the first large-scale ASR system designed for extensibility. Omnilingual ASR enables communities to introduce unserved languages with only a handful of data samples. It scales self-supervised pre-training to 7B parameters to learn robust speech representations and introduces an encoder-decoder architecture designed for zero-shot generalization, leveraging a LLM-inspired decoder. This capability is grounded in a massive and diverse training corpus; by combining breadth of coverage with linguistic variety, the model learns representations robust enough to adapt to unseen languages. Incorporating public resources with community-sourced recordings gathered through compensated local partnerships, Omnilingual ASR expands coverage to over 1,600 languages, the largest such effort to date--including over 500 never before served by ASR. Automatic evaluations show substantial gains over prior systems, especially in low-resource conditions, and strong generalization. We release Omnilingual ASR as a family of models, from 300M variants for low-power devices to 7B for maximum accuracy. We reflect on the ethical considerations shaping this design and conclude by discussing its societal impact. In particular, we highlight how open-sourcing models and tools can lower barriers for researchers and communities, inviting new forms of participation. Open-source artifacts are available at github.com.

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and Language

Similar

Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian LanguagesDvoice : An open source dataset for Automatic Speech Recognition on African Languages and DialectsMultilingual Speech Recognition Initiative for African LanguagesMultilingual Automatic Speech Recognition for Kinyarwanda, Swahili, and Luganda: Advancing ASR in Select East African LanguagesSagalee: an Open Source Automatic Speech Recognition Dataset for Oromo LanguageDARTS-ASR: Differentiable Architecture Search for Multilingual Speech Recognition and Adaptation

Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian Languages

We present Ethio-ASR, a suite of multilingual CTC-based automatic speech recognition (ASR) models jo

Dvoice : An open source dataset for Automatic Speech Recognition on African Languages and Dialects

DVoice is a community initiative that aims to provide African languages and dialects with d

Multilingual Speech Recognition Initiative for African Languages

Abstract This paper summarizes a speech recognition initiative for African languages. More

Multilingual Automatic Speech Recognition for Kinyarwanda, Swahili, and Luganda: Advancing ASR in Select East African Languages

Multilingual Automatic Speech Recognition for Kinyarwanda, Swahili, and Luganda: Advancing ASR in Select East African Languages

Poster presented at the Deep Learning Indaba 2023 by Samuel Rutunda

Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language

We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoke

DARTS-ASR: Differentiable Architecture Search for Multilingual Speech Recognition and Adaptation

In previous works, only parameter weights of ASR models are optimized under fixed-topology architect