Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

An (unhelpful) guide to selecting the best ASR architecture for your under-resourced language

Domain:

natural language processing

Record type:

paper
Creator:
AssJimLiuPru
Publisher:
Und
Host:avatar
Advances in deep neural models for automatic speech recognition (ASR) have yielded dramatic improvements in ASR quality for resource-rich languages, with English ASR now achieving word error rates comparable to that of human transcribers. The vast majority of the world's languages, however, lack the quantity of data necessary to approach this level of accuracy. In this paper we use four of the most popular ASR toolkits to train ASR models for eleven languages with limited ASR training resources: eleven widely spoken languages of Africa, Asia, and South America, one endangered language of Central America, and three critically endangered languages of North America. We find that no single architecture consistently outperforms any other. These differences in performance so far do not appear to be related to any particular feature of the datasets or characteristics of the languages. These findings have important implications for future research in ASR for under-resourced languages. ASR systems for languages with abundant existing media and available speakers may derive the most benefit simply by collecting large amounts of additional acoustic and textual training data. Communities using ASR to support endangered language documentation efforts, who cannot easily collect more data, might instead focus on exploring multiple architectures and hyperparameterizations to optimize performance within the constraints of their available data and resources.

Visit

doi.orgunderline.io

Tasks

automatic speech recognitionspeech processing

Tags

Natural Language ProcessingLanguage ModelsText Summarization

Similar

Using different acoustic, lexical and language modeling units for ASR of an under-resourced language - AmharicAutomatic speech recognition for an under-resourced language - amharicEnriching the NArabizi Treebank: A Multifaceted Approach to Supporting an Under-Resourced LanguageSub-word Based End-to-End Speech Recognition for an Under-Resourced Language: AmharicDocument Classification for the Under-resourced Amharic LanguageShona spaCy: A Morphological Analyzer for an Under-Resourced Bantu Language

Using different acoustic, lexical and language modeling units for ASR of an under-resourced language - Amharic

State-of-the-art large vocabulary continuous speech recognition systems use mostly phone based acoustic models (AMs) and word based lexical and language models. However, phone based AMs are not efficient in modeling long-term temporal dependencies and the use of wo

Automatic speech recognition for an under-resourced language - amharic

Enriching the NArabizi Treebank: A Multifaceted Approach to Supporting an Under-Resourced Language

In this paper we address the scarcity of annotated data for NArabizi, a Romanized form of N

Sub-word Based End-to-End Speech Recognition for an Under-Resourced Language: Amharic

In this work, we focused on end-to-end speech recognition for less-resourced language, Amharic. The result can be integrated with other tasks such as spoken content retrieval. We explored three models, which consist of Convolutional Neural Networks, Recurrent Neura

Document Classification for the Under-resourced Amharic Language

NLP is severely hampered by a scarcity of digital resources. This is especially true for Amharic, a

Shona spaCy: A Morphological Analyzer for an Under-Resourced Bantu Language

Despite rapid advances in multilingual natural language processing (NLP), the Bantu language Shona r