Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Mai Ho'omāuna i ka 'Ai: Language Models Improve Automatic Speech Recognition in Hawaiian

Domain:

natural language processing

Record type:

papermodel
Creator:
ChaZarFisKim
Host:avatar
In this paper we address the challenge of improving Automatic Speech Recognition (ASR) for a low-resource language, Hawaiian, by incorporating large amounts of independent text data into an ASR foundation model, Whisper. To do this, we train an external language model (LM) on ~1.5M words of Hawaiian text. We then use the LM to rescore Whisper and compute word error rates (WERs) on a manually curated test set of labeled Hawaiian data. As a baseline, we use Whisper without an external LM. Experimental results reveal a small but significant improvement in WER when ASR outputs are rescored with a Hawaiian LM. The results support leveraging all available data in the development of ASR systems for underrepresented languages.

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageMachine LearningSoundAudio and Speech Processing

Similar

Using Songs to Improve Kazakh Automatic Speech RecognitionFirst automatic fongbe continuous speech recognition system: Development of acoustic models and language modelsBenchmarking Automatic Speech Recognition Models for African LanguagesAutomatic Speech Recognition for the Ika LanguageAutomatic speech recognition of the isiZulu languageLanguage variation, automatic speech recognition and algorithmic bias

Using Songs to Improve Kazakh Automatic Speech Recognition

Developing automatic speech recognition (ASR) systems for low-resource languages is hindered by the

First automatic fongbe continuous speech recognition system: Development of acoustic models and language models

This paper reports our efforts toward an ASR system for a new under-resourced language (Fongbe). The aim of this work is to build acoustic models and language models for continuous speech decoding in Fongbe. The problem encountered with Fongbe (an African language

Benchmarking Automatic Speech Recognition Models for African Languages

Automatic speech recognition (ASR) for African languages remains constrained by limited labeled data

Automatic Speech Recognition for the Ika Language

We present a cost-effective approach for developing Automatic Speech Recognition (ASR) models for lo

Automatic speech recognition of the isiZulu language

A key component of artificial intelligence is human-to-machine communication. Such communication has

Language variation, automatic speech recognition and algorithmic bias

In this thesis, I situate the impacts of automatic speech recognition systems in relation to socioli