Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models

Domain:

natural language processing

Record type:

paper
Creator:
DasLasAhmKal
Publisher:
arXiv
Host:avatar
Developing Automatic Speech Recognition (ASR) for morphologically rich, low-resource languages such as Assamese is challenging due to insufficient annotated speech data. The pretrained Whisper model performs poorly on Assamese speech recognition tasks. This paper presents a controlled, fine-tuned Whisper-based Assamese ASR system trained on the Mozilla Common Voice 24.0-Assamese corpus. A hardware-aware optimized training pipeline is implemented for resource-constrained environments, employing mixed-precision training and gradient accumulation on Tesla 4 Graphics Processing Units (T4 GPUs). The proposed fine-tuned model significantly outperformed the Zero-shot baseline, yielding Word Error Rate (WER), Character Error Rate (CER), Match Error Rate (MER), and Word Infomation Loss (WIL) of 43.17\%, 13.18\%, 43\%, and 64.81\%, respectively, achieving significant relative improvements of 78.26\%, 93.10\%, 57.0\%, and 35.19\% over the baseline. Semantic evaluation of the fine-tuned model also demonstrates notable improvement over a zero baseline, attaining Bilingual Evaluation Understudy (BLEU) and Metric for Evaluation of Translation with Explicit ORdering (METEOR) scores of 30.81 and 0.5262, respectively. Additionally, the predicted hallucination rate and Real-Time Factor (RTF) are substantially improved by 96.70\% and 32.38\%, compared to the zero-shot baseline. 31 pages, 11 figures

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Tags

Machine Learning (cs.LG)FOS: Computer and information sciences

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/

Similar

sykeb21/Fine-Tuning-OpenAI-Whisper-for-Amharic-Speech-RecognitionAfrikaans Speech Dataset for Whisper Fine-TuningEfficient fine Tuning of Whisper for Algerian Dialect Automatic Speech Recognition using Low Rank AdaptationFine-tuning Whisper Tiny for Swahili ASR: Challenges and Recommendations for Low-Resource Speech RecognitionWitty-Kitty/Phoneme-Recognition-through-Fine-Tuning-of-Phonetic-RepresentationsAV55CS/Speech-to-Text-Fine-tuning-Whisper-Tiny-for-Swahili-ASR

sykeb21/Fine-Tuning-OpenAI-Whisper-for-Amharic-Speech-Recognition

This repository contains a complete, memory-optimized pipeline for fine-tuning OpenAI’s Whisper auto

Afrikaans Speech Dataset for Whisper Fine-Tuning

Dataset Card This dataset consists of approximately 56 hours of Afrikaans speech extracted from chu

Efficient fine Tuning of Whisper for Algerian Dialect Automatic Speech Recognition using Low Rank Adaptation

Fine-tuning Whisper Tiny for Swahili ASR: Challenges and Recommendations for Low-Resource Speech Recognition

Witty-Kitty/Phoneme-Recognition-through-Fine-Tuning-of-Phonetic-Representations

This repository makes available code for running expetiments to finetune Allosaurus on Bukusu and Sa

AV55CS/Speech-to-Text-Fine-tuning-Whisper-Tiny-for-Swahili-ASR

ASR-Low resource Language-Swahili # Swahili Speech Recognition - Whisper Fine-tuning This reposito