Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Post-ASR Correction in Hindi: Comparing Language Models and Large Language Models in Low-Resource Scenarios

Domaine:

natural language processing

Type de record:

paper
Créateur:
AssJyoKriKum
Éditeur:
Und
Hôte:avatar
Automatic Speech Recognition (ASR) systems for low-resource languages like Hindi often produce erroneous transcripts due to limited annotated data and linguistic complexity. Post-ASR correction using language models (LMs) and large language models (LLMs) offers a promising approach to improve transcription quality. In this work, we compare fine-tuned LMs (mT5, ByT5), fine-tuned LLMs (Nanda 10B), and instruction-tuned LLMs (GPT-4o-mini, LLaMA variants) for post-ASR correction in Hindi. Our findings reveal that smaller, fine-tuned models consistently outperform larger LLMs in both fine-tuning and in-context learning (ICL) settings. We observe a U-shaped inverse scaling trend under zero-shot ICL, where mid-sized LLMs degrade performance before marginal recovery at extreme scales, yet still fall short of fine-tuned models. ByT5 is more effective for character-level corrections such as transliteration and word segmentation, while mT5 handles broader semantic inconsistencies. We also identify performance drops in out-of-domain settings and propose mitigation strategies to preserve domain fidelity. In particular, we observe similar trends in Marathi and Telugu, indicating the broader applicability of our findings across low-resource Indian languages.

Visit

doi.orgunderline.io

Tasks

automatic speech recognitionspeech processing

Tags

Computational LinguisticsNatural Language ProcessingArtificial Intelligence

Similaires

Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and FongbeEmploying large language models in Swahili, a low-resource languageHuman review for post-training improvement of low-resource language performance in large language modelsCan Character-based Language Models Improve Downstream Task Performances In Low-Resource And Noisy Language Scenarios?Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?Large language models for frontline healthcare support in low-resource settings

Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe

Large language models (LLMs) are trained on data contributed by low-resource language communities, y

Employing large language models in Swahili, a low-resource language

Human review for post-training improvement of low-resource language performance in large language models

Large language models (LLMs) have significantly improved natural language processing, holding the p

Can Character-based Language Models Improve Downstream Task Performances In Low-Resource And Noisy Language Scenarios?

Recent impressive improvements in NLP, largely based on the success of contextual neural language models, have been mostly demonstrated on at most a couple dozen high- resource languages. Building language mod- els and, more generally, NLP systems for non- standard

Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?

Recent impressive improvements in NLP, largely based on the success of contextual neural la

Large language models for frontline healthcare support in low-resource settings

Abstract Large language models (LLMs) have demonstrated str