Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Enhancing Automatic Speech Recognition for Child Speech in Low-Resource Languages

Domain:

natural language processing

Record type:

paper
Creator:
HasZar
Publisher:
Sha
Host:avatar

Automatic speech recognition (ASR) for children is demanding because their speech differs considerably from adult speech. Shorter vocal tracts, higher fundamental frequencies ($F_0$), and differences in language development can all affect recognition accuracy. This problem is even more difficult for low-resource languages such as Persian (Farsi), where standardized speech datasets for children are scarce. In this study, we investigate several approaches to improve Persian child speech recognition, including a text-based post-processing pipeline and Parameter-Efficient Fine-Tuning (PEFT/LoRA).

Initial experiments with off-the-shelf Vosk and Whisper-Base models resulted in relatively high Word Error Rates (WER). Many of the errors were related to phonetic ambiguity and incorrect or hallucinated words in the generated transcripts. To address these issues, we developed a multi-stage post-processing pipeline that combines exact lexicon matching, Levenshtein-distance-based fuzzy matching, and contextual text correction with large language models (LLMs). We also evaluated an audio-enabled multimodal model that processes the speech directly. This approach produced the lowest WER among the models and methods evaluated.

We further examined LoRA fine-tuning using both real recordings of children's speech and synthesized samples. Fine-tuning improved the performance of the smaller Whisper-Base model, but the same approach caused overfitting and catastrophic forgetting when applied to Whisper-Large-v3. The difference appears to be related to the limited amount of target-domain data relative to the model's capacity. Overall, the results suggest that, for Persian child speech and similarly low-resource settings, text-based correction and direct multimodal speech processing can be more practical than fine-tuning large acoustic models.

بازشناسی خودکار گفتار (ASR) برای خردسالان به دلیل ویژگی‌های خاص آناتومیکی مثل کوتاه بودن مجرای صوتی و بالا بودن فرکانس پایه و نیز نارسایی‌های شناختی، با چالش عدم تطابق دامنه مواجه است. این چالش در زبان‌های کم‌منبع مانند فارسی به دلیل کمبود شدید پایگاه داده گفتاری استاندارد کودکان دوچندان می‌شود. پژوهش حاضر با هدف بهبود کیفیت بازشناسی گفتار کودکان فارسی‌زبان، به پیاده‌سازی و ارزیابی یک خط لوله ترکیبی پس‌پردازش متنی و ارزیابی روش تنظیم دقیق کارا (PEFT/LoRA) می‌پردازد. در بخش ارزیابی مدل‌های پایه، مدل‌های خام Vosk و Whisper Base به دلیل توهم آکوستیکی و ابهام آوایی با نرخ خطای کلمه (WER) بسیار بالایی مواجه شدند. برای بهبود نتایج، یک خط لوله اصلاح متنی چندمرحله‌ای (تطابق قطعی واژگان، تطابق فازی مبتنی بر فاصله لون‌اشتاین و روان‌سازی متنی با مدل زبانی بزرگ) توسعه یافت که عملکرد مدل پایه را بهبود بخشید. همچنین، تغذیه مستقیم سیگنال صوتی به مدل چندوجهی بهترین عملکرد کلی پژوهش و کمترین میزان نرخ خطا را به ثبت رساند. ارزیابی فرآیند تنظیم دقیق لورا روی نمونه‌های واقعی گفتار کودکان و جملات صوتی شبیه‌سازی‌شده نشان داد که این روش موجب ارتقای کارایی مدل کوچک Whisper Base می‌شود؛ اما در مدل بزرگ Whisper Large-v3 به دلیل ظرفیت بالای پارامترها و محدودیت شدید داده‌های آموزشی، به بیش‌برازش و فراموشی فاجعه‌بار منجر شده و عملکرد آن را تخریب می‌کند. یافته‌های این پژوهش نشان می‌دهد در شرایط کم‌منبع، استفاده از خط لوله‌های پس‌پردازش زبانی یا مدل‌های چندوجهی مستقیم، رویکرد عملیاتی بسیار موفق‌تری نسبت به تنظیم دقیق آکوستیک است.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Tags

Automatic Speech Recognition (ASR)Children's Speech RecognitionPersian LanguageTextual Post-ProcessingLarge Language Models (LLMs)Parameter-Efficient Fine-Tuning (PEFT)LoRAWhisper

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode