Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Fine-Tuning SLAM-ASR for Low-Resource Language Speech Recognition with High-Resource Alignment

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art performance in various tasks. However, their applicability is still less explored in low-resource settings. This work investigates the use of Speech LLMs for low-resource Automatic Speech Recognition using the SLAM-ASR framework, where a trainable lightweight projector connects a speech encoder and a LLM. Firstly, we assess training data volume requirements to match Whisper-only performance, re-emphasizing the challenges of limited data. Secondly, we show th Research goal: What is the impact of fine-tuning SLAM-ASR on low-resource language speech recognition when the LLM is first aligned with high-resource language speech data vs. text data? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.5/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.5/10.

Visit

doi.orgzenodo.org

Tasks

automatic speech recognitionspeech processing

Tags

impactfine-tuningSLAM-ASRlow-resourcelanguagespeechrecognitionLLM

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Fine-tuning Whisper Tiny for Swahili ASR: Challenges and Recommendations for Low-Resource Speech RecognitionInstructAlign: High-and-Low Resource Language Alignment via Continual CrosslingualInstruction TuningInstructAlign: High-and-Low Resource Language Alignment via Continual Crosslingual Instruction TuningSpeech Encoder Size and Projector Capacity Trade-offs in SLAM-ASR for High- and Low-Resource LanguagesKrishnateja244/Fine-tuning-of-ASR-models-on-low-resource-languagesWhispering in Amharic: Fine-tuning Whisper for Low-resource Language

Fine-tuning Whisper Tiny for Swahili ASR: Challenges and Recommendations for Low-Resource Speech Recognition

InstructAlign: High-and-Low Resource Language Alignment via Continual CrosslingualInstruction Tuning

Large language models (LLMs) that are tuned with instructions have demonstrated remarkable capabilit

InstructAlign: High-and-Low Resource Language Alignment via Continual Crosslingual Instruction Tuning

Large language models (LLMs) that are tuned with instructions have demonstrated remarkable capabilit

Speech Encoder Size and Projector Capacity Trade-offs in SLAM-ASR for High- and Low-Resource Languages

Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource

Krishnateja244/Fine-tuning-of-ASR-models-on-low-resource-languages

Fine tuning ASR models such as Wave2Vec2.0, Whisper, Nemo and MMS models on low-resource languages

Whispering in Amharic: Fine-tuning Whisper for Low-resource Language

This work explores fine-tuning OpenAI's Whisper automatic speech recognition (ASR) model for Amharic