Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
ZheCheMeiXu,
Hôte:avatar
Large-scale multilingual ASR (mASR) models such as Whisper achieve strong performance but incur high computational and latency costs, limiting their deployment on resource-constrained edge devices. In this study, we propose a lightweight and language-agnostic multilingual ASR system based on a CTC architecture with domain adaptation. Specifically, we introduce a Language-agnostic Hierarchical LoRA-MoE (HLoRA) framework integrated into an mHuBERT-CTC model, enabling end-to-end decoding via LID-posterior-driven LoRA routing. The hierarchical design consists of a multilingual shared LoRA for learning language-invariant acoustic representations and language-specific LoRA experts for modeling language-dependent characteristics. The proposed routing mechanism removes the need for prior language identity information or explicit language labels during inference, achieving true language-agnostic decoding. Experiments on MSR-86K and the MLC-SLM 2025 Challenge datasets demonstrate that HLoRA achieves comparable performance to two-stage inference approaches while reducing RTF by 11.7% and 8.2%, respectively, leading to improved decoding efficiency for low-resource mASR applications. 5 pages, submitted to IEEE Communications Letters

Visit

arxiv.org

Tasks

automatic speech recognitionlanguage identificationspeech processing

Tags

Computation and LanguageSoundAudio and Speech Processing

Similaires

Language-Agnostic Chatbot: A Cross-Lingual Conversational AI Framework for Multilingual Support Using Transformer-Based ArchitectureAllSpeak: A Language-Agnostic Runtime for Computational Literacy in Multilingual CommunitiesUnveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork AdaptationEYEDOL/parakeet-ctc-1.1b-yoruba-loraDARTS-ASR: Differentiable Architecture Search for Multilingual Speech Recognition and AdaptationA Blended Attention-CTC Network Architecture for Amharic Text-image Recognition

Language-Agnostic Chatbot: A Cross-Lingual Conversational AI Framework for Multilingual Support Using Transformer-Based Architecture

Language barriers remain a significant obstacle to equitable access to information and services in a

AllSpeak: A Language-Agnostic Runtime for Computational Literacy in Multilingual Communities

Language is the foundation of human development. It is what enables us to convey ideas, share knowle

Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation

Mixture-of-Experts (MoE) models exhibit striking performance disparities across languages, yet the i

EYEDOL/parakeet-ctc-1.1b-yoruba-lora

DARTS-ASR: Differentiable Architecture Search for Multilingual Speech Recognition and Adaptation

In previous works, only parameter weights of ASR models are optimized under fixed-topology architect

A Blended Attention-CTC Network Architecture for Amharic Text-image Recognition