Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion

Domaine:

natural language processing

Type de record:

paper
Créateur:
KumAmaGho
Éditeur:
arXiv
Hôte:avatar
Automatic Speech Recognition (ASR) and Dialect Identification (DID) are crucial for Indian languages, many of which are low-resource and exhibit significant dialectal differences. Existing methods often optimize ASR or DID individually, resulting in performance trade-offs. In this work, we propose a multimodal framework that jointly improves ASR and DID. Our method employs a Bottleneck Encoder to extract dialectal features from Conformer-based speech representations and a RoBERTa encoder to process ASR-generated CTC embeddings. A gating mechanism merges these features, followed by an attention encoder to refine the representations. The learned embeddings are concatenated with Conformer outputs to enhance ASR features. Evaluated on eight Indian languages with thirty-three dialects, our method achieves an average DID accuracy of 81.63% and average CER and WER of 4.65% and 17.73%, respectively. These results highlight the effectiveness of our method for joint ASR-DID modeling.

Visit

doi.org

Tasks

automatic speech recognitionlanguage identificationspeech processing

Tags

Computation and Language (cs.CL)Audio and Speech Processing (eess.AS)FOS: Computer and information sciencesFOS: Electrical engineering, electronic engineering, information engineering

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Arabic Dialect Identification Using iVectors and ASR TranscriptsEnhancing Biometric Security Using Artificial Neural Network-Based Multimodal Fusion of Facial Recognition and Fingerprint IdentificationFeature Fusion Using GSA for Multi-Instance Authentication SystemAmharic Fake News Detection on Social Media Using Feature FusionMultimodal In-context Learning for ASR of Low-resource LanguagesMultimodal Depression Detection from Speech and Text Using a Fusion Neural Network

Arabic Dialect Identification Using iVectors and ASR Transcripts

This paper presents the systems submitted by the MAZA team to the Arabic Dialect Identification (ADI) shared task at the VarDial Evaluation Campaign 2017. The goal of the task is to evaluate computational models to identify the dialect of Arabic utterances using bo

Enhancing Biometric Security Using Artificial Neural Network-Based Multimodal Fusion of Facial Recognition and Fingerprint Identification

International audience Biometric authentication systems based on a single modality re

Feature Fusion Using GSA for Multi-Instance Authentication System

International audience Multi-instance fusion of fingerprint authentication system at

Amharic Fake News Detection on Social Media Using Feature Fusion

These days, many people use social media as a source of information and medium of communication due

Multimodal In-context Learning for ASR of Low-resource Languages

Automatic speech recognition (ASR) still covers only a small fraction of the world's languages, main

Multimodal Depression Detection from Speech and Text Using a Fusion Neural Network

Depression is among the leading causes of disability worldwide, yet its detection continues to rely