Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

SAGE: Spliced-Audio Generated Data for Enhancing Foundational Models in Low-Resource Arabic-English Code-Switched Speech Recognition

Domaine:

natural language processing

Type de record:

paper
Créateur:
FarSaz
Hôte:avatar
This paper investigates the performance of various speech SSL models on dialectal Arabic (DA) and Arabic-English code-switched (CS) speech. To address data scarcity, a modified audio-splicing approach is introduced to generate artificial CS speech data. Fine-tuning an already fine-tuned SSL model with the proposed Spliced-Audio Generated (SAGE) data results in an absolute improvement on Word Error Rate (WER) of 7.8% on Arabic and English CS benchmarks. Additionally, an Experience Replay (ER) inspired approach is proposed to enhance generalisation across DA and CS speech while mitigating catastrophic forgetting. Integrating an out-of-domain 3-gram language model reduces the overall mean WER from 31.7% to 26.6%. Few-shot fine-tuning for code-switching benchmarks further improves WER by 4.9%. A WER of 31.1% on Arabic-English CS benchmarks surpasses large-scale multilingual models, including USM and Whisper-large-v2 (both over ten times larger) by an absolute margin of 5.5% and 8.4%, respectively. Accepted for IEEE MLSP 2025

Visit

arxiv.org

Tasks

automatic speech recognitioncode switchingspeech processing

Tags

Computation and LanguageSoundAudio and Speech Processing

Similaires

Enhancing multilingual automatic speech recognition for low-resource code-switched languages: a scalable data augmentation strategyLarge language models (LLMs)-generated Afrikaans-English code-switched dataInvestigations on Speech Recognition Systems for Low-Resource Dialectal Arabic-English Code-Switching SpeechEnglish-IsiZulu Code-Switched Speech Recognition DatasetAutomatic Code-switched Academic Tunisian Arabic Speech RecognitionLeveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition

Enhancing multilingual automatic speech recognition for low-resource code-switched languages: a scalable data augmentation strategy

This research addresses the lack of annotated code-switched (CS) speech data for low-resource langua

Large language models (LLMs)-generated Afrikaans-English code-switched data

As highlighted in recent surveys, one of the biggest barriers to progress in code-switc

Investigations on Speech Recognition Systems for Low-Resource Dialectal Arabic-English Code-Switching Speech

Code-switching (CS), defined as the mixing of languages in conversations, has become a worldwide phe

English-IsiZulu Code-Switched Speech Recognition Dataset

Dataset for semi-supervised acoustic and language model training for English-isiZulu code-switched s

Automatic Code-switched Academic Tunisian Arabic Speech Recognition

Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition

Crafting an effective Automatic Speech Recognition (ASR) solution for dialects demands innovative ap