Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Bridging the SEA Gap: An Initial Benchmark for Neural Audio Codec-Synthesized Speech Deepfakes in South-East Asian Languages

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
PhuGirAkhBud
Hôte:avatar
Codecfakes (CFs) are a type of speech deepfakes generated through Audio Language Models (ALMs), with Neural Audio Codecs (NACs) forming the core mechanism for speech encoding and generation. CFs exhibit distributional characteristics that differ from vocoder-based deepfakes, causing detectors trained on vocoder data to generalize poorly to CFs detection. Although this has led to the development of CF detection benchmarks, existing resources are largely confined to English -- and to a limited extent Chinese -- leaving South-East Asian (SEA) languages unexplored. To bridge this gap, we introduce SEA-CF, the first large-scale benchmark for CF detection spanning multiple SEA languages, diverse speaker profiles, and a wide range of NAC architectures. SEA-CF is constructed by synthesizing publicly available real speech corpora. Our experiments show that state-of-the-art (SOTA) CF detectors trained on English-centric datasets fail to generalize to SEA speech due to language-specific phonetic structures, tonal variations, and rich prosodic diversity. We further conduct a comprehensive zero-shot and fine-tuned evaluation of recent SOTA ALMs on SEA-CF. Fine-tuning the ALMs improves performance, however, these are very large being impractical for real-world application due to their scale, particularly in low-resource and latency-constrained settings. To address this limitation, we propose a novel small-ALM, GARUDA tailored for CF detection, which delivers strong performance while remaining lightweight. Extensive evaluations demonstrate that the proposed Small-ALM outperforms strong end-to-end and ALM-based baselines, establishing a new, practical direction for robust CF detection in SEA languages and beyond. Accepted to IJCAI-ECAI 2026

Visit

arxiv.org

Tasks

speech processing

Tags

Audio and Speech ProcessingSound

Similaires

InstituteforDiseaseModeling/Bridging-the-Gap-Low-Resource-African-LanguagesBridging the edge–cloud gap: adaptive AI for robust image and audio wildlife monitoringBridging the Data Provenance Gap Across Text, Speech and VideoSouth-East Asian Features in the Munda Languages: Evidence for the Analytic-to-Synthetic Drift of MundaDevelopment in South Africa: Bridging The Rural-Urban GapBRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages

InstituteforDiseaseModeling/Bridging-the-Gap-Low-Resource-African-Languages

Code and data corresponding to paper Bridging the Gap: Enhancing LLM Performance for Low-Resource Af

Bridging the edge–cloud gap: adaptive AI for robust image and audio wildlife monitoring

Artificial intelligence (AI) is transforming wildlife monitoring through automated analysis of image

Bridging the Data Provenance Gap Across Text, Speech and Video

Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established datasets beyond text. In this work we conduct the largest and first-of-its-kind longit

South-East Asian Features in the Munda Languages: Evidence for the Analytic-to-Synthetic Drift of Munda

Proceedings of the Twenty-Eighth Annual Meeting of the Berkeley Linguistics Society: Special Session

Development in South Africa: Bridging The Rural-Urban Gap

South Africa is a highly unequal society with high levels of inequality, poverty, unemployment and o

BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages

People worldwide use language in subtle and complex ways to express emotions. While emotion recognition -- an umbrella term for several NLP tasks -- significantly impacts different applications in NLP and other fields, most work in the area is focused on high-resou