Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Multilingual Dataset Integration Strategies for Robust Audio Deepfake Detection: A SAFE Challenge System

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
AliSubBolAdu
Hôte:avatar
The SAFE Challenge evaluates synthetic speech detection across three tasks: unmodified audio, processed audio with compression artifacts, and laundered audio designed to evade detection. We systematically explore self-supervised learning (SSL) front-ends, training data compositions, and audio length configurations for robust deepfake detection. Our AASIST-based approach incorporates WavLM large frontend with RawBoost augmentation, trained on a multilingual dataset of 256,600 samples spanning 9 languages and over 70 TTS systems from CodecFake, MLAAD v5, SpoofCeleb, Famous Figures, and MAILABS. Through extensive experimentation with different SSL front-ends, three training data versions, and two audio lengths, we achieved second place in both Task 1 (unmodified audio detection) and Task 3 (laundered audio detection), demonstrating strong generalization and robustness. Accepted @ IEEE ASRU 2025

Visit

arxiv.org

Tasks

speech processing

Tags

Audio and Speech ProcessingMachine Learning

Similaires

Deepfake Africa ChallengeEnglish-Centric Multilingual Audio DatasetEnglish-Centric Multilingual Audio DatasetDeepfake-Synthetic-20K Datasetregak/Swahili-Deepfake-datasetSCDF: A Speaker Characteristics DeepFake Speech Dataset for Bias Analysis

Deepfake Africa Challenge

Shining a light on deepfake media & tools in Africa
There is no data for this challenge. You can use any data / tools and platforms you like.
The only restriction is that the tools, techniques, programming languages, libraries, and/or online platform

English-Centric Multilingual Audio Dataset

This dataset contains generated article and summary audio for English-centric multilingual direction

English-Centric Multilingual Audio Dataset

This dataset contains generated article and summary audio for English-centric multilingual direction

Deepfake-Synthetic-20K Dataset

The Deepfake-Synthetic-20K dataset is a significant contribution to the field of digital forensics a

regak/Swahili-Deepfake-dataset

Creation of Real and fake Swahili voices # Swahili-Deepfake-dataset A multi-corpus, multi-generato

SCDF: A Speaker Characteristics DeepFake Speech Dataset for Bias Analysis

Despite growing attention to deepfake speech detection, the aspects of bias and fairness remain unde