Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

SCDF: A Speaker Characteristics DeepFake Speech Dataset for Bias Analysis

Domain:

natural language processing

Record type:

paperdataset
Creator:
StaSrnFirMal
Host:avatar
Despite growing attention to deepfake speech detection, the aspects of bias and fairness remain underexplored in the speech domain. To address this gap, we introduce the Speaker Characteristics Deepfake (SCDF) dataset: a novel, richly annotated resource enabling systematic evaluation of demographic biases in deepfake speech detection. SCDF contains over 237,000 utterances in a balanced representation of both male and female speakers spanning five languages and a wide age range. We evaluate several state-of-the-art detectors and show that speaker characteristics significantly influence detection performance, revealing disparities across sex, language, age, and synthesizer type. These findings highlight the need for bias-aware development and provide a foundation for building non-discriminatory deepfake detection systems aligned with ethical and regulatory standards.

Visit

arxiv.org

Tasks

speech processing

Tags

SoundArtificial IntelligenceCryptography and Security

Similar

"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"Dendi of Parakou multi-speaker speech datasetJamiil92/Dendi-of-Parakou-multi-speaker-speech-datasetDeepfake-Synthetic-20K Datasetregak/Swahili-Deepfake-datasetBias in Automated Speaker Recognition

"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"

"This dataset is a 20.03-hour Amharic speech corpus recorded from 100 native speakers and containing

Dendi of Parakou multi-speaker speech dataset

This dataset was created for speech research purposes and contains about 676 recordings of participa

Jamiil92/Dendi-of-Parakou-multi-speaker-speech-dataset

:dart: :benin: This dataset was created for speech research purposes and contains about 676 recordin

Deepfake-Synthetic-20K Dataset

The Deepfake-Synthetic-20K dataset is a significant contribution to the field of digital forensics a

regak/Swahili-Deepfake-dataset

Creation of Real and fake Swahili voices # Swahili-Deepfake-dataset A multi-corpus, multi-generato

Bias in Automated Speaker Recognition

Automated speaker recognition uses data processing to identify speakers by their voice. Today, automated speaker recognition technologies are deployed on billions of smart devices and in services such as call centres. Despite their wide-scale deployment and known s