Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

<b>A Pilot Speech Corpus for Studying Device and Environmental Variability in Voice Biometrics</b>

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Oye
Hôte:avatar

This dataset provides a curated pilot corpus for studying device and environmental variability in voice biometrics. It contains 480 speech recordings from 12 participants (Japan, Nigeria, Ivory Coast, France, Germany, and Indonesia), each contributing 40 utterances recorded across multiple devices and environments.

Recordings were made using the Samsung A04s, OnePlus Nord (both direct and in-call), iPhone 15 Pro, and a USB condenser microphone (connected to a MacBook), under both indoor (semi-controlled lobby) and outdoor (campus) conditions. All files are stored in WAV format (8–16 kHz, 16-bit PCM), accompanied by a metadata file (CSV/Excel) with anonymized attributes such as nationality, gender, age, and English proficiency.

The dataset supports research in speech enhancement (spectral subtraction, Wiener filtering, adaptive filtering), speaker identification and verification, spoofing resilience, and liveness detection. Validation experiments confirmed that adaptive filtering achieved the highest accuracy (97%), highlighting both the challenges of cross-device variability and the potential for robust enhancement methods.

This corpus provides a valuable benchmark for developing secure and consistent voice biometric systems, particularly in real-world applications such as mobile banking authentication and low-resource environments.

Visit

figshare.com

Tasks

speaker verificationspeech processing

Tags

Natural language processingSpeech productionSpeech recognitionAudio processingImage processingDeep learningNeural networksVoice biometricsSpeaker verificationSpeaker identification+7

Licenses

CC BY 4.0