Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

<b>A Pilot Speech Corpus for Studying Device and Environmental Variability in Voice Biometrics</b>

Domain:

natural language processing

Record type:

dataset
Creator:
Oye
Host:avatar

This dataset provides a curated pilot corpus for studying device and environmental variability in voice biometrics. It contains 480 speech recordings from 12 participants (Japan, Nigeria, Ivory Coast, France, Germany, and Indonesia), each contributing 40 utterances recorded across multiple devices and environments.

Recordings were made using the Samsung A04s, OnePlus Nord (both direct and in-call), iPhone 15 Pro, and a USB condenser microphone (connected to a MacBook), under both indoor (semi-controlled lobby) and outdoor (campus) conditions. All files are stored in WAV format (8–16 kHz, 16-bit PCM), accompanied by a metadata file (CSV/Excel) with anonymized attributes such as nationality, gender, age, and English proficiency.

The dataset supports research in speech enhancement (spectral subtraction, Wiener filtering, adaptive filtering), speaker identification and verification, spoofing resilience, and liveness detection. Validation experiments confirmed that adaptive filtering achieved the highest accuracy (97%), highlighting both the challenges of cross-device variability and the potential for robust enhancement methods.

This corpus provides a valuable benchmark for developing secure and consistent voice biometric systems, particularly in real-world applications such as mobile banking authentication and low-resource environments.

Visit

figshare.com

Tasks

speaker verificationspeech processing

Tags

Natural language processingSpeech productionSpeech recognitionAudio processingImage processingDeep learningNeural networksVoice biometricsSpeaker verificationSpeaker identification+7

Licenses

CC BY 4.0

Similar

Methods for studying memory B-cell immunity against malariaZambezi Voice: A Multilingual Speech Corpus for Zambian Languages<b>Nutritional and anti-nutritional variability in leaf and grain of cowpea (</b><b><i>Vigna unguiculata</i></b><b> (L). Walp.) genotypes</b>Common Voice: A Massively-Multilingual Speech Corpuscielo-b/Kinyarwanda-Voice-Assistant<b>BBC Igbo–Pidgin Gold-Standard NLP Corpus</b>

Methods for studying memory B-cell immunity against malaria

Plasmodium falciparum malaria remains one of the world’s deadliest infectious diseases and the se

Zambezi Voice: A Multilingual Speech Corpus for Zambian Languages

This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages. It contains two collections of datasets: unlabelled audio recordings of radio news and talk shows programs (160 hours) and labelled data (over 80 hours) consistin

<b>Nutritional and anti-nutritional variability in leaf and grain of cowpea (</b><b><i>Vigna unguiculata</i></b><b> (L). Walp.) genotypes</b>

This dataset presents grain and leaf nutritional and antinutritional trait data for 40 diverse cowpe

Common Voice: A Massively-Multilingual Speech Corpus

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identi

cielo-b/Kinyarwanda-Voice-Assistant

# Kinyarwanda Voice Assistant 🤖🗣️ **Developer**: Cielo B. (cielo-b) **Email**: irumvaregisdmc@gm

<b>BBC Igbo–Pidgin Gold-Standard NLP Corpus</b>

This corpus is a high-quality, manually annotated collection of BBC Igbo and BBC Pidgin