Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Creation of a Nigerian Voice Corpus for Indigenous Speaker Recognition

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AdeEmmJokAde
Éditeur:
IOP
Hôte:
Abstract One of the goals of Word Bank’s Identification for Development (ID4D) is the realization of robust digital identification systems as a means of sustainable development priority. ID4D’s most recent report shows about 1.1 billion of the world’s population are yet to be identified for development. Africa represents about half of that number while Nigeria represents about a quarter of Africa’s share. Biometrics is the state-of-the-art approach for identification using human behavioral and/or physiological digitally calibrated traits and one such trait is the voice. The backbone of biometric research is the database employed in the design of biometric systems. Although many voice databases are publicly available such as the THCHS-30 for Chinese and Microsoft Indian language Speech Corpus for Indians, none is currently publicly available or free for Nigerians. The creation of such an indigenous database (or corpus) can open doors to Nigerian automatic speaker recognition as well as for indigenous language, ethnicity, gender, age group and emotion classification amongst others. This work is a first step in the direction of creating a Nigerian Voice Corpus (NVC) to aid indigenous voice biometric research. A voice corpus of popular Nigerians was created by curation of audio samples of 14 women and 23 men from YouTube. The corpus contains 10 different samples of 5 seconds duration for each individual resulting in a total of 370 samples. The created corpus was used to carry out speaker recognition experiment by dividing the audio samples into 25ms non-overlapping frame durations. Silent frames were excluded using short-term spectral energy threshold for Voice Activity Detection (VAD). This was followed by extraction of Mel Frequency Cepstral Coefficient (MFCC) as descriptors to discriminate different speakers using Support Vector Machine (SVM) with median Gaussian function. An overall recognition accuracy of 93.24% was achieved demonstrating the feasibility and research potential in this direction.

Visit

doi.org

Tasks

speaker verificationspeech processing

Licenses

http://creativecommons.org/licenses/by/3.0/https://iopscience.iop.org/info/page/text-and-data-mining

Similaires

"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"Corpus Voice Dataset Creation in Low-Resource Contexts: A systematic reviewPashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource LanguageSPEAKER RECOGNITION OF MAGHREB DIALECTSA Comprehensive Kurdish Speech Corpus for Speaker Identification and VerificationVoice Conversion Based Speaker Normalization for Acoustic Unit Discovery

"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"

"This dataset is a 20.03-hour Amharic speech corpus recorded from 100 native speakers and containing

Corpus Voice Dataset Creation in Low-Resource Contexts: A systematic review

Voice corpora are fundamental resources for developing speech technologies, such as automatic speech

Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language

We present the Pashto Common Voice corpus -- the first large-scale, openly licensed speech resource

SPEAKER RECOGNITION OF MAGHREB DIALECTS

A few studies have focused on the west Arabic (Maghreb) dialects for which resources are ra

A Comprehensive Kurdish Speech Corpus for Speaker Identification and Verification

Abstract / General Description: This dataset comprises a proprietary acoustic corpus specifically de

Voice Conversion Based Speaker Normalization for Acoustic Unit Discovery

Discovering speaker independent acoustic units purely from spoken input is known to be a hard proble