Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

KARAKALPAK SPEECH CORPUS: THE FIRST BENCHMARK DATASET FOR AUTOMATIC SPEECH RECOGNITION

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
NiyKabJabTan
Éditeur:
Zenodo
Hôte:avatar
While large-scale pre-trained models have significantly advanced multilingual Automatic Speech Recognition (ASR), many low-resource languages remain under-served due to the scarcity of high-quality annotated speech corpora. This paper introduces the Karakalpak Speech Corpus (KSC), the first publicly available benchmark dataset for Karakalpak, a Turkic language spoken by over two million people primarily in Karakalpakstan. The corpus encompasses 50 hours of predominantly read speech. The data was collected from 25 native speakers with a balanced gender distribution. To establish a performance benchmark, we fine-tuned the Wav2Vec 2.0 architecture, specifically evaluating the efficacy of transfer learning from multilingual pre-trained models.

Visit

doi.orgzenodo.org

Tasks

automatic speech recognitionspeech processing

Tags

speech datasetspeech recognitionspeech-to-texttransfer learningWav2Vec 2.0 modelMachine LearningDeep Learning

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Speech Representation System For The Karakalpak Automatic Speech RecognitionAfriVox: An African benchmark dataset for Automatic Speech Translation and Speech RecognitionA New Tunisian Arabic Corpus and Benchmark for Automatic Speech Recognition"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech RecognitionThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

Speech Representation System For The Karakalpak Automatic Speech Recognition

Abstract Recently, automatic speech recognition (ASR) has moved from deeplearning

AfriVox: An African benchmark dataset for Automatic Speech Translation and Speech Recognition

This project creates a benchmark dataset for evaluating Automatic Speech Translation and Speech reco

A New Tunisian Arabic Corpus and Benchmark for Automatic Speech Recognition

"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"

"This dataset is a 20.03-hour Amharic speech corpus recorded from 100 native speakers and containing

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

The Makerere AI Lab has built an end-to-end CTC Luganda ASR model using radio data. Having encountered data challenges in working with low resource languages, we take the initiative together with our partners to release the first radio corpus for Luganda. The corp

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

Building a usable radio monitoring automatic speech recognition (ASR) system is a challenging task for under-resourced languages and yet this is paramount in societies where radio is the main medium of public communication and discussions. Initial efforts by the Un