Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

DynAg Open Voice Dataset for Low-Resource Bihari Languages

Record type:

dataset
Creator:
Dex
Editor:
DexSum
Publisher:
Zenodo
Host:avatar
jjkhkjhkjh

Visit

doi.orgzenodo.org

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

TWB Voice Playbook for voice data collection for low-resource languagesAfrican Voices: Multilingual Speech Dataset for Low-Resource African LanguagesCorpus Voice Dataset Creation in Low-Resource Contexts: A systematic reviewTowards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource LanguagesA Tri-Class Multilingual Phishing Email Dataset for Low-Resource Languages EvaluationPashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language

TWB Voice Playbook for voice data collection for low-resource languages

This playbook will help you to plan and manage projects to collect voice data for low-resource languages. It is aimed at both new and experienced teams and covers the full process, from setting up the project to publishing your dataset. We draw on CLEAR Global’s

African Voices: Multilingual Speech Dataset for Low-Resource African Languages

A large-scale multilingual speech dataset developed by Data Science Nigeria. Contains more than 3,000 hours of transcribed audio across four Nigerian languages: Hausa, Igbo, Nigerian Pidgin, and Yorùbá. The dataset supports Automatic Speech Recognition (ASR) and sp

Corpus Voice Dataset Creation in Low-Resource Contexts: A systematic review

Voice corpora are fundamental resources for developing speech technologies, such as automatic speech

Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages

Automatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, p

A Tri-Class Multilingual Phishing Email Dataset for Low-Resource Languages Evaluation

Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language

We present the Pashto Common Voice corpus -- the first large-scale, openly licensed speech resource