Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

TWB Voice Playbook for voice data collection for low-resource languages

Domaine:

natural language processing

Type de record:

projectmedia
Créateur:
CLEAR Global

This playbook will help you to plan and manage projects to collect voice data for low-resource languages. It is aimed at both new and experienced teams and covers the full process, from setting up the project to publishing your dataset. We draw on CLEAR Global’s experience with our data collection platform “[TWB Voice]”. The playbook also covers aspects of data collection that apply to organizations and communities who want to collect voice data through other platforms or initiatives.

Around four billion people lack access to voice technologies like speech recognition and conversational AI. This is because their languages don’t have the data available to build these tools. This playbook aims to help address this gap by outlining the key steps, challenges, and best practices for collecting voice data in low-resource languages in an effective and ethical way.

In this chapter, we introduce TWB Voice, CLEAR Global’s new platform for collecting voice data. We refer to TWB Voice throughout this playbook. We also explain key terms and concepts in voice technology, and show you how to use the playbook.

Visit

twbvoiceplaybook.clearglobal.org

Tasks

speech processing

Tags

data collectionplaybookTWB Voiceclear global

Similaires

DynAg Open Voice Dataset for Low-Resource Bihari LanguagesTWB Voice 1.0 - KanuriTWB Voice 1.0 - HausaTWB Voice Platform: Donate your voice, make a differenceErwinSchillack/Comon-Voice-Extraction-Data-Collection-South-African-LanguagesCLEAR-Global/TWB-Voice-1.0

DynAg Open Voice Dataset for Low-Resource Bihari Languages

jjkhkjhkjh

TWB Voice 1.0 - Kanuri

TWB Voice 1.0 - Kanuri is the Kanuri language portion of the TWB Voice 1.0 multilingual speech corpu

TWB Voice 1.0 - Hausa

TWB Voice 1.0 - Hausa is the Hausa language portion of the TWB Voice 1.0 multilingual speech corpus,

TWB Voice Platform: Donate your voice, make a difference

TWB Voice is a platform owned by CLEAR Global/Translators without Borders, which is designed to collect voice recordings from many users. These recordings are used to create large voice datasets, which are essential for developing language models for language te

ErwinSchillack/Comon-Voice-Extraction-Data-Collection-South-African-Languages

Collection and extraction of data from South African languages # Comon-Voice-Extraction-Data-Collec

CLEAR-Global/TWB-Voice-1.0

TWB Voice 1.0 is a multilingual speech corpus containing read speech data in three languages from Ni