
This playbook will help you to plan and manage projects to collect voice data for low-resource languages. It is aimed at both new and experienced teams and covers the full process, from setting up the project to publishing your dataset. We draw on CLEAR Global’s experience with our data collection platform “[TWB Voice]”. The playbook also covers aspects of data collection that apply to organizations and communities who want to collect voice data through other platforms or initiatives.
Around four billion people lack access to voice technologies like speech recognition and conversational AI. This is because their languages don’t have the data available to build these tools. This playbook aims to help address this gap by outlining the key steps, challenges, and best practices for collecting voice data in low-resource languages in an effective and ethical way.
In this chapter, we introduce TWB Voice, CLEAR Global’s new platform for collecting voice data. We refer to TWB Voice throughout this playbook. We also explain key terms and concepts in voice technology, and show you how to use the playbook.