Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

CMU Wilderness Multilingual Speech Dataset

Domain:

natural language processing

Record type:

dataset
A dataset of over 700 different languages providing audio, aligned text and word pronunciations. On average each language provides around 20 hours of sentence-lengthed transcriptions. Data is mined from read New Testaments from bible.is

Visit

github.com

Connected records

paper

Tasks

speech translation

Languages

AbidjiAcholiAdeleAdioukrouAkaAkebuAkooseAlurAmazighAnufo+236

Tags

speech alignment

Licenses

Similar

CMU Wilderness Multilingual Speech Dataset (Paper)

CMU Wilderness Multilingual Speech Dataset (Paper)

This paper describes the CMU Wilderness Multilingual Speech Dataset. A dataset of over 700 different languages providing audio, aligned text and word pronunciations. On average each language provides around 20 hours of sentence-lengthed transcriptions. We describe