Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Common Voice Corpus 16

Domain:

natural language processing

Record type:

dataset
Creator:
eld
Host:
The Common Voice dataset consists of a unique MP3 and corresponding text file. Many of the 30328 recorded hours in the dataset also include demographic metadata like age, sex, and accent that can help improve the accuracy of speech recognition engines. The dataset currently consists of 19673 validated hours in 120 languages, but more voices and languages are always added.

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

AfrikaansAmazighAmharicBasaaGandaHausaIgboJulaKinyarwandaSwahili+4

Licenses

cc0-1.0