Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

VoxLingua107

Domain:

natural language processing

Record type:

dataset
VoxLingua107 is a speech dataset for training spoken language identification models. The dataset consists of short speech segments automatically extracted from YouTube videos and labeled according the language of the video title and description, with some post-processing steps to filter out false positives. VoxLingua107 contains data for 107 languages. The total amount of speech in the training set is 6628 hours. The average amount of data per language is 62 hours. However, the real amount per language varies a lot. There is also a seperate development set containing 1609 speech segments from 33 languages, validated by at least two volunteers to really contain the given language.

Visit

bark.phon.ioc.eecs.taltech.ee

Connected records

papermodel

Tasks

keywordsautomatic speech recognitionspeech processinglanguage identification

Languages

AfrikaansAmharicHausaLingalaMalagasyShonaSomaliSwahiliYoruba

Tags

VoxLingua107

Licenses

Creative Commons Attribution 4.0 International License

Similar

VoxLingua107 ECAPA-TDNN Spoken Language Identification ModelVOXLINGUA107: A DATASET FOR SPOKEN LANGUAGE RECOGNITION

VoxLingua107 ECAPA-TDNN Spoken Language Identification Model

This is a spoken language recognition model trained on the VoxLingua107 dataset using SpeechBrain. The model uses the ECAPA-TDNN architecture that has previously been used for speaker recognition. The model can classify a speech utterance according to the language

VOXLINGUA107: A DATASET FOR SPOKEN LANGUAGE RECOGNITION

This paper investigates the use of automatically collected web audio data for the task of spoken language recognition. We generate semirandom search phrases from language-specific Wikipedia data that are then used to retrieve videos from YouTube for 107 languages.