Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

VOXLINGUA107: A DATASET FOR SPOKEN LANGUAGE RECOGNITION

Domaine:

natural language processing

Type de record:

paper
This paper investigates the use of automatically collected web audio data for the task of spoken language recognition. We generate semirandom search phrases from language-specific Wikipedia data that are then used to retrieve videos from YouTube for 107 languages. Speech activity detection and speaker diarization are used to extract segments from the videos that contain speech. Post-filtering is used to remove segments from the database that are likely not in the given language, increasing the proportion of correctly labeled segments to 98%, based on crowd-sourced verification. The size of the resulting training set (VoxLingua107) is 6628 hours (62 hours per language on the average) and it is accompanied by an evaluation set of 1609 verified utterances. We use the data to build language recognition models for several spoken language identification tasks. Experiments show that using the automatically retrieved training data gives competitive results to using hand-labeled proprietary datasets.

Visit

arxiv.org

Connected records

modeldataset

Tasks

keywordsautomatic speech recognitionspeech processinglanguage identification

Languages

AfrikaansAmharicHausaLingalaMalagasyShonaSomaliSwahiliYoruba

Tags

VoxLingua107

Similaires

VoxLingua107 ECAPA-TDNN Spoken Language Identification ModelTARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language UnderstandingA dataset for Moroccan sign language recognition and translationSwahili Speech Dataset Development and Improved Pre-training Method for Spoken Digit RecognitionVoxLingua107Alabib-65: A Realistic Dataset for Algerian Sign Language Recognition

VoxLingua107 ECAPA-TDNN Spoken Language Identification Model

This is a spoken language recognition model trained on the VoxLingua107 dataset using SpeechBrain. The model uses the ECAPA-TDNN architecture that has previously been used for speaker recognition. The model can classify a speech utterance according to the language

TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language Understanding

A dataset for Moroccan sign language recognition and translation

Swahili Speech Dataset Development and Improved Pre-training Method for Spoken Digit Recognition

Speech dataset is an essential component in building commercial speech applications. However, low-re

VoxLingua107

VoxLingua107 is a speech dataset for training spoken language identification models. The dataset consists of short speech segments automatically extracted from YouTube videos and labeled according the language of the video title and description, with some post-proc

Alabib-65: A Realistic Dataset for Algerian Sign Language Recognition

Sign language recognition (SLR) is a promising research field that aims to blur boundaries between D