Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
JinYanShiKan
Hôte:avatar
The evolving speech processing landscape is increasingly focused on complex scenarios like meetings or cocktail parties with multiple simultaneous speakers and far-field conditions. Existing methodologies for addressing these challenges fall into two categories: multi-channel and single-channel solutions. Single-channel approaches, notable for their generality and convenience, do not require specific information about microphone arrays. This paper presents a large-scale far-field overlapping speech dataset, crafted to advance research in speech separation, recognition, and speaker diarization. This dataset is a critical resource for decoding ``Who said What and When'' in multi-talker, reverberant environments, a daunting challenge in the field. Additionally, we introduce a pipeline system encompassing speech separation, recognition, and diarization as a foundational benchmark. Evaluations on the WHAMR! dataset validate the broad applicability of the proposed data. InterSpeech 2024

Visit

arxiv.org

Tasks

speech processing

Tags

SoundComputation and LanguageAudio and Speech Processing

Similaires

"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"Dendi of Parakou multi-speaker speech datasetJamiil92/Dendi-of-Parakou-multi-speaker-speech-datasetAn Investigation Into Various Approaches For Bengali Long-Form Speech Transcription and Bengali Speaker Diarizationlamhaourii/darija-speaker-diarization-finetuningYoruba Multi-Speaker Speech Corpus

"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"

"This dataset is a 20.03-hour Amharic speech corpus recorded from 100 native speakers and containing

Dendi of Parakou multi-speaker speech dataset

This dataset was created for speech research purposes and contains about 676 recordings of participa

Jamiil92/Dendi-of-Parakou-multi-speaker-speech-dataset

:dart: :benin: This dataset was created for speech research purposes and contains about 676 recordin

An Investigation Into Various Approaches For Bengali Long-Form Speech Transcription and Bengali Speaker Diarization

Bengali remains a low-resource language in speech technology, especially for complex tasks like long

lamhaourii/darija-speaker-diarization-finetuning

Fine-tuned pyannote speaker diarization model specifically for Darija (Moroccan Arabic). This projec

Yoruba Multi-Speaker Speech Corpus

Yoruba tts notebook and data