Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Speech Denoising Without Clean Training Data: A Noise2Noise Approach

Domaine:

natural language processing

Type de record:

paper
Créateur:
KasTamManNat
Hôte:avatar
This paper tackles the problem of the heavy dependence of clean speech data required by deep learning based audio-denoising methods by showing that it is possible to train deep speech denoising networks using only noisy speech samples. Conventional wisdom dictates that in order to achieve good speech denoising performance, there is a requirement for a large quantity of both noisy speech samples and perfectly clean speech samples, resulting in a need for expensive audio recording equipment and extremely controlled soundproof recording studios. These requirements pose significant challenges in data collection, especially in economically disadvantaged regions and for low resource languages. This work shows that speech denoising deep neural networks can be successfully trained utilizing only noisy training audio. Furthermore it is revealed that such training regimes achieve superior denoising performance over conventional training regimes utilizing clean training audio targets, in cases involving complex noise distributions and low Signal-to-Noise ratios (high noise environments). This is demonstrated through experiments studying the efficacy of our proposed approach over both real-world noises and synthetic noises using the 20 layered Deep Complex U-Net architecture. Published in Interspeech 2021 ( See isca-speech.org ). 5 pages, 2 figures, 1 table

Visit

arxiv.org

Tasks

speech processing

Tags

SoundMachine LearningAudio and Speech ProcessingI.2.6; I.5.4

Similaires

Speechless: Speech Instruction Training Without Speech for Low Resource LanguagesLLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Modelskalilouisangare/bambara-speech-kis-clean-splitA hybrid denoising and artificial neural network approach for diesel fuel price predictionMonolingual and cross-lingual intent detection without training data in target languagesTranslators without Borders data

Speechless: Speech Instruction Training Without Speech for Low Resource Languages

The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need f

LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models

Despite the advances in neural text to speech (TTS), many Arabic dialectal varieties remain marginal

kalilouisangare/bambara-speech-kis-clean-split

Dataset de reconnaissance vocale en bambara, nettoyé et splitté pour le fine-tuning de modèles ASR (

A hybrid denoising and artificial neural network approach for diesel fuel price prediction

Diesel fuel price (DFP) modeling and prediction are important to an economy since fuel price has dir

Monolingual and cross-lingual intent detection without training data in target languages

International audience Due to recent DNN advancements, many NLP problems can be effec

Translators without Borders data

Speech, parallel text corpora, models, papers for various African languages