Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network

Domain:

natural language processing

Record type:

papermodeldataset
Creator:
ChaParLeeZhang, Yu
Host:avatar
We present SpeechStew, a speech recognition model that is trained on a combination of various publicly available speech recognition datasets: AMI, Broadcast News, Common Voice, LibriSpeech, Switchboard/Fisher, Tedlium, and Wall Street Journal. SpeechStew simply mixes all of these datasets together, without any special re-weighting or re-balancing of the datasets. SpeechStew achieves SoTA or near SoTA results across a variety of tasks, without the use of an external language model. Our results include 9.0\% WER on AMI-IHM, 4.7\% WER on Switchboard, 8.3\% WER on CallHome, and 1.3\% on WSJ, which significantly outperforms prior work with strong external language models. We also demonstrate that SpeechStew learns powerful transfer learning representations. We fine-tune SpeechStew on a noisy low resource speech dataset, CHiME-6. We achieve 38.9\% WER without a language model, which compares to 38.6\% WER to a strong HMM baseline with a language model. submitted to INTERSPEECH

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageMachine Learning

Similar

Neural Network Based Hausa Language Speech RecognitionNEURAL NETWORK BASED ARCHITECTURE FOR AUTOMATIC SPEECH RECOGNITION IN YORUBAThe Application of Probabilistic Neural Network in Speech Recognition Based on Partition ClusteringAdvanced Convolutional Neural Network-Based Hybrid Acoustic Models for Low-Resource Speech RecognitionConvolutional neural network for speech emotion recognition in the Moroccan Arabic dialect languageExploring data augmentation for Amazigh speech recognition with convolutional neural networks

Neural Network Based Hausa Language Speech Recognition

NEURAL NETWORK BASED ARCHITECTURE FOR AUTOMATIC SPEECH RECOGNITION IN YORUBA

The Application of Probabilistic Neural Network in Speech Recognition Based on Partition Clustering

A probabilistic neural network (PNN) speech recognition model based on the partition clustering algo

Advanced Convolutional Neural Network-Based Hybrid Acoustic Models for Low-Resource Speech Recognition

Deep neural networks (DNNs) have shown a great achievement in acoustic modeling for speech recogniti

Convolutional neural network for speech emotion recognition in the Moroccan Arabic dialect language

Extracting the speaker's emotional state has become an active research topic lately due to the deman

Exploring data augmentation for Amazigh speech recognition with convolutional neural networks