Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Swahili Speech Dataset Development and Improved Pre-training Method for Spoken Digit Recognition

Domain:

natural language processing

Record type:

datasetpaper
Creator:
AleQinJim
Publisher:
Ass
Host:
Speech dataset is an essential component in building commercial speech applications. However, low-resource languages such as Swahili lack such a resource that is vital for spoken digit recognition. For languages where such resources exist, they are usually insufficient. Thus, pre-training methods have been used with external resources to improve continuous speech recognition. However, to the best of our knowledge, no study has investigated the effect of pre-training methods specifically for spoken digit recognition. This study aimed at addressing these problems. First, we developed a Swahili spoken digit dataset for Swahili spoken digit recognition. Then, we investigated the effect of cross-lingual and multi-lingual pre-training methods on spoken digit recognition. Finally, we proposed an effective language-independent pre-training method for spoken digit recognition. The proposed method has the advantage of incorporating target language data during the pre-training stage that leads to an optimal solution when using less training data. Experiments on Swahili (being developed), English, and Gujarati datasets show that our method achieves better performance compared with all the baselines listed in this study.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Languages

Swahili

Licenses

https://www.acm.org/publications/policies/copyright_policy#Background

Similar

regak/Spoken-Swahili-Digit-DatasetFine-Tuning VGG19 with Mel Spectrograms for Amazigh Spoken Digit RecognitionAfroDigits: A Community-Driven Spoken Digit Dataset for African LanguagesAfroDigits: A Community-Driven Spoken Digit Dataset for African LanguagesImproved Speech Pre-Training with Supervision-Enhanced Acoustic UnitStable Distillation: Regularizing Continued Pre-training for Low-Resource Automatic Speech Recognition

regak/Spoken-Swahili-Digit-Dataset

# Spoken Swahili Digit Dataset Spoken Swahili Digit Dataset (SSDD) is a spoken digit dataset for Swa

Fine-Tuning VGG19 with Mel Spectrograms for Amazigh Spoken Digit Recognition

This paper shows an improvement in speech recognition performance for the Amazigh language using the

AfroDigits: A Community-Driven Spoken Digit Dataset for African Languages

The advancement of speech technologies has been remarkable, yet its integration with African languages remains limited due to the scarcity of African speech corpora. To address this issue, we present AfroDigits, a minimalist, community-driven dataset of spoken digi

AfroDigits: A Community-Driven Spoken Digit Dataset for African Languages

The advancement of speech technologies has been remarkable, yet its integration with African languages remains limited due to the scarcity of African speech corpora. To address this issue, we present AfroDigits, a minimalist, community-driven dataset of spoken digi

Improved Speech Pre-Training with Supervision-Enhanced Acoustic Unit

Speech pre-training has shown great success in learning useful and general latent representations fr

Stable Distillation: Regularizing Continued Pre-training for Low-Resource Automatic Speech Recognition

Continued self-supervised (SSL) pre-training for adapting existing SSL models to the target domain h