Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Somali Automatic Speech Recognition Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
sky
Host:
This dataset contains audio recordings and corresponding transcriptions in Somali, designed for automatic speech recognition (ASR) tasks. Language: Somali (so) Size: 10K - 100K samples Format: Parquet Modalities: Audio + Text License: CC-BY 4.0 Task: Automatic Speech Recognition Usage

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

Somali

Licenses

cc-by-4.0

Similar

Automatic Speech Recognition for Humanitarian Applications in SomaliVoxMg: An Automatic Speech Recognition Dataset for MalagasyKARAKALPAK SPEECH CORPUS: THE FIRST BENCHMARK DATASET FOR AUTOMATIC SPEECH RECOGNITIONAfriVox: An African benchmark dataset for Automatic Speech Translation and Speech RecognitionReproducible Automatic Speech RecognitionEnabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan Language

Automatic Speech Recognition for Humanitarian Applications in Somali

We present our first efforts in building an automatic speech recognition system for Somali, an under-resourced language, using 1.57 hrs of annotated speech for acoustic model training. The system is part of an ongoing effort by the United Nations (UN) to implement

VoxMg: An Automatic Speech Recognition Dataset for Malagasy

African languages are not well-represented in Natural Language Processing (NLP). The main reason is a lack of resources for training models. Low-resource languages, such as Malagasy, cannot benefit from modern NLP methods if no datasets are available. This paper pr

KARAKALPAK SPEECH CORPUS: THE FIRST BENCHMARK DATASET FOR AUTOMATIC SPEECH RECOGNITION

While large-scale pre-trained models have significantly advanced multilingual Automatic Speech Recog

AfriVox: An African benchmark dataset for Automatic Speech Translation and Speech Recognition

This project creates a benchmark dataset for evaluating Automatic Speech Translation and Speech reco

Reproducible Automatic Speech Recognition

The poster describes the architecture of one of the Sci-GaIA project "champion" use cases proposed f

Enabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan Language

The lack of impaired speech data hinders advancements in the development of inclusive speech technol