Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AfriVox: An African benchmark dataset for Automatic Speech Translation and Speech Recognition

Domain:

natural language processing

Record type:

dataset
Creator:
int
Host:
This project creates a benchmark dataset for evaluating Automatic Speech Translation and Speech recognition models on African languages. This benchmark dataset covers 18 African languages. See language details below. This work is licensed under a

Visit

huggingface.co

Tasks

automatic speech recognitionmachine translationspeech processingspeech translation

Languages

AfrikaansAkanAmharicGaHausaIgboKinyarwandaSetswanaShonaSotho, Northern+5

Similar

AfriVox-Translate: An African benchmark dataset for Automatic Speech TranslationKARAKALPAK SPEECH CORPUS: THE FIRST BENCHMARK DATASET FOR AUTOMATIC SPEECH RECOGNITIONAfriVox-v2: A Domain-Verticalized Benchmark for In-the-Wild African Speech RecognitionVoxMg: An Automatic Speech Recognition Dataset for MalagasySomali Automatic Speech Recognition DatasetDvoice : An open source dataset for Automatic Speech Recognition on African Languages and Dialects

AfriVox-Translate: An African benchmark dataset for Automatic Speech Translation

This project creates a benchmark dataset for evaluating Automatic Speech Translation models on Afric

KARAKALPAK SPEECH CORPUS: THE FIRST BENCHMARK DATASET FOR AUTOMATIC SPEECH RECOGNITION

While large-scale pre-trained models have significantly advanced multilingual Automatic Speech Recog

AfriVox-v2: A Domain-Verticalized Benchmark for In-the-Wild African Speech Recognition

Recent large language models (LLMs) show strong speech recognition and translation capabilities for

VoxMg: An Automatic Speech Recognition Dataset for Malagasy

African languages are not well-represented in Natural Language Processing (NLP). The main reason is a lack of resources for training models. Low-resource languages, such as Malagasy, cannot benefit from modern NLP methods if no datasets are available. This paper pr

Somali Automatic Speech Recognition Dataset

This dataset contains audio recordings and corresponding transcriptions in Somali, designed for auto

Dvoice : An open source dataset for Automatic Speech Recognition on African Languages and Dialects

DVoice is a community initiative that aims to provide African languages and dialects with d