Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

djelia/bambara-audio

Domain:

natural language processing

Record type:

dataset
Creator:
dje
Host:
The Djelia Bambara Audio Dataset is a comprehensive resource aimed at supporting research and development in Bambara language processing. This dataset consists of audio extracted from YouTube videos, denoised and diarized to ensure high-quality segments. Additionally, it features a semi-annotated subset with transcriptions generated using the Djelia Whisper v1 model. Features

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

BamanankanLame

Similar

djelia/bambara-synthetic-audiodjelia/bambara-audio-bdjelia/bambara-textsdjelia/bambara-asrdjelia/bambara-asr-v2djelia/bambara-mt-v2

djelia/bambara-synthetic-audio

djelia/bambara-audio-b

djelia/bambara-texts

The Bambara-Texts dataset is a collection of monolingual Bambara text designed for pretraining langu

djelia/bambara-asr

djelia/bambara-asr-v2

Configuration Train (hours) Dev (hours) Test (hours) Total (hours) bm-to-bm-weak 136.37 23.40 11.28

djelia/bambara-mt-v2