Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

djelia/bambara-texts

Domain:

natural language processing

Record type:

dataset
Creator:
dje
Host:
The Bambara-Texts dataset is a collection of monolingual Bambara text designed for pretraining language models. It provides a diverse set of textual data to improve natural language processing (NLP) applications for the Bambara language. This dataset can be used for: Pretraining large language models (LLMs) Building word embeddings for Bambara Language modeling tasks such as masked language modeling (MLM) and autoregressive modeling

Visit

huggingface.co

Tasks

language modeling

Languages

BamanankanLame

Similar

djelia/bambara-asrdjelia/bambara-audiodjelia/bambara-mt-v2djelia/bambara-tts-waxaldjelia/bambara-lm-qadjelia/bambara-asr-v2

djelia/bambara-asr

djelia/bambara-audio

The Djelia Bambara Audio Dataset is a comprehensive resource aimed at supporting research and develo

djelia/bambara-mt-v2

djelia/bambara-tts-waxal

djelia/bambara-lm-qa

The Bambara-LM-QA dataset is designed to support the fine-tuning of large language models (LLMs) for

djelia/bambara-asr-v2

Configuration Train (hours) Dev (hours) Test (hours) Total (hours) bm-to-bm-weak 136.37 23.40 11.28