Logo Lanfrica

djelia/bambara-texts

Domaine:

natural language processing

Type de record:

dataset
Créateur:
dje
Hôte:
The Bambara-Texts dataset is a collection of monolingual Bambara text designed for pretraining language models. It provides a diverse set of textual data to improve natural language processing (NLP) applications for the Bambara language. This dataset can be used for: Pretraining large language models (LLMs) Building word embeddings for Bambara Language modeling tasks such as masked language modeling (MLM) and autoregressive modeling