Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The-African-Research-Collective/african-llm-datasets

Domaine:

natural language processing

Type de record:

dataset
Créateur:
The
Hôte:
A curated list of LLM datasets for African languages. # African LLM Datasets 🌍 > “More data beats clever algorithms, but better data beats more data” — Peter Norvig This repository is intended to serve as a practical resource listing available LLM training (pretraining and post-training) datasets that include one or more African languages. Data is arguably the most critical ingredient in training language models, yet for African languages it is often difficult to determine what datasets actually exist, where to find them. The repository covers pretraining data, instruction-tuning datasets, and evaluation datasets, with detailed metadata for each dataset. Where possible, we provide per-language breakdowns as well as train/validation/test splits, making it easier to understand the true data coverage and applicability of each dataset. This resource is actively evolving. Our long-term goal is to turn it into a comprehensive dataset discovery and utility tool, fully integrated into the **africanlanguages** package, to support researchers and practitioners working on African language technologies. --- ## 📚 Table of Contents - Contributing - Table Schema \& Column Definitions - Pre-training Datasets - Instruction Tuning Datasets (SFT) - Evaluation Datasets - Dataset Details --- ## Contributing Want to contribute? Here are two ways: - **Add new datasets:** Submit a Pull Request with complete information (see Table Schema). - **Update or correct entries:** Open an issue or submit a Pull Request! --- ## Table Schema & Column Definitions All dataset tables use the following columns: | Column | Description | |------|------------| | **Dataset** | Dataset name | | **Link** | Primary hosting location (Hugging Face, GitHub, etc.) | | **Total Size** | Total number of samples / examples | | **Language Breakdown** | Per-language coverage and approximate counts | | **Splits** | Available splits (train / dev / test) | | **Domain** | Dataset domain (General, QA, Math, Safety, etc.) | | **Type** | Data origin: Human, Synthet …

Visit

github.com

Tasks

language modeling

Similaires

The-African-Research-Collective/iteweThe-African-Research-Collective/translation_benchThe-African-Research-Collective/ogbufoThe-African-Research-Collective/africanlanguagesThe-African-Research-Collective/post-trainingThe-African-Research-Collective/karanta-ocr

The-African-Research-Collective/itewe

This code contains all the processing and scripting for processing and running OCR on documents usin

The-African-Research-Collective/translation_bench

# Translation Experiments A suite of translation experiments across a variety of models and dataset

The-African-Research-Collective/ogbufo

Translation seems to be a key component of our work, The goal of this repository is to create a set

The-African-Research-Collective/africanlanguages

# africanlanguages ## Setup You'll need to install uv. * Install `africanlanguages` by running `u

The-African-Research-Collective/post-training

Language Model Adapation via Direct Policy Optimization of Existing LLMs # Exploring Post Training

The-African-Research-Collective/karanta-ocr

karantaOCR -- Efficient Document Processing for African Languages Karanta OCR # KarantaOCR: Effic