Logo Lanfrica

Afri Code Datasets (A collection of datasets for code generation in African languages)

Domain:

natural language processing

Record type:

dataset
Training and evaluating Large Language Models (LLMs) for code generation, building AI-powered coding assistants that understand African languages, research in low-resource NLP, improving developer productivity for non-English speakers. Notes / challenges: This is a collection, not a single dataset, so users must navigate to individual datasets for use. The primary challenge it addresses is the extreme scarcity of parallel code-text data for low-resource African languages. The quality, size, and language coverage may vary significantly between the individual datasets in the collection.