Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Remithefirst/yoruba-academic-wikipedia-qa

Domain:

natural language processing

Record type:

dataset
Creator:
Rem
Host:
This dataset contains ~5,000 instruction-style academic question–answer pairs in pure Yorùbá with tone marks, derived from Yorùbá Wikipedia. Each sample follows: instruction: academic question in Yorùbá input: empty output: academic answer in Yorùbá Instruction tuning Low-resource language modeling Educational assistants Yorùbá Wikipedia (cleaned & deduplicated)

Visit

huggingface.co

Tasks

question answering

Languages

Yoruba

Similar

hammamwahab/qa-wikipedia-sudanWikipedia: wikipedia-af (Afrikaans)Somaliska Wikipedia Somali WikipediaUnveiling the veiled: Wikipedia collaborating with academic libraries in Africa in creating visibility for African women through Art+Feminism Wikipedia edit-a-thonWikipediaWikipedia

hammamwahab/qa-wikipedia-sudan

This is a synthetic dataset based on "neuml/txtai-wikipedia" embedding index. The generation of stat

Wikipedia: wikipedia-af (Afrikaans)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia

Somaliska Wikipedia Somali Wikipedia

Korpus av somaliska Wikipedia Corpus of Somali Wikipedia

Unveiling the veiled: Wikipedia collaborating with academic libraries in Africa in creating visibility for African women through Art+Feminism Wikipedia edit-a-thon

Purpose This study aims to show that digital literacy can serve as a tool for effecting social cha

Wikipedia

Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdow

Wikipedia

Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wiki