Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

aloginfile/wikipedia

Domain:

natural language processing

Record type:

dataset
Creator:
alo
Host:
Wikipedia dataset containing cleaned articles of all languages. The dataset is built from the Wikipedia dumps (Wikipedia) with one subset per language, each containing a single train split. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.).

Visit

huggingface.co

Tasks

language modeling

Languages

AfrikaansAmazighAmharicArabic, Egyptian SpokenArabic, Moroccan SpokenBamanankanChichewaDagaare, SouthernDagbaniÉwé+38

Licenses

cc-by-sa-3.0gfdl

Similar

Wikipedia: wikipedia-af (Afrikaans)Somaliska Wikipedia Somali WikipediaWikipediaWikipediaWikipediaWikipedia

Wikipedia: wikipedia-af (Afrikaans)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia

Somaliska Wikipedia Somali Wikipedia

Korpus av somaliska Wikipedia Corpus of Somali Wikipedia

Wikipedia

Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wiki

Wikipedia

Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wiki

Wikipedia

Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wiki

Wikipedia

Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdow