NOTE: I am not affiliated with Wikimedia or Wikipedia.
Wikipedia dataset containing cleaned articles of all languages.
The dataset is built from the Wikipedia dumps (Wikipedia)
with one subset per language, each containing a single train split.
Each example contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).