Wikipedia dataset containing cleaned articles of all languages.
The datasets are built from the Wikipedia dump
(Wikipedia) with one split per language. Each example
contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).