This dataset contains articles from the Somali Wikipedia, processed and formatted in JSON Lines (.js
Cleaned Somali Wikipedia corpus (~9,500 articles) for NLP, LLM training, and linguistic research #
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia
Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdow
Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wiki