WURA is a document-level dataset covering 16 African Languages and 4 high-resource languages widely spoken in Africa (English, French, Arabic and Portuguese). This dataset was created by auditing mC4 and crawling additional verified news sources. It was first used to train AfriTeVa V2.
>>> from datasets import load_dataset