SIB-200 is the largest publicly available topic classification dataset based on Flores-200 covering 205 languages and dialects.
The train/validation/test sets are available for all the 205 languages.
topic classification: categorize wikipedia sentences into topics e.g science/technology, sports or politics.
There are 205 languages available :