Logo Lanfrica

sib200_14classes

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Dav
Hôte:
SIB-200 is the largest publicly available topic classification dataset based on Flores-200 covering 205 languages and dialects. The train/validation/test sets are available for all the 205 languages. This is another version with 14 classes, more idea for few-shot evaluation, it has 5 examples for few-shot, and larger test set (1225) Supported Tasks and Leaderboards