Logo Lanfrica

Tigrinya Abusive Language Detection (TiALD) Dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
fga
Hôte:
TiALD is a large-scale, multi-task benchmark dataset for abusive language detection in the Tigrinya language. It consists of 13,717 YouTube comments annotated for abusiveness, sentiment, and topic tasks. The dataset includes comments written in both the Ge’ez script and prevalent non-standard Latin transliterations to mirror real-world usage.