Yor-Sarc is a gold-standard dataset for sarcasm detection in Yorùbá, a tonal and morphologically rich low-resource African language spoken by over 50 million people. The dataset was created to address the scarcity of high-quality annotated resources for figurative language understanding in African NLP.
It contains 436 manually annotated Yorùbá text instances labeled for binary sarcasm classification.
Dataset Details