SARCSenti is a tone-marked Yoruba dataset developed for affective natural language processing, specifically binary sarcasm detection and ternary sentiment classification. The analytical release contains 1,507 Yoruba headline records, annotated for both sarcasm and sentiment. Sarcasm labels are non-sarcastic (0) and sarcastic (1), while sentiment labels are negative (0), neutral (1), and positive (2).
The dataset contains 1,334 BBC Yoruba-derived records and 173 Yoruba translations derived from the Misra/Kaggle News Headlines Dataset for Sarcasm Detection. Record-level provenance is provided in the source field. Three records from the original research workbook were excluded from the analytical release because of missing or invalid numeric sentiment labels.
SARCSenti was developed as part of research conducted in the Department of Computer Science, Lead City University, Ibadan, Nigeria. The dataset is intended to support research on Yoruba and African low-resource NLP, sarcasm detection, sentiment analysis, affective computing, cross-task learning, and the processing of tonal languages.
Yoruba text is provided in UTF-8 and normalized using Unicode NFC. The accompanying README and data dictionary provide information on dataset structure, labels, provenance, and usage.
Access to the full-text dataset is restricted because a substantial portion contains source-derived BBC Yoruba headline text for which open redistribution rights have not yet been confirmed. Research access may be considered subject to applicable source terms and copyright restrictions.