Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Yor-Sarc: A gold-standard dataset for sarcasm detection in a low-resource African language

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
JimDe Nik
Hôte:avatar
Sarcasm detection poses a fundamental challenge in computational semantics, requiring models to resolve disparities between literal and intended meaning. The challenge is amplified in low-resource languages where annotated datasets are scarce or nonexistent. We present \textbf{Yor-Sarc}, the first gold-standard dataset for sarcasm detection in Yorùbá, a tonal Niger-Congo language spoken by over $50$ million people. The dataset comprises 436 instances annotated by three native speakers from diverse dialectal backgrounds using an annotation protocol specifically designed for Yorùbá sarcasm by taking culture into account. This protocol incorporates context-sensitive interpretation and community-informed guidelines and is accompanied by a comprehensive analysis of inter-annotator agreement to support replication in other African languages. Substantial to almost perfect agreement was achieved (Fleiss' $κ= 0.7660$; pairwise Cohen's $κ= 0.6732$--$0.8743$), with $83.3\%$ unanimous consensus. One annotator pair achieved almost perfect agreement ($κ= 0.8743$; $93.8\%$ raw agreement), exceeding a number of reported benchmarks for English sarcasm research works. The remaining $16.7\%$ majority-agreement cases are preserved as soft labels for uncertainty-aware modelling. Yor-Sarc\footnote{github.com is expected to facilitate research on semantic interpretation and culturally informed NLP for low-resource African languages.

Visit

arxiv.org

Tasks

sentiment analysistext classification

Languages

Yoruba

Tags

Computation and Language

Similaires

Multimodal Sarcasm Dataset Generation for a Low-Resource Language: Swahiliyor-sarcBuilding a Dataset for Misinformation Detection in the Low-Resource LanguageA Forensic Linguistic Dataset for Offensive Content Detection in Low-Resource Language: SetswanaHausaMovieReview: A Benchmark Dataset for Sentiment Analysis in Low-Resource African LanguageA Gold-Standard Dataset for Benchmarking Balinese Extractive and Abstractive Text Summarization

Multimodal Sarcasm Dataset Generation for a Low-Resource Language: Swahili

yor-sarc

Yor-Sarc is a gold-standard dataset for sarcasm detection in Yorùbá, a tonal and morphologically ric

Building a Dataset for Misinformation Detection in the Low-Resource Language

A Forensic Linguistic Dataset for Offensive Content Detection in Low-Resource Language: Setswana

Developing Monolingual Setswana Datasets for Offensive Content Detection Reproducibility Package, Me

HausaMovieReview: A Benchmark Dataset for Sentiment Analysis in Low-Resource African Language

The development of Natural Language Processing (NLP) tools for low-resource languages is critically

A Gold-Standard Dataset for Benchmarking Balinese Extractive and Abstractive Text Summarization

1. Research Hypothesis and Data Scope The central hypothesis guiding the creation of the BaliSummari