Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Sarcasm detection in Tamil and Malayalam YouTube comments

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
Cha
Éditeur:
UniUni
Éditeur:
Springer
Hôte:avatar
The expression of sarcasm is a standard literary device where individuals deliberately convey the opposite of what is intended. Accurately identifying sarcasm in the text can aid in comprehending a speaker’s genuine intentions and facilitate other natural language processing activities, particularly sentiment analysis and offensive language identification tasks. We created a dataset for sarcasm from YouTube comments in Dravidian languages and manually annotated them for sarcasm in two Dravidian languages: Tamil (42,244 comments) and Malayalam (18,840 comments). Subsequently, we benchmarked the dataset by comparing it with different text classifiers. Among these approaches, pre-trained transformer models performed well, achieving an accuracy of 0.798 with TamBERT for Tamil and 0.852 with the MuRIL model for Malayalam. Furthermore, we employed SHAP values (explainable AI) to help understand how individual model inputs influence predictions. We also released the dataset on CodaLab and analyzed the participants’ systems. We have presented the results of shared task participants.

Visit

doi.orgresearchrepository.universityofgalway.ie

Tasks

sentiment analysistext classification

Tags

Sarcasm detectionUnder-resourced languagesMultilingualExplainability AIShared task

Licenses

CC BYCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode