Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Sarcasm detection in Tamil and Malayalam YouTube comments

Domain:

natural language processing

Record type:

datasetpaper
Creator:
Cha
Editor:
UniUni
Publisher:
Springer
Host:avatar
The expression of sarcasm is a standard literary device where individuals deliberately convey the opposite of what is intended. Accurately identifying sarcasm in the text can aid in comprehending a speaker’s genuine intentions and facilitate other natural language processing activities, particularly sentiment analysis and offensive language identification tasks. We created a dataset for sarcasm from YouTube comments in Dravidian languages and manually annotated them for sarcasm in two Dravidian languages: Tamil (42,244 comments) and Malayalam (18,840 comments). Subsequently, we benchmarked the dataset by comparing it with different text classifiers. Among these approaches, pre-trained transformer models performed well, achieving an accuracy of 0.798 with TamBERT for Tamil and 0.852 with the MuRIL model for Malayalam. Furthermore, we employed SHAP values (explainable AI) to help understand how individual model inputs influence predictions. We also released the dataset on CodaLab and analyzed the participants’ systems. We have presented the results of shared task participants.

Visit

doi.orgresearchrepository.universityofgalway.ie

Tasks

sentiment analysistext classification

Tags

Sarcasm detectionUnder-resourced languagesMultilingualExplainability AIShared task

Licenses

CC BYCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Indonesian YouTube Comments DatasetAmharic Youtube Comments SentimentTriVector@DravidianLangTech 2026: Depression Detection from Tamil and Malayalam Speech with Speaker-Independent Evaluation using MFCC and Wav2Vec2Samuelogeno/sarcasm-detectionAbdelhamidElKrem/Darija-Sentiment-Analysis-of-YouTube-Commentssarcasm detection and quantification in arabic tweets

Indonesian YouTube Comments Dataset

Dataset ini berisi 5.291 komentar YouTube berbahasa Indonesia yang telah melalui tahapan preprocessi

Amharic Youtube Comments Sentiment

Movie Review Comments

TriVector@DravidianLangTech 2026: Depression Detection from Tamil and Malayalam Speech with Speaker-Independent Evaluation using MFCC and Wav2Vec2

Depression is a major mental health concern that can be reflected through subtle changes in speech p

Samuelogeno/sarcasm-detection

Language evolves rapidly. In Kenya, Gen Z and Gen Alpha have largely moved from "Habari yako" to "Ni

AbdelhamidElKrem/Darija-Sentiment-Analysis-of-YouTube-Comments

REAL-TIME DATA PROCESSING FOR SENTIMENT ANALYSIS OF MOROCCAN DIALECT IN YOUTUBE # Big-Data

sarcasm detection and quantification in arabic tweets

The role of predicting sarcasm in the text is known as automatic sarcasm detection. Given the preval