Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A Corpus of English-Hindi Code-Mixed Tweets for Sarcasm Detection

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
SwaKhaSinAkh
Hôte:avatar
Social media platforms like twitter and facebook have be- come two of the largest mediums used by people to express their views to- wards different topics. Generation of such large user data has made NLP tasks like sentiment analysis and opinion mining much more important. Using sarcasm in texts on social media has become a popular trend lately. Using sarcasm reverses the meaning and polarity of what is implied by the text which poses challenge for many NLP tasks. The task of sarcasm detection in text is gaining more and more importance for both commer- cial and security services. We present the first English-Hindi code-mixed dataset of tweets marked for presence of sarcasm and irony where each token is also annotated with a language tag. We present a baseline su- pervised classification system developed using the same dataset which achieves an average F-score of 78.4 after using random forest classifier and performing 10-fold cross validation. 9 pages, CICLing 2018

Visit

arxiv.org

Tasks

code switching

Tags

Computation and Language

Similaires

sarcasm detection and quantification in arabic tweetsMedSumm: A Multimodal Approach to Summarizing Code-Mixed Hindi-English Clinical QueriesHate speech detection in Hausa code-mixed tweets using machine learningPreparing Bengali-English Code-Mixed Corpus for Sentiment Analysis of Indian LanguagesMisinformation detection in Luganda-English code-mixed social media textSamuelogeno/sarcasm-detection

sarcasm detection and quantification in arabic tweets

The role of predicting sarcasm in the text is known as automatic sarcasm detection. Given the preval

MedSumm: A Multimodal Approach to Summarizing Code-Mixed Hindi-English Clinical Queries

In the healthcare domain, summarizing medical questions posed by patients is critical for improving

Hate speech detection in Hausa code-mixed tweets using machine learning

Code-mixed communication in Nigeria, involving English, Nigerian Pidgin, Hausa, Yoruba, and Igbo, po

Preparing Bengali-English Code-Mixed Corpus for Sentiment Analysis of Indian Languages

Analysis of informative contents and sentiments of social users has been attempted quite intensively

Misinformation detection in Luganda-English code-mixed social media text

The increasing occurrence, forms, and negative effects of misinformation on social media platforms h

Samuelogeno/sarcasm-detection

Language evolves rapidly. In Kenya, Gen Z and Gen Alpha have largely moved from "Habari yako" to "Ni