Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Dataset for "Enhancing multi-label emotion analysis in Indonesian social media with emoji-aware text representations"

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AmaLydRanMar
Éditeur:
Zenodo
Hôte:avatar
Creator Amalia Amalia1*, Maya Silvi Lydia1, Rahmi Putri Rangkuti2, Farhan Purwanto Marulitua3, Fikri Hanif3, Mardanan Fitra3   1 Department of Computer Science, Universitas Sumatera Utara, Indonesia 2 Department of Psychology, Universitas Sumatera Utara, Indonesia 3 Department of Data Science and Artificial Intelligence, Universitas Sumatera Utara, Indonesia   Funding This research was supported by the Directorate of Research, Technology, and Community Service (DRTPM), Ministry of Education, Culture, Research, and Technology of the Republic of Indonesia, under the Fundamental Research Grant Scheme (Regular) 2025, based on Decree No. 0419/C3/DT.05.00/2025 and Contract/Agreement No. 112/C3/DT.05.00/PL/2025.   Description The dataset was constructed from Indonesian social media content, integrating both textual data and paralinguistic signals in the form of emojis. In total, it consists of 139,414 instances annotated for Text Emotion Analysis (TEA) based on Plutchik’s emotion model. Unlike single-label corpora, this dataset supports multi-label classification, where a single post may express more than one emotion simultaneously. For example, text accompanied by multiple emojis (e.g., 😍 and 😢) can convey both joy and sadness, resulting in overlapping emotion categories. To enrich the emotional representation, a vocabulary of 1,040 unique emojis was incorporated, capturing supportive, contrastive, or even sarcastic emotional cues. This makes the dataset a valuable resource for exploring multimodal and multi-label emotion analysis in Indonesian, a low-resource language where high-quality annotated datasets are scarce. The dataset is specifically designed to benchmark models that integrate paralinguistic signals and to evaluate the robustness of LLM-based TEA systems.

Visit

doi.orgzenodo.org

Tasks

emotion identificationtext classification

Tags

Text Emotion AnalysisMultimodal FusionEmoji-aware RepresentationMulti-label ClassificationLow-resource LanguageArtificial IntelligenceComputational LinguisticsNatural Language and SpeechSentiment AnalysisNeural Networks

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcodeCopyright (C) 2025 The Authors.http://rightsstatements.org/vocab/InC/1.0/

Similaires

An explainable AfroXLMR approach for multi-label emotion classification of Amharic social media text with dataset releaseEnhancing Multi-Label Emotion Analysis and Corresponding Intensities for Ethiopian LanguagesHausa Emotion-Tagged Tweets Dataset for Multi-Label Emotion ClassificationMulti-label Emotion Classification on Social Media Comments using Deep learningBESCAD: A Multi-Label Bangla Emotion and Social Context Annotation DatasetEthiopicEmotion: Multi-label Emotion Dataset with Large Language Models Evaluation

An explainable AfroXLMR approach for multi-label emotion classification of Amharic social media text with dataset release

Abstract Emotion detection from social media is crucial for understanding human

Enhancing Multi-Label Emotion Analysis and Corresponding Intensities for Ethiopian Languages

Developing and integrating emotion-understanding models are essential for a wide range of human-comp

Hausa Emotion-Tagged Tweets Dataset for Multi-Label Emotion Classification

The dataset comprises 19,757 Hausa tweets, each annotated with 11 distinct emotions: anger, sadness,

Multi-label Emotion Classification on Social Media Comments using Deep learning

Abstract Social media is an online platform that people use to develop social networks or

BESCAD: A Multi-Label Bangla Emotion and Social Context Annotation Dataset

BESCAD (Bangla Emotion and Social Context Annotation Dataset) is a large-scale multi-label dataset o

EthiopicEmotion: Multi-label Emotion Dataset with Large Language Models Evaluation

Large Language Models (LLMs) show promising learning and reasoning abilities. Compared to other NLP