Logo Lanfrica

Dataset for "Enhancing multi-label emotion analysis in Indonesian social media with emoji-aware text representations"

Domain:

natural language processing

Record type:

dataset
Creator:
AmaLydRanMar
Publisher:
Zenodo
Host:avatar
Creator Amalia Amalia1*, Maya Silvi Lydia1, Rahmi Putri Rangkuti2, Farhan Purwanto Marulitua3, Fikri Hanif3, Mardanan Fitra3   1 Department of Computer Science, Universitas Sumatera Utara, Indonesia 2 Department of Psychology, Universitas Sumatera Utara, Indonesia 3 Department of Data Science and Artificial Intelligence, Universitas Sumatera Utara, Indonesia   Funding This research was supported by the Directorate of Research, Technology, and Community Service (DRTPM), Ministry of Education, Culture, Research, and Technology of the Republic of Indonesia, under the Fundamental Research Grant Scheme (Regular) 2025, based on Decree No. 0419/C3/DT.05.00/2025 and Contract/Agreement No. 112/C3/DT.05.00/PL/2025.   Description The dataset was constructed from Indonesian social media content, integrating both textual data and paralinguistic signals in the form of emojis. In total, it consists of 139,414 instances annotated for Text Emotion Analysis (TEA) based on Plutchik’s emotion model. Unlike single-label corpora, this dataset supports multi-label classification, where a single post may express more than one emotion simultaneously. For example, text accompanied by multiple emojis (e.g., 😍 and 😢) can convey both joy and sadness, resulting in overlapping emotion categories. To enrich the emotional representation, a vocabulary of 1,040 unique emojis was incorporated, capturing supportive, contrastive, or even sarcastic emotional cues. This makes the dataset a valuable resource for exploring multimodal and multi-label emotion analysis in Indonesian, a low-resource language where high-quality annotated datasets are scarce. The dataset is specifically designed to benchmark models that integrate paralinguistic signals and to evaluate the robustness of LLM-based TEA systems.

Similar