A multi-label emotion classification dataset for Amharic language, designed for research and development in natural language processing and emotion detection
Amharic Multi-Label Emotion Dataset
Description
This dataset contains Amharic text samples collected from YouTube, Facebook, and Twitter using APIs and comment extractors. The text data is annotated for multiple emotions by professional psychologists. Each emotion label is binary: '1' indicates the presence of that emotion in the text, while '0' denotes its absence. The dataset is intended for multi-label emotion classification research in the Amharic language.
Data Collection
• Text data sourced from social media platforms (YouTube, Facebook, Twitter) via API and comment extraction tools.
• Annotation performed by trained psychology professionals to ensure accurate emotion labeling.
• Multi-label setting allows each text to be associated with multiple emotions simultaneously.
Dataset Structure
• Format: CSV file
• Columns: Text content and Multi-label emotion labels(sad,happy,anger,fear,disgust,surprise,contempt,neutral)
• Number of samples: 17983
• Label encoding: 1 = emotion present, 0 = emotion absent
Usage
This dataset is useful for training and evaluating machine learning models for emotion detection and natural language understanding specific to the Amharic language. Researchers and practitioners can utilize it for tasks like sentiment analysis, social media monitoring, and psychological studies.
Licensed by
CC BY 4.0
Contact
For any questions or collaboration opportunities, contact Yeshimebet Bayu at yeshi8590@gmail.com/ yeshimebet_baye@dmu.edu.et