This dataset provides a curated and research-ready collection of 9,804 Afaan Oromo social-media records for rumor detection, cross-lingual learning, and cross-platform misinformation analysis. The corpus contains content from Twitter and Facebook and is designed to support research in low-resource natural language processing, multilingual misinformation detection, domain generalization, and cross-platform rumor classification.
The dataset contains 5,019 Rumor instances (51.19%) and 4,785 Non-Rumor instances (48.81%), providing a nearly balanced binary classification setting. In terms of platform distribution, 5,400 records (55.08%) originate from Twitter and 4,404 records (44.92%) from Facebook, allowing bidirectional Twitter-to-Facebook and Facebook-to-Twitter generalization studies.
Each record was independently evaluated by three annotators. Among the 9,804 instances, 9,583 records (97.75%) received unanimous annotations, while 221 records (2.25%) were resolved through majority voting. Inter-annotator reliability was high, with an overall Fleiss’ κ of 0.9699 (approximately 0.97). Pairwise Cohen’s κ values were approximately 0.97, indicating near-perfect agreement among annotators.
The dataset includes predefined training, validation, and test partitions of approximately 70%, 10%, and 20%, respectively, to facilitate reproducible experimentation. Records are additionally organized across topical categories including Public Information, Politics, Social Event, Education, Economy, and Health.
For reproducible content-based experiments, the release includes both the complete preprocessed text representation and a claim_text field. The claim_text representation captures the central proposition while removing source- or verification-status framing and is intended as the primary model input for content-only rumor-detection experiments. The accompanying package also provides annotation fields, final labels, platform information, split information, category metadata, a data dictionary, annotation guidelines, preprocessing documentation, label descriptions, and reproducibility notes.
The dataset was developed as part of the AMTE-DIAL (Adaptive Multi-Transformer Ensemble with Domain-Invariant Attention Learning) research framework and supports evaluation of multilingual and cross-platform rumor-detection systems, particularly in the low-resource Afaan Oromo setting.
The released research representation contains no retained direct personally identifying information and is intended for academic research, reproducibility, benchmarking, and development of responsible misinformation-detection methods. Users should respect applicable social-media platform policies, ethical research practices, and the licensing conditions accompanying the dataset.
Keywords: Afaan Oromo; Rumor Detection; Misinformation Detection; Low-Resource NLP; Cross-Lingual Learning; Cross-Platform Rumor Detection; Twitter; Facebook; Multilingual NLP; Domain Generalization; AMTE-DIAL.