Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Afaan Oromo Twitter–Facebook Rumor Dataset for Cross-Lingual and Cross-Platform Rumor Detection

Domain:

natural language processing

Record type:

dataset
Creator:
G, Gan
Publisher:
Zenodo
Host:avatar
This dataset provides a curated and research-ready collection of 9,804 Afaan Oromo social-media records for rumor detection, cross-lingual learning, and cross-platform misinformation analysis. The corpus contains content from Twitter and Facebook and is designed to support research in low-resource natural language processing, multilingual misinformation detection, domain generalization, and cross-platform rumor classification. The dataset contains 5,019 Rumor instances (51.19%) and 4,785 Non-Rumor instances (48.81%), providing a nearly balanced binary classification setting. In terms of platform distribution, 5,400 records (55.08%) originate from Twitter and 4,404 records (44.92%) from Facebook, allowing bidirectional Twitter-to-Facebook and Facebook-to-Twitter generalization studies. Each record was independently evaluated by three annotators. Among the 9,804 instances, 9,583 records (97.75%) received unanimous annotations, while 221 records (2.25%) were resolved through majority voting. Inter-annotator reliability was high, with an overall Fleiss’ κ of 0.9699 (approximately 0.97). Pairwise Cohen’s κ values were approximately 0.97, indicating near-perfect agreement among annotators. The dataset includes predefined training, validation, and test partitions of approximately 70%, 10%, and 20%, respectively, to facilitate reproducible experimentation. Records are additionally organized across topical categories including Public Information, Politics, Social Event, Education, Economy, and Health. For reproducible content-based experiments, the release includes both the complete preprocessed text representation and a claim_text field. The claim_text representation captures the central proposition while removing source- or verification-status framing and is intended as the primary model input for content-only rumor-detection experiments. The accompanying package also provides annotation fields, final labels, platform information, split information, category metadata, a data dictionary, annotation guidelines, preprocessing documentation, label descriptions, and reproducibility notes. The dataset was developed as part of the AMTE-DIAL (Adaptive Multi-Transformer Ensemble with Domain-Invariant Attention Learning) research framework and supports evaluation of multilingual and cross-platform rumor-detection systems, particularly in the low-resource Afaan Oromo setting. The released research representation contains no retained direct personally identifying information and is intended for academic research, reproducibility, benchmarking, and development of responsible misinformation-detection methods. Users should respect applicable social-media platform policies, ethical research practices, and the licensing conditions accompanying the dataset. Keywords: Afaan Oromo; Rumor Detection; Misinformation Detection; Low-Resource NLP; Cross-Lingual Learning; Cross-Platform Rumor Detection; Twitter; Facebook; Multilingual NLP; Domain Generalization; AMTE-DIAL.

Visit

doi.org

Tasks

text classification

Languages

OromoOromo, Borana-Arsi-Guji

Tags

Afaan Oromo Oromo language rumor detection misinformation detection fake news detection low-resource language natural language processing social media Twitter Facebook cross-lingual learning cross-platform learning domain adaptation multilingual transformers human annotation

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Curated Afaan Oromo Rumor Detection Dataset for Cross-Platform and Cross-Lingual Rumor Detection

Curated Afaan Oromo Rumor Detection Dataset for Cross-Platform and Cross-Lingual Rumor Detection

The dataset includes 9,680 records, consisting of 4,956 Rumor records and 4,724 Non-Rumor records. T