Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Classification of Debunking in Indonesian Fact-Checking Platforms Using NLP and Machine Learning : A Mixed-Methods Approach with Corpus Analysis and IndoBERT

Domaine:

natural language processing

Type de record:

paper
Créateur:
BayRidRudHon
Éditeur:
LPP
Hôte:
The rapid spread of disinformation through digital platforms constitutes a serious threat to social cohesion and public health. Debunking—the systematic refutation of false information using verified evidence—has emerged as a key countermeasure, yet manual identification and classification of debunking strategies is labor-intensive and difficult to scale. This study addresses this gap through a mixed-methods design integrating qualitative corpus analysis with automated machine learning (ML) classification. A corpus of 120 debunking articles published by three leading Indonesian fact-checking institutions (Kominfo AIS, Mafindo, and Cek Fakta Kompas, 2022–2024) was first manually annotated by two trained coders (Cohen's κ = 0.82) to identify four dominant debunking strategies: (1) contextual correction with emotional narrative framing; (2) source authority endorsement; (3) visual verification and reverse image search; and (4) myth-versus-fact inoculation format. This annotated corpus was subsequently used as a training dataset to develop and benchmark five NLP-based text classification models: TF-IDF + Support Vector Machine (SVM), TF-IDF + Random Forest, IndoBERT fine-tuned, IndoBERT with data augmentation (IndoBERT-Aug), and XGBoost with linguistic features. The IndoBERT-Aug model achieved the highest overall performance (macro-averaged F1 = 0.847, Precision = 0.851, Recall = 0.843), substantially outperforming the SVM baseline (F1 = 0.612). Logistic regression analysis further identified three significant moderators of debunking effectiveness: correction timeliness within 6 hours (OR=2.80, p<0.01), content readability (OR=0.68, p<0.01), and multi-platform distribution (OR=1.84, p<0.05), with the full model explaining 41% of variance (Nagelkerke R²=0.41). These contributions are formalized into the Indonesian Debunking Effectiveness Model (IDEM), a framework integrating automated strategy detection with evidence-based deployment guidelines for scalable counter-disinformation operations.

Visit

doi.org

Tasks

text classification

Languages

Kono

Licenses

https://creativecommons.org/licenses/by-sa/4.0

Similaires

A Mixed-Methods Analysis of Repression and Mobilization in Bangladesh's July Revolution Using Machine Learning and Statistical ModelingA clinical support system for classification and prediction of depression using machine learning methodsModeling random events using a mixed approach Machine Learning techniquesCholera Epidemic Prediction Using Internet News Data in West Africa: A NLP and Machine Learning Approach with Transformer-BasedEthnicity Classification: A Machine Learning ApproachFact-Checking, Epistemic Identity, and Rationality (Replication with the Nigeria subjects)

A Mixed-Methods Analysis of Repression and Mobilization in Bangladesh's July Revolution Using Machine Learning and Statistical Modeling

The 2024 July Revolution in Bangladesh represents a landmark event in the study of civil resistance.

A clinical support system for classification and prediction of depression using machine learning methods

Abstract The health sector collects a very large amount of data, hence the diagnostic process proce

Modeling random events using a mixed approach Machine Learning techniques

Modeling random events using a mixed approach Machine Learning techniques 

Poster presented at the Deep Learning Indaba 2023 by KABEYA MWEPU Simon Isaac

Cholera Epidemic Prediction Using Internet News Data in West Africa: A NLP and Machine Learning Approach with Transformer-Based

Cholera remains a persistent public health threat in West Africa, where formal surveillance pipeline

Ethnicity Classification: A Machine Learning Approach

Abstract-Recently, researchers in the field of Machine Learnin

Fact-Checking, Epistemic Identity, and Rationality (Replication with the Nigeria subjects)

The purpose of this Study is to replicate the findings of Study 2 (Citizen Fact-Checkers) with parti