Logo Lanfrica

malayalam-factcheck-corpus

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Bha
Éditeur:
Zenodo
Hôte:avatar
A reproducible batch pipeline that ingests Malayalam fact-check articles from IFCN-certified Indian portals and produces a versioned, deduplicated Parquet corpus of (claim, verdict, evidence) triples for training claim-verification and misinformation-detection models. Source-attribution and verdict provenance are preserved on every record. If this corpus or pipeline is useful in your work, a citation using the metadata below would be appreciated. Where it fits, a citation to the underlying fact-check publishers (mapped via the source_id field on each record) alongside this one is appreciated too, since the verdicts and evidence originate with their reporting.