The scraper programmatically accesses the public-facing archives of the selected fact-checking organisations and extracts report-level metadata for a specified date range.
The collected fields are:
organisation
report title
publication date
verdict label, where available
canonical URL
originating archive URL
retrieval timestamp
SHA-256 digest of the retrieved page
scraper version
Because the organisations use different website and archive structures, the script applies source-specific retrieval and parsing procedures.
PesaCheck is processed through its Medium publication archive. Congo Check and Debunk Media are processed through their respective paginated web archives.