The Nigerian Salafi Discourse Corpus (NSDC) is a purpose-built corpus of 5,791
Facebook posts (2,485,963 words; 2,518,453 tagged tokens; 149,431 sentences)
published by Nigerian pages and profiles between 2015 and 2025. The
corpus was assembled from eight keyword-filtered exports of the Meta Content
Library (Salafis, Salafism, Salafiyyah, Ahlus-Sunnah, Sunni, Wahabism,
Wahabiyyah, Salafool) and processed through a fully documented, machine-verified
pipeline: schema validation, three-pass deduplication, Unicode normalisation and
script-based noise removal, SGML-annotated corpus assembly, cross-keyword
merging, and calendar-year segmentation.
This release contains three artefacts: (1) the cleaned corpus (one UTF-8 file per
post, SGML metadata tags, one sentence per line); (2) the tagged vertical corpus
(Penn Treebank part-of-speech and lemma, 15 semantic fields, 15 discourse
strategies), in the vertical format used by Sketch Engine and CQP; and (3) the
corpus segmented by calendar year for diachronic analysis.