Logo Lanfrica

Nigerian Salafi Discourse Corpus (NSDC): Construction, Methods, and Use

Domain:

natural language processing

Record type:

datasetpaper
Creator:
Oke
Publisher:
Zenodo
Host:avatar
This paper documents the construction of the Nigerian Salafi Discourse Corpus (NSDC), a purpose-built corpus of 5,791 Facebook posts (2,485,963 words; 2,518,453 tagged tokens) published by Nigerian pages and profiles over a ten-year frame (2015-2025). The NSDC was assembled from eight keyword-filtered exports of the Meta Content Library and processed through a fully documented, machine-verified pipeline executed by Hermes Agent, an agentic artificial-intelligence research assistant. The pipeline performs schema validation, three-pass deduplication, Unicode normalisation and noise removal, SGML-annotated corpus assembly, cross-keyword merging, and calendar-year segmentation. The corpus is annotated in three layers: Penn Treebank part-of-speech and lemma tagging; assignment of tokens to fifteen purpose-built semantic fields; and sentence-level coding of fifteen discourse strategies under the discourse-historical approach. This paper reports aggregate statistics, data-quality challenges and solutions, verification outcomes, a manual reliability audit of the annotation layers, and research applications.