Developing Monolingual Setswana Datasets for Offensive Content Detection Reproducibility Package, Metadata, Tools, and Documentation Description
This Zenodo release archives the full reproducibility package for the manuscript "Developing Monolingual Setswana Datasets for Offensive Content Detection." The package includes documentation, code, annotation schema, evaluation scripts, and dataset-format specifications supporting the construction, validation, and intended reuse of a manually curated Setswana offensive-language dataset.
The work addresses the scarcity of publicly documented resources for low-resource African languages, focusing on the development of a structured Setswana corpus suitable for offensive-language detection research, digital forensic applications, and explainable AI investigations.
This release is compliant with FAIR principles—ensuring all non-sensitive artefacts are reusable, accessible, and transparently documented.
📁 Contents of the Release
Dataset Schema & Documentation (non-sensitive version)
Dataset description following OLID Zampierietal.,2019 Zampierietal.,2019 and HateCheck Ro¨ttgeretal.,2021 Rottgeretal.,2021 formats.
Column definitions (TEXT, TARGET, ID, SOURCE)
Annotation guidelines and labelling protocol
Ethical considerations and masking procedures
Dataset card (HuggingFace style)
⚠️ Note: The raw offensive text dataset is not included in this Zenodo release to ensure ethical and legal compliance. Instead, we provide:
metadata describing the dataset,
annotation schema,
sampling notes,
cleaning pipeline, and
usage guidelines.
Annotation Framework & Instructions
Includes:
Complete annotation manual
Examples of label categories (masked versions)
Decision rubrics for offensive vs non-offensive distinctions
Inter-annotator agreement (IAA) computation script (Cohen's kappa)
Annotator training notes
This allows reproducibility of annotation methodology without revealing harmful content.
Preprocessing, Cleaning & Structuring Scripts
Folder: scripts/preprocessing/
Includes:
Text cleaning pipeline
Normalisation for Setswana orthography
Token masking functions
Duplicate removal and source metadata stripping
Tools for exporting to OLID/HateCheck CSV structure
These scripts enable researchers to structure their own Setswana datasets using the same methodology.
Evaluation Scripts
Folder: scripts/evaluation/
Includes:
Binary-class evaluation (Accuracy, F1, MCC, ROC-AUC)
Stratified splitting tools
Error analysis helper functions
Label distribution checkers
These enable reproducibility of experimental analysis described in the paper.
Explainability & Counterfactuals (Sanitized)
Folder: outputs/explainability/
Contains:
LIME/S-LIME interpretation outputs (masked)
Counterfactual examples (masked)
explanation_notes.md describing methodology and sanitization policies
No harmful text is included.
Figures & Tables Used in the Manuscript
Folder: outputs/figures/ and outputs/tables/
Includes:
Dataset distribution plots
Annotator agreement figures
Offensive vs non-offensive category diagrams
Tokenisation breakdowns
Summary statistics
Table files (CSV + LaTeX)
These are all reproducible using provided scripts.
Supporting Notebooks
Folder: notebooks/
Provides:
Dataset structuring and cleaning notebook (safe version)
Annotation guidelines demonstration notebook
Explainability notebook (masked text)
Dataset export notebook
All notebooks are sanitised and suitable for public release.
🔒 Ethical & Safety Considerations
This release does not include:
raw offensive text
user-identifiable content
social media data
platform metadata
detailed examples containing sensitive content
All samples shown in figures or notebooks are masked using ***.
The dataset described in the manuscript should be accessed only through controlled academic channels and under strict ethical governance. This Zenodo package is designed to enable reproducibility without distributing harmful language.
🎯 Intended Use
This release supports:
NLP research in low-resource languages
Cyberbullying and hate speech detection studies
Digital forensics investigations
Linguistic analysis of Setswana online discourse
Educational purposes
Reproduction of the experimental methodology described in the manuscript
It is not intended for:
Real-time automated moderation
Law-enforcement without expert oversight
Deployment in production systems without additional safeguards
📖 Citation
If using this release, please cite the manuscript:
Kekgathetse, B. (2025). Developing Monolingual Setswana Datasets for Offensive Content Detection. Zenodo.
doi.org
📝 Licensing
Code: MIT / Apache-2.0
Documentation & annotation guidelines: CC-BY 4.0
Sanitized examples: CC-BY-SA 4.0
No raw data included
🌍 Audience
NLP researchers
Digital forensics experts
Computer security practitioners
Linguists working with Bantu languages
If you use this resource, please cite this work.