Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Forensic Linguistic Dataset for Offensive Content Detection in Low-Resource Language: Setswana

Domain:

natural language processing

Record type:

dataset
Creator:
Kek
Publisher:
Zenodo
Host:avatar
Developing Monolingual Setswana Datasets for Offensive Content Detection Reproducibility Package, Metadata, Tools, and Documentation Description This Zenodo release archives the full reproducibility package for the manuscript "Developing Monolingual Setswana Datasets for Offensive Content Detection." The package includes documentation, code, annotation schema, evaluation scripts, and dataset-format specifications supporting the construction, validation, and intended reuse of a manually curated Setswana offensive-language dataset. The work addresses the scarcity of publicly documented resources for low-resource African languages, focusing on the development of a structured Setswana corpus suitable for offensive-language detection research, digital forensic applications, and explainable AI investigations. This release is compliant with FAIR principles—ensuring all non-sensitive artefacts are reusable, accessible, and transparently documented. 📁 Contents of the Release Dataset Schema & Documentation (non-sensitive version) Dataset description following OLID Zampierietal.,2019 Zampierietal.,2019 and HateCheck Ro¨ttgeretal.,2021 Rottgeretal.,2021 formats. Column definitions (TEXT, TARGET, ID, SOURCE) Annotation guidelines and labelling protocol Ethical considerations and masking procedures Dataset card (HuggingFace style) ⚠️ Note: The raw offensive text dataset is not included in this Zenodo release to ensure ethical and legal compliance. Instead, we provide: metadata describing the dataset, annotation schema, sampling notes, cleaning pipeline, and usage guidelines. Annotation Framework & Instructions Includes: Complete annotation manual Examples of label categories (masked versions) Decision rubrics for offensive vs non-offensive distinctions Inter-annotator agreement (IAA) computation script (Cohen's kappa) Annotator training notes This allows reproducibility of annotation methodology without revealing harmful content. Preprocessing, Cleaning & Structuring Scripts Folder: scripts/preprocessing/ Includes: Text cleaning pipeline Normalisation for Setswana orthography Token masking functions Duplicate removal and source metadata stripping Tools for exporting to OLID/HateCheck CSV structure These scripts enable researchers to structure their own Setswana datasets using the same methodology. Evaluation Scripts Folder: scripts/evaluation/ Includes: Binary-class evaluation (Accuracy, F1, MCC, ROC-AUC) Stratified splitting tools Error analysis helper functions Label distribution checkers These enable reproducibility of experimental analysis described in the paper. Explainability & Counterfactuals (Sanitized) Folder: outputs/explainability/ Contains: LIME/S-LIME interpretation outputs (masked) Counterfactual examples (masked) explanation_notes.md describing methodology and sanitization policies No harmful text is included. Figures & Tables Used in the Manuscript Folder: outputs/figures/ and outputs/tables/ Includes: Dataset distribution plots Annotator agreement figures Offensive vs non-offensive category diagrams Tokenisation breakdowns Summary statistics Table files (CSV + LaTeX) These are all reproducible using provided scripts. Supporting Notebooks Folder: notebooks/ Provides: Dataset structuring and cleaning notebook (safe version) Annotation guidelines demonstration notebook Explainability notebook (masked text) Dataset export notebook All notebooks are sanitised and suitable for public release. 🔒 Ethical & Safety Considerations This release does not include: raw offensive text user-identifiable content social media data platform metadata detailed examples containing sensitive content All samples shown in figures or notebooks are masked using ***. The dataset described in the manuscript should be accessed only through controlled academic channels and under strict ethical governance. This Zenodo package is designed to enable reproducibility without distributing harmful language. 🎯 Intended Use This release supports: NLP research in low-resource languages Cyberbullying and hate speech detection studies Digital forensics investigations Linguistic analysis of Setswana online discourse Educational purposes Reproduction of the experimental methodology described in the manuscript It is not intended for: Real-time automated moderation Law-enforcement without expert oversight Deployment in production systems without additional safeguards 📖 Citation If using this release, please cite the manuscript: Kekgathetse, B. (2025). Developing Monolingual Setswana Datasets for Offensive Content Detection. Zenodo. doi.org 📝 Licensing Code: MIT / Apache-2.0 Documentation & annotation guidelines: CC-BY 4.0 Sanitized examples: CC-BY-SA 4.0 No raw data included 🌍 Audience NLP researchers Digital forensics experts Computer security practitioners Linguists working with Bantu languages   If you use this resource, please cite this work.

Visit

doi.orgzenodo.org

Tasks

hate speech detectiontext classification

Languages

Setswana

Tags

Setswanaoffensive languagelow-resource NLPforensic NLPexplainability

Licenses

Apache License 2.0http://www.apache.org/licenses/LICENSE-2.0