Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Sheng Phishing Corpus for Low-Resource Cybersecurity NLP

Domain:

natural language processing

Record type:

dataset
Creator:
MwaKimKim
Publisher:
Zenodo
Host:avatar

 

This dataset, the Sheng-English Phishing Corpus (SEPC), was created to address the critical lack of resources for phishing detection in low-resource, code-mixed languages. It contains 9,970 curated samples of legitimate (ham) and phishing messages.

The data is specialized for Sheng, a dynamic Swahili-English sociolect spoken in Kenya. The corpus was constructed using a novel pipeline, including automated transcription of online video content (from YouTube, TikTok, FB and X) using OpenAI's Whisper, and augmented with data from social media, Sheng dictionaries, and web scraping. The full data collection methodology is detailed in our accompanying paper, "Hybrid Knowledge Distillation and Federated Learning for Real-Time, On-Device Phishing Detection in Low-Resource Languages."

The final dataset is provided in CSV format with two columns: 'text' (the message content) and 'label' (0 for legitimate, 1 for phishing).

Visit

doi.org

Tasks

text classification

Languages

Swahili

Tags

Sheng, Swahili-English, Code-Mixed NLP, Phishing Detection, Cybersecurity, Low-Resource Languages, Multilingual NLP, African Languages, Explainable AI, XLM-RoBERTa, Text Classification, Natural Language Processing, Machine Learning, Multilingual Cybersecurity, Social Engineering, Language Resources, Corpus, Dataset, African NLP, AI Security, Linguistic Diversity, Low-Resource NLP

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution–NonCommercial 4.0 International (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/

Similar

amandatheuri/NLP-for-English-Sheng-SwahiliA Tri-Class Multilingual Phishing Email Dataset for Low-Resource Languages EvaluationDecolonizing NLP for “Low-resource Languages”A 10-Million-Row Sinhala Narrative Corpus for Low-Resource NLP: Dataset Construction, Statistical Characterisation, and Ethical ConsiderationsNjambi-M/NLP-text-classification-for-a-low-resource-languageImproving Resource Creation for Low-Resource Languages using NLP Methods

amandatheuri/NLP-for-English-Sheng-Swahili

This project is a Natural Language Processing (NLP) application that automatically identifies the la

A Tri-Class Multilingual Phishing Email Dataset for Low-Resource Languages Evaluation

Decolonizing NLP for “Low-resource Languages”

Today African languages are spoken by more than a billion people, yet in the world of machine transl

A 10-Million-Row Sinhala Narrative Corpus for Low-Resource NLP: Dataset Construction, Statistical Characterisation, and Ethical Considerations

sinhala_stories is a crowdsourced corpus of Sinhala-language narrative text comprising 10,949,004 ro

Njambi-M/NLP-text-classification-for-a-low-resource-language

This an NLP model on Text Classification for Kiswahili text using a dataset from Hugging Face # NLP

Improving Resource Creation for Low-Resource Languages using NLP Methods

To digitize existing high-quality text belonging to a certain low-resource language, we are often fa