Logo Lanfrica

Roman Urdu Phishing SMS Dataset for Smishing Detection Research

Domain:

natural language processing

Record type:

dataset
Creator:
Sid
Publisher:
Zenodo
Host:avatar
This dataset contains labeled Roman Urdu/English SMS messages for phishing (smishing) detection research. It includes: - 1,000 unique messages used in the primary experiments (595 legitimate, 405 phishing) -  - Template IDs and Source_Type categories - Both raw and normalized text versions The dataset was created to study the effect of text normalization on machine learning-based smishing detection in Roman Urdu, a low-resource, code-mixed language. A template-aware evaluation protocol was used to prevent template leakage between train and test sets. Related paper: "Does Text Normalization Improve Machine Learning Detection of Roman Urdu Phishing SMS Messages? An Empirical Study with a Template-Aware Evaluation Protocol"