This dataset contains labeled Roman Urdu/English SMS messages for phishing (smishing) detection research. It includes:
- 1,000 unique messages used in the primary experiments (595 legitimate, 405 phishing) - - Template IDs and Source_Type categories - Both raw and normalized text versions
The dataset was created to study the effect of text normalization on machine learning-based smishing detection in Roman Urdu, a low-resource, code-mixed language. A template-aware evaluation protocol was used to prevent template leakage between train and test sets.
Related paper: "Does Text Normalization Improve Machine Learning Detection of Roman Urdu Phishing SMS Messages? An Empirical Study with a Template-Aware Evaluation Protocol"