Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Roman Urdu Phishing SMS Dataset for Smishing Detection Research

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Sid
Éditeur:
Zenodo
Hôte:avatar
This dataset contains labeled Roman Urdu/English SMS messages for phishing (smishing) detection research. It includes: - 1,000 unique messages used in the primary experiments (595 legitimate, 405 phishing) -  - Template IDs and Source_Type categories - Both raw and normalized text versions The dataset was created to study the effect of text normalization on machine learning-based smishing detection in Roman Urdu, a low-resource, code-mixed language. A template-aware evaluation protocol was used to prevent template leakage between train and test sets. Related paper: "Does Text Normalization Improve Machine Learning Detection of Roman Urdu Phishing SMS Messages? An Empirical Study with a Template-Aware Evaluation Protocol" 

Visit

doi.org

Tasks

text classification

Tags

roman urduSmishingSMS phishingText normalizationLow-resource NLPPhishing detectionCode-mixed languageMachine learningTF-IDFPakistan

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode