Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

SMS Fraud Classification dataset for Chichewa

Domaine:

natural language processing

Type de record:

dataset
Créateur:
TayRob
Éditeur:
Zenodo
Hôte:avatar
The dataset contains 676 SMSs in Chichewa and it was used to experiment with machine learning models for fraud classifciation. There are in total six version of the dataset: D-CHI contains SMSs in Chichewa, D-HT contains a human translated version of D-CHI, and D-MT is a machine translation using google translation of D-CHI. These datasets are all balanced: they contain an equal number of fraudulent and normal SMSs. Three extended datasets of 148 SMSs each was also used that contained only normal SMSs. When added to the three datasets we obtained extended unbalance versions demoted as D-CHIe, D-HTe and D-MTe.  The attached paper explains the methodology used. Please note that the github repo and this dataset are private and  is made public with the publication of the results. Please cite: Taylor, A., Robert, A. (2025). Using Machine Learning to Detect Fraudulent SMSs in Chichewa. In: Sinha, G.R., Fan, C.P., Bajaj, V., Nisar, H., Ullo, S.L. (eds) Integrating AI in Science, Management, and Technology. AISMT 2025. Communications in Computer and Information Science, vol 2699. Springer, Cham. doi.org  

Visit

doi.orgzenodo.org

Tasks

text classification

Languages

Chichewa

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode