Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Amharic Phishing Dataset for Ethiopian Fintech Context

Domaine:

natural language processingdigital infrastructure

Type de record:

dataset
Créateur:
Mur
Éditeur:
Zenodo
Hôte:avatar
This dataset comprises 2,540 Amharic SMS, each divided into equal sets of actual ('ham', 1270 instances) and phishing/spam ('spam', 1270 instances). It was specifically designed for the Ethiopian scenario with a focus on threats to the country's emerging Fintech sector, such as products like TeleBirr and CBE. The data was collected from anonymized real user messages, social media incidents, and institutional security teams. The data is for training and testing machine learning models on problems such as phishing detection, spam filtering, and adversarial AI research, which are most useful for low-resource languages such as Amharic. The data is provided in plain CSV format with 'label' and 'message' columns. Messages are similar to actual interaction, such as normal scam behavior (spurious notifications, OTP requests, job offers) and real interaction. All personally identifiable information has been anonymized. This dataset was developed under the research of the QuantumShield cybersecurity framework.

Visit

doi.orgzenodo.org

Tasks

text classification

Languages

Amharic

Tags

Amharic, SMS, Phishing, Spam Detection, Cybersecurity, Ethiopia, Fintech, TeleBirr, CBE, Natural Language Processing, Machine Learning Dataset, Low-resource language, Text Classification, Adversarial AI

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode