Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Kanjo-Elkamira-Ndi/Data-Science-Operations-Pipeline-Lab

Domain:

socioeconomic

Record type:

datasetproject
Creator:
Kan
Host:
A complete, end-to-end data science pipeline applied to a survey dataset investigating mobile money scam prevalence, victim demographics, and loss patterns in Cameroon. The project covers Exploratory Data Analysis, Data Preprocessing, Feature Engineering, Predictive Modelling, and Evaluation, culminating in a fully formatted Word report. # 📱 Mobile Money Scam — Data Science Analysis > **Course:** Data Science in Python  |  **Assessment:** Continuous Assessment (CA)  |  **Year:** 2025/2026 A complete, end-to-end data science pipeline applied to a survey dataset investigating mobile money scam prevalence, victim demographics, and loss patterns in Cameroon. The project covers Exploratory Data Analysis, Data Preprocessing, Feature Engineering, Predictive Modelling, and Evaluation — culminating in a fully formatted Word report with embedded visualisations. --- ## 📋 Table of Contents 1. Project Overview 2. Dataset Description 3. Project Structure 4. Requirements & Installation 5. How to Run 6. Pipeline Walkthrough - Step 0 — Load Data - Step 1 — Exploratory Data Analysis - Step 2 — Data Preprocessing - Step 3 — Feature Engineering & Scaling - Step 4 — Predictive Modelling - Step 5 — Evaluation 7. Results Summary 8. Generated Outputs 9. Key Findings 10. Limitations & Future Work 11. References --- ## 🔍 Project Overview Mobile money services such as MTN Mobile Money and Orange Money have transformed financial access across sub-Saharan Africa. However, their rapid adoption has been accompanied by a parallel surge in mobile money scams — a threat that disproportionately affects young, digitally active populations with limited fraud awareness. This project applies a structured data science methodology to a real survey dataset of **800 respondents** collected in Cameroon. The goal is twofold: 1. **Descriptive** — Understand who gets scammed, how, and at what cost. 2. **Predictive** — Build machine learning models capable of identifying scam victims based on demographic and behavioural features. The pipeline is implemented entirely in Python and follows industry-standard practices for data cleaning, encoding, scaling, dimensionality reduction, modelling, and evaluation. Results are documented in a professional Word report generated programmatically using the `docx` JavaScript libra …

Visit

github.com