# 🏥 National Health Insurance – Big Data Capstone Project
**Name:** Manzi Delphin
**ID:** 26021
**Group:** E
**Course:** Introduction to Big Data Analytics
**Instructor:** Eric Maniraguha
**Academic Year:** 2024–2025, Semester III
---
## 📁 Project Files & Links
- 🔗 **Dataset Used**: Health Insurance Dataset (CSV)
- 📊 **Power BI Dashboard**: Power BI Project Folder
- 📽️ **PowerPoint Presentation**: Capstone Slides
---
## ✅ PART 1: PROBLEM DEFINITION & PLANNING
### 🏷️ I. Sector Selection
☑ **Health**
### ❓ II. Problem Statement
> This project investigates disparities and patterns in national health insurance coverage across countries using survey-based health indicators. Using big data analytics, it aims to identify trends across demographics, regions, and years to support equitable policy decision-making.
### 📦 III. Dataset Identification
- **Dataset Title:** National Health Insurance – Global DHS Indicators
- **Source Link:** Dataset Download
- **Number of Rows and Columns:** 15,000+ rows × 20 columns
- **Data Structure:** ☑ Structured (CSV)
- **Data Status:** ☑ Requires Preprocessing
---
## 🐍 PART 2: PYTHON ANALYTICS TASKS
### ➡️ 1. Clean the Dataset
We began by inspecting for missing values, inconsistent formats, and extreme outliers. String values were normalized, and outliers in numerical fields were capped using IQR.
### ➡️ 2. Apply Data Transformations
Categorical values were encoded using LabelEncoder, and numerical features such as age, household size, and insurance value were scaled using StandardScaler to ensure modeling fairness.
---
### ➡️ 3. Conduct Exploratory Data Analysis (EDA)
We analyzed central tendencies, distributions, and correlations between variables. This helped us detect relationships and understand variable behavior.
#### Descriptive Statistics
#### Distribution Visuals
We used histograms and boxplots to observe feature spread and detect skews and group variances.
#### Correlation Analysis
This heatmap revea …