# π₯ National Health Insurance β Big Data Capstone Project
**Name:** Manzi Delphin
**ID:** 26021
**Group:** E
**Course:** Introduction to Big Data Analytics
**Instructor:** Eric Maniraguha
**Academic Year:** 2024β2025, Semester III
---
## π Project Files & Links
- π **Dataset Used**: Health Insurance Dataset (CSV)
- π **Power BI Dashboard**: Power BI Project Folder
- π½οΈ **PowerPoint Presentation**: Capstone Slides
---
## β
PART 1: PROBLEM DEFINITION & PLANNING
### π·οΈ I. Sector Selection
β **Health**
### β II. Problem Statement
> This project investigates disparities and patterns in national health insurance coverage across countries using survey-based health indicators. Using big data analytics, it aims to identify trends across demographics, regions, and years to support equitable policy decision-making.
### π¦ III. Dataset Identification
- **Dataset Title:** National Health Insurance β Global DHS Indicators
- **Source Link:** Dataset Download
- **Number of Rows and Columns:** 15,000+ rows Γ 20 columns
- **Data Structure:** β Structured (CSV)
- **Data Status:** β Requires Preprocessing
---
## π PART 2: PYTHON ANALYTICS TASKS
### β‘οΈ 1. Clean the Dataset
We began by inspecting for missing values, inconsistent formats, and extreme outliers. String values were normalized, and outliers in numerical fields were capped using IQR.
### β‘οΈ 2. Apply Data Transformations
Categorical values were encoded using LabelEncoder, and numerical features such as age, household size, and insurance value were scaled using StandardScaler to ensure modeling fairness.
---
### β‘οΈ 3. Conduct Exploratory Data Analysis (EDA)
We analyzed central tendencies, distributions, and correlations between variables. This helped us detect relationships and understand variable behavior.
#### Descriptive Statistics
#### Distribution Visuals
We used histograms and boxplots to observe feature spread and detect skews and group variances.
#### Correlation Analysis
This heatmap revea β¦