Building on my analysis of health facilities in sub-Saharan Africa, where I explored country-level and state-level distribution, facility ownership, and facility types, I aim to further drill down and provide in-depth insights specifically for Nigeria, shedding light on the nuances of healthcare infrastructure within the country.
# Analysis-of-Health-Facilities-in-Sub-Saharan-Africa
Building on my analysis of health facilities in sub-Saharan Africa, where I explored country-level and state-level distribution, facility ownership, and facility types, I aim to further drill down and provide in-depth insights specifically for Nigeria, shedding light on the nuances of healthcare infrastructure within the country.
## Table of Contents
1. Project Overview
2. Tools and Methodology
3. Data Profiling
4. Data Cleaning
5. Univariate Analysis
6. Bivariate Analysis
## Project Overview
This analysis utilizes a dataset compiled by Maina, J., Ouma, P.O., Macharia, P.M., et al., as part of their research titled A spatial database of health facilities managed by the public health sector in sub-Saharan Africa, published in 2019.
The goal is to explore the distribution of health facilities at country and state levels, examining facility types and ownership structures. The analysis will also focus specifically on Nigeria, investigating state-level distribution, ownership, facility types, and healthcare service tiers within the country.
## Tools and Methodology
1. SQL (querying and analyzing data)
2. Data Exploration (understanding the data)
3. Data Cleaning (preparing the data for analysis)
4. Aggregation (grouping and summarizing data)
5. CTE (Common Table Expressions, a powerful SQL feature)
## Data Profiling
This dataset comprises 98,745 rows and 8 columns, featuring 6 categorical variables (Country, Admin1, Facility name, Facility type, Ownership, and LL Source) stored as varchar, and 2 numerical variables (Lat and Long, representing latitude and longitude). Upon inspection, null values are present in several columns: Ownership (30,448 nulls), Lat (6,041 nulls), Long (3,095 nulls), and LL Source (2,350 nulls). Additionally, the dataset contains 1367 duplicate rows. After removing these duplicates, the dataset is refined to 97,378 unique rows.
#### Number of columns
```sql
SELECT COUNT(*) A …