Logo Lanfrica

rikotse/Maji-Ndogo-Data-Architecture

Domain:

environment and energysocioeconomic

Record type:

project
Creator:
rik
Host:
End-to-end data engineering and analytics solution architecting relational SQL databases, Python data pipelines, and interactive Power BI dashboards to optimize public infrastructure and resolve the Maji Ndogo water crisis. # 🌍 Maji Ndogo Infrastructure Revitalization: End-to-End Data Engineering & Analytics **Role:** Lead Data Engineer & BI Developer **Domain:** Public Infrastructure, Environmental Resources, Agribusiness **Tech Stack:** PostgreSQL, Python (Pandas/NumPy), Power BI, DAX, Power Query ## 📖 Executive Summary As a Lead Data Engineer, I architected an end-to-end data solution to combat the severe water crisis in the fictional region of Maji Ndogo. By engineering robust data pipelines, standardizing messy infrastructural data, and modeling intuitive executive dashboards, this project empowered national stakeholders to strategically allocate repair budgets, monitor vendor performance, and ultimately improve water access for thousands of citizens. ## 🏗️ Project Architecture & Methodology ### 🛠️ Phase 1: Database Architecture & Initial EDA (SQL) * **Objective:** Architected a relational database schema to centralize fragmented regional water source data, queue times, and demographic surveys. * **Execution:** Extracted and validated over 60,000 records. Engineered complex SQL joins and aggregations to audit infrastructure failure rates, identify biological/chemical contamination zones, and flag critical bottlenecks in water access. * **Code References:** Navigate to the `/Phase_1_SQL_Architecture` folder to review the database schema and analytical queries. * Engineered multi-table relational joins to track infrastructure bottlenecks, flagging regions with queue times exceeding safety thresholds (>60 mins). * Engineered environmental audit queries using aggregations (GROUP BY, AVG) to categorize chemical and biological contamination across well sources, enabling targeted medical and filtration responses. ### 🐍 Phase 2: Advanced Data Pipelines (Python) * **Objective:** Engineered scalable data pipelines to validate agricultural impacts and standardize disparate datasets. * **Execution:** Wrangled raw environmental and survey data using Pandas and NumPy. Handled null values, …

Languages