Ganga River water-quality forecasting and early-warning system using ML, deep learning, and geospatial analytics.
# 🌊 Aquanga – Predictive Water Monitoring & Early Warning System
**Aquanga** is a production-grade, end-to-end water quality forecasting and environmental early warning system for Central Pollution Control Board (CPCB) monitoring stations along the Ganga River basin. It integrates machine learning and deep learning time-series forecasting, real-time risk assessment, a RESTful FastAPI backend, a PostgreSQL relational datastore, and an interactive Streamlit geospatial dashboard.
---
## 📌 Table of Contents
1. Problem Statement
2. System Architecture
3. Dataset & Preprocessing
4. Feature Engineering
5. Machine Learning & Deep Learning Models
6. Model Evaluation & Benchmark
7. Environmental Risk & Early Warning System
8. FastAPI REST API
9. Database Architecture & Seeding
10. Interactive Streamlit Dashboard
11. Project Structure
12. How to Run Locally
13. Docker & Containerized Deployment
14. Testing Suite
15. Generalization & Future Improvements
---
## 🎯 Problem Statement
The Ganga River supports over 400 million people but faces severe ecological stress from municipal sewage, industrial effluents, and agricultural runoff. Traditional water quality monitoring relies on post-hoc manual laboratory testing, often detecting hypoxic events (low Dissolved Oxygen) and severe microbial surges days after they occur.
**Aquanga** solves this by:
- Forecasting future **Dissolved Oxygen (DO)** levels using historical multi-parameter time-series observations.
- Calculating environmental risk scores based on statutory CPCB water quality criteria.
- Automatically generating actionable early warnings to alert environmental authorities before ecological thresholds are breached.
---
## 🏛️ System Architecture
```mermaid
flowchart TD
A[Raw CPCB Water Quality Data] --> B[Data Preprocessing & Imputation]
B --> C[Feature Engineering & Compliance Flags]
C --> D[Chronological Train / Test Split]
D --> E[6 ML & DL Models Training]
E --> F[Evaluation Benchmark & Best Model Selectio …