EY Open Science AI & Data Challenge 2026 (France finals). Predicting water quality (alkalinity, conductance, phosphorus) of South African rivers using Landsat, TerraClimate and spatiotemporal features. Hybrid Random Forest + LightGBM model with spatial cross-validation.
# EY Open Science AI & Data Challenge 2026 — Water Quality Prediction
> **3rd place — France finals** · Team **Data4Decision** (Hermann Banzouzi Miampassi & Neville Tchatchou Njatcha)
>
> Predicting three water quality indicators across South African rivers using satellite and climate data, with a focus on **spatial generalization** to regions never seen during training.
---
## Table of Contents
- Context
- The Data
- Approach
- Results
- Key Insights
- Repository Structure
- Installation & Reproduction
- Team
- Acknowledgments
---
## Context
Access to safe drinking water remains a challenge for **2.1 billion people** worldwide. In South Africa, around **70% of large river systems are eutrophic or hypereutrophic**, and field measurements remain slow and expensive.
The EY Open Science AI & Data Challenge 2026 asked us to predict three water quality parameters on South African rivers from satellite and climate data:
| Parameter | Name | Unit | Meaning |
|-----------|------|------|---------|
| **TA** | Total Alkalinity | mg/L | Buffering capacity against acidification |
| **EC** | Electrical Conductance | µS/cm | Proxy for mineralization / salinity |
| **DRP** | Dissolved Reactive Phosphorus | mg/L | Key eutrophication indicator |
**Evaluation metric**: mean R² across the three targets, on **24 validation sites not seen during training**.
---
## The Data
| Dataset | Size | Source |
|---------|------|--------|
| Training | 9,319 observations · 162 stations · 2011–2015 | EY Challenge |
| Validation | 200 points · 24 unseen stations | EY Challenge |
| Landsat 7/8 | Spectral reflectances (NIR, Green, SWIR) | Google Earth Archive |
| TerraClimate | PET, precipitation, temperature, runoff | UCAR / Climatology Lab |
**The core difficulty**: the 24 validation stations are **80–280 km** away from their nearest training neighbor (median: 190 km). Random cross-validation splits massively overestimate generalization performance on this kind of spatial transfer p …