Logo Lanfrica

dralexbevan/water-quality-prediction-catboost-time-series

Domain:

environment and energy

Record type:

project
Creator:
dra
Host:
Part of EY Challenge 2026 to create ML water quality predictor for South Africa. Involved 9K samples across 162 stations. I worked with time series models for EDA and XGBoost and CatBoost for the final model. Achieved an average Rsq 26% locally and 12% in comp across three target parameters (domain specific norms hover between 20 to 30). **Water Quality Prediction (Machine Learning & Time Series)** South Africa Water Quality Challenge Dr Alex Bevan **Project Overview** This project develops machine learning and time series models to predict key water quality indicators across South Africa using spatial and temporal data. The goal is both predictive accuracy and interpretability, with a particular focus on identifying environmental and climate drivers of water quality. **Objectives** 1. Predict three core water quality measures: Alkalinity – water’s natural buffer against contaminants Electrical Conductance (EC) – proxy for dissolved solids and overall water “load” Dissolved Reactive Phosphorus (DRP) – nutrient levels that indicate potential pollution 2. Identify the most important features (≈20 variables), with emphasis on: climate drivers, hydrological conditions, geographic variation 3. Develop a deeper understanding of spatiotemporal dynamics in water systems **Data Preparation** The dataset was constructed through a series of: • outer merges across multiple data sources • aggregation by time and station location This process produced a unified panel dataset capturing: • environmental variables • satellite-derived features • geographic identifiers • time-based observations **Missingness** was investigated and found to be: • Non-random • Highly correlated with summer months and extreme weather conditions This supports the conclusion that missing values are driven by cloud coverage affecting satellite measurements. Rather than imputing blindly, this insight informs: targeted imputation strategies and modeling approaches robust to missingness **Exploratory Data Analysis (EDA)** Several key diagnostic steps were performed which provided flexibility for: feature selection, dimensionality reduction and model simplification Extensive Feature Engineering Examined all variables in their raw (unscaled) form Then experimented with a substantial number of custom made features including derivin …