
This 10-year dataset contains hydrogeochemical and hydrogeological variables used for the development and evaluation of machine learning models for seawater intrusion prediction in a coastal alluvial aquifer system in eastern India, covering West Bengal and Odisha. The data are organized into pre-monsoon (PRM) and post-monsoon (POM) seasons and are intended for predictive modeling using Random Forest (RF), Long Short-Term Memory (LSTM), and Support Vector Machine (SVM) algorithms.
The dataset is divided into training (2012–2018) and testing (2019–2021) periods to facilitate model calibration and independent validation. Separate folders are provided for input and target datasets required for machine learning model development.
The 'input' variables include:
A – Aquifer hydraulic conductivity,
L – Groundwater elevation,
D – Distance from the coastline,
I – Extent of seawater intrusion,
T – Aquifer thickness.
The 'target' dataset contains the 'estimated' vulnerability computed from AHP-based modified GALDIT framework, which is used for supervised machine learning model training and prediction. In addition, the outputs generated by the trained RF, LSTM, and SVM models are provided as 'predicted' vulnerability datasets for both pre-monsoon and post-monsoon seasons.
This dataset is intended to support reproducible research in groundwater quality assessment, seawater intrusion modeling, and the application of artificial intelligence techniques in hydrogeological studies. The dataset has been curated to enable independent verification of the study results and to facilitate future research on data-driven coastal groundwater management.
------
Published as a journal paper on 18 July 2026, and interested researchers can cite the published article as:
Ghosh, S., Jha, M.K. and Pandey, V.M. (2026). Application of machine learning techniques for predicting seawater intrusion vulnerability in a coastal alluvial aquifer of eastern India. Marine Pollution Bulletin, 233(1): 120101. doi.org