High-resolution surface PM2.5 concentration for India, built from satellite products, surface monitors, and machine learning pipelines. This dataset is produced using a reimplementation of the Improved daily PM2.5 estima… (Kawano, et al., 2025) built by the Centre for Research on Ener… (CREA). See southasia-pm2-5.energyandcl… for more details.
You can download daily PM2.5 data from the original authors for 1 January 2005 to 30 September 2023 from https://zenodo.org/records/….
The dataset is provided in CF-1.8–compliant NetCDF format, containing daily surface PM2.5 concentrations (μg/m3) at 10 km resolution over India. It includes time, x, and y dimensions, with the primary variable pm25(time, y, x). The data uses the WGS 84 / India NSF Lambert Conformal Conic projection (EPSG: 7755) to define spatial coordinates.
We provide the results of the spatial cross-validation in a CSV file alongside the results. This has the headers r2 and rmse (in ug/m3). We use the same cross-validation metrics and methodology for daily values as for the final full model by Kawano et al. (2025).
When using this data, please reference:
Kawano, Ayako, et al. "Improved daily PM2.5 estimates in India reveal inequalities in recent enhancement of air quality." Science Advances 11.4 (2025): eadq1071. https://doi.org/10.1126/sci…
The version of the data used from Zenodo.
For each release, we provide versioning info to ensure every file name encodes when it was released, which schema it follows, and which model generated it. Major versions indicate breaking or significant changes in the schema or model.
Filename pattern: pm25-___.nc
: spatial–temporal packagingThe file’s spatial–temporal packaging.
Format: lowercase letters and hyphens only. Options available: full only
: releaseRelease month of the file and regeneration version. Independent of s and mb. This does not describe the maximum date available in the dataset, which is usually delayed by up to 2 months.
Format: rel-YYYY-MM-rT:
YYYY: four-digit year.
MM: two-digit month.
T zero-based monthly regenerate counter. Resets to 0 when YYYY-MM changes.
: schemaThe file’s schema.
Format: s-X.Y
X: major schema version. Breaking or significant change in file schema or variable layout.
Y: minor/patch schema version. Backward-compatible metadata or layout additions.
: model bundleThe model bundle used to generate the dataset. The models at every stage are versioned using the same number and changes to any of these can change the model bundle version. We version the features, ingestion, and model code under this version.
Format: mb-X.Y-rT
X: major model version: Changes to model features or model code that alters the feature set or modeling approach.
Y: minor model version. Backward-compatible code changes that keep the feature set identical.
T: zero-based retrain counter. Resets to 0 when X.Y changes.
Disclaimer: The designations employed and the presentation of the material on maps contained in this dataset do not imply the expression of any opinion whatsoever concerning the legal status of any country, territory, city or area or of its authorities, or concerning the delimitation of its frontiers or boundaries.