Comparison of K-Means Clustering and Density-Based Spatial Clustering of Applications with Noise (DBSCAN) in Identifying Particulate Matter (PM2.5) Pollution Hotspots using the Global High Air Pollutants (GHAP) Dataset from September to December 2022
# Comparison of K-Means Clustering and Density-Based Spatial Clustering of Applications with Noise (DBSCAN) in Identifying Particulate Matter (PM2.5) Pollution Hotspots using the Global High Air Pollutants (GHAP) Dataset from September to December 2022
The dataset was collected from zenodo.org, a free, open-access digital repository that allows researchers to share, preserve, and cite datasets, software code, research papers, and other digital artifacts across all scientific disciplines. The original dataset consists of 36.1 GB of daily, monthly, and yearly 1 km (i.e., D1K, M1K, and Y1K) global ground-level PM2.5 data over land for 2022. However, the researchers used only 11.9 GB of data from the D1K for September to December 2022. It is a compressed zip file containing Scientific Data Files, specifically– Network Common Data Form (.nc) files. The columns include: Latitude, the measurement of a location north or south of the Equator; Longitude, the measurement of a location east or west of the prime meridian; and PM2.5 for each month, the monthly 6 measurement of the particulate matter with a diameter of 2.5 micrometers or less.