Spatiotemporal Analysis and Machine Learning-Based Profiling of Atmospheric Pollutants Using Sentinel-5P TROPOMI Satellite Data at Tier-Two Administrative Units in Bangladesh.
DOI:
https://doi.org/10.59185/jos.v46i1.425Keywords:
Air pollution analysis,, K-means clustering,, Multi-pollutant assessment,, District-level air quality mapping,, Google Earth Engine (GEE)Abstract
Air pollution poses a critical environmental and public health threat in Bangladesh, yet a comprehensive, multi pollutant analysis at a high administrative resolution has been lacking. This study bridges this gap by leveraging the high-resolution capabilities of Sentinel-5P TROPOMI satellite data and an unsupervised machine learning framework to conduct a district-level assessment of five key atmospheric pollutants—SO₂, NO₂, O₃, HCHO, and CH₄—from 2020 to 2024. Annual mean concentrations were extracted for all 64 districts using Google Earth Engine, followed by rigorous spatiotemporal trend analysis using the Mann-Kendall test. The core novelty of this research lies in the application of K-means clustering to classify districts based on their holistic, multi-annual pollutant profiles. The analysis revealed distinct national trends: a significant increase in O₃ (54 districts) and CH₄ (61 districts), a volatile trend for SO₂, and a notable decrease in NO₂ in 21 urban districts. Critically, the K-means algorithm (k=5) segmented the country into five discrete pollution clusters, ranked from best to worst: 1) Cleanest (16 districts, e.g., coastal and hilly regions), 2) Moderate (7 districts, e.g., northeastern regions), 3) High O₃ & CH₄ (16 northern agricultural districts), 4) High Multi-Pollutant (23 southwestern and central industrial districts), and 5) Extreme Urban/Industrial (4 districts: Dhaka, Gazipur, Munshiganj, Narayanganj). This final cluster exhibited an extreme concentration of NO₂ and CH₄, over 50% higher than the national average. The findings demonstrate that Bangladesh's air pollution landscape is not monolithic but a mosaic of distinct regional typologies, each driven by different dominant sources—from vehicular and industrial combustion in urban centers to agricultural emissions in the north. This study provides a novel, data-driven framework for "pollution zonation," moving beyond one-size-fits-all policies towards targeted, cluster-specific mitigation strategies. The methodology establishes a replicable model for evidence-based air quality management in data-scarce regions globally.