Advances in Air Pollution Data Analysis: From Classical Geostatistics to Big Data and Artificial Intelligence

A Special Issue of Atmosphere (ISSN 2073-4433) belonging to the section "Air Pollution Control".

Deadline for manuscript submissions: closed (30 June 2026) | Viewed by 9165

Editors


E-Mail Website
Guest Editor
Department of Geoinformatics and Applied Computer Science, Faculty of Geology, Geophysics and Environmental Protection, AGH University of Krakow, 30-059 Krakow, Poland
Interests: air pollution measurements; air quality monitoring; artificial intelligence, anthropogenic emission; spatio-temporal geostatistics; geophysics; smart cities

E-Mail Website
Guest Editor
Department of Geoinformatics and Applied Computer Science, Faculty of Geology, Geophysics and Environmental Protection, AGH University of Krakow, 30-059 Krakow, Poland
Interests: geostatistics; spatial data analysis; machine learning; air pollution measurements; air quality monitoring; geophysics

E-Mail Website
Guest Editor
Yale-NUIST Center on Atmospheric Environment, Nanjing University of Information Science and Technology, Nanjing 210044, China
Interests: atmospheric chemistry; reactive nitrogen; ammonia; isotopic analysis; haze; secondary aerosol formation
Special Issues, Collections and Topics in MDPI journals

Special Issue Information

Dear Colleagues,

We invite researchers and practitioners to contribute to this Special Issue focusing on the evolving landscape of air pollution data analysis, bridging traditional geostatistics with groundbreaking advancements in big data and artificial intelligence (AI). This Special Issue aims to capture the latest innovations and foster interdisciplinary dialog on leveraging advanced analytical techniques to address air quality challenges. Air pollution, a critical environmental and public health concern, demands increasingly sophisticated tools to handle the complexity and scale of modern data. From satellite imagery and ground-based sensors to citizen science and IoT networks, the availability of vast, high-resolution datasets opens new frontiers for exploration. However, extracting actionable insights from these data sources requires a fusion of traditional methods and emerging technologies.

This Special Issue seeks contributions across a broad spectrum of topics, including, but not limited to, the following:

  • Applications of geostatistics for spatial and temporal modeling of air quality;
  • Big data techniques for managing and analyzing large-scale pollution datasets;
  • AI and machine learning models for predictive analysis, anomaly detection, and source apportionment;
  • Integrating heterogeneous data sources (satellite, sensor, and citizen science) for comprehensive air quality assessments;
  • Uncertainty quantification, explainable AI, and ethical considerations in air pollution analysis;
  • Real-time applications in pollution forecasting, urban planning, and policymaking.

By submitting to this Special Issue, you will showcase your research at the forefront of this dynamic field, contributing to innovative solutions for global air quality management. Together, let us push the boundaries of air pollution science and technology.

Dr. Mateusz Zareba
Dr. Elżbieta Węglińska
Prof. Dr. Yunhua Chang
Guest Editors

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Atmosphere is an international peer-reviewed open access monthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2400 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • air pollution
  • air quality monitoring
  • machine learning
  • big data
  • spatial analysis
  • artificial intelligence

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (7 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

Jump to: Review

32 pages, 30808 KB  
Article
Leveraging Remote Traffic Data for Local Air Pollutant Estimation: A Scenario-Based Machine Learning Study Across London Monitoring Sites
by Valeria Legaria-Santiago, Amadeo Arguelles, Magdalena Saldana-Perez, Jocelyn Richardson and Marcella Bona
Atmosphere 2026, 17(8), 806; https://doi.org/10.3390/atmos17080806 - 21 Aug 2026
Viewed by 279
Abstract
Vehicular traffic is a major source of air pollution; however, the contribution of remotely acquired traffic information to local machine-learning (ML) air-pollution models remains insufficiently characterised. This study evaluates four interpretable tree-based ML models (Random Forest, Extra Trees, LightGBM, and XGBoost) under six [...] Read more.
Vehicular traffic is a major source of air pollution; however, the contribution of remotely acquired traffic information to local machine-learning (ML) air-pollution models remains insufficiently characterised. This study evaluates four interpretable tree-based ML models (Random Forest, Extra Trees, LightGBM, and XGBoost) under six predictor scenarios combining progressively larger predictor sets, ranging from remotely acquired traffic, meteorological, and temporal variables alone to the inclusion of measurements from one and four neighbouring monitoring stations, to estimate NO2, PM10, PM2.5, and O3 concentrations across several sites in London. ML model performance was compared with a ridge linear regression model as a baseline, with spatial interpolation methods and with a cross-site validation experiment. When modelling without data from neighbouring stations, the RMSE for NO2 ranged from 9.73 to 11.66 μg/m3 without traffic information, compared with 8.72 to 11.52 μg/m3 when traffic information was included. Additionally, for NO2, SHAP analyses indicate that traffic-related variables can contribute at levels comparable to pollutant measurements from neighbouring monitoring stations in traffic-dominated environments. Full article
Show Figures

Graphical abstract

26 pages, 23869 KB  
Article
Combined O3 and NO2 Pollution Reveals Widespread Nonlinear Impacts on Net Primary Productivity Across China’s Terrestrial Ecosystems
by Zhaosheng Wang and Mei Huang
Atmosphere 2026, 17(8), 799; https://doi.org/10.3390/atmos17080799 - 19 Aug 2026
Viewed by 253
Abstract
Quantifying the large-scale impact of combined ozone (O3) and nitrogen dioxide (NO2) pollution on terrestrial carbon sinks remains a major challenge. Here, we develop a parsimonious yet robust empirical framework that leverages high-resolution remote sensing datasets (CHAP O3 [...] Read more.
Quantifying the large-scale impact of combined ozone (O3) and nitrogen dioxide (NO2) pollution on terrestrial carbon sinks remains a major challenge. Here, we develop a parsimonious yet robust empirical framework that leverages high-resolution remote sensing datasets (CHAP O3/NO2 and MODIS NPP, 2008–2021) to characterize nonlinear threshold responses of terrestrial net primary productivity (NPP) across China’s diverse ecosystems. Our observational analysis identifies only associative temporal relationships between annual NPP variability and pollutant concentrations, with NPP positively correlated with O3 (Pearson’s r = 0.714, p < 0.01) and negatively correlated with NO2 (r = −0.599, p < 0.05). Notably, the ecosystem-specific threshold values (O3: 28,324–34,391 μg m−3 yr−1; NO2: 3646–4968 μg m−3 yr−1) are statistically derived from spatially aggregated pixel-level records across the full 14-year period, independent of the national annual time-series correlation analyses. Distinct from previous single-pollutant national evaluations, our study advances a novel analytical framework focusing on the interactive and combined impacts of O3 and NO2 co-exposure. The results demonstrate that NPP displays an increasing trend under low-level pollutant exposure but declines substantially once pollutant loads exceed the identified threshold ranges. Based on K-means clustering and segmented regression analyses, we estimate a national average NPP reduction of 17.4% per year (−0.68 Pg C yr−1), resulting in a cumulative carbon loss of −9.48 Pg C over the 14-year study period—equivalent to 2.45 years of China’s total terrestrial carbon uptake. Among all ecosystem types, forestlands experience the largest cumulative carbon loss (−4.22 Pg C), with prominent loss hotspots concentrated on the Tibetan Plateau and Northwest China. This refined national-scale assessment of dual-pollutant impacts provides observation-based evidence of substantial terrestrial carbon sink degradation, underscoring the necessity of combined air pollution mitigation strategies to sustain ecosystem stability and climate mitigation targets. Full article
Show Figures

Figure 1

28 pages, 11722 KB  
Article
Symbolic Artificial Intelligence for Ground-Level Ozone Prediction Through Association Rule Mining
by David Camarazo, Aengus Ball, Agnieszka Rorat, Idriss Jairi, Nathalie Pujol-Söhne, Ludivine Canivet and Hayfa Zgaya-Biau
Atmosphere 2026, 17(8), 798; https://doi.org/10.3390/atmos17080798 - 19 Aug 2026
Viewed by 367
Abstract
Accurate air quality prediction is essential for environmental monitoring and public health protection. Among atmospheric pollutants, ground-level ozone remains particularly difficult to predict because of the complex and nonlinear interactions governing its formation. Although recent advances have achieved promising predictive performance using machine [...] Read more.
Accurate air quality prediction is essential for environmental monitoring and public health protection. Among atmospheric pollutants, ground-level ozone remains particularly difficult to predict because of the complex and nonlinear interactions governing its formation. Although recent advances have achieved promising predictive performance using machine learning and deep learning, most existing approaches rely on black-box models whose explanations are provided only through post hoc explainability techniques. This work presents an alternative symbolic artificial intelligence framework based on association rule mining for intrinsically explainable ozone prediction. Hourly atmospheric observations collected from ground-level monitoring stations in the Hauts-de-France region (France) are preprocessed through cleaning, discretization, and class balancing before rule extraction. Two complementary symbolic AI approaches, Formal Concept Analysis (FCA) and a Genetic Algorithm (GA), are employed to automatically discover human-readable association rules linking meteorological and atmospheric variables to ozone concentration classes. The extracted rules provide transparent and directly interpretable decision mechanisms that can be readily validated by air quality experts. The experimental results show that both rule-mining approaches produce substantially more precise rule sets than decision trees, with average rule precisions of 0.76 for FCA and 0.79 for GA. Furthermore, the resulting rule-based classifier achieves an accuracy of approximately 0.79, outperforming the evaluated machine learning baselines while preserving intrinsic interpretability. These results demonstrate that symbolic AI constitutes a promising alternative for trustworthy air quality prediction and knowledge discovery. Full article
Show Figures

Graphical abstract

30 pages, 14210 KB  
Article
Characterising Multivariate Air Pollution State Evolution in an Urban Atmosphere Using Deep-Learned Baseline Representations: London
by Arda Eraslan, David Topping, Dudley E. Shallcross, M. A. H. Khan and Aşan Bacak
Atmosphere 2026, 17(6), 589; https://doi.org/10.3390/atmos17060589 - 8 Jun 2026
Viewed by 1032
Abstract
Urban air quality management has been playing a significant role due to its effects on public health and pollution characteristics of countries with constantly changing policies. Traditional approaches capture how much pollution is present but are unable to detect changes in the chemical [...] Read more.
Urban air quality management has been playing a significant role due to its effects on public health and pollution characteristics of countries with constantly changing policies. Traditional approaches capture how much pollution is present but are unable to detect changes in the chemical character of the atmosphere, the relationships between co-emitted species, the balance of photochemical processing, and the combustion fingerprint of emission sources. This study introduces a framework that identifies and diagnoses such evolutions within the pollutants of the atmosphere. A chemistry-aware Variational Autoencoder is trained on 19 multivariate pollution features (7 raw concentrations, 5 chemical ratios, 7 temporal gradients) at London Marylebone Road (urban roadside) and North Kensington (urban background) from 2015 to 2019, and tested on 2022–2025. A four-method ensemble framework (VAE reconstruction error, reconstruction probability, Isolation Forest, and statistical Z-score) requires ≥3 agreement to identify high-confidence departed pollution states. Per-feature decomposition of the reconstruction probability diagnoses the chemical character of each departure. At the roadside site, 14.5% of post-COVID hours fall within departed states, dominated by the CO/NOx combustion ratio (513.2) and the photostationary state proxy (391.4), chemical relationships rather than individual concentrations. This indicates that at the point of emission, London’s fleet modernisation and Ultra Low Emission Zone (ULEZ) have changed the combustion fingerprint and photochemical equilibrium. The same structural indicators are carried over during the COVID-19 lockdown; however, O3 rises 3.2× during the pandemic period, reflecting suppressed NO titration. Conversely, at the urban background site, where the departures are driven by concentrations and boundary-layer trapping (r=0.659), the combustion fingerprint of the atmosphere is invisible to detect (CO/NOx=45.0). These findings indicate that London’s emission landscape has undergone fundamental transformations over the past decade, and the consequences of ULEZ and similar interventions or greater impacts of pandemic-related events are non-homogeneously distributed across the relevant region. Full article
Show Figures

Graphical abstract

23 pages, 3742 KB  
Article
Emergency Medical Interventions in Areas with High Air Pollution: A Case Study from Małopolska Voivodeship, Poland
by Ewa Szewczyk, Michał Lupa, Mateusz Zaręba, Elżbieta Węglińska, Tomasz Danek and Amit Kumar Mishra
Atmosphere 2025, 16(8), 983; https://doi.org/10.3390/atmos16080983 - 18 Aug 2025
Cited by 3 | Viewed by 3318
Abstract
Air pollution poses a significant threat to public health, particularly in urban and industrialized regions. This study investigates the relationship between air quality and the frequency of Emergency Medical Service (EMS) calls in the Małopolska Voivodeship of Poland between 2020 and 2023. Data [...] Read more.
Air pollution poses a significant threat to public health, particularly in urban and industrialized regions. This study investigates the relationship between air quality and the frequency of Emergency Medical Service (EMS) calls in the Małopolska Voivodeship of Poland between 2020 and 2023. Data from over 190 air quality sensors (PM10) were spatially aggregated using both hexagonal grids and administrative boundaries, while EMS call records were filtered to focus on cardiovascular and respiratory incidents. During 2020–2023, a total of 305,142 EMS calls were analyzed, and months with PM10 exceedances showed an average of 1.50 respiratory calls per 1000 residents compared to 1.19 in months without exceedances. Statistical analyses, including Kolmogorov-Smirnov tests and Pearson correlation, were applied to explore temporal and spatial associations. Results indicate a statistically significant increase in EMS calls during periods of elevated air pollution, with the strongest correlation observed for respiratory-related incidents. Comparative analyses between high- and low-pollution municipalities supported the observed relationships. Further analysis indicated that the COVID-19 pandemic may have partially confounded these associations, particularly for respiratory cases, though significant patterns remained even after accounting for pandemic peaks. While limitations related to data gaps and seasonal biases exist, the findings suggest that real-time air pollution data could inform better EMS resource allocation. This research highlights the potential of integrating environmental data into public health strategies to improve emergency response and reduce health risks in polluted regions. Full article
Show Figures

Figure 1

27 pages, 15404 KB  
Article
Machine-Learning Models for Surface Ozone Forecast in Mexico City
by Mateen Ahmad, Bernhard Rappenglück, Olabosipo O. Osibanjo and Armando Retama
Atmosphere 2025, 16(8), 931; https://doi.org/10.3390/atmos16080931 - 1 Aug 2025
Cited by 5 | Viewed by 2502
Abstract
Mexico City frequently experiences high near-surface ozone concentrations, and exposure to elevated near-surface ozone causes harmful effects to the inhabitants and the environment of Mexico City. This necessitates developing models for Mexico City that predict near-surface ozone levels in advance. Such models are [...] Read more.
Mexico City frequently experiences high near-surface ozone concentrations, and exposure to elevated near-surface ozone causes harmful effects to the inhabitants and the environment of Mexico City. This necessitates developing models for Mexico City that predict near-surface ozone levels in advance. Such models are crucial for regulatory procedures and can save a great deal of near-surface ozone detrimental effects by serving as early warning systems. We utilize three machine-learning models, trained on seven-year data (2015–2021) and tested on one-year data (2022), to forecast the near-surface ozone concentrations. The trained models predict the next day’s 24-h near-surface ozone concentrations for up to one month; before forecasting the following months, the models are trained again and updated. Based on prediction results, the convolutional neural network outperforms the rest of the models on a yearly scale with an index of agreement of 0.93 for three stations, 0.92 for nine stations, and 0.91 for one station. Full article
Show Figures

Figure 1

Review

Jump to: Research

31 pages, 454 KB  
Review
Multi-Model Ensemble Approaches in Air Quality Prediction: A Comprehensive Review from Chemical Transport Models to Hybrid Machine Learning
by Elena Chianese and Angelo Riccio
Atmosphere 2026, 17(7), 689; https://doi.org/10.3390/atmos17070689 - 14 Jul 2026
Cited by 2 | Viewed by 598
Abstract
Over the past two decades, air-quality prediction has moved from a mainly single-model paradigm toward ensemble systems that make explicit use of diversity across models, observations, and data streams. This review connects developments that are often treated separately: chemical transport model (CTM) ensembles, [...] Read more.
Over the past two decades, air-quality prediction has moved from a mainly single-model paradigm toward ensemble systems that make explicit use of diversity across models, observations, and data streams. This review connects developments that are often treated separately: chemical transport model (CTM) ensembles, tree-based and hybrid machine learning ensembles, deep learning architectures, physics-informed neural networks, and distributed approaches such as federated learning. Evidence summarized from recent systematic reviews and coordinated modeling initiatives indicates that, within comparable validation settings, ensembles often outperform individual models for PM2.5, PM10, O3, NO2, CO, and SO2 across a broad range of spatial scales and standard error metrics, including RMSE, MAE, and correlation. Operational CTM ensembles, such as the Copernicus Atmosphere Monitoring Service (CAMS) European system with eleven regional models, improve both forecast skill and uncertainty characterization for ozone and particulate matter. In data-driven applications, tree-based ensembles (Random Forest, gradient boosting, XGBoost, LightGBM) and hybrid deep architectures (CNN–LSTM models, attention-based multi-branch networks, graph neural networks) now form a core part of the state of the art for AQI (Air Quality Index) and particulate-matter estimation from structured and multi-source data. Reported performance can be very high on well-structured tabular datasets, with R2 values above 0.99 in selected benchmarks and RMSE reductions of 23–45% relative to classical statistical baselines in multi-modal studies; however, these values are not directly interchangeable because pollutant type, prediction horizon, monitoring density, and validation design differ among studies. This review proposes a practical taxonomy of ensemble strategies and uses it to explain why diversity, rather than model count alone, is central to reliable air-quality prediction. Drawing on coordinated European and North American model-evaluation initiatives (AQMEII, HTAP) and on case studies in topographically and meteorologically complex Italian regions (the Po Valley, the Naples metropolitan area, and Campania), we show that effective ensemble design requires a balance among diversity, redundancy, computational feasibility, and interpretability. On the basis of a structured narrative synthesis, the main research gaps concern physics-informed and explainable ensemble frameworks, transferable and adaptive models, standardized benchmarks, severe-pollution-episode forecasting, and scalable distributed architectures. Open questions include how to design compact non-redundant CTM sub-ensembles and how to couple deep learning with chemical-transport physics in next-generation operational systems. Full article
Show Figures

Graphical abstract

Back to TopTop