Air Pollution Forecasting Using Autoencoders: A Classification-Based Prediction of NO2, PM10, and SO2 Concentrations
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsDeep-leaning based forecast of the concentration levels of various air pollutants is performed. AE and ACE are used for the prediction, focusing on the case the Bay of Algeciras. The prediction accuracy of the two algorithms is obtained. While the model introduction is extensive, analysis and presentation of the model predictions appear to be premature, as no figure is given for the main research content. The manuscript requires major revision to improve the paper’s presentation, data analyses, and readability.
- First paragraph of the introduction focuses on the situations in Spain, however, your title suggests a very general application/topic. Either constrain the paper scope by changing the title, or improve the introduction to a more global perspective.
- “There are several main reasons why predicting future values of these air pollutants can be important”, however, authors do not clarify what is missing or challenged for such prediction in the current study, and how your work can help to advance/improve the understanding or help to resolve the issue.
- A more extensive review of the deep learning-based methods for air pollution prediction is definitely needed, it should cover most recent literatures, and authors are supposed to comment on the pros and cons of these studies, and how their approaches are different from yours.
- Source of figure 1 should be provided.
- Section 2.2.1, Quality Measurements, the definition of the variables should be more specific. What physical and chemical quantities do “results” represent. Some examples should be given.
- Section 2.2.2 and 2.2.3, there is no need to repeat well-known knowledge, put what is new in your approach which is relevant to the pollutant concentration prediction.
- For the Results section, the data analysis is rather limited and lacks sufficient visualization to support the reported findings. Some performance indicators (accuracy, sensitivity, specificity, and precision) are computed, the results are only presented in large numerical tables without graphical summaries or statistical interpretation. No figures are provided to illustrate model behavior, such as performance comparison between AE and SAE. The absence of visual or exploratory analysis limits the understanding of model performance. The section remains largely descriptive, and requires more analytical and visual validation.
- There is no experimental validation in the sense of comparing model predictions to newly measured or independently collected field data. This severely limits the reliability of the reported accuracy.
- The discussion is limited to the special case in the Bay of Algeciras. This lacks general interest. Authors should expand the discussion on how the proposed model can be applied to other cases.
Author Response
Comments and Suggestions for Authors
Deep-leaning based forecast of the concentration levels of various air pollutants is performed. AE and ACE are used for the prediction, focusing on the case the Bay of Algeciras. The prediction accuracy of the two algorithms is obtained. While the model introduction is extensive, analysis and presentation of the model predictions appear to be premature, as no figure is given for the main research content. The manuscript requires major revision to improve the paper’s presentation, data analyses, and readability.
- First paragraph of the introduction focuses on the situations in Spain, however, your title suggests a very general application/topic. Either constrain the paper scope by changing the title, or improve the introduction to a more global perspective.
The authors agree with this recommendation and a new paragraph has been added in the introduction section highlighted in yellow:
Air pollution forecasting has become a crucial environmental and public health issue worldwide. Urban and industrial regions face increasing challenges in predicting pollutant concentration dynamics due to complex emission sources and meteorological interactions. Among the various forecasting methods, deep learning (DL) models have shown promise for capturing nonlinear dependencies between variables across different geographic contexts.
Besides, some new references have been added to the introduction section [1-3] and in References section:
- Bekkar, A., Hssina, B., Douzi, S. et al.Air-pollution prediction in smart city, deep learning approach. J Big Data 8, 161 (2021)
- Agbehadji, I.E.; Obagbuwa, I.C. Systematic Review of Machine Learning and Deep Learning Techniques for Spatiotemporal Air Quality Prediction. Atmosphere2024, 15, 1352.
- Sun, H., Fung, J. C. H., Chen, Y., Li, Z., Yuan, D., Chen, W., and Lu, X.: Development of an LSTM broadcasting deep-learning framework for regional air pollution forecast improvement. Model Dev., 2022, 15, 8439–8452
- “There are several main reasons why predicting future values of these air pollutants can be important”, however, authors do not clarify what is missing or challenged for such prediction in the current study, and how your work can help to advance/improve the understanding or help to resolve the issue.
The authors agree with this recommendation and a new paragraph has been added in the introduction section highlighted in yellow:
Despite the progress in machine learning for air quality forecasting, most existing studies rely on either traditional regression models or deep recurrent networks such as Long Short-Term Memory (LSTM), which primarily capture temporal dependencies but neglect latent spatial correlations among variables. Recent developments have demonstrated the strong potential of hybrid Deep Learning architectures (e.g., Autoencoders-Long Short-Term Memory (AE–LSTM), Convolutional Neural Network – Long Short-Term Memory (CNN–LSTM), Transformer-based models) to improve pollutant forecasting accuracy. Studies such as [1-3] [37-38] highlight that autoencoders enable feature compression and denoising, while LSTM networks capture temporal dependencies. However, the combination or comparative performance of AE versus Sparse AE in port environments remains largely unexplored, which constitutes the main novelty of this work. Furthermore, few works focus on highly heterogeneous port–industrial regions, where emissions are influenced by both maritime and industrial sources [22-24][33]. This study contributes to filling this gap by applying and comparing Autoencoder (AE) and Sparse Autoencoder (SAE) architectures to model pollutant concentration quartiles, capturing intrinsic features from both meteorological and port activity data.
- A more extensive review of the deep learning-based methods for air pollution prediction is definitely needed, it should cover most recent literatures, and authors are supposed to comment on the pros and cons of these studies, and how their approaches are different from yours.
The authors agree with this recommendation and a new paragraph has been added in the introduction section highlighted in yellow:
Despite the progress in machine learning for air quality forecasting, most existing studies rely on either traditional regression models or deep recurrent networks such as LSTM, which primarily capture temporal dependencies but neglect latent spatial correlations among variables.
Recent developments have demonstrated the strong potential of hybrid DL architectures (e.g., AE–LSTM, CNN–LSTM, Transformer-based models) to improve pollutant forecasting accuracy. Studies such as [1-3] [37-38] highlight that autoencoders enable feature compression and denoising, while LSTM networks capture temporal dependencies. However, the combination or comparative performance of AE versus Sparse AE in port environments remains largely unexplored, which constitutes the main novelty of this work.
Furthermore, few works focus on highly heterogeneous port–industrial regions, where emissions are influenced by both maritime and industrial sources [22-24][33]. This study contributes to filling this gap by applying and comparing Autoencoder (AE) and Sparse Autoencoder (SAE) architectures to model pollutant concentration quartiles, capturing intrinsic features from both meteorological and port activity data.
Besides, two more new deep learning references have been added:
- Mengara Mengara, A.G.; Kim, Y.; Yoo, Y.; Ahn, J. Distributed Deep Features Extraction Model for Air Quality Forecasting. Sustainability 2020, 12, 8014.
- Basir, N. I., Tan, K. K., Djarum, D. H., Ahmad, Z., Vo, D.-V. N., & Jie, Z. Autoencoder Artificial Neural Network Model for Air Pollution Index Prediction. IIUM Engineering Journal, 2025, 26(1), 1–21.
- Source of figure 1 should be provided.
The authors agree with this recommendation and the source has been added in the Figure 1 highlighted in yellow:
Source: Google Earth and authors’ own elaboration
Figures 2 and 3 have been modified to a better design:
- Section 2.2.1, Quality Measurements, the definition of the variables should be more specific. What physical and chemical quantities do “results” represent. Some examples should be given.
The authors agree with this recommendation and a new paragraph has been added in the 2.2.1 Quality Measurements highlighted in yellow:
In this study, the results obtained refer to the hourly concentration predictions (µg/m³) of NO₂, PM₁₀ and SO₂, obtained from the AE/SAE models using as inputs meteorological variables (temperature, wind speed, humidity, etc.), vessel activity (gross tonnage per hour), and pollutant levels at surrounding monitoring stations.
- Section 2.2.2 and 2.2.3, there is no need to repeat well-known knowledge, put what is new in your approach which is relevant to the pollutant concentration prediction.
The authors agree with this recommendation, the well-known knowledge has been reduced,
Autoencoders (AEs) are neural networks designed to replicate input data at the output with minimal distortion and play an important role in machine learning. They were first introduced in the 1980s by Hinton and the Parallel Distributed Processing (PDP) group [48], using the input data as "supervision". AEs are a fundamental paradigm of unsupervised learning, where local synaptic changes can lead to coordinated global learning [48]. An AE consists of an encoder, which transforms the input into an internal representation, and a decoder, which reconstructs the original data from this representation. During training, the network adjusts its weights and biases to minimise the difference between the original and reconstructed data. The hidden layer produces the encoded data: if the number of hidden neurons (NH) is smaller than the input dimension (D), the code is compressed; if NH > D, a sparse representation is obtained [50-51]. The loss function (Eq. 5) measures the reconstruction error, combining mean squared error, L2 regularisation , and sparsity regularisation , with L2 helping to set the parameters and [52].
The L2 regularisation term in Equation 6 sums the squared elements of the weight matrices for each layer, while the sparsity regulariser in Equation 7 encourages sparse representations in the hidden layer, with average neuron activation and target ?. Autoencoders (AEs) are unsupervised networks that replicate input at the output, learning an intermediate representation in a different dimensional space. This study compares Sparse Autoencoder (SAE) and AE performance, showing that SAE can extract specific features from the input, whereas AE cannot, though both reproduce the input at the output. Both networks were trained independently until the validation error reached a minimum. A stacked configuration with two autoencoder layers followed by a supervised layer was used to predict air pollutant concentration quartiles. Preprocessing included imputing missing values and normalising variables. Two autoencoders of dimensions NH1 and NH2 were trained and combined into a stacked AE (Fig. 2), then used in testing with the supervised layer to predict the future signal (Fig. 3).
Grid search is a hyperparameter optimisation method [53] that exhaustively evaluates models for all combinations within a predefined hyperparameter space. Since hyperparameters are set before training, the method systematically explores the grid until the best combination is found. It is most effective when the number of hyperparameters is below seven (M < 7) and the search limits are well defined. Although simple and effective, grid search can be computationally expensive [53].
Figure 4 illustrates movement within 1D, 2D, and 3D hyperparameter grids, where unevaluated neighboring cells are explored to improve performance metrics such as MSE. Previously analysed cells are skipped, and to avoid local minimum, the process is repeated 20 times.
and a new paragraph has been added, highlighted in yellow.
Unlike conventional autoencoders used for dimensionality reduction, in this work the stacked AE/SAE is configured to classify pollutant concentration quartiles, enabling the model to focus on distinct pollution intensity regimes. Furthermore, the grid search was tailored to optimise sparsity parameters jointly with layer size, which has not been previously applied in environmental forecasting models.
- For the Results section, the data analysis is rather limited and lacks sufficient visualization to support the reported findings. Some performance indicators (accuracy, sensitivity, specificity, and precision) are computed, the results are only presented in large numerical tables without graphical summaries or statistical interpretation. No figures are provided to illustrate model behavior, such as performance comparison between AE and SAE. The absence of visual or exploratory analysis limits the understanding of model performance. The section remains largely descriptive, and requires more analytical and visual validation.
The authors agree with this recommendation and the Figures have been added to the Results section:
The differences between Autoencoder (AE) and Sparse Autoencoder (SAE) in accuracy, sensitivity, precision, and specificity for each pollutant (SO₂, PM₁₀, NO₂) will be visually shown using the best model in each case (see Figure 5a,b,c).
|
Figure 5.a. Comparative Performance: AE vs SAE for SO2. |
|
Figure 5.b. Comparative Performance: AE vs SAE for PM10. |
|
Figure 5.c. Comparative Performance: AE vs SAE for NO2. |
Figure 5a,b,c compares the classification performance of Autoencoder (AE) and Sparse Autoencoder (SAE) models across pollutants. AE configurations generally achieve higher accuracy for moderate concentration levels (Q2–Q3), while SAE models perform slightly better at extreme quartiles (Q4), particularly for NO₂. This suggests that sparse architectures may better capture the non-linear behavior of pollutants at peak concentrations.
- There is no experimental validation in the sense of comparing model predictions to newly measured or independently collected field data. This severely limits the reliability of the reported accuracy.
The authors agree with this recommendation, a limitation of the current study is the lack of external validation using independent or newly collected field data.
Although an external validation using newly collected field data has not yet been performed, the proposed models were rigorously validated through a classification-based evaluation using sensitivity, specificity, precision, and accuracy metrics. These results confirmed that the selected configurations significantly outperform the benchmark Multiple Linear Regression (MLR) models across all pollutants and quartiles. Moreover, the Bonferroni and Friedman tests provided statistical validation for the selection of optimal model architectures, supporting the principle of model simplicity (Occam’s Razor) where multiple configurations yielded similar results. Future work will include an additional validation phase using a new dataset to be provided by the Andalusian Regional Government, which will allow for assessing the model’s generalization capability under unseen meteorological conditions.
The following sentence has been added to the conclusion section:
Moreover, future work will be carried out using newly collected data by deploying portable air quality sensors in the Bay of Algeciras to validate the model predictions in real time and to assess the model’s generalization under unseen meteorological conditions.
- The discussion is limited to the special case in the Bay of Algeciras. This lacks general interest. Authors should expand the discussion on how the proposed model can be applied to other cases.
The authors agree with this recommendation, a new paragraph has been added to Discussion Section:
Although the present study focuses on the Bay of Algeciras, the proposed AE/SAE framework can be generalized to other port or industrial regions with similar multivariate data structures. By retraining the network with local meteorological and emission inputs, the model could be adapted to predict air quality in other coastal zones such as Rotterdam, Hamburg, or Singapore. This adaptability demonstrates the potential for scalable implementation in port environmental management systems.
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsIn their paper the author show the results of a modelling study on the prediction of pollutant concentration in a specific region. In general the topic is interesting and the methods are well described, however some small issues should be addressed before publication:
1. in section 3 (lines 385-419) several acronyms are introduced (e.g. SVM,...) which might not be familiar to all readers. In general this section is not well connected to the rest of the article. Maybe it is better placed in the introduction section? Or the comparison with the current findings should be made more specific.
2. The analysis of the results is a bit on the weak side. The paper discusses results solely at a classification level. Maybe it would be useful to check if the model tends to over or underestimate the actual pollutant levels? Or is this balanced? Does misclassification also correlate with very extreme values of the input parameters or is there a seasonal trend that would indicate a hidden parameter (e.g. like additional tourist traffic unaccounted for, or regional differences like wind shadow zones where pollutants can accumulate)?
3. Why did the data set cover only years 2017-2019? Are data no longer recorded? This should be commented in the text.
4. The blocks for Q1 (the first five lines) in tables 4 and 5 are exactly identical. This seems to be very unlikely for real data. Is this a copy&paste error?
Author Response
Comments and Suggestions for Authors
In their paper the author show the results of a modelling study on the prediction of pollutant concentration in a specific region. In general the topic is interesting and the methods are well described, however some small issues should be addressed before publication:
- in section 3 (lines 385-419) several acronyms are introduced (e.g. SVM,...) which might not be familiar to all readers. In general this section is not well connected to the rest of the article. Maybe it is better placed in the introduction section? Or the comparison with the current findings should be made more specific.
The authors are grateful for pointing this out. The authors have revised Section 3 to improve readability and consistency. The acronyms (e.g., SVM, GRU, LSTM, DWT) are now explicitly defined when first mentioned.
…methods (Support Vector Machined (SVM), Gated Recurrent Units (GRU), LSTM, Discrete Wavelet Transform (DWT-LSTM)).
- The analysis of the results is a bit on the weak side. The paper discusses results solely at a classification level. Maybe it would be useful to check if the model tends to over or underestimate the actual pollutant levels? Or is this balanced? Does misclassification also correlate with very extreme values of the input parameters or is there a seasonal trend that would indicate a hidden parameter (e.g. like additional tourist traffic unaccounted for, or regional differences like wind shadow zones where pollutants can accumulate)?
The authors agree that a more detailed interpretation of the model behavior would strengthen the paper. We have therefore added an additional analysis discussing possible over- or under-estimation patterns and their relationship to seasonal or meteorological variations. Although the current model predicts quartiles (classification framework), we have re-examined the confusion matrices to determine whether misclassifications are biased toward higher or lower concentration classes. We found that minor misclassifications tend to occur between adjacent quartiles (e.g., Q2→Q3), which suggests balanced predictions rather than systematic bias.
A brief paragraph has been added to the Discussion section describing this and discussing potential seasonal influences such as wind patterns or increased summer traffic:
An examination of the confusion matrices revealed that most errors occur between adjacent quartiles (Q2–Q3), indicating that the model neither systematically overestimates nor underestimates pollutant levels. Misclassifications were slightly more frequent during summer months, coinciding with higher port activity and temperature inversions, which could suggest hidden exogenous factors (e.g., increased tourist traffic or wind-shadow accumulation zones).
- Why did the data set cover only years 2017-2019? Are data no longer recorded? This should be commented in the text.
The authors are grateful with this observation. The selected period (2017–2019) corresponds to the last continuous and complete dataset available from the Andalusian Government and the Port Authority before interruptions caused by the COVID-19 pandemic. After 2020, some monitoring stations in the Bay of Algeciras had data gaps or maintenance outages, which would have compromised model homogeneity. A short explanation has now been added in Section 2.1 (Materials).
The following paragraph has been added to 2.1 section Materials:
The dataset spans from January 2017 to December 2019, which corresponds to the last period of uninterrupted and homogeneous data availability across all monitoring stations. Subsequent years include several data gaps due to equipment maintenance and disruptions during the COVID-19 pandemic, hence they were excluded to preserve temporal consistency.
- The blocks for Q1 (the first five lines) in tables 4 and 5 are exactly identical. This seems to be very unlikely for real data. Is this a copy&paste error?
The authors appreciate the reviewer’s attention to this detail. The identical Q1 block between Tables 4 and 5 was not a data issue but a formatting error that occurred when adapting the manuscript to a different journal template. The table has now been carefully revised, and the correct SO2 results are presented in the updated version.
Author Response File:
Author Response.pdf
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsMy comments are addressed and I recommend acceptance.
Reviewer 2 Report
Comments and Suggestions for AuthorsAll issues have been addressed in the new version of the manuscript. In my opinion it can be published in the current form.
