3.1. Experimental Comparison
In the systematic evaluation of classification performance on the FY4A+FY4B fused dataset, we first summarized the Precision, Recall, and F1-score values of LF-Transformer, Random Forest, and Saint across the four visibility levels and compiled them into a performance comparison table, as shown in
Table 2. This table provides a quantitative foundation for the subsequent analysis and makes the differences among the three models across various metrics and categories more explicit. On this basis, we further plotted the line charts of Precision, Recall, and F1-score to illustrate the performance variation and overall trends of LF-Transformer, Random Forest, and Saint across different visibility levels, as shown in
Figure 3.
Figure 3 compares the classification performance of LF-Transformer, Saint, and Random Forest on the FY4A+FY4B fused dataset, covering three core evaluation metrics: Precision, Recall, and F1-score. Overall, LF-Transformer significantly outperforms the other two models across all visibility categories and all metrics, while the performance of Saint falls between LF-Transformer and Random Forest.
Taking the F1-score as an example, LF-Transformer achieves an average score of 0.59 across the four categories, representing an increase of approximately 13% over the value of 0.47 obtained by Random Forest and also exceeding the value of 0.48 achieved by Saint. LF-Transformer maintains a high recall while effectively controlling the false-positive rate, demonstrating a more desirable balance between precision and recall. In contrast, although Saint performs similarly to LF-Transformer in Class 0 and Class 3, its F1-score drops sharply to 0.27 in Class 1, indicating that it still struggles to recognize intermediate visibility levels.
For Precision and Recall, the LF-Transformer achieves average values of 0.59 and 0.60, respectively, which are markedly higher than the values of 0.47 obtained by Random Forest and also superior to the value of 0.48 achieved by Saint. Moreover, LF-Transformer exhibits smaller variations across categories (standard deviations of 0.10 and 0.11 for Precision and Recall, respectively; compared with 0.13 and 0.12 for Random Forest, and 0.10 and 0.20 for Saint), suggesting stronger robustness and adaptability under complex or extreme conditions.
This performance advantage is particularly critical for visibility-level classification. Due to the continuity and fuzzy boundaries between visibility levels, traditional methods often struggle in transition regions. By leveraging deep feature modeling and multi-source information fusion, LF-Transformer shows enhanced adaptability and recognition ability when dealing with complex and ambiguous scenarios. Meanwhile, Saint outperforms Random Forest in most categories, further indicating that Transformer-based architectures are more suitable for high-dimensional, multi-channel remote-sensing visibility retrieval tasks.
In summary, the results in
Figure 3 provide a clear and systematic demonstration of the performance superiority of LF-Transformer, offering strong support for its practical application and broader adoption in remote-sensing-based visibility classification.
To further reveal the prediction distribution and misclassification patterns of each model across different categories,
Figure 4 presents the visualized confusion matrices of the three models. These matrices intuitively reflect the prediction distribution for each class and provide useful insight into overall classification performance.
Overall, the LF-Transformer shows noticeably higher values along the main diagonal compared with both Random Forest and Saint, indicating its stronger ability in correct classification. For example, in the dense or thick fog category (Class 0), the LF-Transformer correctly identified 525 out of 700 samples, achieving a precision of 75%, which is 11% higher than Random Forest and also superior to Saint’s 420 correctly identified samples. In addition, LF-Transformer produced fewer false positives in non-target categories, demonstrating its superior capability in recognizing low-visibility conditions. Although Saint performs significantly better than Random Forest in this class, it still falls slightly behind the LF-Transformer.
In the most challenging heavy fog category (Class 1), the LF-Transformer correctly classified 371 out of 700 samples, reaching a precision of 53%, which is 8% higher than the 245 correctly identified samples of Random Forest. Most of its misclassifications were concentrated in the adjacent categories (Class 0 and Class 2), indicating fewer random errors and better interpretability. In contrast, Saint performed relatively poorly in Class 1, correctly identifying only 280 samples, with more dispersed misclassification patterns. This suggests that it still struggles to handle the fuzzy boundaries characteristic of intermediate low-visibility levels.
In the moderate fog (Class 2) and clear weather (Class 3) categories, the LF-Transformer again demonstrates higher precision and more concentrated misclassification patterns. Compared with Random Forest, it achieves performance improvements across all categories. Saint performs similarly to or slightly better than Random Forest in these two classes (e.g., 420 correct predictions in Class 3), but overall, it remains inferior to the LF-Transformer, reflecting differences in the ability of Transformer-based architectures to model samples with fuzzy category boundaries.
In summary, the LF-Transformer provides more stable and accurate class discrimination and shows significant advantages and strong application potential under complex meteorological conditions. Saint performs better overall than Random Forest but still requires improvement in the more difficult classification categories.
To more intuitively analyze the prediction tendencies and misclassification distributions of the models across different categories,
Figure 5 presents a comparison between the predicted results and the true labels for the Random Forest, LF-Transformer, and Saint models. The side-by-side bar charts show the differences in predicted sample counts for each class, a columnar summation in
Figure 4, reflecting the degree of classification bias exhibited by each model.
Overall, the LF-Transformer’s predictions are closer to the true distribution. It slightly overestimates the dense or thick fog (Class 0) and moderate fog (Class 2) categories, while slightly underestimating the heavy fog (Class 1) and clear weather (Class 3) categories. However, the model shows a minimum-bias in Class 1 and Class 3, compared to the other two methods, suggesting the reliability of its discrimination between frog and non-frog events. Also, because its predictions are more concentrated, the model appears to effectively capture inter-class feature boundaries, demonstrating LF-Transformer’s advantages in maintaining classification consistency and controlling misclassification.
In contrast, the Random Forest model shows significant underestimation and dispersed predictions in the heavy fog (Class 1) and moderate fog (Class 2) categories, consistent with the instability observed in its confusion matrix.
The Saint model also exhibits notable prediction bias. It overestimates Class 1, with the predicted sample count reaching 785, while underestimating Class 0 and Class 3, yielding only 648 and 669 samples, respectively, compared with the true count of 700. Although Saint’s predicted count for Class 2 is close to the true value, the confusion matrix in
Figure 4 shows that its true positives are not high, indicating insufficient prediction concentration and stronger class confusion.
To further validate the practical inversion capabilities of the three models under low-visibility conditions, this study selected a typical persistent low-visibility event that occurred near the Jiaxing–Shaoxing Bridge in Zhejiang Province on 28–30 December 2023, and visualized the predicted visibility levels at observation stations, as shown in
Figure 6. The color shading in the figures represents different visibility classes.
As shown in
Figure 6, the LF-Transformer can well reproduce the horizontal distribution patterns and intensity inversion characteristics observed in real conditions. Specifically, the model exhibits the highest stability in identifying dense or thick fog areas (Class 0), with a misclassification rate of only 27%. The misclassification rate for clear-weather conditions without fog or haze (Class 3) is 35%, while the rates for heavy fog (Class 1) and moderate fog (Class 2) are 46% and 52%, respectively.
These differences may be related to the varying complexity of observational features corresponding to different visibility levels. Dense or thick fog (Class 0) typically exhibits more pronounced radiative characteristics (such as strong absorption signals in the 7.42 μm low-level water vapor channel), making it easier for the model to capture its spatial distribution. In contrast, heavy fog (Class 1) and moderate fog (Class 2) exhibit weaker radiative differences from their surroundings and have more ambiguous boundaries, leading to confusion between adjacent levels.
Compared with the LF-Transformer, the spatial inversion results of the Random Forest model reveal significantly weaker recognition capability in low-visibility regions, with reconstructed spatial structures deviating substantially from the observed distribution. Specifically, the Random Forest shows unstable performance in identifying dense or thick fog (Class 0), with a misclassification rate of 38%, and incorrectly classifies large portions of the core low-visibility region as Class 1 or Class 2, resulting in pronounced omission errors. For heavy fog (Class 1) and moderate fog (Class 2), the Random Forest’s misclassification rates reach 60% and 58%, respectively—substantially higher than those of the LF-Transformer. Its spatial predictions exhibit clear over-smoothing, blurred fog boundaries, inaccurate intensity gradients, and even cross-regional misclassification, indicating limited capability in distinguishing mid-level fog features. Under clear-weather conditions (Class 3), although the Random Forest can roughly capture the large-scale structure of non-fog areas, its misclassification rate remains as high as 45%.
The Saint model demonstrates spatial inversion accuracy between that of the LF-Transformer and Random Forest. Its misclassification rates for Class 0, Class 1, Class 2, and Class 3 are 42%, 56%, 60%, and 43%, respectively. Saint performs better than Random Forest in identifying dense fog (Class 0) but still falls noticeably short of the LF-Transformer, with local fog boundaries remaining blurred. For heavy fog (Class 1) and moderate fog (Class 2), Saint’s misclassification rates are comparable to those of the Random Forest, showing substantial spatial confusion and cross-regional errors, especially in mid-level fog regions where it struggles to capture fine-scale structures. In clear-weather conditions (Class 3), Saint can stably identify large-scale non-fog areas, but still exhibits a non-negligible proportion of misclassification, failing to match the spatial consistency achieved by the LF-Transformer.
To characterize the prediction bias of different models across various visibility levels, the spatial distribution of prediction errors at observation stations was analyzed for a representative Zhejiang low-visibility event on 28 December 2023.
Figure 7 presents the spatial error distribution produced by the LF-Transformer model, while the corresponding results for the Random Forest and Saint models are provided in
Supplementary Figures S1 and S2, respectively. The prediction error is defined as the difference between the predicted visibility class and the true visibility class (Predicted Class−True Class). An error value of zero indicates a correct prediction, while positive and negative values indicate overestimation and underestimation of the visibility level, respectively.
As shown in
Figure 7, the LF-Transformer exhibits relatively small prediction errors across all true visibility classes, with a high degree of spatial consistency. Under dense or thick fog (Class 0) and heavy fog (Class 1) conditions, prediction errors are mainly concentrated around zero, with only localized deviations between adjacent visibility classes, indicating strong prediction stability under low-visibility conditions.
For moderate fog conditions (Class 2), prediction errors increase to some extent; however, the spatial distribution remains relatively continuous, suggesting that the LF-Transformer maintains stable discrimination capability under transitional visibility levels. Under clear-weather conditions (Class 3), the error distribution is generally concentrated, demonstrating reliable performance in identifying high-visibility regions.
In contrast, the Random Forest model shown in
Supplementary Figure S1 exhibits pronounced spatial dispersion in prediction errors. Under low-visibility conditions (Class 0–1), predictions at several stations within core low-visibility regions tend to shift toward higher visibility classes, resulting in an underrepresentation of low-visibility extent.
Under moderate fog conditions (Class 2), the Random Forest displays the most scattered error distribution, with fragmented spatial structures, indicating substantial uncertainty in distinguishing adjacent visibility classes. Even under clear-weather conditions (Class 3), a noticeable proportion of prediction bias remains, reflecting limited prediction stability under complex spatial patterns.
Supplementary Figure S2 presents the spatial distribution of prediction errors for the Saint model, whose overall performance lies between that of the LF-Transformer and the Random Forest. Under dense fog conditions (Class 0), Saint shows slight improvement compared to the Random Forest; however, error points still exhibit evident spatial dispersion.
Under heavy and moderate fog conditions (Class 1–2), the Saint model shows scattered error distributions with localized instability. For clear-weather conditions (Class 3), Saint can identify large-scale high-visibility regions, but the concentration of prediction errors remains weaker than that of the LF-Transformer.
Taken together, the comparison of prediction error distributions in
Figure 7 indicates that the LF-Transformer produces more concentrated errors with greater spatial continuity and stability across all visibility classes. In contrast, the Random Forest and Saint models are more prone to dispersed prediction errors and fragmented spatial structures, particularly under moderate visibility conditions.
Overall, the LF-Transformer demonstrates superior performance not only in quantitative evaluation metrics but also in spatial consistency, showing stronger representational capability and more reliable inversion performance for regional fog recognition and early-warning applications. Saint outperforms Random Forest but still shows clear limitations in distinguishing complex fog structures and capturing fog boundaries.
3.4. Comparison Between LF-Transformer and Ensemble LF-Transformer on the FY4A+FY4B Fused Dataset
To further improve the model’s generalization ability and its recognition performance for complex visibility categories, an ensemble training strategy was designed in this study. A single model, when dealing with class imbalance, sample diversity, and fuzzy category boundaries in satellite data, tends to suffer from overfitting and unstable predictions. The ensemble approach mitigates these limitations by constructing multiple diverse sub-models and integrating their predictions. Specifically, the ensemble method involves using different parameter initializations to ensure diversity among sub-models, applying feature perturbation to enhance robustness against noise and complex samples, and fusing sub-model outputs through weighted averaging or voting to improve overall prediction stability and classification performance.
On this basis, we further summarized the Precision, Recall, and F1-score values of the original LF-Transformer and the ensemble LF-Transformer on the FY4A+FY4B fused dataset and compiled them into a comparison table, as shown in
Table 5, to quantitatively demonstrate the performance improvements introduced by the ensemble strategy. Subsequently, we plotted the Precision, Recall, and F1-score curves for both models across different visibility levels to more intuitively illustrate the performance differences and trends, as shown in
Figure 9. In general,
Figure 9 demonstrates that the ensemble strategy provides a comprehensive and consistent performance improvement, particularly in Recall and F1-score. Specifically, although Precision in the dense or thick fog category (Class 0) slightly decreased (from 0.75 to 0.74), Recall increased from 0.79 to 0.82, and F1-score rose slightly from 0.77 to 0.78, indicating better recall capability and balanced performance.
For other categories, Precision improved across the board—for example: heavy fog (Class 1) increased from 0.53 to 0.55, moderate fog (Class 2) from 0.50 to 0.52, and clear weather without fog or haze (Class 3) from 0.61 to 0.64 —demonstrating enhanced discriminative ability of the ensemble model.
In terms of Recall, all categories showed improvement except for Class 1, which slightly decreased (from 0.54 to 0.52). Notably, moderate fog (Class 2) increased from 0.44 to 0.47, and clear weather (Class 3) from 0.64 to 0.66, suggesting that the ensemble model is more sensitive to minority and boundary samples, effectively reducing the risk of missed detections.
F1-scores also increased for all classes—from 0.53, 0.47, and 0.62 to 0.54, 0.49, and 0.65, respectively—reflecting the model’s balanced improvement in both accuracy and recall.
In summary, the ensemble LF-Transformer achieved robust and consistent improvements across multiple metrics, validating the significant effectiveness of ensemble learning for remote sensing-based visibility classification tasks.
In the spatial prediction task for typical visibility events, the ensembled LF-Transformer model maintains higher classification accuracy over a large spatial domain, with particularly strong performance for dense fog (Class 0) and moderate fog (Class 2). The corresponding spatial prediction results are shown in
Supplementary Figure S3. The corresponding misclassification rates are 24% and 48%, representing reductions of 3% and 4%, respectively, compared with the original model. In contrast, the original model frequently exhibited class boundary misjudgments in transition regions (e.g., misclassifying Class 2 moderate fog as Class 1 heavy fog). The ensembled model significantly reduces both misclassification and omission errors, demonstrating its stronger spatial generalization ability in dealing with ambiguous class boundaries and imbalanced sample distributions in complex visibility scenarios.
In summary, the ensemble training not only improved numerical performance metrics but also enhanced class discrimination and spatial consistency, exhibiting greater robustness and adaptability. Compared with traditional machine learning methods, the LF-Transformer—leveraging its deep architecture and multi-head attention mechanism—can effectively capture multi-scale nonlinear features in remote sensing data. By introducing the ensemble strategy, the model further mitigates class confusion, improves stability and reliability in complex remote sensing environments, and demonstrates strong potential for intelligent visibility classification applications.
3.5. Computational Efficiency and Operational Deployability Analysis
Following the spatial inversion analysis of a typical low-visibility case, this study further evaluates the computational efficiency and operational deployability of the three models to assess their potential for real-time visibility monitoring. All experiments were conducted on a unified hardware platform equipped with an NVIDIA A30 GPU.
In terms of computational cost, the LF-Transformer used in this study contains 36,598,616 trainable parameters, representing a mid-sized Transformer architecture. Although its parameter scale is considerably larger than that of the lightweight Random Forest and Saint models, the LF-Transformer still demonstrates favorable efficiency during inference. Its single-sample inference latency ranges between 8–15 ms, with a peak GPU memory usage of approximately 0.8–1.5 GB. For the complete test set of 2800 samples, batch inference with a batch size of 50 requires only 4–6 s, indicating strong throughput performance sufficient to meet the real-time, minute-level update requirements of operational visibility monitoring systems.
Regarding training cost, the LF-Transformer required 34 min and 59 s to complete 100 training epochs on the A30 GPU, which remains within an acceptable range for operational applications. Although its training time is noticeably longer than that of Random Forest and Saint, the LF-Transformer exhibits substantially stronger multi-channel feature modeling capacity and better generalization stability. On the test set, the model achieved a 60% classification accuracy, with F1-scores of 0.77 (Class 0), 0.53 (Class 1), 0.47 (Class 2), and 0.62 (Class 3), indicating superior capability in identifying extremely low-visibility regions.
Random Forest, as a traditional statistical learning method, provides the fastest inference speed with the lowest memory consumption, while the Saint model—owing to its lightweight architecture—also delivers higher inference efficiency compared with the LF-Transformer. However, these computational advantages come at the cost of significantly reduced predictive performance. Both Random Forest and Saint exhibit low quantitative accuracy, blurred fog boundaries, and poor spatial consistency, making them unsuitable for fine-grained visibility classification under complex low-visibility conditions. In comparison, the LF-Transformer achieves improvements of approximately 12% in precision, 13% in recall, and 12% in F1-score, and shows markedly better performance in preserving spatial structures, identifying extreme low-visibility events, and maintaining cross-regional stability, thereby striking a more effective balance between computational cost and predictive performance.
From a deployment perspective, the inference latency, memory usage, and throughput of the LF-Transformer all fall within the acceptable range for current operational platforms such as traffic monitoring and meteorological early-warning systems that commonly utilize A30 GPUs. Its high stability and accuracy under low-visibility conditions provide strong support for fog monitoring and early-warning applications. Additionally, further optimizations—such as model pruning, knowledge distillation, and mixed-precision inference—may reduce computational overhead and enhance adaptability in broader operational environments.
In summary, considering both training and inference efficiency, the LF-Transformer achieves significant performance gains in multi-source remote sensing visibility classification while maintaining manageable computational requirements. It represents a highly accurate, scalable, and operationally valuable solution. Although Random Forest and Saint offer advantages in computational efficiency, their limitations in recognizing complex fog structures restrict their applicability in high-precision operational scenarios.