Next Article in Journal
Enhanced Natural Remediation of Nitrate by Pumping Groundwater from Active Denitrification Depth
Previous Article in Journal
Regional Differences in the Potential Drivers of Grassland Degradation from the Perspective of Partial-Order Theory: A Case Study of Ordos
Previous Article in Special Issue
The Tropical Challenge in Solar Energy Modelling: Spatial and Seasonal Breakdown of Semi-Empirical Approaches Under Topographic Heterogeneity
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Wildfire Susceptibility Mapping in China Combining Machine Learning, Deep Learning, and Transformer-Based Models

by
Uroš Durlević
1,*,
Velibor Ilić
2,
Milan M. Radovanović
1,
Ana Milanović Pešić
1,3,
Marko D. Petrović
1,4,
Milan Milenković
1,
Jasmina M. Jovanović
5 and
Emin Atasoy
6
1
Geographical Institute “Jovan Cvijić”, Serbian Academy of Sciences and Arts, Đure Jakšića 9, 11000 Belgrade, Serbia
2
Institute for Artificial Intelligence Research and Development of Serbia, Fruškogorska 1, 21000 Novi Sad, Serbia
3
Department of Physical and Economic Geography, Institute of Natural Sciences, L.N. Gumilyov Eurasian National University, Satbayev 2 Str., 010008 Astana, Kazakhstan
4
Department of Regional Economics and Geography, Faculty of Economics, RUDN University (Peoples’ Friendship University of Russia), Miklukho-Maklaya 6 Str., 117198 Moscow, Russia
5
Faculty of Geography, University of Belgrade, Studentski trg 3/3, 11000 Belgrade, Serbia
6
Department of Social Sciences, Faculty of Education, Uludag University, 16000 Bursa, Türkiye
*
Author to whom correspondence should be addressed.
Earth 2026, 7(4), 119; https://doi.org/10.3390/earth7040119
Submission received: 8 June 2026 / Revised: 9 July 2026 / Accepted: 10 July 2026 / Published: 13 July 2026
(This article belongs to the Special Issue Special Issue Series: Young Investigators in Earth Science)

Abstract

Long-term wildfire susceptibility mapping represents a significant component of disaster prevention and the protection of human communities, public health, and local ecosystems. In this study, a wildfire inventory was developed through multi-sensor fusion of satellite data (MODIS and VIIRS), comprising 153,305 fire events across China for the period 2001–2024. In addition to historical incidents, 14 predictive variables were processed, representing geomorphological, climatological, hydrological, vegetative, and anthropogenic conditions. This study evaluates long-term spatial wildfire susceptibility based on long-term mean environmental and climatic conditions. Methodologically, the research applies six models from machine learning (ML), deep learning (DL), and transformer-based approaches: Random Forest (RF), Extreme Gradient Boosting (XGBoost), Deep Neural Network (DNN), Fourier Multi-Layer Perceptron (F-MLP), Kolmogorov–Arnold Network (KAN), and Feature Tokenizer (FT) Transformer. The results were integrated into an ensemble susceptibility map with a spatial resolution of 500 m using Geographic Information Systems (GIS), indicating that 7.4% of China’s territory is classified as having a very high wildfire susceptibility. In addition to the national-scale assessment, a local differentiation was conducted across 34 province-level divisions, revealing that Fujian Province (86.8%) and the Guangxi Zhuang Autonomous Region (82.9%) had the largest shares of areas classified as high and very high wildfire susceptibility. Performance evaluation under spatial block-based validation demonstrated that the Random Forest model achieved the highest predictive power, with an area under the curve (AUC) of 87.8%, followed by XGBoost (87.3%) and Fourier MLP (86.6%). Based on the combined SHAP (Shapley additive explanations) analysis of all applied models, soil moisture, elevation, and terrain slope were identified as the most influential factors affecting wildfire occurrence in China. Overall, the findings contribute to more effective wildfire prevention and risk management strategies at both the local and national levels.

1. Introduction

In 2025, forests covered approximately 31.8% of the world’s total land area [1], playing an important role in various aspects of the environment, maintaining biodiversity and regulating water cycles [2]. It is also important to emphasize that, each year, the world’s forests absorb 2.4 billion tons of CO2, accounting for one-third of the CO2 emitted by the combustion of fossil fuels [3].
In addition, forests play an important role in reducing soil erosion and the occurrence of flash floods, landslides, droughts, and other natural disasters. However, in the context of increasing global warming, manifested through frequent occurrences of extremely high air temperatures and long-lasting heat waves, wildfires represent a significant challenge and are gradually becoming a global issue [4,5]. According to Oom and Pereira (2013), approximately 3% of forests worldwide experience fires each year [6]. The consequences of wildfires are multiple and range from direct damage, such as the destruction of forest resources, degradation of biodiversity, release of harmful gases into the atmosphere, potential casualties, and population displacement, to significant economic losses in agricultural production, as well as damage to buildings and infrastructure [7]. The increasing frequency and severity of wildfires have increased interest in researching this issue at a global scale in order to prevent and control these natural disasters.
According to the EM-DAT database (2026) (which records natural disasters that meet one of the following criteria: 10 or more fatalities, 100 or more injuries, a declared state of emergency due to a natural disaster, or request for international aid), 336 wildfires were recorded worldwide between 2001 and 2025 [8]. These fires resulted in 2341 fatalities and 17,573 injuries, with economic losses estimated at approximately USD 170.3 billion. During this period, the highest number of wildfires occurred in North America (93), followed by Europe (73), South America (49), Asia (46), Africa (26), and Australia and Oceania (25). According to the Food and Agriculture Organization (2025), wildfires remain the greatest threat to the world’s forests, affecting an average of 127 million hectares of forests per year [9]. The Global Fire Atlas identified approximately 13.3 million individual wildfires worldwide between 2003 and 2013 [10]. At the global scale, the temporal and spatial distribution of wildfires is uneven. Data from the Fire Information for Resource Management System and Global Wildfire Information System indicate that the highest numbers of wildfires occur in August and October worldwide, while in Europe specifically, this is in August, followed by April [11,12]. The lowest numbers of wildfires occur during December, January, and February, which are the winter months in the northern hemisphere. In terms of location, the highest numbers of wildfires are registered in areas where global forests are widespread, such as Central Africa, the Amazon Basin in South America, the northern part of North America, Southeast and East Asia, and the Mediterranean area. Recent research has indicated that wildfires in the Mediterranean area account for more than 80% of the total forest area burned on the European continent [13].
Contemporary research on wildfires often has a multidisciplinary approach, from investigating causes and patterns of occurrence and spread to analyzing trends [14,15,16,17]. It also includes research on biological, ecological, and socioeconomic aspects, including assessment of impact on public health [18,19,20,21,22,23]. From a geographical perspective, many studies focus on how climatic, soil and anthropogenic factors, air pollutants, climate change and solar activity influence the occurrence of wildfires [24,25,26]. The utilization of Geographic Information Systems (GIS) and remote sensing has become essential to wildfire risk assessment, wildfire management, zoning risk areas, and modeling the spread of fires [27,28,29,30,31,32,33,34]. Moreover, significant attention has been paid to cascading natural hazards, including the occurrence of floods and flash floods after wildfires.

2. Literature Review

Research on wildfires has been the subject of numerous studies that have examined various aspects of this pressing issue worldwide. Based on satellite wildfire data from 2003 to 2023, Chen et al. (2026) indicate a continued global decline in fire activity, especially at low-to-mid latitudes in both hemispheres, although some regions are evidencing an accelerated increase [35]. Notably, increased wildfire activity has been observed in the northern part of North America, the northeastern part of Eurasia, and central Africa, primarily associated with a warmer climate, and in South Asia and China due to intensive agricultural burning. Conversely, fire activity has declined in the southeastern part of South America, some areas of Africa and Southeastern Asia, and the northwestern part of Eurasia. In Europe, research on wildfires is based on the application of modern spatiotemporal models to analyze the frequency and patterns of fire spread, particularly in the Mediterranean Basin [13] and southeastern Europe [16]. Certain studies have focused on specific countries, such as Portugal [36], Greece [37], and Serbia [38,39]. Studies focusing on Asia indicate an increasing trend in wildfire occurrence across South and Southeast Asia. Based on data from the Moderate Resolution Imaging Spectroradiometer (MODIS) (2003–2016) and Visible Infrared Imaging Radiometer Suite (VIIRS) (2012–2016), it was found that the highest number of recorded wildfires in South Asia occurred in India, followed by Pakistan and other countries. In Southeast Asia, the highest number of wildfires was recorded in Indonesia, followed by Myanmar and Laos [40]. In addition, it is important to highlight research on the impact of climate change, projected for the period 2019–2100, on future wildfire regimes in two large boreal study areas in central Russia and western Canada using three global climate models. Projections indicate that future fire weather severity will increase significantly in these boreal regions, with some climate change scenarios estimating a 400–500% increase in daily severity ratings [41]. Recent studies of wildfires in North and South America indicate an increase in their frequency, and highlight the March 2024–February 2025 season, when fires in the Canadian boreal forests [42] and in the Amazonia and the Pantanal–Chiquitano basins [43] led to record global carbon emissions. Research based on a combination of various satellite images for assessing burned areas and terrestrial ecosystem models to simulate fuel quantities and the impact of fires on ecosystem dynamics indicate that Africa as a continent represents a global hotspot for wildfires [44]. Geographical studies highlight Central Africa in particular as a critical region for biomass burning and greenhouse gas emissions from wildfires [45], with over 60% of such emissions occurring in Angola, DR Congo, Sudan, and the Central African Republic [46]. According to Haque et al. (2023), interest in wildfire research in Australia has increased following the so-called “Black Summer” of 2019–2020, during which 12 million hectares were destroyed and biodiversity was severely impacted [47].
Statistical studies of wildfires implement various quantitative indices to assess wildfire danger levels and environmental impacts, as well as sophisticated spatiotemporal models to analyze the patterns of wildfire occurrence. Conventional indices, such as the Canadian Fire Weather Index (FWI) and the differenced normalized burn ratio (dNBR) provide daily risk estimates [48,49,50]. Jolly et al. (2015) calculated global wildfire danger from 1979 to 2013 by applying the US Burning Index (BI), Canadian Fire Weather Index (FWI), and Australian (or McArthur) Forest Fire Danger Index (FFDI), using National Centers for Environmental Prediction (NCEP) and European Centre for Medium-Range Weather Forecasts (ECMWF) data [51]. In addition, the Drought Code (DC), part of the FWI system, can also be applied, based on an approach that balances daily precipitation and evaporation [52]. However, it is important to emphasize that many indices have limitations regarding spatial accuracy and the holistic integration of non-meteorological covariates, which restricts their applicability [53]. One of the primary methodologies employed in a number of studies is the point process approach, which treats wildfires as discrete events in the spatiotemporal domain in order to assess the impact of climatic and ecological covariates [10]. More recent studies have linked spatial statistics and machine learning components. Techniques such as Bayesian network and log-Gaussian Cox process models are applied, which allow the capture of spatiotemporal aggregation structures through random effects [54,55,56,57,58]. In addition, researchers have also utilized Gibbs point process models, such as the Geyer saturation process and the Strauss hardcore process, to examine multiscale clustering and inhibition in fire phenomena [58]. Furthermore, the application of multi-criteria decision analysis, particularly the analytic hierarchy process, has been widespread in wildfire risk assessment studies, including those evaluating wildfire susceptibility in northern Türkiye [59] and the prioritization of factors contributing to wildfire risk in Brazil [32].
In recent years, machine-learning-based approaches to studying wildfires have also gained significant support in the development of predictive models. Systematic reviews indicate that machine learning now accounts for approximately 25% of all methodologies used in wildfire risk assessments, with the Random Forest (RF) algorithm emerging as the dominant technique due to its accuracy [60,61]. In addition, support vector machines and artificial neural networks are often used due to their robustness in handling high-dimensional datasets (such as topography, meteorological parameters, and vegetation), while convolutional neural networks enable automated processing of visual data from ground-based sensors, drones, and satellites [61]. The widespread application of machine learning and deep learning in wildfire research is also confirmed by numerous studies. For instance, Symeonidis et al. (2025) used four machine learning models (Extreme Gradient Boosting (XGBoost), Gradient Boosting Machines (GBM), Light Gradient Boosting Machine (LightGBM), and Categorical Boosting (CatBoost)) to create wildfire susceptibility maps in Greece, analyzing data from wildfires that occurred between 2000 and 2024 [37]. In Serbia, recent studies have utilized three deep and machine learning models (Deep Neural Network (DNN), Kolmogorov–Arnold Networks (KANs), and XGBoost) for spatial prediction of wildfire vulnerability [39] and two models (Random Forest (RF) and Logistic Regression (LR)) for mapping the probability of wildfire occurrence in Eastern Serbia [62]. Durlević et al. (2026) applied two machine learning models (RF and XGBoost) and two deep learning models (DNN and KAN) with Sentinel-2 and VIIRS images, in combination with GIS, to distinguish countries in Southeast Europe based on their degree of susceptibility to wildfires [38]. Additionally, Safariallahkheili et al. (2025) developed a web-based geospatial artificial intelligence system (GeoXAI) to investigate wildfire susceptibility in Berlin and Brandenburg in Germany, by combining ecological, topographic, and meteorological features derived from high-resolution geospatial data to train an RF model [63]. In order to study wildfires in Iran, Bahadori et al. (2023) generated wildfire susceptibility maps using deep learning (Recurrent Neural Network and Long Short-Term Memory (LSTM)) with MODIS and Landsat-8 satellite images for areas in western Iran affected by fires in 2021 [64]. In addition. Noroozi et al. (2024) identified wildfire-prone areas in the Firuzabad region of Fars province using Bayesian and RF methodologies [31]. Akıncı et al. (2024) conducted a comprehensive evaluation of six different machine learning algorithms (K-Nearest Neighbors (KNN), Support Vector Machines (SVM), tree-based Conditional Inference Trees (CTREE), RF, GBM, and XGBoost) to map wildfire susceptibility in Antalya Province in Türkiye [65]. Jamshed et al. (2022) used an LSTM model to make weekly wildfire forecasts and calculate the associated fire area in hectares in Pakistan using historical wildfire data provided by Global Forest Watch [66]. Larsen et al. (2021) implemented a Fully Convolutional Network (FCN) for real-time fire smoke detection in satellite imagery for Australia, which is important for health risk assessment [67]. In North America, Moghim and Mehrabi (2024) used two machine learning algorithms (LR and RF) to predict wildfires in the United States and Canada in 2024 [68]. Lastly, in Africa, Seddouki et al. (2023) compared three machine learning algorithms (XGBoost, RF, and SVM) to predict wildfire susceptibility in Tetouan Province in northern Morocco [69].
According to Lu et al. (2023), there were a total of 1317 wildfires recorded in China between 2011 and 2020, which burned a total of 155,700 ha of forest and caused 543 fatalities [70]. China’s forest resources are primarily located in the northeast (the Daxinganling, Xiaoxinganling, and Changbai mountain ranges), southwest (Hengduan range), and the great bend of the Yarlung Zangbo River. These regions are significant areas for wildfire risk research due to the widespread distribution of coniferous and broadleaf mixed forests, as well as subtropical evergreen broadleaf forests [70]. In recent years, studies for predicting wildfires in China have focused on employing machine learning in creating probability models. Thus, Shao et al. (2022) analyzed 96,594 wildfire samples collected from 2001 to 2019 to produce wildfire zonation maps that clearly identified monthly trends using a Fully Convolutional Network (FCN) model [71]. They also applied spatial autocorrelation to examine the spatial aggregation of active wildfire hotspots throughout China. Chen et al. (2023) focused their research on improving wildfire probability modeling in the northwest region of Sichuan by integrating variables such as weather, fuel, topography, infrastructure, and other factors using two machine learning methods: RF and XGBoost [72]. He et al. (2024) forecasted wildfire risk in six eastern Chinese provinces by optimizing a ConvLSTM model using various data sources (satellite-monitored wildfire products, terrestrial and human activities, simulated meteorological elements, and high-resolution vegetation imagery) for the period 2012–2022 [73]. In addition, Jiang et al. (2024) developed a Convolutional Neural Network (CNN) model based on over 11,000 wildfires recorded in Guangdong Province (2011–2021), incorporating four categories of wildfire-driven factors: topography, vegetation, weather, and human activity [74].
In addition to the aforementioned studies, which mainly deal with wildfire research in specific areas of China, this study aims to provide a comprehensive survey of the entire territory of China.
The main objectives and contributions of this study are:
  • Development of a national wildfire inventory;
  • Application of machine learning, deep learning, and transformer-based models for spatial wildfire prediction;
  • Production of wildfire susceptibility maps at a 500 m spatial resolution;
  • Identification of the provinces most susceptible to wildfires across China;
  • SHAP-based analysis of predictor importance and nonlinear relationships under the pronounced geographical asymmetry between eastern and western China.
This is the first study that analyzes wildfire patterns at the national scale in China using machine learning, deep learning, and transformer-based models. Specific spatial patterns have been identified as a result of the heterogeneity of natural and anthropogenic conditions across all regions of China.

3. Materials and Methods

3.1. Study Area

The study area covers most of the territory of the People’s Republic of China, the fourth-largest country in the world by area (Figure 1). Its total surface area is 9.6 million km2, while the estimated population is approximately 1.416 billion [75]. In terms of natural conditions, China is characterized by high geodiversity and biodiversity. Elevation ranges from −154 m (Turpan Depression) to the highest point on Earth, Mount Everest, at 8848 m [76].
Climate characteristics are influenced by the distance from the sea [77]. Eastern China experiences a typical monsoon climate with high levels of precipitation, whereas the northwestern regions are dominated by an arid continental climate [78]. Hydrological conditions are unevenly distributed, although more than 50,000 rivers flow through China, the most important being the Yangtze, the Yellow River, and the Pearl River [79]. In addition to geomorphological, climatic, and hydrological conditions, vegetation plays a key role in wildfire occurrence [80]. Forests, meadows, grasslands, and agricultural areas in northern, eastern, and southern China are more prone to wildfires than those in the northwestern and southwestern regions, where deserts and glaciers dominate [81].
Due to the pronounced asymmetry between eastern and western China, a number of studies have focused exclusively on the eastern part, i.e., areas east of the Hu Huanyong Line, which separates the densely populated eastern region from the less developed western region [82]. Owing to the combination of human and natural factors, areas east of the Hu Huanyong Line exhibit a higher frequency of wildfires [83].

3.2. Data Sources and Processing

For the purposes of this research, large geospatial databases from various European and global geoportals were utilized. The application of a large number of satellite sensors from different programs (Sentinel-2, Landsat 8/9, VIIRS, MODIS), along with the interpretation of their data, enables the creation of extensive geospatial databases for various geographic processes [84,85,86]. The study area polygon was derived from spatial datasets available from the United Nations Office for the Coordination of Humanitarian Affairs [87]. Disputed border areas are shown as “no data” because reliable data are unavailable for these zones due to their complex status, in accordance with the Agreement between the Government of the People’s Republic of China and the Government of the Republic of India on Border Defense Cooperation [88]. Additionally, there are no available data for certain island territories in the East China Sea and the South China Sea due to their complex status. The Copernicus Global Digital Elevation Model was used to obtain elevation data [89,90].

3.2.1. Spatial Wildfire Inventory

Spatial wildfire data for China were obtained by downloading vector datasets from the FIRMS platform [91]. The analysis included samples from two satellite sensors: the Moderate Resolution Imaging Spectroradiometer (MODIS) and the Visible Infrared Imaging Radiometer Suite (VIIRS). To avoid data duplication, MODIS wildfire samples were analyzed for 2001–2011, while VIIRS data were used for 2012–2024 (Figure 2).
To avoid false fire events and reduce weaker heat sources, only historical fire detections with a Fire Radiative Power (FRP) greater than 20 MW and classified as nominal or high confidence were retained. Although satellite sensors detect many thermal anomalies, not all detections correspond to actual wildfire events. Therefore, this filtering procedure served as an additional quality-control step, increasing the likelihood that the retained detections represented genuine wildfire occurrences. Consequently, although VIIRS provides higher spatial resolution and improved fire-detection capability compared to MODIS, many low-intensity VIIRS detections (<20 MW) were excluded during filtering, resulting in fewer VIIRS samples.
This procedure increases the reliability and accuracy of the wildfire data [92,93]. Multi-sensor data fusion (MODIS and VIIRS) was performed in QGIS v3.40.9 software, resulting in a wildfire inventory comprising 153,305 historical samples (Table 1) [94].
Although the original spatial resolutions of the datasets differ (MODIS—1 km; VIIRS—375 m), the pixels were resampled to 500 m in order to ensure methodological consistency. For this purpose, the nearest neighbor resampling method was applied, as this preserves the original binary categories without generating interpolated or artificial values [38].

3.2.2. Topographic Conditions

For the purpose of wildfire prediction, topographic variables including elevation, slope, aspect, and wind exposure were generated (Figure 3). All maps were created in QGIS, and the data were obtained from the Copernicus Global Digital Elevation Model [90]. Elevation, slope, and aspect were derived through raster-based terrain analysis, while wind exposure was calculated using a 50 km search radius, a 30° angular resolution, and a 1.5 acceleration factor to better represent terrain-driven wind effects. The initial spatial resolution of the dataset was 90 m.
For the sake of methodological consistency, the TIFF datasets were resampled to 500 m, applying different resampling approaches according to the data type. For continuous variables such as elevation and slope, the average method was applied, as it provides a representative depiction of terrain conditions at a broader spatial scale [95]. For the aspect, as a categorical variable, the mode method was used to preserve the dominant slope direction. The nearest-neighbor method was applied to wind exposure to retain the original values without interpolation and to avoid altering the physical meaning of the data [96].

3.2.3. Climate Characteristics

Based on the climatic conditions, data were collected for global horizontal irradiance, air temperature, precipitation, and wind speed. The TIFF file of global horizontal irradiance was obtained from the Global Solar Atlas geoportal for 2024 [97]. The pixels were resampled from the original 240 m resolution to 500 m using the cubic 4 × 4 kernel method. This approach enables smoother and more precise interpolation of values when changing spatial resolution, while preserving continuity and minimizing information loss in the global horizontal irradiance dataset [95].
Air temperature represents the mean annual values across the territory of China (Figure 4). Although the average annual maximum temperature reaches approximately 26 °C, daily maximum temperatures frequently exceed 30 °C, under which the ignition of combustible fuels becomes highly likely. Precipitation refers to the total annual sum of all types of precipitation, while wind speed represents the mean annual value. Data for these three factors were generated from the Climatology Lab platform for the period 2001–2024 to ensure temporal consistency with the wildfire inventory data [98,99]. The pixels were resampled from 4 km to 500 m resolution using bilinear interpolation, as these are continuous climatic variables for which it is important to preserve smooth spatial transitions without abrupt discontinuities. This method computes a weighted average of the four nearest cells, thereby reducing abrupt changes in values and ensuring an appropriate spatial distribution when adjusting the spatial resolution [100].

3.2.4. Hydrological Conditions

Hydrological conditions are determined through the geospatial analysis of three criteria: evapotranspiration, distance from water surfaces, and soil moisture (Figure 5). Data on the mean annual values of evapotranspiration and soil moisture for the period 2001–2024 were obtained from the Climatology Lab geoportal, while the distance from water surfaces was derived from dynamic land use data for 2022 [99,101].
The spatial resolution for evapotranspiration and soil moisture was resampled from 4 km to 500 m, while the original 30 m resolution for distance from water surfaces was resampled to 500 m. In both cases, bilinear interpolation was applied. Bilinear interpolation was selected due to the fact that all considered criteria represent continuous spatial variables, whose values change gradually across space rather than abruptly [100].

3.2.5. Vegetation Characteristics

Land use data for China were derived from Landsat imagery from 2022 processed in Google Earth Engine, with a spatial resolution of 30 m [101]. A total of eight land use classes were identified: agricultural plots, meadows, and pastures; forests; wet bare land; water surfaces; snow and glaciers; bare dry land; settlements, and seasonally flooded areas (Figure 6). To maintain methodological consistency, the pixels were resampled to a spatial resolution of 500 m. For this purpose, the nearest neighbor method was applied, as the data represent categorical values rather than continuous numerical variables. Unlike bilinear interpolation, which computes averaged values and is used for continuous data, the nearest neighbor method preserves each pixel’s original class value without altering or mixing it with neighboring classes [96].

3.2.6. Anthropogenic Conditions

From the anthropogenic factors relevant for wildfire prediction, distance from settlements and distance from roads were identified as being of greatest importance (Figure 7). To determine distance from settlements, an existing land use database was used, from which only built-up areas were extracted. Subsequently, rasterization and the proximity method in QGIS were applied to generate the distance map [101]. For the creation of the road network, the OpenStreetMap geospatial database was used [102]. All roads were subsequently rasterized, and the proximity method was applied to calculate the distance from roads.
Regarding spatial resolution, for distance from settlements, the pixels were resampled from 30 m to 500 m using bilinear interpolation, while the original resolution of both the vector and raster data used for distance from roads was 500 m. All criteria, together with their spatial resolutions, dataset year, and data sources, are presented in Table 2.
The selection of 14 variables was based on the availability of open-access spatial datasets and on previous studies on wildfire spatial prediction, which identified similar environmental factors as important predictors of wildfire occurrence [38,39]. In the study, all spatial data were harmonized both temporally and spatially to ensure consistency with the wildfire inventory for the period 2001–2024. QGIS v3.40.09 with GRASS was used for complete cartographic visualization and raster generation [94].

3.3. Methodology

This section describes the methodological framework used for wildfire susceptibility modeling, encompassing dataset construction, model development, validation strategy, threshold optimization, and interpretability analysis. The study integrates multiple modeling paradigms, including conventional machine learning algorithms, DNNs, and transformer-based architectures, enabling a comprehensive evaluation of predictive behavior.
Particular emphasis is placed on probability calibration, decision threshold selection, and cross-model consistency, as these factors critically influence wildfire prediction reliability. The adopted framework ensures methodological comparability across models, while preserving realistic spatial prediction conditions [103].

3.3.1. Machine Learning, Deep Learning, and Transformer Frameworks

In order to model wildfire occurrence and generate hazard prediction maps, a diverse set of machine learning paradigms was employed, encompassing ensemble tree-based algorithms, deep neural architectures, spectral learning approaches, and transformer-based models. The selection of models was designed to capture different representational biases, learning mechanisms, and nonlinear modeling capabilities. All models were trained and evaluated using an identical dataset derived from harmonized geospatial predictors, ensuring methodological consistency and enabling a fair comparative analysis [104]. The overall workflow is illustrated in Figure 8.
Extreme Gradient Boosting (XGBoost) is a gradient boosting ensemble algorithm that constructs an additive predictive model through sequentially optimized decision trees [105]. The method effectively captures nonlinear relationships and complex feature interactions, making it well suited to structured geospatial datasets [106]. In this study, XGBoost was implemented as a probabilistic classifier using a binary logistic objective function, with performance evaluated via Precision–Recall Area Under the Curve (PR-AUC), Receiver Operating Characteristic Area Under the Curve (ROC-AUC), and Logloss metrics.
Equation (1) defines the binary cross-entropy (log-loss), where N is the number of samples, yi is the true label, and pi is the predicted probability:
L = 1 N i = 1 N y i · log p i + 1 y i · l o g 1 p i .
The loss quantifies the discrepancy between predicted probabilities and actual outcomes and is minimized during model training. Tree construction employed a histogram-based method with a maximum bin count of 256 to ensure computational efficiency. Class imbalance was addressed through automatic class weighting based on the training sample distribution. Model complexity was controlled using parameters governing maximum tree depth and minimum child weight, while stochastic regularization was introduced via row and column subsampling. The learning process utilized a small learning rate with an adaptive schedule. Training was conducted for up to 6000 boosting rounds with validation-based early stopping after 200 rounds. To improve probability calibration, Platt scaling was employed, while the final decision threshold was chosen according to optimization results on the validation set.
Random Forest (RF) is an ensemble-based learning algorithm that builds multiple decision trees and integrates their outputs to generate the final classification [107]. The method is well suited to structured tabular data, particularly when complex nonlinear relationships exist between environmental predictors [108]. In this study, the model was trained using a staged growth strategy to analyze performance stability, whereby the number of trees was increased progressively from 50 to 1200. Key hyperparameters included a maximum tree depth of 20, which limits model complexity, and a minimum leaf size of five samples to improve generalization. Feature subsampling was controlled by setting the fraction of features per split to 0.5, promoting tree diversity. Class imbalance was addressed through balanced class weights derived from the training data. Bootstrap sampling was enabled, allowing the computation of out-of-bag (OOB) estimates used as an internal validation proxy. The warm-start mechanism allowed incremental forest expansion without retraining from scratch. Model predictions were interpreted probabilistically, with an initial classification threshold of 0.5. This configuration provides a balance between predictive stability, variance reduction, and overfitting control [38].
The Deep Neural Network (DNN) was implemented as a residual multilayer perceptron for probabilistic wildfire susceptibility prediction [39]. The architecture consists of stacked residual blocks with progressively decreasing widths (1024–512–256–128), enabling stable gradient flow and hierarchical feature learning. Training was performed using the AdamW optimizer, with the learning rate set to 1 × 10−3 and weight decay to 1 × 10−4. Learning dynamics were controlled using a OneCycleLR scheduler over a training budget of 180 epochs. Training was performed with a batch size of 4096 and validation-based early stopping with a patience of 10 epochs. Class imbalance was handled through balanced class weights incorporated into the Binary Cross-Entropy with Logits (BCEWithLogits) loss, defined in Equation (2):
B C E x , y = max x , 0 x · y + log 1 + e x ,
where x denotes the raw model output (logit) and y ∈ {0,1} represents the ground-truth label. This formulation combines the sigmoid activation and cross-entropy loss into a single numerically stable expression, ensuring stable optimization even for extreme probability values.
Mild label smoothing (0.08) was applied to reduce prediction overconfidence. To improve computational efficiency, automatic mixed precision (AMP) was used, while stochastic weight averaging (SWA) was introduced from epoch 90 onward to enhance the model’s generalization ability. The final classification threshold was optimized on the validation set, and probability calibration was performed using temperature scaling [109].
The Kolmogorov–Arnold Network (KAN) was implemented as a residual multilayer perceptron for probabilistic wildfire susceptibility prediction [110]. The architecture consists of stacked residual blocks with layer widths of 512, 512, 256, 128, and 64 neurons, incorporating batch normalization, SiLU activations, and dropout regularization (p = 0.20). Input features were normalized using StandardScaler to stabilize optimization dynamics. The model was trained using the AdamW optimizer with an initial learning rate of 3 × 10−3 and weight decay of 1 × 10−4. Learning dynamics were governed by a CosineAnnealingLR scheduler over a maximum of 100 epochs. Training employed large batch sizes (4096) and balanced class weights within the BCEWithLogits loss to address class imbalance. Early stopping with a patience of 40 epochs was applied to prevent overfitting. To enhance generalization stability, exponential moving average (EMA) weight smoothing was used, followed by validation-based threshold optimization and probability calibration via temperature scaling [111].
The Fourier MLP (Spectral Neural Model) model was implemented as a spectral neural architecture designed for probabilistic wildfire susceptibility prediction. Although Fourier MLP is not yet widely used in geoinformatics, it was included as an exploratory spectral representation approach to assess whether deterministic Fourier feature mapping can improve the representation of continuous environmental gradients. The model applies deterministic Fourier feature mapping exclusively to continuous predictors, while one-hot encoded variables are passed directly to the network [112]. The model architecture includes three fully connected hidden layers of 256 neurons each, combined with ReLU activations and dropout regularization (p = 0.15). The Fourier mapping employed six frequency components per continuous feature, enabling enhanced representation of smooth and periodic patterns. Training was carried out with the AdamW optimizer, using a learning rate of 2 × 10−4 and weight decay set to 1 × 10−4. The process lasted up to 80 epochs, while early stopping based on validation loss was applied with a patience of 10 epochs. Class imbalance was addressed through a dynamically computed positive class weight within the BCEWithLogits loss. The final classification threshold was determined via validation-based F1 score optimization, followed by probability calibration using temperature scaling.
The Feature Tokenizer Transformer (FT Transformer) model was implemented as a transformer-based neural architecture for probabilistic wildfire susceptibility prediction. The model employs a feature tokenizer layer that converts each input predictor into a learnable embedding, enabling attention-based modeling of feature interactions [113].
The architecture was configured with an embedding dimension of 64, four attention heads, and three stacked transformer encoder layers incorporating GELU activations and dropout regularization (p = 0.1). Optimization relied on the AdamW optimizer with a learning rate of 2 × 10−4 and a weight decay of 1 × 10−4. The model was trained for up to 100 epochs using batches of 512 samples, with validation-based early stopping implemented using a patience of 10 epochs. Class imbalance was addressed through dynamically computed positive class weighting within the BCEWithLogits loss function.
Model outputs were interpreted probabilistically, with the final classification threshold determined via validation-based F1 score optimization. Feature importance was approximated using tokenizer weight magnitudes, followed by probability calibration using temperature scaling [114].
Given the heterogeneity of the evaluated models, hyperparameter selection and model configuration were performed using a validation-based strategy rather than a single uniform tuning procedure. For tree-based models, complexity-related parameters such as tree depth, number of estimators, subsampling, and class weighting were controlled to reduce overfitting and improve generalization. For neural, spectral, and transformer-based models, compact architectures were adopted and regularized using dropout, weight decay, and early stopping, followed by probability calibration. The advanced neural architectures, including Fourier MLP, KAN, and FT Transformer, were included to examine alternative representational mechanisms rather than to assume superior performance over conventional ensemble models. A summary of the main hyperparameters, validation-based tuning choices, and regularization mechanisms is provided in Table 3.

3.3.2. Dataset Construction

The modeling dataset was constructed by integrating wildfire inventory data with environmental predictor variables derived from raster layers. Predictor values were extracted at sampled spatial locations to create a structured tabular dataset suitable for machine learning analysis. The final modeling dataset contained 38,432 spatial sample points, including 15,373 fire-presence samples and 23,059 background/no-recorded-fire samples. These samples were not generated for every 500 m raster cell. Instead, fire-presence locations were combined with background/no-recorded-fire locations, and predictor values were extracted from the corresponding 500 m raster cells at each sample location. Thus, each row of the dataset represents one spatial sample location with an associated wildfire label and a vector of environmental attributes. The same sample set was used for all models to ensure direct comparability.
To ensure stable model training and reduce bias toward the majority class, a controlled class sampling strategy was adopted. Fire-presence samples were combined with a controlled number of background/no-recorded-fire samples, enabling the models to learn discriminative patterns rather than prevalence-driven decision rules. Samples representing the background/no-recorded-fire class were taken from spatial locations without documented wildfire occurrence during the 2001–2024 study period, thereby providing representative coverage of environmental conditions [39]. Accordingly, the response variable should be interpreted as a long-term spatial occurrence indicator rather than as a temporally explicit fire/non-fire state.
All predictor variables were spatially aligned and resampled to a common reference grid prior to extraction. This preprocessing step guarantees predictor consistency and prevents geometric misalignment effects. The output wildfire susceptibility maps were generated on the same common reference grid used for the predictor raster layers. The spatial resolution of this grid was 500 m × 500 m, meaning that each output cell represents an area of 0.25 km2. No spatial interpolation method was used to generate the susceptibility maps. Instead, the trained models were applied directly to each valid raster cell of the aligned predictor stack, producing one predicted susceptibility value per cell. Continuous variables were normalized where required by model architecture, while categorical predictors were encoded using appropriate numerical representations. This dataset construction framework isolates class discrimination learning from spatial event rarity, ensuring methodological stability during model training.

3.3.3. Validation Procedure

Model validation was conducted using a structured data partitioning strategy designed to ensure statistical reliability and prevent information leakage.
The complete dataset was partitioned using a spatial block-based strategy into three non-overlapping subsets designated for training (80%), validation (10%), and testing (10%). Instead of assigning individual pixels randomly, spatially contiguous blocks were assigned exclusively to one of the three subsets. This procedure reduces the risk of spatial data leakage caused by spatial autocorrelation among neighboring raster cells. The resulting spatial partition is illustrated in Figure 9, which shows the block-level allocation of samples across the study area and confirms the absence of overlap between the training, validation, and test subsets.
The study area was divided into a 25 × 25 spatial grid, and each occupied block was assigned exclusively to the training, validation, or test subset. The final split contained 370 occupied blocks, including 296 training blocks (30,831 samples), 37 validation blocks (3749 samples), and 37 test blocks (3852 samples). This spatial partitioning reduces data leakage by preventing neighboring samples from being shared across subsets.
To maintain methodological consistency across modeling paradigms, identical data splits were applied to all evaluated algorithms. The training subset was exclusively used for parameter learning, while the validation subset supported threshold optimization and calibration analysis. The test subset remained strictly isolated and was only used for final performance assessment [115]. Given the inherent class imbalance characteristic of wildfire occurrence data, dataset construction incorporated a controlled sampling strategy. A controlled ratio of fire-presence and background/no-recorded-fire samples was used to stabilize the learning process and mitigate classification bias toward the dominant class. This controlled sampling procedure improves model sensitivity to wildfire-related patterns without altering the spatial predictor distributions.
Importantly, performance evaluation was undertaken from two complementary perspectives: classification metrics and confusion matrices were derived from the spatially independent block-based test subset to examine model discrimination behavior, while ROC and precision–recall analyses were used to assess threshold-independent ranking performance. This dual evaluation framework provides both algorithmic comparability and realistic performance interpretation [116]. The adopted validation strategy ensures that model performance reflects spatial generalization capability rather than dataset artifacts. By separating training, threshold optimization, and testing stages, and by preventing spatial blocks from being shared between subsets, the procedure minimizes overfitting risks and enables consistent cross-model comparison.

4. Results and Discussion

This section provides an evaluation of model performance and a comparative analysis of the applied machine learning, deep learning, and transformer-based approaches for wildfire susceptibility prediction. The results are organized to provide a structured examination of model behavior, beginning with threshold-free performance evaluation of the model probability maps under real spatial distribution conditions, followed by threshold optimization, error analysis, predictor importance assessment, and cross-model consistency evaluation.
Model performance is analyzed from both probabilistic and classification perspectives to ensure a comprehensive understanding of predictive capability. Special attention is given to the effects of class imbalance, probability calibration, spatial generalization, and decision threshold selection, which are critical factors in wildfire susceptibility modeling applications.
In addition, this section includes a spatial modeling analysis of wildfire susceptibility, enabling the visualization of spatial distribution patterns across the study area. Based on the modeled probabilities, the most susceptible administrative units within the territory of China are identified and ranked according to their wildfire susceptibility levels.

4.1. Classification Performance

Model classification performance was evaluated using conventional binary classification metrics, namely accuracy, precision, recall, F1 score, ROC-AUC, and PR-AUC. These indicators offer complementary perspectives on predictive behavior, especially in the presence of class imbalance and spatial autocorrelation, which are characteristic of wildfire datasets [117].
Figure 10 illustrates the ROC curves obtained from the threshold-free evaluation of the model probability maps against the wildfire inventory raster under real spatial distribution conditions. The results indicate that all models retained meaningful discriminative capability under the more conservative validation setting. Among the evaluated approaches, the tree-based ensemble models, particularly RF and XGBoost, achieved the strongest ROC-AUC performance, while Fourier MLP also showed competitive performance, and the DNN, FT Transformer, and KAN models exhibited somewhat lower but still informative discrimination ability.
Precision–recall (PR-AUC) analysis, shown in Figure 11, provides additional insight into model effectiveness under class-imbalanced conditions. Compared with ROC-AUC, PR-AUC is more sensitive to the proportion of wildfire samples and therefore provides a stricter assessment of model behavior for the fire class. The results confirm that the models maintain useful precision–recall trade-offs under the original highly imbalanced spatial distribution, although the obtained values are more conservative than those expected from a random pixel-level split.
Table 4 presents the comparative threshold-free performance of the evaluated model probability maps under real spatial distribution conditions. The evaluation was performed by comparing the predicted probability rasters against the wildfire inventory raster. The results indicate that all models retained meaningful discriminative capability, with ROC-AUC values ranging from 0.871 to 0.910. RF achieved the highest ROC-AUC, followed by XGBoost and Fourier MLP. In contrast, PR-AUC values were substantially lower because they were computed under the original highly imbalanced spatial distribution, where wildfire pixels represent only a very small fraction of the full raster domain.
Overall, the results indicate that all evaluated models demonstrate stable ranking and discrimination capability under real spatial distribution conditions, while the low PR-AUC values reflect the extreme rarity of wildfire pixels in the full raster domain rather than poor discrimination alone.

4.2. Threshold Optimization

In addition to probabilistic predictions, classification performance is influenced by the selection of an appropriate decision threshold. Since wildfire datasets are characterized by class imbalance and spatial heterogeneity, the conventional threshold value of 0.5 does not necessarily provide optimal classification behavior. To address this issue, threshold optimization was performed on the validation subset by evaluating model outputs across a range of threshold values. The optimal threshold was selected using the F1 score criterion, which balances precision and recall and is commonly applied in imbalanced classification problems [118]. The F1 score used for threshold optimization is defined in Equation (3):
F 1 = 2 T P 2 T P   +   F P   +   F N ,
where TP, FP, and FN denote true positives, false positives, and false negatives, respectively. The F1 score represents the harmonic mean of precision and recall, providing a balanced measure of classification performance under imbalanced conditions.
The final class label ŷi is obtained by applying a decision threshold t to the predicted probability pi, as defined in Equation (4):
u ^ i = 1 ,   p t 0 ,   p < t .
The thresholds reported in Table 5 were used exclusively to derive binary predictions for the confusion matrix analysis. They were not used to define the final wildfire susceptibility classes, which were based on calibrated probability outputs.
It should be emphasized that the optimized thresholds were used only for comparative binary classification and confusion matrix analysis. They were not used as direct operational cut-off values for defining the final wildfire susceptibility classes across the full spatial domain. The final susceptibility maps were interpreted as calibrated probability surfaces and subsequently classified into susceptibility categories for spatial analysis.
The results confirmed that optimal decision thresholds are model-dependent and reflect differences in probability calibration and score distribution across algorithms. Therefore, identical probability values should not be interpreted equivalently across different modeling paradigms. This confirms that threshold optimization is useful for fair model comparison, while final wildfire susceptibility mapping should rely primarily on calibrated probability outputs rather than binary thresholded predictions.

4.3. Confusion Matrix Analysis

To provide a detailed assessment of classification behavior, model performance was further examined using confusion matrices. It is important to emphasize that evaluation was not conducted on the complete spatial dataset, but on the curated machine learning dataset constructed for model training and validation [119]. The dataset was generated through a controlled sampling procedure designed to ensure a balanced representation of classes. Specifically, an approximately equal number of samples corresponding to wildfire presence (Fire) and wildfire absence (No Fire) were selected. This balancing strategy was intentionally applied to mitigate the severe class imbalance inherent to wildfire occurrence data, where non-fire pixels dominate the spatial domain.
Consequently, the confusion matrices reflect model behavior under balanced classification conditions rather than real-world fire frequency. This evaluation perspective isolates class discrimination capability from spatial event rarity. Under such conditions, performance metrics primarily characterize class separation behavior rather than event prevalence. Across models, results indicate consistent patterns of predictive behavior. Most algorithms demonstrate high recall for the Fire class, indicating strong sensitivity to wildfire-prone conditions. This behavior is to be expected, given the threshold optimization procedure, which prioritizes detection performance and F1 score maximization. Tree-based ensemble models, particularly XGBoost and RF, exhibit a favorable balance between false positives and false negatives. Deep learning models display slightly higher variance, while Fourier MLP and FT Transformer models demonstrate competitive detection characteristics with distinct error distributions.
Overall, confusion matrix analysis confirms that the models achieve stable class separation, with errors primarily arising from transitional environmental gradients where environmental predictors exhibit overlapping characteristics between fire and non-fire samples [39]. Figure 12 presents the confusion matrices obtained using the optimized decision thresholds.
The observed misclassification patterns motivate further examination of predictor contributions and model decision mechanisms. To better understand the drivers of model behavior, feature importance analysis was conducted.

4.4. Spatial Modeling of Wildfire Susceptibility

Spatial modeling of wildfire susceptibility is the key component of modern approaches to the risk assessment and management of natural disasters. The incorporation of GIS, remote sensing, and artificial intelligence methods facilitates the verification of spatial patterns in wildfire occurrence and the detection of highly susceptible areas at high spatial resolutions [77,120].
In this research, spatial wildfire occurrence probability was modeled using multiple machine learning, deep learning, and transformer-based models (Figure 13). The continuous susceptibility values (0–1) were classified into five classes using the equal interval method—very low (0.0–0.2), low (0.2–0.4), medium (0.4–0.6), high (0.6–0.8), and very high (0.8–1.0)—to ensure consistent interpretation and comparison of the outputs from all applied models.
The models were applied to the integrated set of geospatial predictors, which include topographic, climatic, hydrological, vegetational, and anthropogenic factors. The modeling results were presented in the form of raster maps with a spatial resolution of 500 × 500 m, which enabled detailed spatial analysis of wildfire susceptibility across the territory of China. The modeling results were divided into five types of sensitivity: very low, low, medium, high, and very high wildfire occurrence probability [115]. Such a classification approach is often applied in research on the assessment of hazard zones due to the fact that it allows for simpler clarification of the findings and their use in spatial planning. The individual models indicate a relatively stable distribution of susceptibility categories. For example, the RF model classified approximately 62.1% of the territory as an area with very low wildfire susceptibility, while around 5.8% was identified as an area with very high risk (Table 6). Similar spatial patterns were also identified in other models, including XGBoost, DNN, and the transformer architecture.
The variations between the models are primarily related to their different abilities regarding the modeling of complex nonlinear relationships between predictor variables. Algorithms like RF and XGBoost proved to be highly efficient in the analysis of structured geospatial data, due to their ability to detect complex interactions between factors that affect wildfire occurrence [105]. On the other hand, DNNs and transformers enabled the identification of latent patterns in the data but are also able to model complex relationships between climatic, geomorphological, and anthropogenic factors [120].
Ensemble modeling, which combines the predictions of all the models, was used to lower the variability of individual models and produce a more stable evaluation of wildfire susceptibility. Equation (5) shows that the ensemble probability is computed as the arithmetic mean of the predicted probabilities produced by the individual models:
p e n s = 1 M m = 1 M p m   ,
where M denotes the number of models and pm denotes the predicted probability from model m.
The ensemble model was also evaluated on the spatially independent test subset used for model validation. For each test sample, the ensemble probability was calculated as the arithmetic mean of the calibrated probabilities produced by the six individual models. The resulting ensemble predictions were then assessed using the same performance metrics as the individual models, including ROC-AUC and PR-AUC. This procedure ensured that the final ensemble susceptibility map was supported by an independent spatial generalization assessment rather than being interpreted only as a post-processing product.
Ensemble modeling represents one of the most reliable strategies in modern geospatial risk analysis, since it enables a combination of the advantages of various algorithms and reduces the impact of single-model errors [118]. According to the ensemble model results, around 56.8% of the territory of China was classified as an area with very low wildfire probability, whereas 14.4% belonged to the category of low probability. Approximately 10.7% of the territory was determined to be of medium susceptibility, while areas of high and very high probability together make up around 18% of the total land area (Figure 14).
The spatial distribution of the zones with increased wildfire susceptibility shows noticeable regional differentiation. China’s eastern and southern regions were found to have the highest concentrations of high-risk zones. These areas include large population densities, ideal climatic conditions for vegetation growth, and intense human activity, all of which raise the risk of wildfire ignition [80,83]. Conversely, because of the predominance of desert and high mountain ecosystems with little vegetation cover, the western and northwestern regions of China, such as Xinjiang and Tibet, exhibit lower susceptibility (Table 7). A more detailed analysis was carried out on the level of 34 administrative units, enabling a more precise identification of regional susceptibility patterns.
The results show that the provinces of Fujian and Guangxi Zhuang are the most susceptible regions; the areas of high and very high wildfire susceptibility comprise 86.8% and 82.9% of their territory, respectively. A high level of susceptibility was also identified in the provinces of Guangdong, Jiangxi, and Heilongjiang, confirming that regions with developed forest ecosystems and intensive human activity are especially susceptible to wildfire occurrence.
The interpretation of modeling results was additionally improved using the methods of interpretable artificial intelligence, such as SHAP (Shapley additive explanations), which enabled a quantification of the contribution of individual predictor variables to the modeling predictions [109]. To obtain an objective ranking of predictor importance, SHAP values were converted into mean absolute SHAP values for each variable across all samples. Predictors were then ranked in descending order according to this mean absolute contribution, rather than by visual inspection of the SHAP summary plots. The feature importance analysis showed that soil moisture was the highest-ranked predictor, followed by elevation, slope, global horizontal irradiance, air temperature, and wind speed. These variables therefore represented the dominant model-based contributors to wildfire susceptibility predictions. The SHAP results were interpreted as model-based feature contribution patterns rather than direct causal effects. Higher mean absolute SHAP importance indicates that a predictor contributed more strongly to model predictions, while positive and negative SHAP values were interpreted cautiously because some predictors showed mixed contributions across samples. These results suggest that moisture availability, terrain conditions, solar radiation, temperature, and wind-related factors played an important role in shaping the predicted susceptibility patterns, which is consistent with previous wildfire studies across different ecosystems [77]. A high degree of spatial agreement among the models indicates the robustness of the identified susceptibility patterns from a meteorological perspective. Analysis of correlations between model predictions showed high values of correlation, indicating that the various algorithms detect similar spatial trends of wildfire probability [121]. Such consistency implies that integration of multiple algorithms improves the reliability of the results and enables more stable mapping of hazard zones.

4.5. Cross-Model Prediction Consistency

Although the evaluated models differ substantially in terms of both their mathematical foundations and learning mechanisms, it is important to assess whether they produce coherent susceptibility patterns. High classification performance alone does not guarantee structural agreement between predicted probability distributions; models may achieve similar discrimination capability while exhibiting different calibration characteristics or localized prediction behavior.
To quantify model agreement, cross-model prediction consistency was evaluated using pairwise error metrics and correlation analysis [121]. Unlike classification metrics, this analysis focuses on the similarity of continuous wildfire probability estimates across the full spatial domain.
Table 8 summarizes the pairwise comparison results, including mean absolute error (MAE), root mean square error (RMSE), and Pearson correlation coefficients (PCC) computed between model predictions. The results reveal a generally high level of agreement among models, with correlation values consistently exceeding 0.88 across all pairwise combinations. The strongest consistency was observed between the RF and XGBoost models (corr = 0.9788), accompanied by the lowest MAE and RMSE values. This behavior indicates that tree-based ensemble models produce highly similar probability distributions and capture comparable susceptibility gradients. High degrees of correlation were also observed between neural network models and ensemble learners, particularly for DNN, KAN, and XGBoost combinations. The number of validation samples was N = 31,340,053 for all pairwise model comparisons.
Moderate reductions in agreement are evident for comparisons involving the Fourier MLP model. While correlations remain high, slightly elevated error metrics suggest that spectral feature mapping introduces variations in probability scaling and local sensitivity patterns. Similarly, the FT Transformer exhibits strong correlations with other models, indicating stable detection of dominant spatial trends.
Overall, cross-model consistency analysis confirms that the evaluated algorithms identify coherent large-scale wildfire susceptibility structures. Observed differences primarily reflect model-specific probability calibration characteristics as opposed to fundamental discrepancies in spatial prediction patterns. The high correlation structure indicates robust detection of underlying environmental drivers while preserving meaningful model diversity. Figure 15 presents the cross-model Spearman rank correlations, highlighting the degree of agreement among probabilistic predictions.

4.6. Environmental Analysis and Importance of Predictive Features

The significance and trends of the predictive variables were examined using SHAP analysis. Points on the right side of the SHAP plot increase susceptibility to wildfires, while those on the left side reduce the risk of occurrence. The color represents the feature value, with blue indicating low values and red indicating high values (Figure 16).
The ranking of environmental criteria was determined through a comparative analysis of the importance of predictive factors across all six models. The factor that most strongly influences wildfire susceptibility was soil moisture (Table 9). The drier the soil, the higher the risk of fire, as combustible material becomes more flammable [122]. However, in China, a positive correlation between soil moisture and wildfire probability was observed, which could be explained by the nonlinear relationships between dryness, fuel availability, and anthropogenic influences. Although desert areas are characterized by extremely low soil moisture, the absence of human activity and extremely sparse vegetation limit the presence of combustible biomass, thereby reducing the actual probability of fire occurrence. Therefore, at the national level, high wildfire susceptibility occurs in zones with moderate soil moisture, where sufficient vegetation is present and, under drought conditions, becomes highly flammable.
Other factors that are highly influential include topographic conditions, namely elevation and slope. Elevation and slope showed mixed SHAP contributions, indicating that their effects were context-dependent rather than strictly negative or positive. The synergistic effects of climatic conditions, spatial vegetation distribution, and anthropogenic influence could explain this trend. Lowland areas are characterized by higher air temperatures, which lead to faster drying of fuel material and increased flammability [123]. At the same time, flat terrains tend to have greater levels of continuous vegetation cover, enabling more efficient horizontal fire spread. In addition, these zones are spatially associated with more intensive human activities, thereby increasing the probability of fire ignition. Following geomorphological factors, significant criteria include climatic conditions: global horizontal irradiance (GHI), air temperature, and wind speed. Higher values of GHI and air temperature increase surface and vegetation heating, accelerate evaporation processes, and reduce soil moisture content [124,125]. Although global horizontal irradiance generally increases vegetation flammability, the highest GHI values in China occur in high-altitude regions (the Himalayas), where wildfire frequency is very low. This indicates a nonlinear relationship between GHI and wildfire probability, where the effect of radiation depends on biomass availability and specific geoecological conditions.
A similar phenomenon is observed with air temperature. An exception appears when desert regions are included in the analysis, as, simultaneously, they are the warmest zones but exhibit very low wildfire frequency. Therefore, the relationship between temperature and wildfire susceptibility is not strictly linear. Extremely high temperatures in the absence of available fuel do not increase fire risk. The positive correlation between wind speed and wildfire probability indicates that higher wind speeds increase the likelihood of fire occurrence and spread. This is due to the fact that stronger winds increase the oxygen supply to the combustion zone, intensifying the oxidation process and raising the flame temperature [126].
From a land use perspective, agricultural plots, meadows, and pastures also have notable significance. Agricultural areas represent zones of increased anthropogenic activity, thereby raising the probability of fire ignition [49]. After harvest, crop residues and dry straw form a highly flammable fuel layer. During dry periods, meadows and pastures are characterized by grassy biomass with low moisture content, and a relatively low ignition temperature [127]. In the arid and semi-arid regions of northern and western China, grassland ecosystems can contribute significantly to overall fire frequency. However, the influence of these land use types depends on the regional climatic context. In areas with sufficient moisture and intensive land management, the risk may be reduced due to vegetation fragmentation and the presence of firebreaks. Therefore, their effect on wildfire probability is not universal, but rather results from the interaction between fuel structure, seasonality, and anthropogenic factors.
Precipitation trends in relation to wildfire probability generally showed a negative correlation, whereby increased rainfall contributes to a lower likelihood of fire occurrence. Higher rainfall amounts increase soil and vegetation moisture, reduce fuel flammability, and prolong drying time. However, in the context of China, a nonlinear relationship arises due to extensive desert areas in the northwestern part of the country. These regions are characterized by extremely low levels of precipitation and minimal vegetation, which limit fuel availability and result in low fire frequency in spite of the arid conditions.
Distance from water surfaces also evidences a nonlinear relationship with wildfire probability, reflecting the pronounced geoecological asymmetry between the western and eastern parts of China. Desert regions are characterized by large distances from water bodies and extremely low fire frequency. Although these areas are far from rivers and lakes, the lack of vegetation limits fuel availability, thereby reducing the likelihood of fire. The effect of distance from water can therefore be seen to be dependent on its interaction with the spatial distribution of biomass. The negative relationship between distance from roads and wildfire probability suggests that areas closer to roads are generally more susceptible to wildfire. However, this pattern primarily reflects roads intersecting or bordering forested and natural landscapes rather than densely urbanized areas, where limited fuel availability reduces wildfire potential despite a high road density.
Criteria with very low importance included evapotranspiration, distance from settlements, and (presence of) forests, bare dry land (deserts), and wet barren land. Evapotranspiration shows a positive correlation with wildfire probability, indicating that higher values of this parameter are associated with a greater likelihood of fire occurrence. This relationship is influenced by the spatial distribution of desert areas in northwestern China, where evapotranspiration values are lowest due to extreme climatic and biogeographical conditions, and fire frequency is also low because of the absence of vegetation and human activity.
Shorter distances from settlements are associated with a higher probability of fire occurrence due to stronger anthropogenic influence and a greater number of potential ignition sources. A higher proportion of forest ecosystems increases risk due to greater quantities of combustible biomass [128]. In contrast, desert areas, despite their pronounced aridity, are characterized by low wildfire risk due to the limited availability of fuel and minimal human activity. Wet barren lands also show reduced susceptibility to fires, as the combination of higher moisture levels and sparse vegetation limits flammability and the potential for spatial fire spread.

4.7. Comparison with Wildfire Studies in China and Model Performance Evaluation

Based on the created wildfire inventory (153,305 samples), it was determined that 54.2% of fire events occurred in forests, 35% in meadows and pastures, and 5.2% in the vicinity of settlements. This distribution clearly indicates that a substantial proportion of fires occur outside forest ecosystems, particularly in rangelands and areas influenced by anthropogenic activities. Therefore, wildfire research should not be limited solely to forest fire prediction, but should also address the broader phenomenon of wildfires, which includes all types of vegetation and land cover where fires may occur and spread. Machine learning models enable the enhancement of forest ecosystem management and wildfire prevention strategies, particularly in regions where climatic and anthropogenic factors are strongly associated with wildfire dynamics [120].
Numerous studies addressing forest fires and wildfires have been conducted at the local level across the territory of the People’s Republic of China. Zhang et al. (2019) applied Convolutional Neural Networks (CNN) in order to predict forest fires in Yunnan Province for the period 2002–2010. By processing 14 environmental and anthropogenic factors, the results indicated that the CNN model achieved an AUC value of 86% [129]. For the same province, Zhang et al. (2025b) applied six machine learning models to identify the optimal model and analyze the dominant factors influencing fire occurrence across different seasons [120]. The LightGBM model achieved the highest AUC (89.5%), thus confirming the strong predictive performance of gradient-boosting ensemble methods.
Li et al. (2024b) investigated forest fire prediction in the eastern part of China by combining kernel density analysis, global and local spatial autocorrelation, deep learning techniques, the standard deviation ellipse method, and Geographic Information Systems (GIS) [116]. The results showed that high-fire-frequency zones are primarily concentrated in Guangdong, Fujian, and Zhejiang provinces, with the model demonstrating strong predictive performance (88.2%). Pang et al. (2022) analyzed forest fire prediction at the national level using several machine learning models, including Artificial Neural Networks (ANN), Radial Basis Function Networks (RBFN), Support Vector Machines (SVM), and RF [130]. The results indicated that AUC values ranged from 84% to 96%, with the RF model achieving the best performance (accuracy of 89.2% and AUC of 0.96).
In addition to studies focusing exclusively on forest fire prediction, several investigations in China have addressed wildfires more broadly, encompassing all types of vegetation fires. Yan et al. (2026) integrated multi-sensor data using five machine learning classifiers—RF, Gradient-Boosting Decision Trees (GBDT), XGBoost, Light Gradient Boosting Machine (LightGBM), and Logistic Regression (LR)—to predict wildfires within the “Three-North Shelterbelt Program” region [131]. Model performance comparisons showed that tree-based ensemble models significantly outperform linear models, with LightGBM achieving the best results (AUC > 0.98). Jiang et al. (2024) developed a database comprising 11,507 historical wildfire events in Guangdong Province for the period 2011–2021 [74]. They applied a Convolutional Neural Network, which demonstrated superior performance (AUC = 96.2%) compared to classical machine learning models. Finally, He et al. (2024) investigated six provinces in China and improved a Convolutional Neural Networks–Long Short-Term Memory (ConvLSTM) model, achieving AUC > 80% [73]. These studies demonstrate that advanced machine learning and deep learning models can effectively capture complex relationships between environmental factors and wildfire occurrence, enabling accurate prediction across different spatial and temporal scales.
In our study, several advantages and novel contributions can be highlighted:
  • Development of a national wildfire inventory containing 153,305 filtered historical records of fire incidents;
  • Combined application of machine learning, deep learning, and transformer-based models for spatial wildfire prediction;
  • A spatial resolution of 500 m for both input parameters and final susceptibility maps, enabling adequate spatial differentiation of results from the national to the local level;
  • Identification of the most susceptible provinces across the territory of China;
  • SHAP analysis and the identification of nonlinear relationships among certain predictors, caused by the significant geographical asymmetry between eastern and western China.
In addition to its advantages, the study also has certain limitations. In this research, the spatial resolution of the input data was harmonized through resampling, as the predictor variables were derived from datasets with different native spatial resolutions. Another limitation of this study concerns the absence of seasonal wildfire prediction. While annual datasets for the predictor variables were available and incorporated into the modeling framework, consistent monthly data covering the entire study area were not available, preventing the development of season-specific wildfire prediction models.
Future studies should incorporate land use datasets spanning multiple time periods to improve the accuracy and reliability of the final results, particularly given the intensive land use changes occurring in China, such as increasing urbanization and large-scale afforestation. In addition, future research should place greater emphasis on incorporating temporal components and seasonal variability into wildfire prediction models. The use of monthly or seasonal environmental and climatic datasets would enable a more detailed analysis of intra-annual dynamics and improve the understanding of how seasonal changes influence wildfire occurrence. The integration of such temporal variability could significantly enhance the accuracy of wildfire susceptibility models, and support more effective early warning and fire risk management strategies. Future studies could perform separate wildfire susceptibility analyses for eastern and western China to better account for the pronounced geographical and environmental differences between these regions and potentially improve model performance.
Another limitation of this approach is that the dataset does not represent a fully spatiotemporal panel of fire and non-fire states. Non-fire/background samples correspond to locations with no recorded wildfire occurrence during the full study period, while locations that burned at least once are treated as fire-presence samples. Consequently, the model is intended for long-term spatial susceptibility mapping rather than short-term temporal fire forecasting. Future research could extend this framework by constructing annual or monthly fire/non-fire samples with time-varying climatic, vegetation, and human-activity predictors. In addition, future research could investigate the contribution of lagged climatic effects by incorporating the Standardized Precipitation Index (SPI) and the Standardized Precipitation Evapotranspiration Index (SPEI) to assess their potential for improving nationwide wildfire susceptibility mapping.

5. Conclusions

Due to increasingly frequent climate extremes, wildfires in China pose a major threat to the environment, socio-economic development, and biodiversity. Effective assessment and mapping of fire-prone areas are key approaches to preventive action and wildfire risk management. This study focuses on large-scale wildfire susceptibility modeling at the national level, with a spatial resolution of 500 m. Because the model is based on long-term mean climatic and environmental variables, it provides a long-term spatial assessment of wildfire susceptibility rather than an operational wildfire forecasting system. The development of the geospatial database began with the filtering and integration of satellite data from MODIS and VIIRS products, based on which a fire inventory containing 153,305 wildfire samples was generated for the period 2001–2024. A total of 14 topographic, climatic, hydrological, vegetation, and anthropogenic variables were processed as spatial predictors.
Regarding the methodological framework, six artificial intelligence models were applied: RF, XGBoost, DNN, Fourier MLP, KAN, and FT Transformer. The application of multiple ML, DL, and transformer-based models with different architectures enables the capture of nonlinear relationships and higher-order interactions among wildfire-related variables. This approach improves the robustness of the modeling framework while allowing comparative performance evaluation and identification of the most efficient model.
The results indicate that all models retained meaningful ranking capability under real spatial distribution conditions, with ROC-AUC values ranging from 0.871 to 0.910. RF achieved the strongest performance (ROC-AUC = 0.910), followed by XGBoost (ROC-AUC = 0.902) and Fourier MLP (ROC-AUC = 0.891). At the national scale, 10.6% of China’s territory was identified as highly susceptible and 7.4% as very highly susceptible to wildfires, with the most vulnerable provinces being Fujian (86.8%) and Guangxi Zhuang (82.9%). Due to the pronounced geographical and environmental asymmetry between eastern and western China, it was necessary to determine the importance and trends of predictive variables. SHAP analysis revealed that the three most influential factors affecting wildfire occurrence are soil moisture, elevation, and terrain slope.
This study provides relevant insights, as the results obtained can be applied in emergency situations from the perspective of wildfire management at both the national and local levels. The susceptibility maps can help responsible authorities better understand the spatial characteristics and risk trends of wildfires. Furthermore, the final outputs can contribute to the development of more efficient targeted wildfire prevention and control strategies.

Author Contributions

Conceptualization, U.D.; methodology, V.I.; software, V.I.; validation, V.I. and U.D.; formal analysis, M.M.R. and J.M.J.; investigation, M.M. and A.M.P.; resources, M.D.P.; data curation, M.M.R. and A.M.P.; writing—original draft preparation, U.D. and V.I.; writing—review and editing, M.D.P. and E.A.; visualization, U.D.; supervision, M.M.R.; project administration, U.D.; funding acquisition, M.M.R. and U.D. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Ministry of Science, Technological Development and Innovation of the Republic of Serbia (Contract number 451-03-33/2026-03/200172 and 451-03-33/2026-03/200091). Marko Petrović acknowledges the support of RUDN University (grant number 060510-0-000) for conducting this research.

Data Availability Statement

The data sources for the wildfire inventory are presented in Table 1. The sources of input data on natural and anthropogenic conditions are presented in Table 2. The remaining datasets, due to their large file sizes, will be made available upon reasonable request to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Food and Agriculture Organization. Forest Area as a Percentage of Total Land Area. 2026. Available online: https://de-public-statsuite.fao.org/vis?fs [0]=Sustainable%20Development%20Goals%20%28SDGs%29%2C1%7CGoal%2015%20Life%20on%20Land%23SDG_G15%23%7C15.1.1%20Forest%20area%23SDG_G15_1511%23&pg=0&fc=Sustainable%20Development%20Goals%20%28SDGs%29&bp=true&snb=1&vw=ov&df[ds]=ds-release&df[id]=DF_SDG_15_1_1&df[ag]=FAO&df[vs]=1.0 (accessed on 30 March 2026).
  2. Milenković, M.; Ducić, V.; Obradović, D.; Dedić, A.; Burić, D. Climatic and anthropogenic impacts on forest fires in conditions of extreme fire danger on sandy soils. J. Geogr. Inst. Jovan Cvijic SASA 2023, 73, 155–168. [Google Scholar] [CrossRef] [Scilit]
  3. Bhatt, R.P. Achievement of SDGs globally in biodiversity conservation and reduction of greenhouse gas emissions by using green energy and maintaining forest cover. GSC Adv. Res. Rev. 2023, 17, 001–021. [Google Scholar] [CrossRef] [Scilit]
  4. Strader, S.M. Spatiotemporal changes in conterminous US wildfire exposure from 1940 to 2010. Nat. Hazards 2018, 92, 543–565. [Google Scholar] [CrossRef] [Scilit]
  5. Jones, M.W.; Abatzoglou, J.T.; Veraverbeke, S.; Andela, N.; Lasslop, G.; Forkel, M.; Smith, A.J.P.; Burton, C.; Betts, R.; van der Werf, G.R.; et al. Global and regional trends and drivers of fire under climate change. Rev. Geophys. 2022, 60, e2020RG000726. [Google Scholar] [CrossRef] [Scilit]
  6. Oom, D.; Pereira, J.M.C. Exploratory spatial data analysis of global MODIS active fire data. Int. J. Appl. Earth Obs. Geoinf. 2013, 21, 326–340. [Google Scholar] [CrossRef] [Scilit]
  7. Gómez-González, J.L.; Cantizano, A.; Caro-Carretero, R.; Castro, M. Leveraging national forestry data repositories to advocate wildfire modeling towards simulation-driven risk assessment. Ecol. Indic. 2024, 158, 111306. [Google Scholar] [CrossRef] [Scilit]
  8. EM-DAT. Emergency Events Database: Wildfire Data Store in the World from 2001 to 2025. Available online: https://public.emdat.be/data (accessed on 26 March 2026).
  9. Food and Agriculture Organization. Fire Management. 2025. Available online: https://www.fao.org/forestry/firemanagement/en (accessed on 30 March 2026).
  10. Global Wildfire Information System. 2026. Available online: https://gwis.jrc.ec.europa.eu/ (accessed on 30 March 2026).
  11. Fire Information for Resource Management System. 2026. Available online: https://www.earthdata.nasa.gov/data/tools/firms (accessed on 30 March 2026).
  12. San-Miguel-Ayanz, J.; Oom, D.; Artes, T.; Viegas, D.X.; Fernandes, P.; Faivre, N.; Freire, S.; Moore, P.; Rego, F.; Castellnou, M. Forest fires in Portugal in 2017. In Science for Disaster Risk Management 2020: Acting Today, Protecting Tomorrow; Casajus Valles, A., Marin Ferrer, M., Poljansek, K., Clark, I., Eds.; Publications Office of the European Union: Luxembourg, 2020. [Google Scholar] [CrossRef] [PubMed]
  13. De Rivera, Ó.R.; Espinosa, J.; Madrigal, J.; Blangiardo, M.; López-Quílez, A. Spatio-temporal marked point process model to understand forest fires in the Mediterranean Basin. J. Agric. Biol. Environ. Stat. 2025, 30, 700–729. [Google Scholar] [CrossRef] [Scilit]
  14. Sari, F. Identifying anthropogenic and natural causes of wildfires by maximum entropy method-based ignition susceptibility distribution models. J. For. Res. 2023, 34, 355–371. [Google Scholar] [CrossRef] [Scilit]
  15. Ruffault, J.; Mouillot, F. Contribution of human and biophysical factors to the spatial distribution of forest fire ignitions and large wildfires in a French Mediterranean region. Int. J. Wildland Fire 2017, 26, 498–508. [Google Scholar] [CrossRef] [Scilit]
  16. Hysa, A.; Teqja, Z. Counting fuel properties as input in the wildfire spreading capacities of vegetated surfaces: Case of Albania. Not. Bot. Horti Agrobot. Cluj-Napoca 2020, 48, 1667–1682. [Google Scholar] [CrossRef] [Scilit]
  17. Naderpour, M.; Rizeei, H.M.; Ramezani, F. Wildfire prediction: Handling uncertainties using integrated Bayesian networks and fuzzy logic. In Proceedings of the 2020 IEEE International Conference on Fuzzy Systems (FUZZ-IEEE); IEEE: Glasgow, UK, 2020; pp. 1–7. [Google Scholar] [CrossRef] [Scilit]
  18. Lamont, B.B.; He, T. The third dimension: How fire-related research can advance ecology and evolutionary biology. Ideas Ecol. Evol. 2020, 13, 25–58. [Google Scholar] [CrossRef] [Scilit]
  19. Geraskina, A.P.; Tebenkova, D.N.; Ershov, D.V.; Ruchinskaya, E.V.; Sibirtseva, N.V.; Lukina, N.V. Wildfires as a factor of loss of biodiversity and forest ecosystem functions. For. Sci. Issues 2022, 4, 82. [Google Scholar] [CrossRef] [Scilit]
  20. Nelson, A.R.; Narrowe, A.B.; Rhoades, C.C.; Fegel, T.S.; Daly, R.A.; Roth, H.K.; Chu, R.; Amundson, K.K.; Young, R.B.; Steindorff, A.S.; et al. Wildfire-dependent changes in soil microbiome diversity and function. Nat. Microbiol. 2022, 7, 1419–1430. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Lappe, B.; Vargo, J. Disruptions from Wildfire Smoke: Vulnerabilities in Local Economies and Disadvantaged Communities in the U.S.; Federal Reserve Bank of San Francisco Community Development Research Brief 2022-06; Federal Reserve Bank of San Francisco: San Francisco, CA, USA, 2022. [Google Scholar] [CrossRef] [Scilit]
  22. De Diego, J.; Fernández, M.; Rúa, A.; Kline, J.D. Examining socioeconomic factors associated with wildfire occurrence and burned area in Galicia (Spain) using spatial and temporal data. Fire Ecol. 2023, 19, 18. [Google Scholar] [CrossRef] [Scilit]
  23. Ye, X.; Ye, Y.; Huang, X.; Onega, T. Wildfires and public health: A comprehensive review of human-centric studies. GeoHealth 2026, 10, e2025GH001534. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Senande-Rivera, M.; Insúa-Costa, D.; Miguez-Macho, G. Spatial and temporal expansion of global wildland fire activity in response to climate change. Nat. Commun. 2022, 13, 1208. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Vyklyuk, Y.; Radovanović, M.M.; Stanojević, G.; Petrović, M.D.; Ćurčić, N.B.; Milenković, M.; Malinović Milićević, S.; Milovanović, B.; Yamashkin, A.A.; Milanović Pešić, A.; et al. Connection of solar activities and forest fires in 2018: Events in the USA (California), Portugal and Greece. Sustainability 2020, 12, 10261. [Google Scholar] [CrossRef] [Scilit]
  26. Xu, J.; Liang, S.; Ma, H.; Chen, Y.; Li, W.; Ma, Y.; Zhao, X.; Jiang, B.; Zhang, X.; Guan, S. Joint estimation of global daily 1 km surface radiation budget components from MODIS observations (2000–2023) using conservation-constrained deep neural networks. Remote Sens. Environ. 2026, 333, 115135. [Google Scholar] [CrossRef] [Scilit]
  27. Srivanit, M. Community risk assessment: Spatial patterns and GIS-based model for fire risk assessment—A case study of Chiang Mai Municipality. J. Archit. Plan. Res. Stud. 2011, 8, 113–126. [Google Scholar] [CrossRef] [Scilit]
  28. Abdi, O.; Kamkar, B.; Shirvani, Z.; Teixeira da Silva, J.A.; Buchroithner, M.F. Spatial-statistical analysis of factors determining forest fires: A case study from Golestan, Northeast Iran. Geomat. Nat. Hazards Risk 2018, 9, 267–280. [Google Scholar] [CrossRef] [Scilit]
  29. Hwang, S.N.; Meier, K. Associations between wildfire risk and socio-economic-demographic characteristics using GIS technology. J. Geogr. Inf. Syst. 2022, 14, 365–388. [Google Scholar] [CrossRef]
  30. Sakellariou, S.; Parisien, M.-A.; Flannigan, M.; Wang, X.; de Groot, B.; Tampekis, S.; Samara, F.; Sfougaris, A.; Christopoulou, O. Spatial planning of fire-agency stations as a function of wildfire likelihood in Thasos, Greece. Sci. Total Environ. 2020, 729, 139004. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Noroozi, F.; Ghanbarian, G.; Safaeian, R.; Pourghasemi, H.R. Forest fire mapping: A comparison between GIS-based random forest and Bayesian models. Nat. Hazards 2024, 120, 6569–6592. [Google Scholar] [CrossRef] [Scilit]
  32. Coutinho, L.M.; Marchioro, E. Forest fire risk zoning in an administrative division of Cachoeiro de Itapemirim (Brazil/ES). Cuad. Investig. Geogr. 2025, 51, 9–28. [Google Scholar] [CrossRef] [Scilit]
  33. Asori, M. Wildfire hazard and risk modelling in the Northern regions of Ghana using GIS-based multi-criteria decision making analysis. J. Environ. Earth Sci. 2020, 10, 11–28. [Google Scholar] [CrossRef] [Scilit]
  34. Tomar, J.S.; Kanga, S.; Singh, S.K. GIScience for forest fire modeling: New advances in wildland fires with management. Plant Arch. 2021, 21, 473–481. [Google Scholar] [CrossRef] [Scilit]
  35. Chen, X.-T.; Kang, S.-C.; Ji, Z.-M.; Yang, J.-H.; Hu, Y.-L.; Xu, M.; Ma, M.-M.; Zhou, Y. Divergent trajectories of global fires: Overall decline, regional intensification. Adv. Clim. Change Res. 2026, 17, 575–587. [Google Scholar] [CrossRef] [Scilit]
  36. Castellnou, M.; Guiomar, N.; Rego, F.; Fernandes, P.M. Fire growth patterns in the 2017 mega fire episode of October 15, central Portugal. In Advances in Forest Fire Research; Viegas, D.X., Ed.; ADAI & CEIF, University of Coimbra: Coimbra, Portugal, 2018; pp. 447–453. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Symeonidis, P.; Vafeiadis, T.; Ioannidis, D.; Tzovaras, D. Wildfire susceptibility mapping in Greece using ensemble machine learning. Earth 2025, 6, 75. [Google Scholar] [CrossRef] [Scilit]
  38. Durlević, U.; Ilić, V.; Aleksova, B. Wildfire probability mapping in Southeastern Europe using deep learning and machine learning models based on open satellite data. AI 2026, 7, 21. [Google Scholar] [CrossRef] [Scilit]
  39. Durlević, U.; Ilić, V.; Valjarević, A. Wildfire susceptibility mapping using deep learning and machine learning models based on multi-sensor satellite data fusion: A case study of Serbia. Fire 2025, 8, 407. [Google Scholar] [CrossRef] [Scilit]
  40. Vadrevu, K.P.; Lasko, K.; Giglio, L.; Schroeder, W.; Biswas, S.; Justice, C. Trends in vegetation fires in South and Southeast Asian countries. Sci. Rep. 2019, 9, 7422. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. De Groot, W.J.; Flannigan, M.D.; Cantin, A.S. Climate change impacts on future boreal fire regimes. For. Ecol. Manag. 2013, 294, 35–44. [Google Scholar] [CrossRef] [Scilit]
  42. Miller, E.A. A conceptual interpretation of the drought code of the Canadian forest fire weather index system. Fire 2020, 3, 23. [Google Scholar] [CrossRef] [Scilit]
  43. Mataveli, G.; Maure, L.A.; Sanchez, A.; Dutra, D.J.; de Oliveira, G.; Jones, M.W.; Amaral, C.; Artaxo, P.; Aragão, L.E.O.C. Forest degradation is undermining progress on deforestation in the Amazon. Glob. Change Biol. 2025, 31, e70209. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  44. Knorr, W.; Lehsten, V.; Arneth, A. Determinants and predictability of global wildfire emissions. Atmos. Chem. Phys. 2012, 12, 6845–6861. [Google Scholar] [CrossRef] [Scilit]
  45. Xu, R.; Ye, T.; Yue, X.; Yang, Z.; Yu, W.; Zhang, Y.; Bell, M.L.; Morawska, L.; Yu, P.; Zhang, Y.; et al. Global population exposure to landscape fire air pollution from 2000 to 2019. Nature 2023, 621, 521–529. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  46. Roberts, G.; Wooster, M.J.; Lagoudakis, E. Annual and diurnal African biomass burning temporal dynamics. Biogeosciences 2009, 6, 849–866. [Google Scholar] [CrossRef] [Scilit]
  47. Haque, K.M.S.; Uddin, M.; Ampah, J.D.; Haque, M.K.; Hossen, M.S.; Rokonuzzaman, M.; Hossain, M.Y.; Hossain, M.S.; Rahman, M.Z. Wildfires in Australia: A bibliometric analysis and a glimpse on ‘Black Summer’ (2019/2020) disaster. Environ. Sci. Pollut. Res. 2023, 30, 73061–73086. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  48. Silva, F.R.; Martínez, J.R.M.; González-Cabán, A. A methodology for determining operational priorities for prevention and suppression of wildland fires. Int. J. Wildland Fire 2014, 23, 544–554. [Google Scholar] [CrossRef] [Scilit]
  49. Nikolić, N. Assessing wildfire impact on vegetation in protected areas using the dNBR index: Insights from the designated location in Serbia. J. Geogr. Inst. Jovan Cvijic SASA 2025, 75, 453–460. [Google Scholar] [CrossRef] [Scilit]
  50. Key, C.H.; Benson, N.C. Landscape assessment (LA). In FIREMON: Fire Effects Monitoring and Inventory System; Lutes, D.C., Keane, R.E., Caratti, J.F., Key, C.H., Benson, N.C., Sutherland, S., Gangi, L.J., Eds.; U.S. Department of Agriculture, Forest Service, Rocky Mountain Research Station: Fort Collins, CO, USA, 2006; pp. LA-1–LA-55. Available online: https://research.fs.usda.gov/treesearch/24066 (accessed on 5 April 2026).
  51. Jolly, W.M.; Cochrane, M.A.; Freeborn, P.H.; Holden, Z.A.; Brown, T.J.; Williamson, G.J.; Bowman, D.M.J.S. Climate-induced variations in global wildfire danger from 1979 to 2013. Nat. Commun. 2015, 6, 7537. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  52. Di Giuseppe, F.; Vitolo, C.; Krzeminski, B.; Barnard, C.; Miguel, J.S. Drought Code—ERA-Interim. Contract 933710 Between JRC and ECMWF (Copernicus—Fire Danger Forecast Computation). 2019. Available online: https://zenodo.org/records/3250960 (accessed on 10 April 2026).
  53. Durlević, U.; Čegar, N.; Ilić, V.; Kovjanić, A. Machine learning and deep learning approaches for wildfire susceptibility prediction: A case study of the Djerdap Geopark, Serbia. Earth Syst. Environ. 2025, 10, 5999–6020. [Google Scholar] [CrossRef] [Scilit]
  54. Villar, B.J.; Rodríguez, P.P.; de Souza, A. Long temporal trend and seasonal variation analysis of forest fires in Brazilian biomes: A stochastic approach. Rev. Mex. Cienc. For. 2024, 15, 29–53. [Google Scholar] [CrossRef] [Scilit]
  55. Konurhan, Z.; Yucesan, M.; Gul, M. Investigating forest fire causes through an integrated Bayesian network and geographic information system approach. Nat. Hazards 2025, 121, 12933–12958. [Google Scholar] [CrossRef] [Scilit]
  56. Koh, J.; Pimont, F.; Dupuy, J.-L.; Opitz, T. Spatiotemporal wildfire modeling through point processes with moderate and extreme marks. Ann. Appl. Stat. 2023, 17, 560–582. [Google Scholar] [CrossRef] [Scilit]
  57. Gabriel, E.; Opitz, T.; Bonneu, F. Detecting and modeling multi-scale space-time structures: The case of wildfire occurrences. J. Soc. Fr. Stat. 2017, 158, 86–105. [Google Scholar]
  58. Raeisi, M.; Bonneu, F.; Gabriel, E. Spatio-Temporal Hybrid Strauss Hardcore Point Process and Application. HAL (Le Centre pour la Communication Scientifique Directe). 2021. Available online: https://univ-avignon.hal.science/hal-03193464v1 (accessed on 12 April 2026).
  59. Bağcı, H.R.; Cansu Kaya, C. Evaluating forest fire susceptibility levels in Şahinkaya Canyon (Northern Türkiye) using AHP. J. Geogr. Inst. Jovan Cvijic SASA 2026, 76, 173–190. [Google Scholar] [CrossRef] [Scilit]
  60. Mihajlovski, B.; Zhiyanski, M. Global forest fire assessment methods: A comparative analysis of hazard, susceptibility, and vulnerability approaches in different landscapes. Fire 2025, 8, 380. [Google Scholar] [CrossRef] [Scilit]
  61. Sobha, P.; Latifi, S. A survey of the machine learning models for forest fire prediction and detection. Int. J. Commun. Netw. Syst. Sci. 2023, 16, 131–150. [Google Scholar] [CrossRef]
  62. Milanović, S.; Marković, N.; Pamučar, D.; Gigović, L.; Kostić, P.; Milanović, S.D. Forest fire probability mapping in Eastern Serbia: Logistic regression versus random forest method. Forests 2021, 12, 5. [Google Scholar] [CrossRef] [Scilit]
  63. Safariallahkheili, Q.; Schiewe, J.; Meier, S. Post-hoc explanation of AI predictions in wildfire risk mapping through an interactive web-based GeoXAI system. KN J. Cartogr. Geogr. Inf. 2025, 75, 143–158. [Google Scholar] [CrossRef] [Scilit]
  64. Bahadori, N.; Razavi-Termeh, S.V.; Sadeghi-Niaraki, A.; Al-Kindi, K.M.; Abuhmed, T.; Nazeri, B.; Choi, S.-M. Wildfire susceptibility mapping using deep learning algorithms in two satellite imagery datasets. Forests 2023, 14, 1325. [Google Scholar] [CrossRef] [Scilit]
  65. Akıncı, H.A.; Akıncı, H.; Zeybek, M. Comparison of diverse machine learning algorithms for forest fire susceptibility mapping in Antalya, Türkiye. Adv. Space Res. 2024, 74, 647–667. [Google Scholar] [CrossRef] [Scilit]
  66. Jamshed, M.A.; Theodorou, C.; Kalsoom, T.; Anjum, N.; Abbasi, Q.H.; Rehman, M.U. Intelligent computing based forecasting of deforestation using fire alerts: A deep learning approach. Phys. Commun. 2022, 55, 101941. [Google Scholar] [CrossRef] [Scilit]
  67. Larsen, A.; Hanigan, I.; Reich, B.J.; Qin, Y.; Cope, M.; Morgan, G.; Rappold, A.G. A deep learning approach to identify smoke plumes in satellite imagery in near-real time for health risk communication. J. Expo. Sci. Environ. Epidemiol. 2021, 31, 170–176. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  68. Moghim, S.; Mehrabi, M. Wildfire assessment using machine learning algorithms in different regions. Fire Ecol. 2024, 20, 104. [Google Scholar] [CrossRef] [Scilit]
  69. Seddouki, M.; Benayad, M.; Aamir, Z.; Tahiri, M.; Maanan, M.; Rhinane, H. Using machine learning coupled with remote sensing for forest fire susceptibility mapping: Case study—Tetouan Province, Northern Morocco. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2023, XLVIII-4/W6-2022, 333–342. [Google Scholar] [CrossRef] [Scilit]
  70. Lu, Y.; Zhou, Q.; Shao, S.; Wang, W.; Dai, Y.; Wei, X. Influence and prediction of climatic factors on forest fires in China. China Saf. Sci. J. 2023, 33, 53–59. [Google Scholar] [CrossRef]
  71. Shao, Y.; Feng, Z.; Sun, L.; Yang, X.; Li, Y.; Xu, B.; Chen, Y. Mapping China’s forest fire risks with machine learning. Forests 2022, 13, 856. [Google Scholar] [CrossRef] [Scilit]
  72. Chen, R.; He, B.; Quan, X.; Lai, X.; Fan, C. Improving wildfire probability modeling by integrating dynamic-step weather variables over northwestern Sichuan, China. Int. J. Disaster Risk Sci. 2023, 14, 313–325. [Google Scholar] [CrossRef] [Scilit]
  73. He, Z.; Fan, G.; Li, Z.; Li, S.; Gao, L.; Li, X.; Zeng, C. Deep learning modeling of human activity affected wildfire risk by incorporating structural features: A case study in eastern China. Ecol. Indic. 2024, 160, 111946. [Google Scholar] [CrossRef] [Scilit]
  74. Jiang, W.; Qiao, Y.; Zheng, X.; Zhou, J.; Jiang, J.; Meng, Q.; Su, G.; Zhong, S.; Wang, F. Wildfire risk assessment using deep learning in Guangdong Province, China. Int. J. Appl. Earth Obs. Geoinf. 2024, 128, 103750. [Google Scholar] [CrossRef] [Scilit]
  75. United Nations. UN Data: A World of Information. 2025. Available online: https://data.un.org/default.aspx (accessed on 18 April 2026).
  76. Cheng, B.; Sun, G.; Xie, A.; Sun, J.; Bo, X.; Tian, X. Secondary alteration and probable sources of oils and condensates in Jurassic reservoirs of the Turpan Depression, Turpan–Hami Basin, NW China. Org. Geochem. 2025, 208, 105052. [Google Scholar] [CrossRef] [Scilit]
  77. Seitzinger, S.P.; Chuvieco, E.; Di Giuseppe, F.; Bombelli, A.; Cagnazzo, C.; Harris, S.; Tapper, N. Relevance of earth observations of essential climate variables in wildfire adaptation. Remote Sens. Environ. 2026, 332, 115082. [Google Scholar] [CrossRef] [Scilit]
  78. Zhang, G.; Zhang, L.; Li, X.; Feng, X.; Wang, Y.; Guo, J.; Li, P.; Wei, X. Spatiotemporal evolution characteristics and driving mechanisms of wildfires in China under the context of climate change and human activities. Ecol. Indic. 2025, 176, 113694. [Google Scholar] [CrossRef] [Scilit]
  79. Dong, Y.-Y.; Wang, P.; Hua, Z.-L.; Liu, X.-D. River networks evolution under multiple stresses: A geometric and structural fractal perspective. J. Clean. Prod. 2024, 448, 141411. [Google Scholar] [CrossRef] [Scilit]
  80. Wang, W.; Wang, C. Spatiotemporal dynamics of active fire in China (2003–2024): Regional patterns and land cover associations. Fire 2025, 8, 445. [Google Scholar] [CrossRef] [Scilit]
  81. Ren, H.; Wen, Z.; Liu, Y.; Lin, Z.; Han, P.; Shi, H.; Wang, Z.; Su, T. Vegetation response to changes in climate across different climate zones in China. Ecol. Indic. 2023, 155, 110932. [Google Scholar] [CrossRef] [Scilit]
  82. Ouyang, H.; Liu, Y.; Zhao, X. Natural conditions and foundation of the Chinese economy. In Introduction to Chinese Economy; Springer: Singapore, 2025. [Google Scholar] [CrossRef] [Scilit]
  83. Lian, C.; Xiao, C.; Feng, Z.; Ma, Q. Accelerating decline of wildfires in China in the 21st century. Front. For. Glob. Change 2024, 6, 1252587. [Google Scholar] [CrossRef] [Scilit]
  84. Stojković, S.; Marković, D.; Durlević, U. Snow cover estimation using Sentinel-2 high spatial resolution data: A case study of National Park Šar Planina (Serbia). In Advanced Technologies, Systems, and Applications VII (IAT 2022); Ademović, N., Mujčić, E., Mulić, M., Kevrić, J., Akšamija, Z., Eds.; Lecture Notes in Networks and Systems; Springer: Cham, Switzerland, 2023; Volume 539. [Google Scholar] [CrossRef] [Scilit]
  85. Vujović, F.; Valjarević, A.; Durlević, U.; Morar, C.; Grama, V.; Spalević, V.; Milanović, M.; Filipović, D.; Ćulafić, G.; Gazdić, M.; et al. A Comparison of the AHP and BWM Models for the Flash Flood Susceptibility Assessment: A Case Study of the Ibar River Basin in Montenegro. Water 2025, 17, 844. [Google Scholar] [CrossRef] [Scilit]
  86. Deđanski, V.; Durlević, U.; Kovjanić, A.; Lukić, T. GIS-based spatial modeling of landslide susceptibility using BWM-LSI: A case study—City of Smederevo (Serbia). Open Geosci. 2024, 16, 20220688. [Google Scholar] [CrossRef] [Scilit]
  87. United Nations Office for the Coordination of Humanitarian Affairs. China—Subnational Administrative Boundaries. 2020. Available online: https://data.humdata.org/dataset/cod-ab-chn (accessed on 5 March 2026).
  88. Ministry of Foreign Affairs of the People’s Republic of China. Joint Statement Between the People’s Republic of China and the Republic of India. 2013. Available online: https://www.mfa.gov.cn/eng/gjhdq_665435/2675_665437/2711_663426/2712_663428/202406/t20240607_11407512.html (accessed on 7 March 2026).
  89. Google. Google Earth Pro, version 7.3.6; Google: Mountain View, CA, USA, 2024; Available online: https://www.google.com/earth/about/versions/#earth-pro (accessed on 10 March 2026).
  90. European Space Agency. Copernicus Global Digital Elevation Model [Dataset]. Distributed by OpenTopography. 2024. Available online: https://portal.opentopography.org/datasetMetadata?otCollectionID=OT.032021.4326.1 (accessed on 10 March 2026).
  91. Fire Information for Resource Management System (FIRMS). Archive Download. 2025. Available online: https://firms.modaps.eosdis.nasa.gov/download/ (accessed on 13 March 2026).
  92. Laurent, P.; Mouillot, F.; Moreno, M.V.; Yue, C.; Ciais, P. Varying relationships between fire radiative power and fire size at a global scale. Biogeosciences 2019, 16, 275–288. [Google Scholar] [CrossRef] [Scilit]
  93. Schroeder, W.; Oliva, P.; Giglio, L.; Csiszar, I.A. The new VIIRS 375 m active fire detection data product: Algorithm description and initial assessment. Remote Sens. Environ. 2014, 143, 85–96. [Google Scholar] [CrossRef] [Scilit]
  94. QGIS Development Team. QGIS Geographic Information System, version 3.40.09; Open Source Geospatial Foundation: Beaverton, OR, USA, 2025; Available online: http://qgis.osgeo.org (accessed on 15 March 2026).
  95. Minh, N.Q.; Huong, N.T.T.; Khanh, P.Q.; Hien, L.P.; Bui, D.T. Impacts of resampling and downscaling digital elevation model and its morphometric factors: A comparison of Hopfield neural network, bilinear, bicubic, and kriging interpolations. Remote Sens. 2024, 16, 819. [Google Scholar] [CrossRef] [Scilit]
  96. Cappello, G.; Angrisano, A.; Gioia, C.; Maratea, A.; Gaglione, S. On the context-aware GNSS navigation: Test of a k-nearest neighbors classifier in different environments. Eng. Proc. 2026, 126, 10. [Google Scholar] [CrossRef] [Scilit]
  97. World Bank; ESMAP; Solargis. Global Solar Atlas. The World Bank. 2025. Available online: https://globalsolaratlas.info (accessed on 15 March 2026).
  98. Abatzoglou, J.T.; Dobrowski, S.Z.; Parks, S.A.; Hegewisch, K.C. TerraClimate: A high-resolution global dataset of monthly climate and climatic water balance from 1958–2015. Sci. Data 2018, 5, 170191. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  99. Climatology Lab. TerraClimate: Dataset of Monthly Climate and Climatic Water Balance for Global Terrestrial Surfaces from 1950–Present. Available online: https://www.climatologylab.org/terraclimate.html (accessed on 26 March 2026).
  100. Arif, F.; Akbar, M. Resampling air borne sensed data using bilinear interpolation algorithm. In Proceedings of the IEEE International Conference on Mechatronics (ICM 2005); IEEE: Piscataway, NJ, USA, 2005; pp. 62–65. [Google Scholar] [CrossRef] [Scilit]
  101. Yang, J.; Huang, X. The 30 m annual land cover datasets and its dynamics in China from 1985 to 2022 (Version 1.0.2). Earth Syst. Sci. Data 2024, 13, 3907–3925. [Google Scholar] [CrossRef]
  102. United Nations Office for the Coordination of Humanitarian Affairs. China Roads (OpenStreetMap Export). 2025. Available online: https://data.humdata.org/dataset/hotosm_chn_roads (accessed on 17 March 2026).
  103. Ilić, V.; Stojković, M.; Dodevska, Z.; Ilić, S. Machine learning model for prediction of indicative water parameters on the Danube River based on satellite data. In Disruptive Information Technologies for a Smart Society (ICIST 2024); Trajanović, M., Filipović, N., Zdravković, M., Eds.; Lecture Notes in Networks and Systems; Springer: Cham, Switzerland, 2024; Volume 860. [Google Scholar] [CrossRef] [Scilit]
  104. Ilić, V.; Turk Sekulić, M.; Brborić, M.; Radonić, J.; Dmitrašinović, S.; Stojković, M. Enhancing the monitoring system for river water quality: Harnessing the power of satellite data and machine learning. Blue-Green Syst. 2025, 7, 338–352. [Google Scholar] [CrossRef] [Scilit]
  105. Monteiro, T.V.P.; Castor, G.J.B.C.; Castillo Correa, C.G.; Arias, H.R.C.; Ñaupari Huatuco, D.Z.; Molina Rodriguez, Y.P. A hybrid machine learning framework for electricity fraud detection: Integrating isolation forest and XGBoost for real-world utility data. Energies 2025, 18, 6249. [Google Scholar] [CrossRef] [Scilit]
  106. Zhang, J.; Ni, J.; Wang, F.; Huang, H.; Zhang, D. Soil classification from cone penetration test profiles based on XGBoost. Appl. Sci. 2026, 16, 280. [Google Scholar] [CrossRef] [Scilit]
  107. Chen, J.; Hu, R.; Chen, L.; Liao, Z.; Che, L.; Li, T. Multi-sensor integrated mapping of global XCO2 from 2015 to 2021 with a local random forest model. ISPRS J. Photogramm. Remote Sens. 2024, 208, 107–120. [Google Scholar] [CrossRef] [Scilit]
  108. Ren, F.; He, J.; Zhang, Y.; Kong, F. Estimating and projecting forest biomass energy potential in China: A panel and random forest analysis. Land 2026, 15, 152. [Google Scholar] [CrossRef] [Scilit]
  109. Li, Z.; Yuan, Q.; Yang, Q.; Li, J.; Zhao, T. Differentiable modeling for soil moisture retrieval by unifying deep neural networks and water cloud model. Remote Sens. Environ. 2024, 311, 114281. [Google Scholar] [CrossRef] [Scilit]
  110. Luna-Villagómez, E.; Mahalec, V. Exploring Kolmogorov–Arnold networks for unsupervised anomaly detection in industrial processes. Processes 2025, 13, 3672. [Google Scholar] [CrossRef] [Scilit]
  111. Yuan, S.; Liu, Y.; Zhang, X.; Yan, X.; Qin, H.; Akhtar, N. SP-KAN: Sparse-sine perception Kolmogorov–Arnold networks for infrared small target detection. ISPRS J. Photogramm. Remote Sens. 2026, 234, 1–19. [Google Scholar] [CrossRef] [Scilit]
  112. Aravanis, T.; Papadopoulos, P.; Georgikos, D. Fourier feature-enhanced neural networks for wind turbine power modeling. Electricity 2025, 6, 70. [Google Scholar] [CrossRef] [Scilit]
  113. Tran, V.Q.; Byeon, H. Explainable hybrid tabular variational autoencoder and feature tokenizer transformer for depression prediction. Expert Syst. Appl. 2025, 265, 126084. [Google Scholar] [CrossRef] [Scilit]
  114. Aksholak, G.; Bedelbayev, A.; Magazov, R.; Kaplan, K. Transformer tokenization strategies for network intrusion detection: Addressing class imbalance through architecture optimization. Computers 2026, 15, 75. [Google Scholar] [CrossRef] [Scilit]
  115. Caron, N.; Noura, H.N.; Nakache, L.; Guyeux, C.; Aynes, B. AI for wildfire management: From prediction to detection, simulation, and impact analysis—Bridging lab metrics and real-world validation. AI 2025, 6, 253. [Google Scholar] [CrossRef] [Scilit]
  116. Li, J.; Huang, D.; Chen, C.; Liu, Y.; Wang, J.; Shao, Y.; Wang, A.; Li, X. Prediction of forest-fire occurrence in eastern China utilizing deep learning and spatial analysis. Forests 2024, 15, 1672. [Google Scholar] [CrossRef] [Scilit]
  117. Xu, Z.; Li, J.; Cheng, S.; Rui, X.; Zhao, Y.; He, H.; Guan, H.; Sharma, A.; Erxleben, M.; Chang, R.; et al. Deep learning for wildfire risk prediction: Integrating remote sensing and environmental data. ISPRS J. Photogramm. Remote Sens. 2025, 227, 632–677. [Google Scholar] [CrossRef] [Scilit]
  118. Kolokas, N.; Tatsis, V.; Zacharaki, A.; Ioannidis, D.; Tzovaras, D. A unified approach for ensemble function and threshold optimization in anomaly-based failure forecasting. Appl. Sci. 2026, 16, 1452. [Google Scholar] [CrossRef] [Scilit]
  119. Markoulidakis, I.; Markoulidakis, G. Probabilistic confusion matrix: A novel method for machine learning algorithm generalized performance analysis. Technologies 2024, 12, 113. [Google Scholar] [CrossRef] [Scilit]
  120. Zhang, H.; Wang, W.; Ban, Q. Seasonal forest fire risk and key drivers in Yunnan Province: A machine learning approach. npj Nat. Hazards 2025, 2, 59. [Google Scholar] [CrossRef] [Scilit]
  121. Vaiani, L.; Cagliero, L.; Garza, P.; Ravagli, J. Cross-modal consistency types in multimodal social data. Knowl.-Based Syst. 2025, 322, 113705. [Google Scholar] [CrossRef] [Scilit]
  122. Wu, T.; Wang, L.; Wu, J.; Yang, C. Is pre-fire soil moisture an important factor affecting post-fire soil susceptibility to erosion? Int. Soil Water Conserv. Res. 2026, 14, 100573. [Google Scholar] [CrossRef] [Scilit]
  123. Durlević, U.; Srejić, T.; Valjarević, A.; Aleksova, B.; Deđanski, V.; Vujović, F.; Lukić, T. GIS-Based Spatial Modeling of Soil Erosion and Wildfire Susceptibility Using VIIRS and Sentinel-2 Data: A Case Study of Šar Mountains National Park, Serbia. Forests 2025, 16, 484. [Google Scholar] [CrossRef] [Scilit]
  124. Radovanović, M.M. Investigation of solar influence on terrestrial processes: Activities in Serbia. J. Geogr. Inst. Jovan Cvijic SASA 2018, 68, 149–155. [Google Scholar] [CrossRef] [Scilit]
  125. Langović, M.; Srećković, V.A.; Vidović, Z.; Langović, M.; Mijić, Z. The nexus between solar activity and population displacement: The case study of southern Europe. J. Geogr. Inst. Jovan Cvijic SASA 2025, 75, 329–345. [Google Scholar] [CrossRef] [Scilit]
  126. Cruz, M.G.; Alexander, M.E.; Fernandes, P.M.; Kilinc, M.; Sil, Â. Evaluating the 10% wind speed rule of thumb for estimating a wildfire’s forward rate of spread against an extensive independent set of observations. Environ. Model. Softw. 2020, 133, 104818. [Google Scholar] [CrossRef] [Scilit]
  127. Kalfas, D.; Kalogiannidis, S.; Spinthiropoulos, K.; Chatzitheodoridis, F.; Georgitsi, M. The role of traditional fire management practices in mitigating wildfire risk: A case study of Greece. Fire 2025, 8, 389. [Google Scholar] [CrossRef] [Scilit]
  128. Ćurić, V.; Durlević, U.; Ristić, N.; Novković, I.; Čegar, N. GIS application in analysis of threat of forest fires and landslides in the Svrljiški Timok Basin (Serbia). Glas. Srp. Geogr. Drust. 2022, 102, 107–130. [Google Scholar] [CrossRef] [Scilit]
  129. Zhang, G.; Wang, M.; Liu, K. Forest fire susceptibility modeling using a convolutional neural network for Yunnan Province of China. Int. J. Disaster Risk Sci. 2019, 10, 386–403. [Google Scholar] [CrossRef] [Scilit]
  130. Pang, Y.; Li, Y.; Feng, Z.; Feng, Z.; Zhao, Z.; Chen, S.; Zhang, H. Forest fire occurrence prediction in China based on machine learning methods. Remote Sens. 2022, 14, 5546. [Google Scholar] [CrossRef] [Scilit]
  131. Yan, K.; Zhao, F.; Shu, L.; Liu, Y.; Wang, M.; Si, L.; Li, W.; Li, X.; Zhang, S.; Wang, J. A machine learning-based wildfire susceptibility mapping framework for China’s three-north shelterbelt region. Ecol. Inform. 2026, 93, 103539. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Geographical location of the study area.
Figure 1. Geographical location of the study area.
Earth 07 00119 g001
Figure 2. Wildfire inventory for China (2001–2024).
Figure 2. Wildfire inventory for China (2001–2024).
Earth 07 00119 g002
Figure 3. Topography conditions: (a) elevation; (b) slope; (c) aspect; (d) wind exposure.
Figure 3. Topography conditions: (a) elevation; (b) slope; (c) aspect; (d) wind exposure.
Earth 07 00119 g003
Figure 4. Climate characteristics: (a) global horizontal irradiation; (b) air temperature; (c) precipitation; (d) wind speed.
Figure 4. Climate characteristics: (a) global horizontal irradiation; (b) air temperature; (c) precipitation; (d) wind speed.
Earth 07 00119 g004
Figure 5. Hydrological conditions: (a) evapotranspiration; (b) distance from water surfaces; (c) soil moisture.
Figure 5. Hydrological conditions: (a) evapotranspiration; (b) distance from water surfaces; (c) soil moisture.
Earth 07 00119 g005
Figure 6. Land use map.
Figure 6. Land use map.
Earth 07 00119 g006
Figure 7. Anthropogenic conditions: (a) distance from settlements; (b) distance from roads.
Figure 7. Anthropogenic conditions: (a) distance from settlements; (b) distance from roads.
Earth 07 00119 g007
Figure 8. Machine learning workflow for wildfire susceptibility modeling and hazard map generation.
Figure 8. Machine learning workflow for wildfire susceptibility modeling and hazard map generation.
Earth 07 00119 g008
Figure 9. Spatial block-based train/validation/test split used for model validation.
Figure 9. Spatial block-based train/validation/test split used for model validation.
Earth 07 00119 g009
Figure 10. Receiver Operating Characteristic (ROC) curves obtained from raster-level evaluation of the wildfire susceptibility probability maps under real spatial distribution conditions.
Figure 10. Receiver Operating Characteristic (ROC) curves obtained from raster-level evaluation of the wildfire susceptibility probability maps under real spatial distribution conditions.
Earth 07 00119 g010
Figure 11. Precision–recall (PR) curves obtained from raster-level evaluation of the wildfire susceptibility probability maps under class-imbalanced real spatial distribution conditions.
Figure 11. Precision–recall (PR) curves obtained from raster-level evaluation of the wildfire susceptibility probability maps under class-imbalanced real spatial distribution conditions.
Earth 07 00119 g011
Figure 12. Confusion matrices of classification performance for: (a) RF, (b) XGBoost test, (c) XGBoost training, (d) XGBoost validation, (e) DNN, (f) Fourier MLP, (g) KAN, (h) FT Transformer.
Figure 12. Confusion matrices of classification performance for: (a) RF, (b) XGBoost test, (c) XGBoost training, (d) XGBoost validation, (e) DNN, (f) Fourier MLP, (g) KAN, (h) FT Transformer.
Earth 07 00119 g012
Figure 13. Wildfire susceptibility maps generated by: (a) RF; (b) XGBoost; (c) DNN; (d) Fourier MLP; (e) KAN; (f) FT Transformer model.
Figure 13. Wildfire susceptibility maps generated by: (a) RF; (b) XGBoost; (c) DNN; (d) Fourier MLP; (e) KAN; (f) FT Transformer model.
Earth 07 00119 g013
Figure 14. Ensemble-based wildfire susceptibility map with province-level divisions.
Figure 14. Ensemble-based wildfire susceptibility map with province-level divisions.
Earth 07 00119 g014
Figure 15. Cross-model Spearman rank correlation matrix of predicted wildfire susceptibility.
Figure 15. Cross-model Spearman rank correlation matrix of predicted wildfire susceptibility.
Earth 07 00119 g015
Figure 16. SHAP analyses for: (a) RF, (b) XGBoost, (c) DNN, (d) Fourier MLP, (e) KAN, (f) FT Transformer.
Figure 16. SHAP analyses for: (a) RF, (b) XGBoost, (c) DNN, (d) Fourier MLP, (e) KAN, (f) FT Transformer.
Earth 07 00119 g016
Table 1. Characteristics of the multi-sensor wildfire dataset.
Table 1. Characteristics of the multi-sensor wildfire dataset.
SensorsSamplesDateConfidenceFire Radiative PowerResolution
MODIS128,0982001–2011Nominal, High>20 MW500 m (resampled from 1000 m)
VIIRS25,2072012–2024Nominal, High>20 MW500 m (resampled from 375 m)
Table 2. Summary of input variables, resolution, resampling methods, and data sources.
Table 2. Summary of input variables, resolution, resampling methods, and data sources.
CriteriaResolutionResampling MethodYearSource
Elevation (m)90 m to 500 mAverage2024European Space Agency [90]
Slope (°)90 m to 500 mAverage2024
Aspect90 m to 500 mMode2024
Wind exposure90 m to 500 mNearest neighbor2024
Global horizontal irradiation (kWh/m2/year)240 m to 500 mCubic 4 × 4 kernel2024Global Solar Atlas [97]
Air temperature (°C)4 km to 500 mBilinear2001–2024Climatology Lab [99]
Precipitation (mm)4 km to 500 mBilinear2001–2024
Wind speed (m/s)4 km to 500 mBilinear2001–2024
Evapotranspiration (mm/year)4 km to 500 mBilinear2001–2024
Distance from water surfaces (m)30 m to 500 mBilinear2022Yang and Huang [101]
Soil moisture (mm/year)4 km to 500 mBilinear2001–2024Climatology Lab [99]
Land use30 m to 500 mNearest neighbor2022Yang and Huang [101]
Distance from settlements (m)30 m to 500 mBilinear2022Yang and Huang [101]
Distance from roads (m)500 mVector-based2024/25Open Street Map [102]
Table 3. Summary of model configuration and validation-based tuning strategy.
Table 3. Summary of model configuration and validation-based tuning strategy.
ModelMain Tuned/Configured ParametersRegularization/Control Strategy
RFNumber of trees 50–1200, max depth = 20, minimum leaf size = 5, feature fraction = 0.5Bootstrap sampling, balanced class weights, OOB monitoring
XGBoostUp to 6000 boosting rounds, histogram method, max bins = 256, learning-rate schedule, tree-depth and child-weight controlRow/column subsampling, class weighting, early stopping after 200 rounds, Platt scaling
DNNResidual MLP, layer widths 1024–512–256–128, learning rate = 1 × 10−3, batch size = 4096, 180 epochsAdamW, weight decay = 1 × 10−4, label smoothing = 0.08, early stopping, SWA, temperature scaling
KANResidual architecture, widths 512–512–256–128–64, learning rate = 3 × 10−3, batch size = 4096, 100 epochsDropout = 0.20, BatchNorm, AdamW, EMA, early stopping, temperature scaling
Fourier MLPThree hidden layers of 256 neurons, six Fourier frequency components, learning rate = 2 × 10−4, 80 epochsDropout = 0.15, AdamW, weight decay = 1 × 10−4, early stopping, temperature scaling
FT TransformerEmbedding dimension = 64, four attention heads, three encoder layers, learning rate = 2 × 10−4, batch size = 512, 100 epochsDropout = 0.1, AdamW, weight decay = 1 × 10−4, early stopping, temperature scaling
Table 4. Threshold-free raster-level performance metrics computed under real spatial distribution conditions.
Table 4. Threshold-free raster-level performance metrics computed under real spatial distribution conditions.
ModelROC-AUCPR-AUC
RF0.9100.047
XGBoost0.9020.041
Fourier MLP0.8910.034
KAN0.8800.030
FT Transformer0.8740.026
DNN0.8710.025
Table 5. Optimal decision thresholds used for confusion matrix analysis.
Table 5. Optimal decision thresholds used for confusion matrix analysis.
ModelOptimal Threshold
RF0.30
XGBoost0.05
DNN0.05
Fourier MLP0.50
KAN0.05
FT Transformer0.45
Table 6. Model-based distribution of wildfire susceptibility categories.
Table 6. Model-based distribution of wildfire susceptibility categories.
Wildfire Susceptibility (%)
ModelVery LowLowMediumHighVery High
RF62.114.28.79.25.8
XGBoost61.312.79.19.47.5
DNN54.113.710.913.08.3
Fourier MLP61.011.28.29.210.4
KAN56.610.211.911.89.5
FT Transformer53.814.910.410.210.7
Ensemble56.914.410.710.67.4
Table 7. Spatial distribution of wildfire susceptibility (high and very high) in China by province-level divisions.
Table 7. Spatial distribution of wildfire susceptibility (high and very high) in China by province-level divisions.
NumberProvince-Level DivisionTotal Area (km2)Susceptible Area (km2)Susceptibility Within the Province (%)
1Anhui 140,70837,77226.8
2Beijing 16,3170.30.0
3Chongqing 81,1512.60.0
4Fujian 120,747104,84686.8
5Gansu 403,6670.60.0
6Guangdong 174,485140,78080.7
7Guangxi Zhuang 234,136194,17182.9
8Guizhou 175,38050,68428.9
9Hainan 33,71014,87544.1
10Hebei 186,45355543.0
11Heilongjiang 448,245303,49967.7
12Henan 164,78255,12033.5
13Hong Kong 889.1549.661.8
14Hubei 186,19296695.2
15Hunan 211,44997,30646.0
16Inner Mongolia 1,140,435104,3619.2
17Jiangsu 99,89533,40933.4
18Jiangxi 166,486127,27976.5
19Jilin 190,27258,42530.7
20Liaoning 143,69419,84613.8
21Macao4.70.00.0
22Ningxia Hui 50,4040.00.0
23Qinghai 715,9030.00.0
24Shaanxi 205,68395554.7
25Shandong 153,07129,27719.1
26Shanghai5923.00.00.0
27Shanxi 157,013880.60.6
28Sichuan481,40713,6622.8
29Taiwan35,643250.00.7
30Tianjin11,61844.60.4
31Tibet 1,124,20819460.2
32Xinjiang Uygur 1,602,19832430.2
33Yunnan 381,527198,95552.2
34Zhejiang 100,12622,54922.5
Table 8. Cross-model prediction consistency metrics. Pairwise comparison of wildfire susceptibility predictions across models, including mean absolute error (MAE), root mean square error (RMSE), and Pearson correlation coefficients (PCC).
Table 8. Cross-model prediction consistency metrics. Pairwise comparison of wildfire susceptibility predictions across models, including mean absolute error (MAE), root mean square error (RMSE), and Pearson correlation coefficients (PCC).
Model aModel bMAERMSEPCC
DNNFT Transformer0.0910.1130.938
DNNFourier MLP0.1200.1530.896
DNNKAN0.0950.1160.945
DNNRF0.1120.1420.918
DNNXGBoost0.1080.1350.926
FT TransformerFourier MLP0.0830.1350.910
FT TransformerKAN0.0720.1190.927
FT TransformerRF0.0800.1280.926
FT TransformerXGBoost0.0720.1180.934
Fourier MLPKAN0.0890.1520.883
Fourier MLPRF0.0670.1120.937
Fourier MLPXGBoost0.0590.1020.946
KANRF0.0830.1380.908
KANXGBoost0.0760.1290.917
RFXGBoost0.0370.0610.979
Table 9. Predictor importance ranking based on mean absolute SHAP values.
Table 9. Predictor importance ranking based on mean absolute SHAP values.
Predictive VariableRank
Soil moisture1
Elevation2
Slope3
Global horizontal irradiance4
Air temperature5
Wind speed6
Agricultural plots, meadows, and pastures7
Precipitation8
Distance from water surfaces9
Distance from roads10
Evapotranspiration11
Distance from settlements12
Forests13
Bare dry land (deserts)14
Wet barren land15
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Durlević, U.; Ilić, V.; Radovanović, M.M.; Milanović Pešić, A.; Petrović, M.D.; Milenković, M.; Jovanović, J.M.; Atasoy, E. Wildfire Susceptibility Mapping in China Combining Machine Learning, Deep Learning, and Transformer-Based Models. Earth 2026, 7, 119. https://doi.org/10.3390/earth7040119

AMA Style

Durlević U, Ilić V, Radovanović MM, Milanović Pešić A, Petrović MD, Milenković M, Jovanović JM, Atasoy E. Wildfire Susceptibility Mapping in China Combining Machine Learning, Deep Learning, and Transformer-Based Models. Earth. 2026; 7(4):119. https://doi.org/10.3390/earth7040119

Chicago/Turabian Style

Durlević, Uroš, Velibor Ilić, Milan M. Radovanović, Ana Milanović Pešić, Marko D. Petrović, Milan Milenković, Jasmina M. Jovanović, and Emin Atasoy. 2026. "Wildfire Susceptibility Mapping in China Combining Machine Learning, Deep Learning, and Transformer-Based Models" Earth 7, no. 4: 119. https://doi.org/10.3390/earth7040119

APA Style

Durlević, U., Ilić, V., Radovanović, M. M., Milanović Pešić, A., Petrović, M. D., Milenković, M., Jovanović, J. M., & Atasoy, E. (2026). Wildfire Susceptibility Mapping in China Combining Machine Learning, Deep Learning, and Transformer-Based Models. Earth, 7(4), 119. https://doi.org/10.3390/earth7040119

Article Metrics

Back to TopTop