Next Article in Journal
Multi-Hydrological Factor-Driven Attribution and Future Prediction of Vegetation Dynamics on the Qinghai-Tibetan Plateau
Previous Article in Journal
Soil CO2 Flux in Middle-Aged Pedunculate Oak (Quercus robur L.) Stands on Different Chernozem Subtypes
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

FasterNetFire: A Cost-Effective Fast Neural Network for Forest Fire Detection with Partial Convolution

1
International College of Digital Innovation, Chiang Mai University, Chiang Mai 50200, Thailand
2
Sichuan Provincial Key Laboratory of Philosophy and Social Sciences for Mountain Tourism Safety, Chengdu 610041, China
3
School of Information and Engineering, Chengdu University, Chengdu 610106, China
4
School of Electronic Information and Automation, Civil Aviation University of China, Tianjin 300300, China
*
Author to whom correspondence should be addressed.
Forests 2026, 17(6), 672; https://doi.org/10.3390/f17060672
Submission received: 15 April 2026 / Revised: 15 May 2026 / Accepted: 28 May 2026 / Published: 31 May 2026
(This article belongs to the Section Natural Hazards and Risk Management)

Abstract

Forest fires occur frequently around the world due to extreme weather conditions of high temperatures and drought. Vision-based convolutional neural networks (CNNs) have greatly improved forest fire detection accuracy. However, slow inference speed severely restricts real-time deployment in actual forest scenes. Existing models generally adopt group convolution (GConv) or depthwise convolution (DWConv) to reduce computational complexity, which causes frequent memory access and result in a practical inference speed far below theoretical expectations. Therefore, we propose a novel fast neural network named FasterNetFire for forest fire detection, which introduces partial convolution (PConv) as the basic feature extraction operator to reduce redundant computation as well as memory access overhead simultaneously. FasterNetFire is composed of four cascaded stages, each stage contains several stacked FasterNet Blocks, and the core of each FasterNet Block is an inverted residual module built upon PConv. The proposed network significantly improves inference efficiency while maintaining the effectiveness of spatial feature extraction. Experiments conducted on the FD and Foggia’s fire detection dataset demonstrate that our FasterNetFire achieves an impressive inference speed of up to 290 frames per second (FPS) on graphics processing unit (GPU) platforms. Compared with current representative methods, its inference speed is 4.5× and 4× faster than that of EFDNet and DFAN. Furthermore, FasterNetFire achieves the best results among 17 state-of-the-art methods, achieving an excellent balance between detection accuracy and real-time response performance. This advantage fully verifies the high efficiency of PConv in vision forest fire detection tasks and provides a novel lightweight solution for real-time monitoring and early warning of forest fires in resource-constrained environments.

1. Introduction

Driven by extremely high-temperature and drought conditions, forest fuels persistently accumulate under extremely dry environments, resulting in frequent forest fires [1,2]. Forest fire is a major sudden-onset disaster that threatens human life and property and severely damages natural ecosystems [2]. Due to its rapid spread, wide impact range, and destructive power, once a fire occurs, it may cause irreversible casualties, substantial economic losses, and long-term ecological damage [3]. To effectively prevent large-scale harm, early, accurate, and real-time fire detection is essential.
For many years, traditional contact sensors such as those for smoke, temperature, and particles were employed for fire detection. Such sensors are inexpensive and easy to deploy. However, they require proximity to the fire source to be activated and manual intervention to verify the authenticity of alarms. Moreover, they cannot provide information regarding the scale and location of a fire, which may result in missing the optimal time for fire extinguishing. To overcome these limitations, researchers have proposed fire detection systems based on visual sensors [4,5]. Such sensors feature fast response, wide coverage and strong environmental robustness, and have attracted increasing attention [6].
Traditional vision methods depend on handcrafted color, texture, and shape features [5,7,8,9]. In complex scenes, such features are sensitive to illumination changes, occlusion, and cluttered backgrounds. Moreover, flame-like objects (e.g., sunset, red light sources, and vehicle lights) exhibit high visual similarity to real fire, leading to poor generalization and high false-alarm rates.
With the success of deep learning in image classification, object detection, and semantic segmentation, fire detection has also made major progress, with representative models, including ANetFire [10], CNNFire [11], ResNetFire [12], GNetFire [13], EMNFire [14], EFDNet [15], IEFDNet [16], and DFAN [17], as shown in Figure 1. ANetFire and CNNFire [10,11] achieve extremely high inference speed by adopting standard convolutions; nevertheless, their model accuracy is below 90%. Both EFDNet and DFAN [15,17] achieve accuracy exceeding 95%, yet their inference speed on GPUs is merely around 70 FPS. Although DFAN_Comp [17] nearly doubles the inference speed of DFAN to 125 FPS, it employs genetic algorithms with high time complexity for channel pruning. By optimizing multi-scale convolutions, IEFDNet [16] improves inference speed but incurs a certain loss in accuracy. Therefore, achieving a better trade-off among fire detection accuracy, inference speed, and model size remains a major challenge in the current field of fire detection. Most existing state-of-the-art (SOTA) models adopt convolution operators such as group convolution [18] (GConv) and depthwise convolution [19] (DWConv). These convolution operators significantly reduce model floating-point operations (FLOPs) and improve accuracy; nevertheless, frequent memory access severely compromises model inference speed.
Beyond fire detection, recent studies also show that efficient learning and optimization models are being actively applied in trajectory segmentation for urban GPS [20], thermal-property inversion for airport pavement materials [21], stochastic aircraft–pavement vibration analysis [22], multi-UAV path planning [23], flight-arrival prediction [24], and facial-video-based affective computing [25]. These cross-domain deployments further motivate pursuing better accuracy–efficiency trade-offs for real-world intelligent systems.
In this paper, we propose a novel network, FasterNetFire, for forest fire detection based on PConv [26], which applies convolution only on a subset of input channels to reduce both computation and memory access. The network contains four stages, each preceded by an Embedding or Merging layer for spatial downsampling and channel expansion. Specifically, the Embedding layer uses a 4 × 4 convolution with stride 4, while the Merging layer uses a 2 × 2 convolution with stride 2. Each stage is composed of stacked FasterNet Blocks [26]. Each FasterNet Block contains one PConv layer followed by two pointwise convolution (PWConv, i.e., 1 × 1 convolution) layers, where the first PWConv expands channels. A shortcut connection is used to reuse input features, forming an overall inverted residual structure.
The main contributions of this study are as follows:
1.
We point out that improving inference speed is a key challenge for current SOTA methods, and that frequent memory access caused by DWConv/GConv-based operators is a major bottleneck.
2.
We introduce PConv, which applies convolution kernels to only a subset of input channels while keeping the remaining channels unchanged, enabling fast and efficient computation. Compared with standard convolution, PConv has lower FLOPs; compared with DWConv/GConv, it achieves higher computational efficiency.
3.
We propose FasterNetFire and evaluate it on the FD dataset, one of the largest and most challenging datasets in fire detection. Experimental results show that FasterNetFire significantly outperforms mainstream methods in detection accuracy, model size, and inference speed. On GPU, our method reaches an impressive 290 FPS on the FD dataset, which is about 4.5 × and 4 × faster than EFDNet [15] and DFAN [17].
The rest of this paper is organized as follows: Section 2 reviews related work, Section 3 presents the proposed method in detail, Section 4 reports experimental results and analysis, and Section 5 concludes the paper and discusses future work.

2. Related Work

In this section, we briefly review related research on vision-based fire detection, including traditional machine-learning vision detection methods and deep-learning-based detection methods, and clarify the differences between our work and these studies.

2.1. Traditional Machine Learning

Traditional machine-learning fire detection methods mainly rely on extracting color, texture, shape, motion, and other features from input images. Researchers extracted color features in RGB/YCbCr/YUV color spaces to achieve general fire detection. For instance, Healey et al. [7] used class probabilities based on color information for fire detection. Chen et al. [4] proposed extracting flame and smoke pixels using a color model to realize fire detection. Celik et al. [5] proposed a generic fire detection method based on color features in the YCbCr color space. However, this type of detection method, which relies only on color features, has poor robustness and leads to a high false-alarm rate. To solve this problem, Celik et al. [8] introduced a fuzzy-logic system to improve the color model and enhance the model’s ability to distinguish real flames from flame-like targets. Angayarkkani et al. [9] proposed a data-driven forest fire detection method that combines color information with a fuzzy inference system. Some researchers began to apply motion information to fire detection tasks. For example, Toreyin et al. [27] used a hidden Markov model to distinguish real flames from flame-like moving objects, and also used a spatiotemporal wavelet transform to analyze flame dynamics. However, too many heuristic thresholds limited this method in practical applications. Lee et al. [28] proposed a method based on spatiotemporal feature analysis to extract smoke motion features. To extract moving regions more accurately, researchers introduced Gaussian mixture models (GMM) and optical-flow methods. Han et al. [29] and Chen et al. [30] both used GMM to extract moving foreground targets, and then combined other algorithms to determine whether the foreground was flame. Ha et al. [31] used motion vectors to describe dynamic targets, extracted color features in the LAB color space, and combined flame features to identify real fires. However, classical optical-flow algorithms based on the brightness-constancy assumption cannot adequately represent flame appearance. Therefore, Mueller et al. [32] specifically designed two improved optical-flow methods for fire detection: an optimal transport model integrating dynamic texture information and a data-driven optical-flow method. However, due to high computational cost, this type of moving-target extraction method has not been widely used in fire detection. Because heuristic-rule-based methods have limitations in feature extraction, researchers proposed trainable methods with automatic feature-learning capability to avoid subjective factors affecting classification results. For example, Gubbi et al. [33] extracted features through discrete cosine transform and discrete wavelet transform, and then used support vector machines (SVM) for classification. Emmy Prema et al. [34] segmented suspected smoke regions using color filters, extracted features, and then used SVM to classify the extracted features. Byoung et al. [35] extracted high-frequency component magnitudes in horizontal, vertical, and diagonal directions as features, and then used a two-class SVM for final classification. Habiboglu et al. [36] used SVM to classify covariance features extracted from spatiotemporal blocks. Truong et al. [37] used an adaptive Gaussian model to segment moving regions, clustered flame pixels by fuzzy C-means, extracted spatiotemporal features, and then used SVM for feature classification. Yuan [38] combined local binary patterns (LBP) with color features and used AdaBoost to distinguish smoke images from non-smoke images. To enhance discrimination between flames and flame-like targets, Zhang et al. [39] designed a BP neural network based on flame dynamic characteristics for fire detection. Traditional methods based on manually designed features require complex feature engineering and cumbersome pipelines. In addition, maintaining a good balance between detection accuracy and false-alarm rate remains a difficult problem.

2.2. CNN

In recent years, deep learning has made significant progress in computer vision and has been widely applied. In particular, image recognition based on deep convolutional neural networks can automatically perform end-to-end feature extraction and classification, making it more convenient and reliable. To overcome the limitations of traditional vision detection methods, several researchers began to apply existing CNN methods to fire detection tasks. For instance, Frizzi et al. [40] and Mao et al. [41] used LeNet-5 [42] to detect fire, while Sun et al. [43] applied this network to the specific scenario of forest fire detection. Compared with traditional methods based on handcrafted features, CNN-based methods achieved significant improvements in detection accuracy. Lee et al. [44] verified through multiple experiments the performance of GoogLeNet, VGG, and AlexNet for fire detection, among which GoogLeNet performed best. Sharma et al. [12] evaluated VGG16 [45] and ResNet50 [46]; experimental results showed that ResNetFire outperformed VGG16, but the dataset used in the experiments was relatively small. Dunnings et al. [47] compared InceptionV1 [48], VGG16 [45], and AlexNet [18]. On this basis, they simplified the networks to form FireNet and InceptionV1-OnFire. Although the simplified networks reduced classification accuracy by about 1%, inference speed increased by 3 to 4 times. Muhammad et al. [10] proposed AlexNetFire for early fire detection, but the model had too many parameters and was difficult to deploy on resource-constrained devices. To solve this problem, they proposed GNetFire [13] for fire detection in surveillance videos, then designed a lightweight CNNFire [11] network for fire detection in surveillance videos, and proposed EMNFire [14] for fire detection in IoT environments with complex uncertainty. In recent studies, they further expanded related auxiliary tasks, including human detection and counting, evacuation monitoring, combustible-object detection, and target-type recognition [49]. Existing CNN-based fire detection methods have achieved significant progress, but they have not effectively balanced detection accuracy and model complexity. To reduce CNN architectural complexity, Li et al. proposed EFDNet [15], which integrated AlexNet, Inception modules, and an implicit deep-supervision module. DFAN [17] applied a dual fire attention network for efficient and accurate fire detection. MS-Net [50] proposed a fire detection method that can handle variable-scale images. Although these networks improved fire detection accuracy and reduced computational complexity and model size, inference speed is equally important as recognition accuracy in practical deployment scenarios. Existing models have still not achieved a new breakthrough in inference speed, which greatly limits their deployment on resource-constrained devices.

3. Methodology

As illustrated in Figure 2, FasterNetFire adopts a four-stage hierarchical backbone. The input image is first processed by an embedding layer for initial downsampling and channel projection, and then passed through consecutive stages with progressively increased channel width. Each stage stacks several FasterNet blocks, where a PConv layer performs efficient spatial feature extraction and two pointwise ( 1 × 1 ) convolutions complete channel mixing and transformation. This design preserves discriminative fire features while reducing computational overhead.

3.1. Partial Convolution

For fire/non-fire classification, we need operators that preserve discriminative flame cues while limiting per-layer computation and memory traffic. PConv [26] achieves this by applying a k × k spatial kernel only to a fraction of input channels and forwarding the remaining channels without spatial convolution, instead of touching every channel with depthwise or grouped kernels at each step.
Given a feature map X R h × w × c , where h and w are spatial height and width, c is the number of channels, and r ( 0 , 1 ] is a partial ratio, the number of channels that participate in spatial mixing is
c p = max 1 , r c .
We partition X along the channel axis into X 1 R h × w × c p and X 2 R h × w × ( c c p ) , i.e., the partial slice to be convolved and the identity slice that bypasses the k × k operator, respectively. Only X 1 is passed through a standard convolution with kernel size k, stride s, padding p, and parameters θ ; X 2 is carried forward unchanged until the two slices are merged. The forward mapping of PConv is, therefore,
Y 1 = Conv k × k ( X 1 ; θ ) , Y = Concat ( Y 1 , X 2 ) R h × w × c ,
where Conv k × k ( · ; θ ) denotes the k × k convolution on X 1 , Concat ( · , · ) stacks tensors along the channel dimension, and ( h , w ) is the spatial size induced by ( k , s , p ) . Stride and padding are chosen consistently with the backbone so that Y 1 and X 2 stay spatially aligned before concatenation. Intuitively, c p controls how aggressively spatial context is extracted, while the c c p identity channels retain low-cost cues (e.g., color layout of fire-like distractors) without extra k × k multiply–accumulates on the full width [26].

3.2. FasterNetBlock

A FasterNet block is the elementary residual-style unit in our backbone [26]: it performs inexpensive spatial mixing on a channel subset via PConv, then performs full cross-channel fusion with a pointwise ( 1 × 1 ) convolution, batch normalization, and a nonlinearity. Thus, spatial context is first injected on a partial channel slice, and channel interaction is completed afterward at full width—a compact pattern suited to lightweight backbones.
Let X R h × w × c denote the block input. Partial convolution with ratio r and spatial parameters ( k , s , p , θ ) maps X to an intermediate tensor U of the same channel width c, following (2). A pointwise convolution with weights W pw then mixes information across all channels, producing V . Batch normalization with learnable scale and shift ( γ , β ) and an element-wise activation σ ( · ) yields the block output Y . Formally,
U = PConv ( X ; r , k , s , p , θ ) , V = PWConv ( U ; W pw ) , Y = σ BN ( V ; γ , β ) ,
where PWConv ( · ) denotes 1 × 1 convolution over the full c channels, BN ( · ) applies per-channel normalization with running statistics at inference time, and U , V , and Y share the spatial size ( h b , w b ) whenever PConv uses stride one with padding that preserves resolution inside a stage. Any residual wiring used in the concrete FasterNet instantiations [26] is absorbed into the same operator ordering and is not expanded separately in (3).

3.3. FasterNetFire Network

We now assemble PConv and FasterNet blocks [26] into an end-to-end fire–non-fire classifier, denoted FasterNetFire, using the four-stage layout in Figure 2. The architecture follows a standard recipe for image-level binary classification: a compact backbone maps each RGB frame to deep features, and a small head maps pooled features to class logits.
Let I R h in × w in × 3 denote an input RGB image. A stem embedding ϕ emb ( · ) performs an initial 4 × 4 convolution with stride 4 and 64 output channels, reducing spatial resolution while expanding channels so that subsequent stages operate on semantically richer tokens:
Z ( 0 ) = ϕ emb ( I ) = Conv 4 × 4 , s = 4 I ; W emb , W emb R 4 × 4 × 3 × 64 .
The backbone is organized into L = 4 stages indexed by { 1 , , 4 } . Stage stacks N FasterNet blocks (3) with a fixed partial ratio r across the network; we denote the composition of these blocks by B ( l ) ( · ) . Between consecutive stages, a transition operator S ( l ) ( · ) adjusts spatial size and channel width (“merging” in the FasterNet terminology [26]) so that deeper stages see coarser maps with wider representations:
S ( l ) X ( l ) = Conv 2 × 2 , s = 2 X ( l ) ; W merge ( l ) , W merge ( l ) R 2 × 2 × c l 1 × c l .
In our training configuration, ( N 1 , , N 4 ) = ( 1 , 2 , 8 , 2 ) and ( c 1 , , c 4 ) = ( 64 , 128 , 256 , 512 ) channels after each stage transition, consistent with Table 3. Using the embedding and merging definitions in Equations (4) and (5), the recursive feature extraction pipeline can be written as
Z ( 0 ) = ϕ emb ( I ) , Z ( l ) = S ( l ) B ( l ) ( Z ( l 1 ) ) , l = 1 , , 4 ,
where Z ( l ) is the output tensor of stage , Z ( 0 ) is the stem feature, and Z ( 4 ) is the deepest map fed to the classifier.
For binary fire detection, we apply global average pooling (GAP) over the spatial dimensions of Z ( 4 ) and attach a linear layer with weights W and bias b , followed by a softmax to obtain class probabilities p ^ R 2 :
p ^ = softmax W GAP ( Z ( 4 ) ) + b .
Here GAP ( · ) outputs a c 4 -dimensional vector by averaging each channel of Z ( 4 ) over its spatial grid, W R 2 × c 4 , and the j-th entry of p ^ is interpreted as the predicted probability of class j (fire vs. non-fire).

4. Experimental Results

4.1. Experimental Setup

4.1.1. Dataset

In this paper, experiments are conducted on both the FD dataset [15] and Foggia’s dataset [6]. These two benchmarks provide complementary evaluation for image-based and video-sequence-based fire detection scenarios.
The FD dataset is currently a widely used public benchmark dataset in fire detection, with abundant samples and high scene complexity, and is specifically designed to validate flame recognition and anti-interference capability of models in complex environments. This dataset is a balanced binary-class dataset, containing two major categories: Fire and Non-Fire. The samples cover multiple real environments, including indoor, outdoor, forest, building, and traffic scenes, as shown in Figure 3. It also contains a large number of flame-like interference samples that are highly similar to flames in color and brightness, such as sunsets, evening glow, strong light, red lights, red vehicles, red tents, and lanterns, which can fully evaluate the model’s feature discrimination ability and generalization performance. The dataset is split into training, validation, and test sets according to a fixed ratio, and all subsets maintain a balanced class distribution. The detailed sample numbers are shown in Table 1. The Fire category includes flame samples with different scales, different shapes, and different occlusion levels, covering various cases such as large-area open flames at close range, medium-distance flames, small flames in long-range views, and flames accompanied by smoke. The Non-Fire category contains a large number of flame-like interference samples that are prone to misclassification, which is a key component for evaluating model robustness. The dataset has a large scale, balanced distribution, and complex scenes, which ensures the objectivity, reliability, and reproducibility of experimental results, and meets current research and evaluation standards for deep-learning-based fire detection.
Foggia’s dataset is a video-based fire detection dataset developed by Foggia et al. [6]. This dataset consists of 31 videos, where 14 videos contain fire scenes and 17 videos do not contain fire scenes. Following the frame-extraction protocol adopted in prior studies, we extracted 14,036 frames from these videos and organized them into Fire and Non-Fire classes, with 7018 frames per class. The dataset is further divided into training, validation, and test sets with a ratio of 7:2:1.

4.1.2. Implementation Details

To ensure fairness and reproducibility of experimental results, all comparative experiments, ablation experiments, and parameter analysis experiments in this study are conducted under a unified software and hardware environment, as summarized in Table 2. The hardware platform adopts dual Intel Xeon Gold 6330 central processing unit (CPU) processors with a total of 112 threads, equipped with large-capacity L3 cache and a high-speed memory architecture. It is paired with an NVIDIA A100-SXM4-80GB high-performance computing GPU to provide strong computing support for model training and inference. The software environment is based on Ubuntu 22.04.3 LTS (Jammy Jellyfish), uses CUDA 12.2 for GPU acceleration, and adopts PyTorch 1.11.0 as the deep learning framework. The Raspberry Pi benchmark is conducted on a Raspberry Pi 4 Model B with 4 GB LPDDR4 RAM and a quad-core 64-bit Cortex-A72 CPU at 1.5 GHz. In both training and testing stages, a unified batch size of 32 is used, and the model input image size is fixed at 224 × 224 × 3 , ensuring fair evaluation of all models under the same hardware load and computational conditions.
FasterNetFire is based on FasterNet-T0 [26], with fine-tuning of channel width, block number, and downsampling strategy for the fire-detection task. The specific structural configuration is shown in Table 3, which gives the corresponding tensor shapes when the input resolution is 224 × 224 × 3 .

4.1.3. Evaluation Metrics

In fire detection field, the common evaluation metrics include Recall [51], Precision [51], Accuracy [51], F 1 -score [52], area under the curve (AUC) [53], and false alarm rate (FAR) [54]. The specific description is shown as follows.
Let true positive (TP), true negative (TN), false positive (FP), and false negative (FN) denote the four entries in the confusion matrix, respectively.
Accuracy = T P + T N T P + T N + F P + F N
Precision = T P T P + F P
Recall = T P T P + F N
F 1 - score = 2 × Precision × Recall Precision + Recall
FAR = F P F P + T N
where T P is the number of true positives, T N is true negatives, F P is false positives, and F N is false negatives. These six metrics jointly characterize model performance from complementary perspectives. Accuracy reflects overall correctness on all samples, while Precision measures the reliability of positive (fire) predictions by penalizing false alarms. Recall quantifies the ability to capture real fire events and is directly related to missed detections, which is especially critical in safety-oriented applications. The F 1 -score provides a balanced summary of Precision and Recall, enabling robust comparison when the costs of false positives and false negatives must be considered simultaneously. FAR directly measures false alarms over all negative samples, which is important for practical alarm systems. AUC evaluates the ranking quality across all classification thresholds and thus reflects threshold-independent discrimination ability.

4.2. Performance of FasterNetFire

Table 4 reports the comparative performance of different methods on the FD and Foggia’s datasets. For Foggia’s dataset, only the results reported by the original papers are listed, and missing entries are kept as blanks for fair comparison. Traditional machine-learning methods (FD-GCM [5], FFD-ANN [39], and FPC [8]) are overall significantly weaker than deep-learning methods in terms of accuracy-related performance on the FD dataset. Although FPC achieves the highest recall of 99.90%, its precision and accuracy are only 52.00% and 53.90%, respectively, indicating a severe false-alarm problem. Among deep-learning methods, FasterNetFire achieves 94.69% precision, 98.32% recall, 96.47% F1-score, and 96.18 ± 0.50 % accuracy, with all metrics at a leading level. In particular, both precision and accuracy surpass the current SOTA model DFAN [17] (96.00% and 96.17%), demonstrating a clear breakthrough in overall detection accuracy. On the FD test set, using the confusion-matrix statistics ( T P = 2459 , F P = 138 , F N = 42 , T N = 2361 ), FasterNetFire further achieves a FAR of F P / ( F P + T N ) = 138 / ( 138 + 2361 ) = 5.52 % , and the corresponding single-operating-point AUC (computed from TPR/FPR) is approximately 96.40%. Over 10 repeated runs, our accuracy is 96.18 ± 0.50 on FD and 99.80 ± 0.22 on Foggia’s dataset. More importantly, the model keeps competitive and stable performance across two datasets with different data-collection protocols and scene distributions, including complex illumination conditions (e.g., strong light, sunset-like glow, and nighttime low-light) and highly cluttered backgrounds. This cross-dataset consistency indicates that FasterNetFire learns transferable fire semantics rather than dataset-specific shortcuts, and therefore, does not exhibit single-dataset overfitting. In addition, Foggia’s benchmark is constructed from video sequences, and our strong FP/FN and ACC results on this benchmark further verify the model’s generalization ability and temporal-scene robustness for practical video-surveillance applications.

4.3. Analysis of Model Complexity

Table 5 presents a comparison of FLOPs, model size, and inference speed among deep-learning models. While maintaining a high accuracy of 96.18 ± 0.50 %, FasterNetFire achieves 290.54 FPS on GPU and 38.11 FPS on CPU, significantly outperforming all comparison models in inference speed. On edge-device deployment, FasterNetFire also reaches 8.44 FPS on Raspberry Pi, clearly higher than DFAN (0.83 FPS) and DFAN_Comp (3.21 FPS), showing superior practical responsiveness under constrained hardware. Although the absolute accuracy gain over the strongest baseline is only 0.23%, this margin is still meaningful in safety-critical fire detection, where even a small reduction in misclassification can directly reduce missed alarms and false dispatches in long-term deployment. More importantly, the primary contribution of this work is efficiency-oriented optimization: under nearly the same accuracy level, FasterNetFire delivers about 2–4.5× faster inference than representative high-accuracy baselines, which substantially improves real-time responsiveness and practical deployability on resource-constrained platforms. Figure 4 plots the same quantities in a radial-bar form, so that cross-model contrasts in accuracy, FLOPs, model size, and GPU/CPU/Pi FPS can be read at a glance and checked against Table 5. At the same time, FasterNetFire remains lightweight, with only 6.32 MB model size and 0.85 G FLOPs, demonstrating the best overall balance between detection accuracy and computational efficiency.

4.4. Analysis of Partial Ratio r

Table 6 reports an ablation study on different partial-convolution ratios r in PConv. The model achieves the best overall performance at r = 1 / 4 , where all metrics reach their peak values. As r decreases ( 1 / 8 , 1 / 16 , 1 / 32 ), detection performance consistently drops. When r increases to 1 / 2 , model accuracy also degrades notably, verifying that r = 1 / 4 is the optimal configuration for the fire-detection task. To further explain why r = 1 / 4 yields the best trade-off in Table 6, we analyze PConv from a memory-access perspective and compare it with commonly used convolution designs in prior fire-detection backbones.
In previous SOTA models, efficiency is often improved by reducing model FLOPs. For example, ResNetFire [12] mainly uses group convolution, EMNFire [14] mainly uses depthwise separable convolution, and GNetFire/EFDNet [13,15] extensively use Inception-style multi-branch structures. These designs reduce FLOPs to some extent, but from the perspective of memory access, the increased memory-access overhead often dominates runtime and can even prevent practical inference speed from improving. Let the input be X R c × h × w , the output channels be c, the kernel size be k × k , and let group convolution split channels into g groups (each group has c / g channels). Its memory access can be written as follows:
MAC Group = h × w × 2 c + k 2 × c 2 g .
As g increases, per-group computation becomes smaller, but data fragmentation across groups reduces cache hit rate and adds extra data transfer between memory and registers, further amplifying memory-access bottlenecks.
To compensate for accuracy loss, depthwise separable convolution is usually combined with channel expansion ( c c , c > c , e.g., × 6 in MobileNetV2 [59]), and its memory access gradually increases to the following:
MAC DW + PW h × w × 2 c + k 2 × c h × w × 2 c .
Here, the term h × w × 2 c is an I/O bottleneck that is difficult to optimize further and directly increases inference latency.
For Inception-style multi-branch structures, branch outputs are concatenated after parallel computation. Assume there are b branches, and the output channels of each branch are c i with i c i = c . Then memory access is as follows:
MAC Inception = h × w × 2 c + i = 1 b h × w × c i + MAC branch , i .
Compared with standard convolution, the extra term i = 1 b ( h × w × c i + MAC branch , i ) introduces significant overhead, causing obvious inference slowdown and making the accuracy gain costly.
In contrast, the memory access of PConv is as follows:
MAC PConv = h × w × 2 c + k 2 × ( r c ) 2 , 0 < r < 1 ,
where r is the ratio of channels participating in spatial convolution to all channels.

4.5. Analysis of Confusion Matrix

By jointly analyzing Table 4 and Table 5, the core advantage of FasterNetFire is a dual breakthrough in both accuracy and efficiency on the FD dataset. On one hand, with a PConv-based high-compute-density feature extraction design, the model substantially reduces computation and parameter size while preserving the key discriminative fire features, and thus achieves superior accuracy to existing SOTA models. This result validates the effectiveness of PConv for fire detection. On the other hand, the partial-channel convolution mechanism in PConv significantly reduces memory-access cost, enabling much higher inference speed on both GPU and CPU while maintaining a lightweight model size, which makes FasterNetFire highly suitable for edge deployment. The ablation study in Table 6 further shows that fire feature maps contain considerable channel redundancy. When r is too small, the model cannot extract sufficient effective features, leading to accuracy degradation. When r is too large, redundant computation increases, which fails to improve accuracy and may reduce inference efficiency. The optimal setting r = 1 / 4 strongly confirms the rationality of the PConv design in exploiting feature redundancy and balancing efficiency with accuracy. Overall, FasterNetFire demonstrates clear advantages in accuracy, speed, and parameter size, addressing the long-standing trade-off between high accuracy and high speed in fire detection, and providing an effective solution for real-time and lightweight deployment.
As shown in Figure 5, among 2500 true non-fire samples, 138 are misclassified as fire, yielding a false-alarm rate of 5.52%; among 2500 true fire samples, 42 are missed, yielding a miss-detection rate of 1.68%.
As shown in Figure 6, FasterNetFire keeps both false positives (FP) and false negatives (FN) at very low levels, indicating excellent overall classification performance and satisfying the safety-oriented requirement in fire detection that controlling missed detections is critical while false alarms are acceptable to some extent. The clear score boundary, where TP samples are concentrated above 0.8 and TN samples are concentrated below 0.2, further confirms high classification confidence and a stable decision boundary. As shown in Figure 7, false-positive cases are mainly sunset glow, red vehicles, and strong illumination scenes that are highly similar to real flames in visual appearance. This suggests that the model assigns relatively high response weights to color and brightness cues during feature extraction, making it difficult to distinguish semantic differences between flame-like interference and true fire, which is the primary cause of false alarms. In contrast, false-negative cases are concentrated in small-scale, low-light, and occluded flame scenarios. For smoke-only scenes without clearly visible flames, the current flame-centric visual cues are often insufficient, so such samples are more likely to be predicted as non-fire. Nighttime fires and very small fire sources also remain challenging because weak illumination and limited flame pixels reduce discriminative feature responses. For these extreme conditions, the model is less capable of capturing weak discriminative cues and cannot reliably extract subtle flame features from complex backgrounds, which leads to missed detections. In summary, both the number of FP/FN samples and their error distributions are limited and pattern-consistent, with no evidence of large-scale random misclassification, indicating strong classification stability and good generalization on the FD test set. While substantially reducing computation and improving inference speed, the PConv-based design still preserves core fire-discriminative features, resulting in a low miss-detection rate of only 1.68% and high-confidence correct recognition for most fire samples, which validates its effectiveness and rationality for fire-feature extraction. Meanwhile, the 5.52% false-alarm cases are all explainable flame-like interference samples, rather than random errors.

5. Conclusions

In this paper, we investigated a common limitation of deep learning models for forest fire detection, namely, slow inference speed. Our analysis indicates that a major cause is the frequent memory access introduced by DWConv/GConv operators. To address this issue, we introduced PConv as the basic operator and proposed FasterNetFire. Experimental results on the highly challenging FD dataset show that FasterNetFire achieves the best trade-off between detection accuracy, model size, and inference speed. Considering that about 70% of false alarms in the test set come from flame-like samples (e.g., sunset and red clouds), these interference samples are highly similar to real flames in shallow cues such as color, texture, and brightness. Future work will reduce false alarms by integrating temporal and complementary spectral/modal cues, introducing spatial-attention mechanisms, and exploring privacy-preserving cross-site collaborative training [60,61]. This is expected to improve practical deployment in forest watchtowers, UAV patrol systems, and edge-camera early-warning pipelines while reducing environmental and ecological losses through earlier fire discovery. In practical emergency-management workflows, this can support earlier dispatch decisions and reduce unnecessary patrol and suppression costs caused by delayed or unstable alarms. From an environmental perspective, improving early-stage detection reliability can also help limit secondary emissions and post-fire restoration burdens by preventing small ignitions from escalating into large-scale fires.

Author Contributions

Conceptualization, G.C. and A.T.; methodology, G.C. and X.Z.; software, L.J.; formal analysis, G.C. and A.T.; investigation, L.J.; data curation, L.M.; writing—original draft preparation, G.C.; writing—review and editing, X.Z. and W.D.; supervision X.Z. funding acquisition X.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported in part by the National Natural Science Foundation of China under Grant 42271387, in part by the General Project of Sichuan Provincial Science and Technology Education Joint Fund under Grant 2024NSFSC1987, in part by the Science and Technology Plan Project of Longquanyi District in Chengdu under Grant 2024LQRD0049.

Data Availability Statement

We provide the source codes for download at the link: https://github.com/gschen/FasterNetFire (accessed on 27 May 2026).

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
CNNConvolutional Neural Network
PConvPartial Convolution
GConvGroup Convolution
DWConvDepthWise Convolution
PWConvPointWise Convolution
FLOPsFoating-point Operations
FPSFrame per Second
FPFalse Positive
FNFalse Negative
TPTrue Positive
TNTrue Negative
SOTAState-Of-The-Art
GPUGraphics Processing Unit
CPUCentral Processing Unit

References

  1. Jolly, W.M.; Cochrane, M.A.; Freeborn, P.H.; Holden, Z.A.; Brown, T.J.; Williamson, G.J.; Bowman, D.M.J.S. Climate-induced variations in global wildfire danger from 1979 to 2013. Nat. Commun. 2015, 6, 7537. [Google Scholar] [CrossRef] [Scilit]
  2. Bowman, D.M.J.S.; Balch, J.K.; Artaxo, P.; Bond, W.J.; Carlson, J.M.; Cochrane, M.A.; D’Antonio, C.M.; DeFries, R.S.; Doyle, J.C.; Harrison, S.P.; et al. Fire in the Earth System. Science 2009, 324, 481–484. [Google Scholar] [CrossRef] [Scilit]
  3. Johnston, F.H.; Henderson, S.B.; Chen, Y.; Randerson, J.T.; Marlier, M.; DeFries, R.S.; Kinney, P.; Bowman, D.M.J.S.; Brauer, M. Estimated global mortality attributable to smoke from landscape fires. Environ. Health Perspect. 2012, 120, 695–701. [Google Scholar] [CrossRef] [Scilit]
  4. Chen, T.H.; Wu, P.H.; Chiou, Y.C. An early fire-detection method based on image processing. In Proceedings of the 2004 International Conference on Image Processing, 2004. ICIP ’04; IEEE: Piscataway, NJ, USA, 2004; Volume 3, pp. 1707–1710. [Google Scholar] [CrossRef] [Scilit]
  5. Celik, T.; Demirel, H. Fire detection in video sequences using a generic color model. Fire Saf. J. 2009, 44, 147–158. [Google Scholar] [CrossRef] [Scilit]
  6. Foggia, P.; Saggese, A.; Vento, M. Real-time fire detection for video-surveillance applications using a combination of experts based on color, shape, and motion. IEEE Trans. Circuits Syst. Video Technol. 2015, 25, 1545–1556. [Google Scholar] [CrossRef] [Scilit]
  7. Healey, G.; Slater, D.; Lin, T.; Drda, B.; Goedeke, A.D. A system for real-time fire detection. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 1993; pp. 605–606. [Google Scholar] [CrossRef] [Scilit]
  8. Celik, T.; Ozkaramanli, H.; Demirel, H. Fire Pixel Classification using Fuzzy Logic and Statistical Color Model. In Proceedings of the 2007 IEEE International Conference on Acoustics, Speech and Signal Processing—ICASSP ’07; IEEE: Piscataway, NJ, USA, 2007; Volume 1, pp. I-1205–I-1208. [Google Scholar] [CrossRef] [Scilit]
  9. Angayarkkani, K.; Radhakrishnan, N. Efficient forest fire detection system: A spatial data mining and image processing based approach. Int. J. Comput. Sci. Netw. Secur. (IJCSNS) 2009, 9, 100–107. [Google Scholar]
  10. Muhammad, K.; Ahmad, J.; Baik, S.W. Early fire detection using convolutional neural networks during surveillance for effective disaster management. Neurocomputing 2018, 288, 30–42. [Google Scholar] [CrossRef] [Scilit]
  11. Muhammad, K.; Ahmad, J.; Lv, Z.; Bellavista, P.; Yang, P.; Baik, S.W. Efficient Deep CNN-Based fire detection and localization in video surveillance applications. IEEE Trans. Syst. Man Cybern. Syst. 2018, 49, 1419–1434. [Google Scholar] [CrossRef] [Scilit]
  12. Sharma, J.; Granmo, O.C.; Goodwin, M.; Fidje, J.T. Deep convolutional neural networks for fire detection in images. In Proceedings of the Engineering Applications of Neural Networks: 18th International Conference, EANN 2017, Athens, Greece, 25–27 August 2017; Springer: Berlin/Heidelberg, Germany, 2017; pp. 183–193. [Google Scholar] [CrossRef] [Scilit]
  13. Muhammad, K.; Ahmad, J.; Mehmood, I.; Rho, S.; Baik, S.W. Convolutional neural networks based fire detection in surveillance videos. IEEE Access 2018, 6, 18174–18183. [Google Scholar] [CrossRef] [Scilit]
  14. Muhammad, K.; Khan, S.; Elhoseny, M.; Ahmed, S.H.; Baik, S.W. Efficient fire detection for uncertain surveillance environment. IEEE Trans. Ind. Inform. 2019, 15, 3113–3122. [Google Scholar] [CrossRef] [Scilit]
  15. Li, S.; Yan, Q.; Liu, P. An efficient fire detection method based on multiscale feature extraction, implicit deep supervision and channel attention mechanism. IEEE Trans. Image Process. 2020, 29, 8467–8475. [Google Scholar] [CrossRef] [Scilit]
  16. Chen, G.S.; Chandarasupsang, T.; Luo, X.D.; Tananchana, A.; Mu, L. IEFDNet: Detection of Fire Using Partial Convolution Inception. In Proceedings of the 2024 4th International Conference on Big Data Engineering and Education (BDEE); IEEE: Piscataway, NJ, USA, 2024; pp. 32–37. [Google Scholar] [CrossRef] [Scilit]
  17. Yar, H.; Hussain, T.; Agarwal, M.; Khan, Z.A.; Gupta, S.K.; Baik, S.W. Optimized dual fire attention network and medium-scale fire classification benchmark. IEEE Trans. Image Process. 2022, 31, 6331–6343. [Google Scholar] [CrossRef] [Scilit]
  18. Krizhevsky, A.; Sutskever, I.; Hinton, G.E. ImageNet classification with deep convolutional neural networks. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2012; Volume 2, pp. 1097–1105. [Google Scholar]
  19. Sifre, L.; Mallat, S. Rigid-Motion Scattering for Texture Classification. arXiv 2014, arXiv:1403.1687. [Google Scholar] [CrossRef] [Scilit]
  20. Ran, X.; Suyaroj, N.; Tepsan, W.; Lei, M.; Ma, H.; Zhou, X.; Deng, W. A Novel Fuzzy System-Based Genetic Algorithm for Trajectory Segment Generation in Urban Global Positioning System. J. Adv. Res. 2025, 81, 469–480. [Google Scholar] [CrossRef] [Scilit]
  21. Xing, X.; Ling, J.; Liu, S.; Tao, Z. Physics-Informed Neural Network for Thermal Property Inversion of Airport Pavement Multilayer Materials under Icing Conditions. Constr. Build. Mater. 2026, 522, 146164. [Google Scholar] [CrossRef] [Scilit]
  22. Hou, T.; Liu, S.; Mao, W.; Zhao, J.; Ling, J.; Xing, X. Stochastic Vibration Analysis of Aircraft-Rigid Pavement System under Random Aircraft Parameters. Int. J. Pavement Eng. 2026, 27, 2666268. [Google Scholar] [CrossRef] [Scilit]
  23. Zhao, H.; Li, L.; Deng, W. Multi-UAV Path Planning Using Improved Artificial Hummingbird Algorithm Based on Differential Evolution and Gradient Descent. IEEE Trans. Consum. Electron. 2025, 72, 558–569. [Google Scholar] [CrossRef] [Scilit]
  24. Deng, W.; Li, K.; Zhao, H. A Flight Arrival Time Prediction Method Based on Cluster Clustering-Based Modular with Deep Neural Network. IEEE Trans. Intell. Transp. Syst. 2024, 25, 6238–6247. [Google Scholar] [CrossRef] [Scilit]
  25. Li, M.; Chen, Y.; Gao, J.; Dai, Z.; Li, W. Automatic Depression Level Prediction Based on Facial Video with Emotional Cues Using a Hybrid Machine Learning Framework. IEEE Trans. Comput. Soc. Syst. 2025, 1–15. [Google Scholar] [CrossRef] [Scilit]
  26. Chen, J.; Kao, S.; He, H.; Zhuo, W.; Wen, S.; Li, C.L.; Chen, M.S. Run, Don’t Walk: Chasing Higher FLOPS for Faster Neural Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 12021–12031. [Google Scholar] [CrossRef] [Scilit]
  27. Toreyin, B.U.; Dedeoglu, Y.; Cetin, A.E. Flame detection in video using hidden Markov models. In Proceedings of the IEEE International Conference on Image Processing 2005, Genova, Italy, 14–14 September 2005; IEEE: Piscataway, NJ, USA, 2005; Volume 2, p. II-1230. [Google Scholar] [CrossRef] [Scilit]
  28. Lee, C.; Lin, C.; Hong, C.; Su, M. Smoke detection using spatial and temporal analyses. Int. J. Innov. Comput. Inf. Control 2012, 8, 1–11. [Google Scholar]
  29. Han, X.F.; Jin, J.S.; Wang, M.J.; Jiang, W.; Gao, L.; Xiao, L.P. Video fire detection based on Gaussian mixture model and multi-color features. Signal Image Video Process. 2017, 11, 1419–1425. [Google Scholar] [CrossRef] [Scilit]
  30. Chen, J.; He, Y.; Wang, J. Multi-feature fusion based fast video flame detection. Build. Environ. 2010, 45, 1113–1122. [Google Scholar] [CrossRef] [Scilit]
  31. Ha, C.; Hwang, U.; Jeon, G.; Cho, J.; Jeong, J. Vision-based fire detection algorithm using optical flow. In Proceedings of the 2012 Sixth International Conference on Complex, Intelligent, and Software Intensive Systems, Palermo, Italy, 4–6 July 2012; IEEE: Piscataway, NJ, USA, 2012; pp. 526–530. [Google Scholar] [CrossRef] [Scilit]
  32. Mueller, M.; Karasev, P.; Kolesov, I.; Tannenbaum, A. Optical flow estimation for flame detection in videos. IEEE Trans. Image Process. 2013, 22, 2786–2797. [Google Scholar] [CrossRef] [Scilit]
  33. Gubbi, J.; Marusic, S.; Palaniswami, M. Smoke detection in video using wavelets and support vector machines. Fire Saf. J. 2009, 44, 1110–1115. [Google Scholar] [CrossRef] [Scilit]
  34. Prema, C.E.; Vinsley, S.; Suresh, S. Multi feature analysis of smoke in YUV color space for early forest fire detection. Fire Technol. 2016, 52, 1319–1342. [Google Scholar] [CrossRef] [Scilit]
  35. Ko, B.C.; Cheong, K.H.; Nam, J.Y. Fire detection based on vision sensor and support vector machines. Fire Saf. J. 2009, 44, 322–329. [Google Scholar] [CrossRef] [Scilit]
  36. Habiboğlu, Y.H.; Günay, O.; Çetin, A.E. Covariance matrix-based fire and flame detection method in video. Mach. Vis. Appl. 2012, 23, 1103–1113. [Google Scholar] [CrossRef] [Scilit]
  37. Truong, T.X.; Kim, J.M. Fire flame detection in video sequences using multi-stage pattern recognition techniques. Eng. Appl. Artif. Intell. 2012, 25, 1365–1372. [Google Scholar] [CrossRef] [Scilit]
  38. Yuan, F. A double mapping framework for extraction of shape-invariant features based on multi-scale partitions with AdaBoost for video smoke detection. Pattern Recognit. 2012, 45, 4326–4336. [Google Scholar] [CrossRef] [Scilit]
  39. Zhang, D.; Han, S.; Zhao, J.; Zhang, Z.; Qu, C.; Ke, Y.; Chen, X. Image based forest fire detection using dynamic characteristics with artificial neural networks. In Proceedings of the 2009 International Joint Conference on Artificial Intelligence, Haikou, China, 25–26 April 2009; IEEE: Piscataway, NJ, USA, 2009; pp. 290–293. [Google Scholar] [CrossRef] [Scilit]
  40. Frizzi, S.; Kaabi, R.; Bouchouicha, M.; Ginoux, J.M.; Moreau, E.; Fnaiech, F. Convolutional neural network for video fire and smoke detection. In Proceedings of the IECON 2016—42nd Annual Conference of the IEEE Industrial Electronics Society, Florence, Italy, 23–26 October 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 877–882. [Google Scholar] [CrossRef] [Scilit]
  41. Mao, W.; Wang, W.; Dou, Z.; Li, Y. Fire recognition based on multi-channel convolutional neural network. Fire Technol. 2018, 54, 531–554. [Google Scholar] [CrossRef] [Scilit]
  42. Lecun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-based learning applied to document recognition. Proc. IEEE 1998, 86, 2278–2324. [Google Scholar] [CrossRef] [Scilit]
  43. Sun, X.; Sun, L.; Huang, Y. Forest fire smoke recognition based on convolutional neural network. J. For. Res. 2021, 32, 1921–1927. [Google Scholar] [CrossRef] [Scilit]
  44. Lee, W.; Kim, S.; Lee, Y.T.; Lee, H.W.; Choi, M. Deep neural networks for wild fire detection with unmanned aerial vehicle. In Proceedings of the 2017 IEEE International Conference on Consumer Electronics (ICCE), Las Vegas, NV, USA, 8–10 January 2017; IEEE: Piscataway, NJ, USA, 2017; pp. 252–253. [Google Scholar] [CrossRef] [Scilit]
  45. Simonyan, K.; Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv 2014, arXiv:1409.1556v6. [Google Scholar] [CrossRef] [Scilit]
  46. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 770–778. [Google Scholar] [CrossRef] [Scilit]
  47. Dunnings, A.J.; Breckon, T.P. Experimentally defined convolutional neural network architecture variants for non-temporal real-time fire detection. In Proceedings of the 2018 25th IEEE International Conference on Image Processing (ICIP), Athens, Greece, 7–10 October 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 1558–1562. [Google Scholar] [CrossRef] [Scilit]
  48. Szegedy, C.; Liu, W.; Jia, Y.; Sermanet, P.; Reed, S.; Anguelov, D.; Erhan, D.; Vanhoucke, V.; Rabinovich, A. Going Deeper with Convolutions. In Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; IEEE: Piscataway, NJ, USA, 2015; pp. 1–9. [Google Scholar] [CrossRef] [Scilit]
  49. Muhammad, K.; Rodrigues, J.J.P.C.; Kozlov, S.; Piccialli, F.; de Albuquerque, V.H.C. Energy-efficient monitoring of fire scenes for intelligent networks. IEEE Netw. 2020, 34, 108–115. [Google Scholar] [CrossRef] [Scilit]
  50. Feng, J.; Sun, Y. Multiscale network based on feature fusion for fire disaster detection in complex scenes. Expert Syst. Appl. 2024, 240, 122494. [Google Scholar] [CrossRef] [Scilit]
  51. Powers, D.M.W. Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation. Int. J. Mach. Learn. Technol. 2011, 2, 37–63. [Google Scholar]
  52. van Rijsbergen, C.J. Information Retrieval; Butterworth-Heinemann: London, UK, 1979. [Google Scholar]
  53. Hanley, J.A.; McNeil, B.J. The Meaning and Use of the Area under a Receiver Operating Characteristic (ROC) Curve. Radiology 1982, 143, 29–36. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  54. Fawcett, T. An introduction to ROC analysis. Pattern Recognit. Lett. 2006, 27, 861–874. [Google Scholar] [CrossRef] [Scilit]
  55. Khan, S.; Muhammad, K.; Mumtaz, S.; Baik, S.W.; de Albuquerque, V.H.C. Energy-efficient deep CNN for smoke detection in foggy IoT environment. IEEE Internet Things J. 2019, 6, 9237–9245. [Google Scholar] [CrossRef] [Scilit]
  56. Wang, H.; Pan, Z.; Zhang, Z.; Song, H.; Zhang, S.; Zhang, J. Deep Learning Based Fire Detection System for Surveillance Videos. In Proceedings of the Intelligent Robotics and Applications(ICIRA); Springer: Berlin/Heidelberg, Germany, 2019; pp. 318–328. [Google Scholar] [CrossRef] [Scilit]
  57. Zhang, Q.; Xu, J.; Xu, L.; Guo, H. Deep convolutional neural networks for forest fire detection. In Proceedings of the International Forum on Management Education and Information Technology Application (IFMEITA), Guangzhou, China, 30–31 January 2016; pp. 568–575. [Google Scholar] [CrossRef] [Scilit]
  58. Shahid, M.; Hua, K.l. Fire detection using transformer network. In Proceedings of the 2021 International Conference on Multimedia Retrieval (ICMR); Association for Computing Machinery: New York, NY, USA, 2021; pp. 627–630. [Google Scholar] [CrossRef] [Scilit]
  59. Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.C. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR), Salt Lake City, UT, USA, 18–23 June 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 4510–4520. [Google Scholar] [CrossRef] [Scilit]
  60. Deng, W.; Li, X.; Sun, Y.; Zhao, H. Privacy Protection-Enhanced Vertical-Horizontal Federated Learning Secure Sharing for Multisource Heterogeneous Data. IEEE Trans. Ind. Inform. 2026, 22, 3138–3147. [Google Scholar] [CrossRef] [Scilit]
  61. Li, X.; Zhao, H.; Xu, J.; Zhu, G.; Deng, W. APDPFL: Anti-Poisoning Attack Decentralized Privacy Enhanced Federated Learning Scheme for Flight Operation Data Sharing. IEEE Trans. Wirel. Commun. 2024, 23, 19098–19109. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Frame per second (FPS) under varied accuracy on the FD dataset. Many existing neural networks for forest fire detection suffer from low inference speed issues due to frequent memory access, most of them are lower than 150 FPS. Our FasterNetFire, a convolutional neural network with partial convolution, obtains the best tradeoff between accuracy and latency, it is 4.5× and 4× faster than EFDNet and DFAN.
Figure 1. Frame per second (FPS) under varied accuracy on the FD dataset. Many existing neural networks for forest fire detection suffer from low inference speed issues due to frequent memory access, most of them are lower than 150 FPS. Our FasterNetFire, a convolutional neural network with partial convolution, obtains the best tradeoff between accuracy and latency, it is 4.5× and 4× faster than EFDNet and DFAN.
Forests 17 00672 g001
Figure 2. Overall architecture of our FasterNetFire. It has four hierarchical stages, each with a stack of FasterNet blocks and preceded by an embedding or merging layer. Within each FasterNet block, a PConv layer is followed by two PWConv layers. The stage depths are configured as [ 1 , 2 , 8 , 2 ] , and each stage transition uses a 2 × 2 stride/2 merging layer after the initial 4 × 4 stride/4 embedding. Both the embedding layer and merging layer are at the beginning of each stage to downscale. The * symbol denotes convolution operation.
Figure 2. Overall architecture of our FasterNetFire. It has four hierarchical stages, each with a stack of FasterNet blocks and preceded by an embedding or merging layer. Within each FasterNet block, a PConv layer is followed by two PWConv layers. The stage depths are configured as [ 1 , 2 , 8 , 2 ] , and each stage transition uses a 2 × 2 stride/2 merging layer after the initial 4 × 4 stride/4 embedding. Both the embedding layer and merging layer are at the beginning of each stage to downscale. The * symbol denotes convolution operation.
Forests 17 00672 g002
Figure 3. Some representative fire images and fire-like images from the FD dataset. (a) The fire images. (b) The neural and fire-like images.
Figure 3. Some representative fire images and fire-like images from the FD dataset. (a) The fire images. (b) The neural and fire-like images.
Forests 17 00672 g003
Figure 4. Radial-bar comparison of the accuracy, FLOPs, model size, and GPU/CPU inference speed of CNN-based models on the FD dataset.
Figure 4. Radial-bar comparison of the accuracy, FLOPs, model size, and GPU/CPU inference speed of CNN-based models on the FD dataset.
Forests 17 00672 g004
Figure 5. Confusion matrix of our FasterNetFire on the FD dataset. The number of TP is 2459, TN is 2361, FP is 138, and FN is 42.
Figure 5. Confusion matrix of our FasterNetFire on the FD dataset. The number of TP is 2459, TN is 2361, FP is 138, and FN is 42.
Forests 17 00672 g005
Figure 6. Visualized confusion matrix of our FasterNetFire on the FD dataset. The horizontal axis denotes the image index and the vertical axis denotes the fire-class probability score from the softmax output. The prediction scores of TP images are close to 1, as well as the TN are to 0. Most of the prediction scores of FP are very high, which indicates these images are very challenging to classify to the model.
Figure 6. Visualized confusion matrix of our FasterNetFire on the FD dataset. The horizontal axis denotes the image index and the vertical axis denotes the fire-class probability score from the softmax output. The prediction scores of TP images are close to 1, as well as the TN are to 0. Most of the prediction scores of FP are very high, which indicates these images are very challenging to classify to the model.
Forests 17 00672 g006
Figure 7. Representative FP and FN images with prediction scores from the FD test set. The first three rows are FP images and 70% of them are from sunset, burning clouds. The last row is FN images and most of them are the fire at a long distance or blocked by heavy smoke.
Figure 7. Representative FP and FN images with prediction scores from the FD test set. The first three rows are FP images and 70% of them are from sunset, burning clouds. The last row is FN images and most of them are the fire at a long distance or blocked by heavy smoke.
Forests 17 00672 g007
Table 1. The detailed statistics of training, validation and testing set of the FD dataset.
Table 1. The detailed statistics of training, validation and testing set of the FD dataset.
DatasetFireNon FireTotal
Training set17,50017,50035,000
Validation set5000500010,000
Testing set250025005000
Total25,00025,00050,000
Table 2. Experimental software, hardware environments and training hyperparameters.
Table 2. Experimental software, hardware environments and training hyperparameters.
NameValue
Operating systemUbuntu 22.04.3 LTS
Programming languagePython 3.8.10
Deep learning frameworkPyTorch 1.11.0
CUDA Version12.2
GPUNVIDIA A100-SXM4-80 GB
CPUIntel Xeon Gold 6330
RAM80 GB
Video memory12 GB
OptimizerAdam
Initial LR 5 × 10 4
Weight decay 1 × 10 5
LR scheduleCosine annealing
Training epochs300
Table 3. The outline of our FasterNetFire architecture.
Table 3. The outline of our FasterNetFire architecture.
TypePatch Size/Stride or RemarksInput Size
Embedding 4 × 4 / 4 224 × 224 × 3
1 × FasterNet Block 3 × 3 / 1 56 × 56 × 64
Merging 2 × 2 / 2 56 × 56 × 64
2 × FasterNet Block 3 × 3 / 1 28 × 28 × 128
Merging 2 × 2 / 2 28 × 28 × 128
8 × FasterNet Block 3 × 3 / 1 14 × 14 × 256
Merging 2 × 2 / 2 14 × 14 × 256
2 × FasterNet Block 3 × 3 / 1 7 × 7 × 512
AdaptiveAvgPool2d 1 × 1 7 × 7 × 512
Flatten 1 × 1 × 1280
1 × 1 Conv 1 × 1 / 1 1 × 1 × 1280
LinearLogits 1 × 1 × 1280
SoftmaxClassifier 1 × 1 × 2
Table 4. Comparison of the performance of state-of-the-art methods on the FD and Foggia’s datasets. The best and second-best methods are highlighted in bold and underline.
Table 4. Comparison of the performance of state-of-the-art methods on the FD and Foggia’s datasets. The best and second-best methods are highlighted in bold and underline.
NetworksFD DatasetFoggia’s Dataset
PrecisionRecallF1-ScoreAccuracyFPFNACC
FD-GCM [5]63.9090.0074.7069.6029.41083.87
FFD-ANN [39]71.1073.2072.1071.70...
FPC [8]52.0099.9068.4053.90...
SqueezeNet [55]90.5293.6292.0391.90...
ANetFire [10]83.3093.2087.9087.209.072.1394.39
CNNFire [11]84.6091.3087.9087.308.872.1294.50
GoogLeNet [13]88.0098.0092.7392.32...
MobileNet [56]91.4593.7192.5392.40...
GNetFire [13]88.0098.0092.8092.300.0541.5094.43
EMNFire [14]88.3098.7093.2092.8000.1495.86
ResNetFire [12]84.8097.6090.0090.07...
ResNet-50 [57]84.8297.5790.7490.00...
EFDNet [15]93.5097.4095.4095.30...
DFAN [17]96.0097.0096.0096.1700.5899.60
DFAN_Comp [17]95.5096.3095.9095.7000.6399.47
IEFDNet [16]92.9397.8095.3095.18...
ViT-B/32 [58]92.3692.1892.1892.192.151.0294.03
FasterNetFire94.6998.3296.4796.18 ± 0.500.12099.66 ± 0.22
Table 5. Comparison of time complexity, model size and FLOPs on the FD dataset. The best and second-best methods are highlighted in bold and underline.
Table 5. Comparison of time complexity, model size and FLOPs on the FD dataset. The best and second-best methods are highlighted in bold and underline.
ModelAccuracyFLOPs (G)Size (MB)GPU (FPS)CPU (FPS)Pi (FPS)
ANetFire [10]87.20..382.018.1.
CNNFire [14]87.30..135.914.2.
GNetFire [13]92.301.543.3048.24.3.
EMNFire [14]92.800.313.0061.22.4.
ResNetFire [12]90.073.898.0057.32.4.
EFDNet [15]95.301.134.8063.53.0.
DFAN [17]96.170.14180.6370.5512.900.83
DFAN_Comp [17]95.700.07341.09125.3322.733.21
IEFDNet [16]95.181.112.69118.804.90.
FasterNetFire96.18 ± 0.500.856.32290.5438.118.44
Table 6. The performance of different partial ratio r of PConv on the FD dataset. The best and second-best methods are highlighted in bold and underline.
Table 6. The performance of different partial ratio r of PConv on the FD dataset. The best and second-best methods are highlighted in bold and underline.
Partial Ratio rPrecisionRecallF1-ScoreAccuracy
189.1590.4089.7789.70
1 / 2 89.5890.4490.0189.92
1 / 4 94.6998.3296.4796.40
1 / 8 87.4986.3386.9186.67
1 / 16 86.3784.5085.4285.11
1 / 32 83.3080.7782.0281.57
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Chen, G.; Tananchana, A.; Jiang, L.; Zhou, X.; Mu, L.; Deng, W. FasterNetFire: A Cost-Effective Fast Neural Network for Forest Fire Detection with Partial Convolution. Forests 2026, 17, 672. https://doi.org/10.3390/f17060672

AMA Style

Chen G, Tananchana A, Jiang L, Zhou X, Mu L, Deng W. FasterNetFire: A Cost-Effective Fast Neural Network for Forest Fire Detection with Partial Convolution. Forests. 2026; 17(6):672. https://doi.org/10.3390/f17060672

Chicago/Turabian Style

Chen, Gongsuo, Annop Tananchana, Laihong Jiang, Xiangbing Zhou, Lei Mu, and Wu Deng. 2026. "FasterNetFire: A Cost-Effective Fast Neural Network for Forest Fire Detection with Partial Convolution" Forests 17, no. 6: 672. https://doi.org/10.3390/f17060672

APA Style

Chen, G., Tananchana, A., Jiang, L., Zhou, X., Mu, L., & Deng, W. (2026). FasterNetFire: A Cost-Effective Fast Neural Network for Forest Fire Detection with Partial Convolution. Forests, 17(6), 672. https://doi.org/10.3390/f17060672

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop