Next Article in Journal
Effect of Height on the Combustion Rate and Flame of Two Circular Isooctane Pools
Previous Article in Journal
Fire Performance and Leaching Resistance of Poplar Wood Treated with Tannin-Diammonium Phosphate and Ammonia Fuming
Previous Article in Special Issue
Real-Time Tiny Fire-Spot Detection in Farmland Scenes Based on an Improved YOLOv11
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

MFFLNet: Multi-Scale Fire Feature Learning for Fire Video Recognition

by
Shanzheng Yang
1 and
Yun Yi
1,2,*
1
School of Mathematics and Computer Science, Gannan Normal University, Ganzhou 341000, China
2
Key Laboratory of Data Science and Artificial Intelligence of Jiangxi Education Institutes, Gannan Normal University, Ganzhou 341000, China
*
Author to whom correspondence should be addressed.
Fire 2026, 9(10), 427; https://doi.org/10.3390/fire9100427
Submission received: 18 July 2026 / Revised: 19 September 2026 / Accepted: 21 September 2026 / Published: 1 October 2026

Abstract

Accurate and timely recognition of flame and smoke is critical for fire warning systems to mitigate casualties and property damage. Most existing fire datasets are image-based and thus fail to capture the dynamic spatiotemporal information of fire. Furthermore, the lack of large-scale fire video datasets poses a significant challenge to the training of neural networks. To address these limitations, we developed the Flame-Smoke Video Recognition (FSVR) dataset, a large-scale collection of 30,000 video clips that significantly exceeds the scale of prior datasets in the same domain. Since flame and smoke exhibit multiscale characteristics, small-scale targets are often obscured by complex backgrounds. Existing methods lack robust multiscale feature learning capabilities, which limits their ability to achieve precise fire recognition under complex conditions. To address this issue, we proposed the Multi-scale Fire Feature Learning Network (MFFLNet), which integrates multiple Multi-scale Fire Feature Learning (MFFL) blocks into a Transformer backbone. Each MFFL block comprises two key components, i.e., the multiscale fire Conv3D layer and the fire spatiotemporal feature learning layer. The experimental results obtained from the FSVR and LFVR datasets demonstrate that MFFLNet surpasses the baseline model and other comparative methods. When the backbone network is initialized with pre-trained weights from the Kinetics-710 dataset, MFFLNet attained an accuracy of 79.54% and a macro-F1 score of 79.22% on the FSVR dataset, while achieving an accuracy of 95.64% and an F1 score of 95.46% on the LFVR dataset.

1. Introduction

Fire incidents lead to loss of life and property globally [1,2]. Vision-based fire warning has emerged as a significant research direction [3,4]. Sensors and cross-technology communication are also research directions [5,6,7]. Existing public fire datasets are predominantly image-based. Consequently, they lack the dynamic spatiotemporal characteristics of flame and smoke. This scarcity of video data hinders the training of networks. Consequently, the construction of a large-scale, high-quality fire video dataset is of substantial research significance.
The multiscale characteristic of flame and smoke in real-world scenarios presents significant challenges for fire recognition. While large flame and smoke are relatively easy to identify, small ones are often concealed in complex backgrounds and exhibit a similarity to the background textures and colors, thus posing challenges for fire recognition. However, most previous methods lack the ability to effectively learn the multiscale features of fire. Enhancing the model’s ability to accurately recognize multiscale flame and smoke remains a challenge in the field of Fire Video Recognition (FVR).
To address the scarcity of public fire video datasets, we constructed the large-scale Flame-Smoke Video Recognition (FSVR) dataset. This dataset encompasses diverse real-world fire scenarios, and provides ample training data for deep learning models. To overcome the limitations in multiscale feature learning, we proposed the Multi-scale Fire Feature Learning Network (MFFLNet), which integrates multiple Multi-scale Fire Feature Learning (MFFL) blocks within a Vision Transformer (ViT) [8] backbone. This design effectively enhances the learning of multiscale fire features and improves the recognition accuracy of flame and smoke. Each MFFL block comprises two key components: a Multi-scale Fire Conv3D (MFConv3D) layer and a Fire Spatio-Temporal Feature Learning (FSTFL) layer. By leveraging multiple 3D convolutions with varying kernel sizes and a maximum fusion operation, MFConv3D fuses fire features across different scales. Through operations including 3D convolution, average pooling, and channel compression, FSTFL learns the spatiotemporal features of fire. Experimental results demonstrate that MFFLNet outperforms the baseline and other methods on both the FSVR and LFVR [9] datasets. The main contributions are as follows.
  • We constructed the FSVR dataset, comprising 30,000 clips across diverse real-world scenarios. This dataset provides substantial data support for FVR, which is conducive to advancing research in this domain.
  • MFFL was proposed, which is composed of MFConv3D and FSTFL. By innovatively integrating the 3D convolution, average pooling, channel compression and maximum fusion, it enhances the multiscale fire feature learning capability of the neural network.
  • MFFLNet was designed, which integrates the MFFL blocks within a ViT backbone. By learning the multiscale features of fire, MFFLNet attains state-of-the-art performance on two public datasets.

2. Related Work

Notable advancements have been achieved in the development of fire recognition datasets and fire-warning systems [10]. This section focuses primarily on research related to fire recognition datasets and fire recognition methods.

2.1. Fire Recognition Datasets

The construction of high-quality datasets is essential for advancing the fire warning system [11]. Fire incidents in real-world scenarios are diverse, making it challenging to comprehensively represent them using a single data source. Consequently, recently released public fire recognition datasets have provided robust support for algorithm training and validation. Table 1 presents several datasets in this field.
Flame-only datasets focus on identifying visible flame across various scenarios, including building and forest fires. These datasets often include interference objects, such as sunsets and red-colored items. The Corsican Fire dataset [12] is specifically tailored for wildfire monitoring and comprises 500 wildfire images, 100 near-infrared images, and 5 videos. The training set of FireNet [14] comprises 1124 flame-inclusive images and 1301 non-flame images, whereas the test set consists of 46 flame-inclusive videos, 16 non-flame videos, and 160 non-flame images. LFVR [9] contains 11,560 video clips sourced from the internet and real-world recordings, covering different scales, scenarios, and burning objects, thus exhibiting high diversity. A significant limitation of these datasets is their reliance on flame, which precludes early fire warning.
Smoke-only datasets are designed to facilitate early fire warning. These datasets include smoke at varying concentrations and use visually similar entities as primary interference. The DSDF dataset [15] is specifically designed for smoke recognition, containing 18,413 images. Nemo [17] is a dataset designed for fine-grained fire and smoke recognition, comprising 4347 images and 95 test videos. Smoke-only datasets have theoretical advantages in early fire warning. However, methods that solely rely on smoke detection are associated with relatively high false-alarm rates.
To address the aforementioned issue, researchers have constructed datasets for fire and smoke recognition. These datasets are compiled from diverse sources, such as surveillance videos, drone images, satellite images, etc. The FiSmo dataset [13] comprises an image sublibrary and a video sublibrary, totaling 9448 images and 88 videos. The D-Fire [16] dataset contains 21,527 bounding-box-annotated RGB images and 100 test videos. The DFS [18] dataset subdivides 9462 images by fire intensity into three categories. MS-FSDB [19] contains 12,586 images, including 3603 positive and 8983 negative samples.

2.2. Fire Recognition Methods Based on Handcrafted Features

Color features, as the most intuitive visual representation of flame, were central to early fire recognition methods. By exploring the representation capabilities of various color spaces, researchers achieved localization and classification of flame. For instance, Foggia et al. [20] implemented video flame detection by designing a system that integrates the color distribution, dynamic shape variations, and motion trajectory characteristics of flame.
The detailed analysis of motion features has become a significant research direction. Torabian et al. [21] proposed a motion analysis-based method that distinguishes flame from other moving targets by capturing their motion patterns. To improve the performance in dynamic scenes, Dou et al. [22] focused on the features of motion and color of flame. For fire image classification, Harkat et al. [23] employed machine learning techniques to compute higher-order features. For forest fire detection, Zhang et al. [24] proposed a vector machine.
The multi-modal data fusion technology has enhanced the accuracy of fire recognition. To improve fire detection performance, Sharma et al. [25] fused sensor data with environmental data and aggregated model parameters through a federated learning framework. By utilizing multispectral data, Mfondoum [26] proposed a multi-filter active fire mapping method. Panneerselvam et al. [27] proposed a fire recognition method based on federated learning. Liu et al. [28] proposed a fire classification framework that integrates aspect ratio quantitative analysis and point cloud segmentation.

2.3. Fire Recognition Methods Based on Deep Learning

Deep learning-based fire recognition methods have evolved from single-image classification to comprehensive spatiotemporal frameworks. To meet the requirements of diverse scenarios, extensive research has been conducted on model architecture, multimodal fusion, etc.
Regarding architectural innovation, various advanced architectures have been applied to address the diversity and complexity of fire scenarios. Lightweight design is essential for meeting real-time recognition requirements. Muhammad et al. [29] adopted the lightweight MobileNetV2 network to adapt to computational demands in uncertain environments. YOLO-LFD [30] employs depthwise separable convolutions to reduce computational complexity. To enhance the perception of fire, Shahid et al. [31] integrated attention mechanisms into a convolutional neural network. Similarly, CCi-YOLOv8n [32] optimizes the detection of small-fire targets by integrating CARAFE and context-guided blocks. MTL-FFDET [33] is a fire recognition method based on multi-task learning. Bendouma et al. [34] proposed a two-stage method to address complex background interference. Majid et al. [35] leveraged transfer learning and attention mechanism to enhance the generalization ability. LGAE [9] enhances performance by designing the LAME and GAME blocks to effectively capture local and global motion characteristics of fire.
To overcome the limitations of single-source visual data, researchers have focused on integrating multimodal information. Sheng et al. [36] proposed a method enhanced by SLIC-DBSCAN, utilizing clustering algorithms to pre-identify suspected fire and smoke regions. For wildfire monitoring, Dong et al. [37] developed a framework to predict the spatiotemporal dynamic of fire. Han et al. [38] designed a Transformer-based method that addresses image resolution degradation in forest fire. For fire recognition in industrial environments, Deshpande et al. [39] designed a system utilizing transfer learning that operates stably under low visibility and adverse weather conditions.
In contrast to these methods, MFFLNet is a multiscale feature learning approach for FVR. Through the MFFL block, MFFLNet captures feature information of flame and smoke at different scales and enhances key features, thereby improving the recognition capability for fire.

2.4. Methods Based on Multi-Scale or Attention Mechanisms

Spatiotemporal feature modeling methods for videos can be broadly classified into two categories: multi-scale convolution-based approaches and attention-based approaches. While both categories have yielded notable achievements, they exhibit inherent limitations when applied to the FVR task.
Multi-scale 3D Convolution (M3D) [40] captures spatiotemporal information across varying receptive field scales. Building upon 2D convolutions, M3D utilizes multiple parallel temporal convolutional kernels with different dilation rates to achieve multi-scale feature extraction along the temporal dimension. This method fuses features from different scales via equal-weight summation, lacking an adaptive mechanism for scale selection.
Squeeze-and-Excitation (SE) [41] is a lightweight attention module focused on the channel dimension. It compresses the spatial dimensions of features through global average pooling to aggregate global information, then uses two fully connected layers to learn dependencies among channels, thereby selectively enhancing informative channels. The module has low parameter count and computational complexity. However, SE does not model features in the spatial dimension nor process temporal information, making it suitable only for static image tasks and incapable of multi-scale feature extraction.
Convolutional Block Attention Module (CBAM) [42] extends the concept of channel attention by further incorporating spatial attention. It aggregates spatial information using both average and max pooling, generating channel attention maps via a shared multi-layer perceptron. Spatial attention maps are generated through convolutional layers to perform spatial weighting. Nevertheless, CBAM lacks capabilities for multi-scale feature extraction and time-series modeling.
TimeSformer [43] adopts a decoupled spatiotemporal attention strategy, applying self-attention operations independently along the temporal and spatial axes, thus enabling robust global context modeling for long video sequences. While it achieves a full-image receptive field in the spatial dimension, its representation of fine-grained local information is less precise than that of convolutional operations. In terms of computational complexity, TimeSformer exhibits higher parameter counts and time complexity.
As summarized in Table 2, MFFL effectively addresses these limitations by unifying multi-scale temporal feature extraction and spatiotemporal feature learning within a coherent framework. This design allows MFFL to strike a superior balance between accuracy and computational efficiency, making it particularly suitable for the domain of video fire recognition.

3. FSVR Dataset

As shown in Table 1, previous fire datasets primarily focused on recognizing a single category: either fire or smoke. Recently, datasets designed for simultaneous fire and smoke recognition have been predominantly image-based, with few video samples. To address this gap, we constructed the FSVR dataset, comprising 30,000 video clips. To our knowledge, FSVR is currently the largest FVR dataset.

3.1. Videos

The creation of high-quality video datasets depends on the diversity of data sources and the standardization of processing workflows. To comprehensively capture the characteristics of flame and smoke across various scenarios, original videos were acquired via two methods: downloading from the Internet and on-site recording.
First, 2245 long videos containing fire, smoke, and negative videos, were collected from the Internet, with a cumulative duration of 34.37 h. These videos exhibit a wide range of flame and smoke characteristics. Fire intensity is categorized into three levels, i.e., small, medium, and large. Smoke is captured at various scales, reflecting different stages of fire development. Scenarios include residential areas, factories, and wildlands, with burning objects such as buildings, vehicles, electrical appliances, and forests. To validate and enhance the generalization capability of network models, these collected videos include 200 non-fire videos, which contain easily confusing samples such as fog, haze, etc.
Second, 203 long videos completely devoid of fire were recorded on-site, totaling 4.76 h. These videos, which contain objects visually analogous to fire, are utilized for the construction of negative samples. Sourced from diverse origins, these videos provide training data to assist the model in differentiating real flame from analogous distractors.

3.2. Annotations

The primary characteristics of flame and smoke lie in their dynamic visual variations, including changes in the color, shape, and diffusion patterns of smoke, as well as the flickering and expansion of flame. A 2-s video contains sufficient discriminative features for distinguishing between flame and smoke. The collected videos were segmented into non-overlapping 2-s clips, yielding 70,427 clips. Clips with inconsistent annotations, poor quality, or redundancy were removed, resulting in a final set of 30,000 clips for the FSVR dataset.
The annotator instructions are provided as follows. (1) Annotators shall strictly adhere to unified category judgment criteria. The FSVR dataset comprises three categories, namely Fire, Smoke, and Normal. The Fire category is defined as video clips containing flames, irrespective of the presence of smoke. The Smoke category consists of video clips with smoke but no flames, which provides critical data support for early fire detection. The Normal category refers to clips that contain neither smoke nor flames and exhibit no signs of combustion. (2) Annotators shall precisely identify real targets and interfering pseudo-targets. During annotation, distinguish between flames or smoke and easily confused interfering pseudo-targets. Strictly follow the unified category judgment criteria to avoid mislabeling. (3) Annotators shall standardize the execution of frame-by-frame annotation and sample verification processes. Annotation is conducted in a frame-by-frame manner, with each frame in a clip assigned an independent category label. After annotation, check the consistency of labels across frames within the same clip. If label inconsistencies are detected, the clip shall be removed to ensure the consistency and reliability of each sample’s category in the dataset.
Three categories of video clips were excluded, namely low-quality clips, redundant clips, and inconsistent clips. First, low-quality clips refer to video segments that suffer from poor resolution, excessive noise, shaky footage, or other technical flaws that significantly degrade visual clarity. Second, redundant clips are defined as those whose content shares a high degree of overlap with that of other clips; for instance, consecutive video clips extracted from a long video often exhibit such high repetition. Third, inconsistent clips are defined as clips in which the labels assigned to frames within the same clip are inconsistent; for instance, one frame within a given clip may be labeled fire while another frame in the same clip is labeled smoke. The number of clips removed for these three categories was 15,409, 18,954, and 6064, respectively.
To ensure accuracy and objectivity, three annotators independently labeled the video clips. Following this process, 9900 fire clips were annotated, covering various scenarios and intensities. Additionally, 10,040 smoke-without-fire clips were annotated. Due to the scarcity of early-stage smoke-without-fire samples, all clips that exhibited only smoke without flame were included in this category. Moreover, 10,060 non-fire and non-smoke clips were annotated, including various challenging samples. Figure 1 shows some frames from FSVR.

3.3. Dataset Partition

The FSVR dataset is partitioned into training, validation, and test sets using a 2:2:6 ratio. This partition scheme, which accounts for diverse scenarios and dynamic changes, achieves a better balance across training, validation, and testing. One of the research objectives of this dataset is to comprehensively assess the model’s performance across diverse and complex real-world scenarios by leveraging a large-scale test set. Against the backdrop of large-scale video data, allocating 20% of the samples to training furnishes an adequate volume of training instances, enabling the model to fully capture fire-related features. The 20% sample allocation for validation satisfies the requirements for parameter optimization. Moreover, the 60% sample allocation reserved for testing facilitates a more robust assessment of the model’s performance across diverse and complex real-world scenarios. The statistical data of FSVR are presented in Table 3.
FSVR includes 30,000 video clips. Following the 2:2:6 split ratio, the final quantities of the training, validation, and test sets are determined to be 6000, 6000, and 18,000, respectively. To ensure the source-level independence, clips derived from the same long video are assigned to the identical set.
Regarding class distribution, the category ratios in the training and validation sets remain consistent, each comprising 1980 clips with fire, 2008 clips with smoke, and 2012 clips without fire and smoke. The test set contains 5940 clips with fire, 6024 clips with smoke, and 6036 clips without fire and smoke. All experiments in this study were carried out in accordance with the aforementioned partitioning scheme.

4. Method

4.1. MFFL Block

4.1.1. Architecture of MFFL

To enhance the multiscale feature learning capability for fire, we proposed the MFFL block. This block primarily consists of the MFConv3D layer and the FSTFL layer. Figure 2 illustrates the detailed architecture of MFFL. For 3D Convolution (Conv3D), the default kernel size and stride are both 1 × 1 × 1 , with a padding of 0. Non-default convolution parameters are explicitly noted.
Let X ∈ R ( H ∗ W + 1 ) × ( N ∗ T ) × C be the input of MFFL, where N is the batch size, C represents the number of channels, T indicates the temporal dimension, and H and W denote the height and width of the image, respectively. The class token in X is discarded, followed by a reshape operation to yield the output tensor X i n ∈ R N × C × T × H × W . A Conv3D branch is introduced to perform element-wise addition with the weighted features. This preserves the foundational feature information of the original input while enhancing key spatiotemporal features. The function of MFFL is defined as:
f MFFL ( X i n ) = f FSTFL ( f M F C 3 D ( X i n ) ) + f C 3 D ( X i n , 1 , C ) ,
where f C 3 D ( · , · , · ) is the Conv3D function with the second parameter as kernel size and the third as the number of output filters, f FSTFL ( · ) represents the FSTFL function, and f M F C 3 D ( · ) denotes the MFConv3D function. The class token is recombined via the concatenation operation, followed by the application of a reshape operation to yield an output tensor that matches the shape of the input X . MFConv3D and FSTFL are elaborated in the subsequent sections.

4.1.2. MFConv3D

The flame and smoke in the video exhibit multiscale characteristics. To capture the multiscale information, MFConv3D is designed, which comprises three Fire Conv3D (FConv3D) layers and a fusion operation. FConv3D with large convolutional kernels can learn information about the long-term diffusion and spread of smoke and flame, thereby capturing global temporal patterns. Conversely, FConv3D with small convolutional kernels captures the fine-grained details of local smoke density variations and the short-span temporal features of flame.
Specifically, three parallel FConv3D layers extract key features at multiple scales. As shown in Figure 2, each FConv3D comprises a Batch Normalization (BN) layer and three Conv3D layers. Let the input tensor be X i n ∈ R N × C × T × H × W . To enhance the representational capacity, FConv3D first compresses the channels to ⌊ 2 C 3 ⌋ before restoring them to C. The function of FConv3D is presented as follows.
f F C 3 D ( X i n , k ) = f C 3 D f C 3 D f C 3 D f BN ( X i n ) , 1 , ⌊ 2 C 3 ⌋ , ( k , 1 , 1 ) , ⌊ 2 C 3 ⌋ , 1 , C ,
where the parameter k is the kernel size of Conv3D, and f BN ( · ) denotes the function of BN. As shown in Equation (2), the parameter k in f F C 3 D indicates that the kernel size of the intermediate Conv3D is k × 1 × 1 , with a corresponding padding size of ⌊ k / 2 ⌋ × 1 × 1 , where k ∈ { 3 , 5 , 7 } .
By fusing the outputs of these three FConv3D layers, MFConv3D learns multiscale features of fire. MFConv3D is calculated as follows:
X F = f M F C 3 D ( X i n ) = f Max f Stack k ∈ { 3 , 5 , 7 } f F C 3 D ( X i n , k ) , dim = 0 .
where f Stack ( · ) is the stacking operation, and f Max ( · , · ) performs element-wise maximum fusion across the parallel branches. By employing f Stack , multiscale information is concatenated, followed by fusion along the first dimension via f Max . This improves the model’s capacity to learn fire features across multiple scales.

4.1.3. FSTFL

Following multiscale feature fusion, the FSTFL layer is designed to learn the spatiotemporal features of fire. To suppress background noise and focus on critical fire regions, FSTFL utilizes a Conv3D layer with a 7 × 7 × 7 kernel and 3 × 3 × 3 padding, followed by a Sigmoid function. Long-term temporal information is learned by using a 3D Adaptive Average Pooling (AAP3D) layer, two Conv3D layers, a Rectified Linear Unit (ReLU) layer, and a Sigmoid layer, thereby strengthening the key spatiotemporal features of fire.
Within FSTFL, the number of channels is compressed to one via a large-kernel convolution, thus generating a cross-channel-shared spatiotemporal weight tensor. This tensor captures both large-scale spatial features and the small-scale changes of flame and smoke. By using X F from Equation (3) as the input tensor, the first part of FSTFL is expressed as:
X S = X F ⊙ σ f C 3 D ( X F , 7 , 1 ) ,
where ⊙ denotes element-wise tensor multiplication. By applying element-wise weighted enhancement to X F , X S ∈ R N × C × T × H × W is obtained. This assigns high weights to prominent flame or smoke areas, while reducing the weight of irrelevant regions.
Global spatiotemporal context across multiple frames is aggregated via AAP3D, resulting in a feature tensor ∈ R N × C × 1 × 1 × 1 . Cross-frame long-range dependencies are then established using the Conv3D, ReLU, and Sigmoid functions to perform secondary weighted enhancement. To reduce computational complexity, the number of channels is compressed to ⌊ C / 8 ⌋ and subsequently restored to C. By utilizing X S as the input, the function of FSTFL can be defined as below.
f FSTFL ( X S ) = X S ⊙ σ f C 3 D f Relu f C 3 D f A A P 3 D ( X S ) , 1 , ⌊ C 8 ⌋ , 1 , C ,
where f A A P 3 D ( · ) denotes the function of AdaptiveAvgPool3D, and f Relu ( · ) denotes the ReLU function. An element-wise multiplication is applied to X S to enhance key spatiotemporal features. This process learns fire-related spatiotemporal features while mitigating the interference caused by irrelevant features.

4.2. Architecture of MFFLNet

To enhance the capability of multiscale fire feature learning, we propose MFFLNet, which integrates multiple MFFL blocks within a ViT-B [8] backbone. This backbone comprises 12 blocks. To precisely capture multiscale features of flame and smoke, MFFL is embedded before the final 1 to 4 blocks.
As shown in Figure 3, a Conv3D layer with a kernel size of 1 × 16 × 16 is used for downsampling, yielding a tensor in R N × C × T × H × W . This tensor is then transformed into a tensor in R ( N ∗ T ) × ( H ∗ W + 1 ) × C via adding the class token and positional embedding. Layer Normalization (LN) is applied, followed by a dimension transformation to yield X ∈ R ( H ∗ W + 1 ) × ( N ∗ T ) × C , which is then fed into the backbone. Ultimately, a linear classifier is employed for sample classification.

5. Experiments

5.1. Datasets

The FSVR and LFVR [9] datasets are used for experimentation. As shown in Table 3, FSVR is partitioned into training, validation, and test sets, containing 6000, 6000, and 18,000 video clips, respectively. The model checkpoint corresponding to the best validation performance is utilized for evaluation on the test set. Accuracy (Acc) and the macro-F1 score are utilized as the evaluation metrics for FSVR.
As presented in Table 1, LFVR is divided into 2 categories: Fire and Non-fire. It comprises 11,560 video clips, with approximately 70% used for training and 30% for testing. In line with prior studies, LFVR is partitioned into three distinct training and test sets, and 3-fold cross-validation is utilized to derive final results. F1 and Acc are employed as evaluation metrics for LFVR.

5.2. Experimental Setup

Several preprocessing operations are applied to the video data. During training, T frames are randomly and uniformly sampled from each clip. The shorter side of the frames is resized to 256 pixels, while the longer side is adjusted adaptively. The RandAugment [44] and RandomResizedCrop strategies are employed for data augmentation. Then, the cropped frames are scaled to 224 × 224 pixels, and horizontal flipping is applied with a 50% probability to increase data diversity.
During the validation and testing phases, we employed the UniformSample algorithm to uniformly select a fixed set of T frame indices from the videos, thereby ensuring the reproducibility of the test results. In the validation phase, the video is uniformly partitioned into T segments, where the sampling step within each segment is equal to half the length of the segment. Therefore, each sampled frame is positioned at the center of its corresponding segment, achieving deterministic central uniform sampling across these segments. In the testing phase, the video is also partitioned into T equal-length segments, yet the sampling step within each segment is adjusted to one-third of the segment length, thereby generating two distinct samples. These frames are adjusted to a shorter side of 224 pixels while maintaining the aspect ratio. Finally, center cropping is used to extract 224 × 224 pixel regions. For fair comparison, all baseline methods utilize the same sampling strategy described above. In the experiments, T was set to 8, 16, and 32 based on specific requirements.
AdamW [45] is employed for training, with an initial learning rate of 2 × 10 − 5 , β 1 and β 2 of (0.9, 0.999), and a weight decay of 0.01. Weight decay is excluded for normalization layers and bias parameters. All models are trained for 30 epochs. The learning rate is adjusted by using the linear warmup and cosine annealing strategies. Gradient clipping with the L2 norm is applied to stabilize training. All experiments are performed on two NVIDIA Tesla V100 GPUs with a batch size of 2 per GPU, and are implemented by using PyTorch [46] and MMAction2 [47].

5.3. Comparison with Baseline Methods

Owing to their similar network architectures, MViTv2 [48], and UniformerV2 [49] are selected as baseline methods. These methods are compared under two experimental settings: one without pre-trained weights and the other with weights pre-trained on the Kinetics-710 [49] dataset. All approaches use the parameters described in Section 5.2. For the setting without the pre-trained weights, the number of epochs is increased to 50, and the learning rate is set to 1 × 10 − 3 .
In accordance with the aforementioned experimental protocol, all methods were evaluated on the FSVR dataset. The results are shown in Table 4, where “None” indicates that the corresponding method was trained without the use of pre-trained weights.
Under the setting without pre-trained weights, all model parameters are randomly initialized. MViTv2 achieved an accuracy of 56.30% and a macro-F1 score of 54.30%, UniformerV2 improved the accuracy to 59.82% and a macro-F1 score of 58.57%. Under the same configuration, MFFLNet achieved an accuracy of 60.58% and a macro-F1 score of 59.56%. These results demonstrate that, without leveraging pre-training data, MFFLNet can more sufficiently capture fire-related information.
Given that MViTv2 does not provide weights pre-trained on Kinetics-710, only UniformerV2 and MFFLNet are evaluated using weights pre-trained on Kinetics-710. Both methods achieved improvements on the FSVR dataset, which demonstrates that weights pre-trained on Kinetics-710 can improve the FVR performance. Leveraging the pre-trained weights, the accuracy of UniformerV2 was elevated to 78.01%. Under the same experimental setup, MFFLNet achieved an accuracy of 79.54%, outperforming UniformerV2 by 1.53%. This indicates that when pre-trained weights are employed, MFFLNet can better characterize fire-related information that baseline methods fail to capture.

5.4. Ablation Study

To investigate the influence of the position and number of MFFL blocks, ablation experiments were conducted on FSVR. Table 5 presents the results. Given that deeper networks capture high-level abstract features such as semantics and context, embedding the MFFL block in deeper layers can enhance the ability to learn multiscale features of fire.
As shown in Table 5, we progressively increased the number of MFFL blocks, starting from the last layer and embedding them in the final 1 to 2, 3, 4, 5, and 6 layers. Performance declined as the number of layers increased beyond the final 4 layer. To avoid redundancy, subsequent experiments only embedded MFFL in the final 1 to 8, 1 to 10, and all 12 layers. A substantial performance gap exists between the “Final 1 to 3 layers” and “Final 1 to 4 layers” configurations. To validate the independent contribution of the 4th-to-last layer, a supplementary experiment was performed where the MFFL module was embedded exclusively in the 4th-to-last layer.
The best result is achieved when MFFL is embedded in the final 1 to 4 layers. These findings suggest that the position and quantity of MFFL blocks influence the performance. Concentrating the MFFL blocks in a limited number of deep layers improves recognition accuracy while reducing the number of parameters, thereby serving as an optimal configuration strategy. Consequently, MFFLNet embeds MFFL blocks in the final 1 to 4 layers of the ViT-B backbone.
To evaluate the contribution of MFConv3D and FSTFL within the MFFL block, ablation experiments were conducted on FSVR. The results are presented in Table 6, where “Baseline” denotes the method that excludes the utilization of the proposed FSTFL and MFConv3D. All methods in this table were evaluated under identical experimental settings. Moreover, we conducted the experiment three times with distinct random seeds and report the mean ± standard deviation.
As presented in Table 6, the baseline method, with the integration of the FSTFL layer, yields an accuracy of 78.46% and a macro-F1 score of 78.08%, respectively. This finding demonstrates that FSTFL effectively captures the spatiotemporal features of flames and smoke. When the MFConv3D layer is further incorporated into the FSTFL-equipped method, the resulting accuracy and macro-F1 score are 2.14% and 2.27% higher than those of the baseline method, respectively. Notably, the standard deviation of the experimental results gradually decreases with the sequential addition of FSTFL and MFConv3D. This observation indicates that these two layers not only enhance the model’s capability to learn features of flames and smoke but also improve the stability of the experimental results. The ablation experiments demonstrate that the two layers effectively enhance the performance of the baseline method on the FSVR dataset, thereby validating the effectiveness of both layers.

5.5. Comparison with Other Methods

5.5.1. Comparison on FSVR

On the FSVR dataset, we compare the performance of MFFLNet against 4 classic methods, i.e., TimeSformer [43], VideoSwin [50], MViTv2 [48], and UniformerV2 [49]. For TimeSformer, the spatial-only model was used for testing.
Table 7 presents the comparative results across three experimental settings: “None” indicates that networks were not initialized with pre-trained weights; “Kinetics-400” refers to initializing networks with pre-trained weights obtained by the corresponding method on the Kinetics-400 [51] dataset; and “Kinetics-710” refers to initializing networks with pre-trained weights from the Kinetics-710 dataset. To save time, we directly used publicly available pre-trained weights from existing methods.
Under the experimental setup where no pre-trained weights are used and 16 frames are utilized, MFFLNet achieves an accuracy of 60.58% and a macro-F1 score of 59.56%, outperforming all other baseline methods. These experimental results demonstrate that the multiscale fire feature learning architecture of MFFLNet can effectively capture the visual variations and temporal dynamics of flames and smoke. During the training process, MFFLNet exhibits a robust ability to learn fire-related features without relying on large-scale external data for pre-training.
Under the experimental setup employing pre-trained weights from the Kinetics-400 dataset, the recognition performance of all methods in Table 7 was significantly enhanced. Under identical experimental settings, MFFLNet outperformed the other comparative methods. These results demonstrate that MFFLNet can effectively transfer general video features when leveraging the Kinetics-400 pre-trained weights.
When pre-trained weights from the larger-scale Kinetics-710 dataset are utilized, the performance of all methods is further enhanced. When only 8 frames are fed as input to the network, both metrics of MFFLNet surpass those of UniformerV2 that uses 16 frames. Generally, using fewer frames as input reduces the time complexity of the algorithm. This indicates that MFFLNet can improve training and inference speed by using fewer frames as input, while maintaining high recognition performance. In summary, MFFLNet achieves the best results on FSVR, demonstrating superior performance in the FVR domain.

5.5.2. Comparison on LFVR

To further validate MFFLNet, experiments were performed on the LFVR dataset. Table 8 compares the performance of TimeSformer, VideoSwin, I3D [51], LGAE [9], and MFFLNet on LFVR. Following the previous study [9], 32 frames were employed as the input to the neural network. Similar to the experiments conducted on the FSVR dataset, we also present the experimental results of MFFLNet with 16 frames as input.
Generally, the method utilizing 32 frames as training input achieves superior performance compared to the same method employing 16 frames. With 16 frames as input, MFFLNet achieves an accuracy of 95.34% and an F1-score of 95.17%, outperforming other methods that utilize 32 frames. Despite using half the number of input frames compared to other approaches, MFFLNet can still precisely capture the dynamic features of fire via the proposed MFFL block. When 32 frames are employed as input, the performance of MFFLNet on the LFVR dataset is further enhanced. These experimental results demonstrate the superior performance of MFFLNet in the FVR task.

5.6. Time Complexity

Experiments were performed on the FSVR dataset to evaluate computational efficiency. The comparison results are shown in Table 9, where “Parameter” denotes the number of network parameters, “GFLOPs” represents the number of floating-point operations, “Training Time” indicates the average training time per epoch in seconds, and “Testing Time” refers to the average inference time per sample in milliseconds. These experiments were conducted on two NVIDIA Tesla V100 GPUs.
As shown in Table 9, MViTv2 has the smallest parameter count, but suffers from significant shortcomings in performance. VideoSwin achieves the best training and testing speed, yet its accuracy remains limited. Compared to UniformerV2, MFFLNet achieves lower parameter count and computational cost. In terms of model performance, MFFLNet reaches an accuracy of 60.58% with a standard deviation of only 0.18, demonstrating excellent stability. Compared to high-accuracy models such as TimeSformer and UniformerV2, MFFLNet achieves significant performance improvement with comparable computational time. Overall, MFFLNet effectively balances model time complexity, computational cost, and recognition performance.

5.7. Discussion and Analysis

To further investigate the effectiveness of MFFLNet, two approaches are employed to visualize and analyze the experimental results. Figure 4 presents the confusion matrices for the FSVR and LFVR datasets.
As shown in Figure 4a, MFFLNet achieves the highest accuracy for the Fire category, with a value of 93.4%. This indicates that MFFLNet is capable of stably capturing the discriminative information associated with this category. The accuracies for the Normal and Smoke categories are 67.5% and 78.4%, respectively. Although significantly lower than that of the Fire category, these results remain acceptable given the inherent visual similarities between smoke and normal scenes.
While significantly lower than that of the Fire category, these results remain acceptable. The categories with the highest misclassification rates are Normal and Smoke. This is likely attributed to the visual similarity between smoke and certain normal scenes, which may result in misclassification under low-contrast or complex background conditions. As shown in Figure 4b, the accuracies for the Non-fire and Fire categories reach 98.1% and 96.1%, respectively, with cross-misclassification rates below 4%. These results demonstrate the outstanding performance of MFFLNet, exhibiting stable and excellent classification across both datasets. Moreover, Table 10 reports the raw confusion-matrix counts of MFFLNet on FSVR.
The diagonal entries of Table 10 indicate the number of correctly classified samples, which are 4076 for Normal, 5546 for Fire, and 4725 for Smoke. Regarding misclassifications, 651 and 1309 Normal samples were incorrectly predicted as Fire and Smoke, respectively; 138 and 256 Fire samples were misidentified as Normal and Smoke, respectively; and 517 and 782 Smoke samples were erroneously classified as Normal and Fire, respectively. The row sums show that the total number of test samples per class is 6036, 5940, and 6024, indicating a relatively balanced class distribution. Overall, MFFLNet attains the highest recognition accuracy for the Fire class.
Table 11 presents a comparison of MFFLNet’s performance across three categories on FSVR. Experimental results demonstrate that MFFLNet exhibits the optimal detection performance in the Fire category, achieving a recall of 93.37% and an F1-score of 85.86%, with a corresponding Missed-Detection Rate (MDR) of merely 6.63%. This underscores the model’s robust capability in fire recognition, which effectively ensures a high detection rate for fire. The F1-scores for the Smoke and Normal categories are 76.74% and 75.71%, respectively, indicating that their overall performance lags behind that of the Fire category. The Normal category achieves a specificity of 94.53% and a relatively low False Alarm Rate (FAR) of 5.47%, though its MDR remains comparatively high. The Smoke category, by contrast, has the highest FAR among the three at 13.07%. This discrepancy is closely tied to the properties of smoke, including its variable morphology, ambiguous texture features, and proneness to confusion with fog and other interfering factors. A primary direction for future optimization will be to enhance the smoke feature representation capability, thereby reducing the MDR and improving overall detection performance.
Figure 5 presents typical cases where MFFLNet predicts correctly while the baseline method fails, covering the Normal, Fire, and Smoke categories in the FSVR dataset. For the Normal category, the baseline method is susceptible to interference from smoke-like and flame-like backgrounds, such as mountain clouds, mist, dusk glow, night road scenes, etc. For the Fire category, the baseline method fails to recognize small-area flame and early-stage fire. For the Smoke category, the baseline method exhibits missed detection issues when identifying weak and diffuse smoke, and is prone to confusing smoke with fire. The comparative results demonstrate that MFFLNet exhibits superior FVR capability, achieving robust and accurate classification performance across multiple scenarios.
Figure 6 presents representative False-Positive (FP) and False-Negative (FN) instances of MFFLNet on the FSVR dataset. These samples are primarily attributed to environmental interference, and their visual characteristics resemble those of flames and smoke, such as clouds, fog, sunlight at dawn or dusk, dust from construction sites, and warm light sources at night.
Moreover, FP and FN samples also appear in scenarios involving small-scale, distant, or occluded objects. When smoke or flames occupy a relatively small area in the image, are partially occluded by structures such as buildings, or exist under low-contrast lighting conditions, the model struggles to extract sufficiently prominent and discriminative features. These failure cases highlight the model’s limitations in complex real-world scenarios.
To reduce false alarm rates, future research will focus on two aspects. First, the fine-grained feature decoupling and adversarial discriminative learning algorithms will be developed to distinguish fire from environmental interference at the representation level. Second, the multi-resolution local attention and temporal context modeling algorithms will be developed to enhance the recognition capabilities for fires in scenarios involving small targets, long distances, and occlusions.

6. Conclusions

To advance research in the field of FVR, the large-scale FSVR dataset was constructed. To better utilize multiscale information for FVR, we proposed the Transformer-based MFFLNet, which integrates multiple MFFL blocks. Extensive experiments were conducted on the FSVR and LFVR datasets. These results indicate that MFFLNet outperforms the baseline methods and achieves the best results in the FVR field. Future work will explore the integration of motion information from flame and smoke for more precise recognition. Furthermore, we aim to expand the dataset by collecting a more extensive range of fire videos across diverse scenarios.

Author Contributions

Conceptualization, Y.Y.; methodology, Y.Y. and S.Y.; software, S.Y.; validation, S.Y. and Y.Y.; formal analysis, Y.Y.; investigation, S.Y.; resources, Y.Y.; data curation, S.Y. and Y.Y.; writing—original draft preparation, S.Y.; writing—review and editing, S.Y., Y.Y.; visualization, S.Y. and Y.Y.; supervision, Y.Y.; project administration, Y.Y.; funding acquisition, Y.Y. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported in part by the National Natural Science Foundation of China (Grant No. 62362003).

Data Availability Statement

The FSVR dataset and code for MFFLNet have been released at https://github.com/GNNUCV/MFFLNet (accessed on 20 September 2026).

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Danish, S.; Piran, M.J.; Khan, S.U.; Khan, M.A.; Dang, L.M.; Zweiri, Y.; Song, H.K.; Moon, H. Vision-based fire management system using autonomous unmanned aerial vehicles: A comprehensive survey. Artif. Intell. Rev. 2026, 59, 16. [Google Scholar] [CrossRef] [Scilit]
  2. Gragnaniello, D.; Greco, A.; Sansone, C.; Vento, B. Fire and smoke detection from videos: A literature review under a novel taxonomy. Expert Syst. Appl. 2024, 255, 124783. [Google Scholar] [CrossRef] [Scilit]
  3. Cheng, G.; Chen, X.; Wang, C.; Li, X.; Xian, B.; Yu, H. Visual fire detection using deep learning: A survey. Neurocomputing 2024, 596, 127975. [Google Scholar] [CrossRef] [Scilit]
  4. Bugarić, M.; Krstinić, D.; Šerić, L.; Stipaničev, D. Current Trends in Wildfire Detection, Monitoring and Surveillance. Fire 2025, 8, 356. [Google Scholar] [CrossRef] [Scilit]
  5. Wu, S.; Jiang, W.; Gao, D.; He, T. Cross-Technology Signal Detection and Jamming Attack for Heterogeneous Internet of Things. IEEE Trans. Dependable Secur. Comput. 2026, 23, 11134–11150. [Google Scholar] [CrossRef] [Scilit]
  6. Lv, X.; Jiang, W.; Gao, D.; Liu, Y.; He, T. WeRa: LoRa Over Wi-Fi. IEEE Trans. Wirel. Commun. 2026, 25, 18006–18022. [Google Scholar] [CrossRef] [Scilit]
  7. Gao, D.; Li, X.; Wang, W.; Han, Z.; Gadekallu, T.R. Generative AI for Cross-Technology Communication in Consumer Electronics. IEEE Trans. Consum. Electron. 2026, 72, 7026–7038. [Google Scholar] [CrossRef] [Scilit]
  8. Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the International Conference on Learning Representations; OpenReview.net: Online, 2021. [Google Scholar]
  9. Ding, J.; Yi, Y.; Wang, T.; Tian, T. Fire Video Recognition Based on Local and Global Adaptive Enhancement. Algorithms 2025, 18, 8. [Google Scholar] [CrossRef] [Scilit]
  10. Elhanashi, A.; Essahraui, S.; Dini, P.; Saponara, S. Early fire and smoke detection using deep learning: A comprehensive review of models, datasets, and challenges. Appl. Sci. 2025, 15, 10255. [Google Scholar] [CrossRef] [Scilit]
  11. Deng, L.; Wu, S.; Zou, S.; Liu, Q. Large-Space Fire Detection Technology: A Review of Conventional Detector Limitations and Image-Based Target Detection Techniques. Fire 2025, 8, 358. [Google Scholar] [CrossRef] [Scilit]
  12. Toulouse, T.; Rossi, L.; Campana, A.; Celik, T.; Akhloufi, M.A. Computer vision for wildfire research: An evolving image dataset for processing and analysis. Fire Saf. J. 2017, 92, 188–194. [Google Scholar] [CrossRef] [Scilit]
  13. Cazzolato, M.T.; Avalhais, L.; Chino, D.; Ramos, J.S.; de Souza, J.A.; Rodrigues-Jr, J.F.; Traina, A. Fismo: A compilation of datasets from emergency situations for fire and smoke analysis. In Proceedings of the Brazilian Symposium on Databases; SBC: Uberlândia, Brazil, 2017; pp. 213–223. [Google Scholar]
  14. Jadon, A.; Omama, M.; Varshney, A.; Ansari, M.S.; Sharma, R. FireNet: A specialized lightweight fire & smoke detection model for real-time IoT applications. arXiv 2019, arXiv:1905.11922. [Google Scholar]
  15. Gong, X.; Hu, H.; Wu, Z.; He, L.; Yang, L.; Li, F. Dark-channel based attention and classifier retraining for smoke detection in foggy environments. Digit. Signal Process. 2022, 123, 103454. [Google Scholar] [CrossRef] [Scilit]
  16. De Venâncio, P.V.A.; Rezende, T.M.; Lisboa, A.C.; Barbosa, A.V. Fire detection based on a two-dimensional convolutional neural network and temporal analysis. In Proceedings of the IEEE Latin American Conference on Computational Intelligence; IEEE: New York, NY, USA, 2021; pp. 1–6. [Google Scholar]
  17. Yazdi, A.; Qin, H.; Jordan, C.B.; Yang, L.; Yan, F. Nemo: An open-source transformer-supercharged benchmark for fine-grained wildfire smoke detection. Remote Sens. 2022, 14, 3979. [Google Scholar] [CrossRef] [Scilit]
  18. Wu, S.; Zhang, X.; Liu, R.; Li, B. A dataset for fire and smoke object detection. Multimed. Tools Appl. 2023, 82, 6707–6726. [Google Scholar] [CrossRef] [Scilit]
  19. Han, X.; Pu, N.; Feng, Z.; Bei, Y.; Zhang, Q.; Cheng, L.; Xue, L. Benchmarking multi-scene fire and smoke detection. In Proceedings of the Chinese Conference on Pattern Recognition and Computer Vision; Springer: Singapore, 2024; pp. 203–218. [Google Scholar]
  20. Foggia, P.; Saggese, A.; Vento, M. Real-time fire detection for video-surveillance applications using a combination of experts based on color, shape, and motion. IEEE Trans. Circuits Syst. Video Technol. 2015, 25, 1545–1556. [Google Scholar] [CrossRef] [Scilit]
  21. Torabian, M.; Pourghassem, H.; Mahdavi-Nasab, H. Fire detection based on fractal analysis and spatio-temporal features. Fire Technol. 2021, 57, 2583–2614. [Google Scholar] [CrossRef] [Scilit]
  22. Dou, Z.; Ma, X.; Xie, X.; Liu, H.; Guo, C. A hybrid method of detecting flame from video stream. IET Image Process. 2022, 16, 2937–2946. [Google Scholar] [CrossRef] [Scilit]
  23. Harkat, H.; Nascimento, J.M.; Bernardino, A.; Ahmed, H.F.T. Fire images classification based on a handcraft approach. Expert Syst. Appl. 2023, 212, 118594. [Google Scholar] [CrossRef] [Scilit]
  24. Yang, X.; Hua, Z.; Zhang, L.; Fan, X.; Zhang, F.; Ye, Q.; Fu, L. Preferred vector machine for forest fire detection. Pattern Recognit. 2023, 143, 109722. [Google Scholar] [CrossRef] [Scilit]
  25. Sharma, A.; Kumar, R.; Kansal, I.; Popli, R.; Khullar, V.; Verma, J.; Kumar, S. Fire detection in urban areas using multimodal data and federated learning. Fire 2024, 7, 104. [Google Scholar] [CrossRef] [Scilit]
  26. Ngandam Mfondoum, A.H. Sentinel2 image multi-filtering potential for active fire mapping. Remote Sens. Lett. 2025, 16, 1303–1314. [Google Scholar] [CrossRef] [Scilit]
  27. Panneerselvam, S.; Thangavel, S.K.; Ponnam, V.S.; Sengan, S. Federated learning based fire detection method using local MobileNet. Sci. Rep. 2024, 14, 30388. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Liu, P.; Ni, S.; Stanislav, S.; Tang, P. Automated image-based identification and consistent classification of fire patterns with quantitative shape analysis and spatial location identification. Dev. Built Environ. 2025, 21, 100612. [Google Scholar] [CrossRef] [Scilit]
  29. Muhammad, K.; Khan, S.; Elhoseny, M.; Ahmed, S.H.; Baik, S.W. Efficient fire detection for uncertain surveillance environment. IEEE Trans. Ind. Inform. 2019, 15, 3113–3122. [Google Scholar] [CrossRef] [Scilit]
  30. Wang, H.; Zhang, Y.; Zhu, C. YOLO-LFD: A Lightweight and Fast Model for Forest Fire Detection. Comput. Mater. Contin. 2025, 82, 3399–3417. [Google Scholar] [CrossRef] [Scilit]
  31. Shahid, M.; Virtusio, J.J.; Wu, Y.H.; Chen, Y.Y.; Tanveer, M.; Muhammad, K.; Hua, K.L. Spatio-temporal self-attention network for fire detection and segmentation in video surveillance. IEEE Access 2021, 10, 1259–1275. [Google Scholar] [CrossRef] [Scilit]
  32. Lv, K.; Wu, R.; Chen, S.; Lan, P. CCi-YOLOv8n: Enhanced Fire Detection with CARAFE and Context-Guided Modules. In Proceedings of the Advanced Intelligent Computing Technology and Applications; Springer: Singapore, 2025; pp. 128–140. [Google Scholar]
  33. Lu, K.; Huang, J.; Li, J.; Zhou, J.; Chen, X.; Liu, Y. MTL-FFDET: A multi-task learning-based model for forest fire detection. Forests 2022, 13, 1448. [Google Scholar] [CrossRef] [Scilit]
  34. Cheknane, M.; Bendouma, T.; Boudouh, S.S. Advancing fire detection: Two-stage deep learning with hybrid feature extraction using faster R-CNN approach. Signal Image Video Process. 2024, 18, 5503–5510. [Google Scholar] [CrossRef] [Scilit]
  35. Majid, S.; Alenezi, F.; Masood, S.; Ahmad, M.; Gündüz, E.S.; Polat, K. Attention based CNN model for fire detection and localization in real-world images. Expert Syst. Appl. 2022, 189, 116114. [Google Scholar] [CrossRef] [Scilit]
  36. Sheng, D.; Deng, J.; Xiang, J. Automatic smoke detection based on SLIC-DBSCAN enhanced convolutional neural network. IEEE Access 2021, 9, 63933–63942. [Google Scholar] [CrossRef] [Scilit]
  37. Dong, Z.; Zhao, F.; Wang, G.; Tian, Y.; Li, H. A deep learning framework: Predicting fire radiative power from the combination of polar-orbiting and geostationary satellite data during wildfire spread. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 10827–10841. [Google Scholar] [CrossRef] [Scilit]
  38. Han, Y.; Liu, X.; Tian, Y.; Dong, Z. Burned area and burn severity mapping with a transformer-based change detection model. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 13866–13880. [Google Scholar] [CrossRef] [Scilit]
  39. Deshpande, U.U.; Michael, G.K.O.; Srinivasaiah, S.H.; Malawade, H.; Kulkarni, Y.; Desai, Y. Real-time fire and smoke detection system for diverse indoor and outdoor industrial environmental conditions using a vision-based transfer learning approach. Front. Comput. Sci. 2025, 7, 1636758. [Google Scholar] [CrossRef] [Scilit]
  40. Li, J.; Zhang, S.; Huang, T. Multi-Scale 3D Convolution Network for Video Based Person Re-Identification. Proc. Aaai Conf. Artif. Intell. 2019, 33, 8618–8625. [Google Scholar] [CrossRef] [Scilit]
  41. Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2018; pp. 7132–7141. [Google Scholar]
  42. Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision; Ferrari, V., Hebert, M., Sminchisescu, C., Weiss, Y., Eds.; Springer: Cham, Switzerland, 2018; pp. 3–19. [Google Scholar]
  43. Bertasius, G.; Wang, H.; Torresani, L. Is Space-Time Attention All You Need for Video Understanding? In Proceedings of the International Conference on Machine Learning; ACM: New York, NY, USA, 2021; pp. 813–824. [Google Scholar]
  44. Cubuk, E.D.; Zoph, B.; Shlens, J.; Le, Q.V. Randaugment: Practical Automated Data Augmentation With a Reduced Search Space. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 3008–3017. [Google Scholar]
  45. Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. In Proceedings of the International Conference on Learning Representations; OpenReview.net: Online, 2019. [Google Scholar]
  46. Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Proceedings of the Advances in Neural Information Processing Systems; MIT Press: Cambridge, MA, USA, 2019. [Google Scholar]
  47. Contributors, M. OpenMMLab’s Next Generation Video Understanding Toolbox and Benchmark. 2020. Available online: https://github.com/open-mmlab/mmaction2 (accessed on 20 September 2026).
  48. Li, Y.; Wu, C.Y.; Fan, H.; Mangalam, K.; Xiong, B.; Malik, J.; Feichtenhofer, C. MViTv2: Improved multiscale vision transformers for classification and detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 4794–4804. [Google Scholar]
  49. Li, K.; Wang, Y.; He, Y.; Li, Y.; Wang, Y.; Wang, L.; Qiao, Y. UniFormerV2: Unlocking the Potential of Image ViTs for Video Understanding. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2023; pp. 1632–1643. [Google Scholar]
  50. Liu, Z.; Ning, J.; Cao, Y.; Wei, Y.; Zhang, Z.; Lin, S.; Hu, H. Video swin transformer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 3202–3211. [Google Scholar]
  51. Carreira, J.; Zisserman, A. Quo vadis, action recognition? A new model and the kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2017; pp. 6299–6308. [Google Scholar]
Figure 1. Example frames from FSVR.
Figure 1. Example frames from FSVR.
Fire 09 00427 g001
Figure 2. Architecture of MFFL.
Figure 2. Architecture of MFFL.
Fire 09 00427 g002
Figure 3. A diagram of MFFLNet.
Figure 3. A diagram of MFFLNet.
Fire 09 00427 g003
Figure 4. Confusion matrices of MFFLNet.
Figure 4. Confusion matrices of MFFLNet.
Fire 09 00427 g004
Figure 5. Some visualizations of predictions.
Figure 5. Some visualizations of predictions.
Fire 09 00427 g005
Figure 6. Examples of FP and FN samples of MFFLNet on the FSVR dataset.
Figure 6. Examples of FP and FN samples of MFFLNet on the FSVR dataset.
Fire 09 00427 g006
Table 1. Fire recognition datasets.
Table 1. Fire recognition datasets.
DatasetCategoryYearTypeNumber
Corsican fire [12]Fire2017Image/video600/5
FiSmo [13]Fire and smoke2017Image/video9448/88
FireNet [14]Fire2019Image/video2585/62
DSDF [15]Smoke2020Image18,413
D-Fire [16]Fire and smoke2021Image/video21,527/100
Nemo [17]Smoke2022Image/video4347/95
DFS [18]Fire and smoke2023Image9462
LFVR [9]Fire2024Video11,560
MS-FSDB [19]Fire and smoke2024Image12,586
FSVRFire and smoke2026Video30,000
Table 2. Comparison between MFFL and Representative Modules.
Table 2. Comparison between MFFL and Representative Modules.
ModuleSpatial ModelingTemporal ModelingMulti-Scale ModelingComplexity
M3D [40]2D CNN stream extracts appearance featuresMulti-scale 3D convolutions with parallel temporal kernelsMulti-scale only in temporal dimensionLow
SE [41]Global average pooling aggregates spatial information, no spatial attentionNo dedicated temporal modelingNo built-in multi-scale modeling capabilityLow
CBAM [42]2D spatial attention map via channel-wise poolingNo dedicated temporal modelingNo built-in multi-scale modeling capabilityMedium
TimeSformer [43]Spatial self-attention over image patchesTemporal self-attention across framesNo native multi-scale designHigh
MFFLFSTFL and MFConv3D captures multi-scale spatial featuresMFConv3D learns multi-scale temporal informationMulti-receptive-field parallel temporal convolutionsMedium
Table 3. Statistics of the FSVR dataset.
Table 3. Statistics of the FSVR dataset.
SetCategoryNumber of ClipsTotal
TrainFire19806000
Smoke2008
Normal2012
ValidationFire19806000
Smoke2008
Normal2012
TestFire594018,000
Smoke6024
Normal6036
Table 4. Comparison with baseline methods.
Table 4. Comparison with baseline methods.
MethodBackbonePre-Trained WeightsAcc (%)Macro-F1 (%)
MViTv2 [48]MViT-BNone 56.30 ± 0.95 54.30 ± 0.71
UniformerV2 [49]ViT-BNone 59.82 ± 0.13 58.57 ± 0.43
MFFLNetViT-BNone 60.58 ± 0.18 59.56 ± 0.99
UniformerV2 [49]ViT-BKinetics-710 78.01 ± 0.08 77.46 ± 0.10
MFFLNetViT-BKinetics-710 79.54 ± 0.19 79.22 ± 0.15
Table 5. Ablation experiments of positions on the FSVR dataset.
Table 5. Ablation experiments of positions on the FSVR dataset.
PositionParameter (M)Acc (%)
Final 1 to 2 layers92.5677.31
Final 1 to 3 layers95.9477.48
Final 1 to 4 layers99.3279.64
Only 4th-to-last layer89.1878.07
Final 1 to 5 layers102.776.72
Final 1 to 6 layers106.0875.18
Final 1 to 8 layers112.8472.79
Final 1 to 10 layers119.670.37
All 12 layers126.3669.48
Table 6. Ablation experiments on the FSVR dataset.
Table 6. Ablation experiments on the FSVR dataset.
MethodBackbonePre-Trained WeightsAcc (%)Macro-F1 (%)
BaselineViT-BKinetics-710 77.40 ± 0.39 76.95 ± 0.48
FSTFLViT-BKinetics-710 78.46 ± 0.27 78.08 ± 0.27
FSTFL + MFConv3DViT-BKinetics-710 79.54 ± 0.19 79.22 ± 0.15
Table 7. Comparison of other methods on FSVR.
Table 7. Comparison of other methods on FSVR.
MethodBackbonePre-Trained WeightsTraining FramesAcc (%)Macro-F1 (%)
MViTv2 [48]MViT-BNone16 56.30 ± 0.95 54.30 ± 0.71
VideoSwin [50]Swin-BNone16 57.08 ± 0.62 55.68 ± 0.81
TimeSformer [43]ViT-BNone16 58.00 ± 1.55 57.61 ± 2.16
UniformerV2 [49]ViT-BNone16 59.82 ± 0.13 58.57 ± 0.43
MFFLNetViT-BNone16 60.58 ± 0.18 59.56 ± 0.99
TimeSformer [43]ViT-BKinetics-40016 69.89 ± 0.68 69.34 ± 1.11
VideoSwin [50]Swin-BKinetics-40016 71.35 ± 1.90 70.57 ± 2.24
MViTv2 [48]MViT-BKinetics-40016 72.41 ± 0.95 71.30 ± 1.14
UniformerV2 [49]ViT-BKinetics-40016 78.33 ± 0.88 77.86 ± 0.98
MFFLNetViT-BKinetics-40016 78.66 ± 0.89 78.26 ± 0.94
UniformerV2 [49]ViT-BKinetics-7108 77.31 ± 0.18 76.70 ± 0.26
UniformerV2 [49]ViT-BKinetics-71016 78.01 ± 0.08 77.46 ± 0.10
MFFLNetViT-BKinetics-7108 78.70 ± 0.21 78.26 ± 0.31
MFFLNetViT-BKinetics-71016 79.54 ± 0.19 79.22 ± 0.15
Table 8. Comparison of other methods on LFVR.
Table 8. Comparison of other methods on LFVR.
MethodBackboneTraining FramesAcc (%)F1-Score (%)
TimeSformer [43]ViT-B3285.6781.14
VideoSwin [50]Swin-B3289.4086.41
I3D [51]ResNet503289.6486.76
LGAE [9]Swin-B3291.2788.93
MFFLNetViT-B1695.3495.17
MFFLNetViT-B3295.6495.46
Table 9. Comparison of time complexity on FSVR.
Table 9. Comparison of time complexity on FSVR.
MethodTraining FramesParameter (M)GFLOPsTraining Time (s)Testing Time (ms)Acc (%)
MViTv2 [48]1650.9218564160 56.30 ± 0.95
VideoSwin [50]1687.6428238558 57.08 ± 0.62
TimeSformer [43]1685.805631133136 58.00 ± 1.55
UniformerV2 [49]16123.7465273282 59.82 ± 0.13
MFFLNet1699.3264471080 60.58 ± 0.18
Table 10. Raw confusion-matrix counts of MFFLNet on FSVR.
Table 10. Raw confusion-matrix counts of MFFLNet on FSVR.
NormalFireSmokeRow Sum
Normal407665113096036
Fire13855462565940
Smoke51778247256024
Table 11. A comparison of MFFLNet’s performance across three categories on FSVR.
Table 11. A comparison of MFFLNet’s performance across three categories on FSVR.
ClassPrecision (%)Recall (%)F1-Score (%)Specificity (%)FAR (%)MDR (%)
Normal86.1667.5375.7194.535.4732.47
Fire79.4793.3785.8688.1211.886.63
Smoke75.1278.4476.7486.9313.0721.56
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Yang, S.; Yi, Y. MFFLNet: Multi-Scale Fire Feature Learning for Fire Video Recognition. Fire 2026, 9, 427. https://doi.org/10.3390/fire9100427

AMA Style

Yang S, Yi Y. MFFLNet: Multi-Scale Fire Feature Learning for Fire Video Recognition. Fire. 2026; 9(10):427. https://doi.org/10.3390/fire9100427

Chicago/Turabian Style

Yang, Shanzheng, and Yun Yi. 2026. "MFFLNet: Multi-Scale Fire Feature Learning for Fire Video Recognition" Fire 9, no. 10: 427. https://doi.org/10.3390/fire9100427

APA Style

Yang, S., & Yi, Y. (2026). MFFLNet: Multi-Scale Fire Feature Learning for Fire Video Recognition. Fire, 9(10), 427. https://doi.org/10.3390/fire9100427

Article Metrics

Back to TopTop