5.2. Experimental Setup
Several preprocessing operations are applied to the video data. During training,
T frames are randomly and uniformly sampled from each clip. The shorter side of the frames is resized to 256 pixels, while the longer side is adjusted adaptively. The RandAugment [
44] and RandomResizedCrop strategies are employed for data augmentation. Then, the cropped frames are scaled to
pixels, and horizontal flipping is applied with a 50% probability to increase data diversity.
During the validation and testing phases, we employed the UniformSample algorithm to uniformly select a fixed set of T frame indices from the videos, thereby ensuring the reproducibility of the test results. In the validation phase, the video is uniformly partitioned into T segments, where the sampling step within each segment is equal to half the length of the segment. Therefore, each sampled frame is positioned at the center of its corresponding segment, achieving deterministic central uniform sampling across these segments. In the testing phase, the video is also partitioned into T equal-length segments, yet the sampling step within each segment is adjusted to one-third of the segment length, thereby generating two distinct samples. These frames are adjusted to a shorter side of 224 pixels while maintaining the aspect ratio. Finally, center cropping is used to extract pixel regions. For fair comparison, all baseline methods utilize the same sampling strategy described above. In the experiments, T was set to 8, 16, and 32 based on specific requirements.
AdamW [
45] is employed for training, with an initial learning rate of 2 ×
,
1 and
2 of (0.9, 0.999), and a weight decay of 0.01. Weight decay is excluded for normalization layers and bias parameters. All models are trained for 30 epochs. The learning rate is adjusted by using the linear warmup and cosine annealing strategies. Gradient clipping with the L2 norm is applied to stabilize training. All experiments are performed on two NVIDIA Tesla V100 GPUs with a batch size of 2 per GPU, and are implemented by using PyTorch [
46] and MMAction2 [
47].
5.3. Comparison with Baseline Methods
Owing to their similar network architectures, MViTv2 [
48], and UniformerV2 [
49] are selected as baseline methods. These methods are compared under two experimental settings: one without pre-trained weights and the other with weights pre-trained on the Kinetics-710 [
49] dataset. All approaches use the parameters described in
Section 5.2. For the setting without the pre-trained weights, the number of epochs is increased to 50, and the learning rate is set to 1 ×
.
In accordance with the aforementioned experimental protocol, all methods were evaluated on the FSVR dataset. The results are shown in
Table 4, where “None” indicates that the corresponding method was trained without the use of pre-trained weights.
Under the setting without pre-trained weights, all model parameters are randomly initialized. MViTv2 achieved an accuracy of 56.30% and a macro-F1 score of 54.30%, UniformerV2 improved the accuracy to 59.82% and a macro-F1 score of 58.57%. Under the same configuration, MFFLNet achieved an accuracy of 60.58% and a macro-F1 score of 59.56%. These results demonstrate that, without leveraging pre-training data, MFFLNet can more sufficiently capture fire-related information.
Given that MViTv2 does not provide weights pre-trained on Kinetics-710, only UniformerV2 and MFFLNet are evaluated using weights pre-trained on Kinetics-710. Both methods achieved improvements on the FSVR dataset, which demonstrates that weights pre-trained on Kinetics-710 can improve the FVR performance. Leveraging the pre-trained weights, the accuracy of UniformerV2 was elevated to 78.01%. Under the same experimental setup, MFFLNet achieved an accuracy of 79.54%, outperforming UniformerV2 by 1.53%. This indicates that when pre-trained weights are employed, MFFLNet can better characterize fire-related information that baseline methods fail to capture.
5.4. Ablation Study
To investigate the influence of the position and number of MFFL blocks, ablation experiments were conducted on FSVR.
Table 5 presents the results. Given that deeper networks capture high-level abstract features such as semantics and context, embedding the MFFL block in deeper layers can enhance the ability to learn multiscale features of fire.
As shown in
Table 5, we progressively increased the number of MFFL blocks, starting from the last layer and embedding them in the final 1 to 2, 3, 4, 5, and 6 layers. Performance declined as the number of layers increased beyond the final 4 layer. To avoid redundancy, subsequent experiments only embedded MFFL in the final 1 to 8, 1 to 10, and all 12 layers. A substantial performance gap exists between the “Final 1 to 3 layers” and “Final 1 to 4 layers” configurations. To validate the independent contribution of the 4th-to-last layer, a supplementary experiment was performed where the MFFL module was embedded exclusively in the 4th-to-last layer.
The best result is achieved when MFFL is embedded in the final 1 to 4 layers. These findings suggest that the position and quantity of MFFL blocks influence the performance. Concentrating the MFFL blocks in a limited number of deep layers improves recognition accuracy while reducing the number of parameters, thereby serving as an optimal configuration strategy. Consequently, MFFLNet embeds MFFL blocks in the final 1 to 4 layers of the ViT-B backbone.
To evaluate the contribution of MFConv3D and FSTFL within the MFFL block, ablation experiments were conducted on FSVR. The results are presented in
Table 6, where “Baseline” denotes the method that excludes the utilization of the proposed FSTFL and MFConv3D. All methods in this table were evaluated under identical experimental settings. Moreover, we conducted the experiment three times with distinct random seeds and report the mean ± standard deviation.
As presented in
Table 6, the baseline method, with the integration of the FSTFL layer, yields an accuracy of 78.46% and a macro-F1 score of 78.08%, respectively. This finding demonstrates that FSTFL effectively captures the spatiotemporal features of flames and smoke. When the MFConv3D layer is further incorporated into the FSTFL-equipped method, the resulting accuracy and macro-F1 score are 2.14% and 2.27% higher than those of the baseline method, respectively. Notably, the standard deviation of the experimental results gradually decreases with the sequential addition of FSTFL and MFConv3D. This observation indicates that these two layers not only enhance the model’s capability to learn features of flames and smoke but also improve the stability of the experimental results. The ablation experiments demonstrate that the two layers effectively enhance the performance of the baseline method on the FSVR dataset, thereby validating the effectiveness of both layers.
5.5. Comparison with Other Methods
5.5.1. Comparison on FSVR
On the FSVR dataset, we compare the performance of MFFLNet against 4 classic methods, i.e., TimeSformer [
43], VideoSwin [
50], MViTv2 [
48], and UniformerV2 [
49]. For TimeSformer, the spatial-only model was used for testing.
Table 7 presents the comparative results across three experimental settings: “None” indicates that networks were not initialized with pre-trained weights; “Kinetics-400” refers to initializing networks with pre-trained weights obtained by the corresponding method on the Kinetics-400 [
51] dataset; and “Kinetics-710” refers to initializing networks with pre-trained weights from the Kinetics-710 dataset. To save time, we directly used publicly available pre-trained weights from existing methods.
Under the experimental setup where no pre-trained weights are used and 16 frames are utilized, MFFLNet achieves an accuracy of 60.58% and a macro-F1 score of 59.56%, outperforming all other baseline methods. These experimental results demonstrate that the multiscale fire feature learning architecture of MFFLNet can effectively capture the visual variations and temporal dynamics of flames and smoke. During the training process, MFFLNet exhibits a robust ability to learn fire-related features without relying on large-scale external data for pre-training.
Under the experimental setup employing pre-trained weights from the Kinetics-400 dataset, the recognition performance of all methods in
Table 7 was significantly enhanced. Under identical experimental settings, MFFLNet outperformed the other comparative methods. These results demonstrate that MFFLNet can effectively transfer general video features when leveraging the Kinetics-400 pre-trained weights.
When pre-trained weights from the larger-scale Kinetics-710 dataset are utilized, the performance of all methods is further enhanced. When only 8 frames are fed as input to the network, both metrics of MFFLNet surpass those of UniformerV2 that uses 16 frames. Generally, using fewer frames as input reduces the time complexity of the algorithm. This indicates that MFFLNet can improve training and inference speed by using fewer frames as input, while maintaining high recognition performance. In summary, MFFLNet achieves the best results on FSVR, demonstrating superior performance in the FVR domain.
5.5.2. Comparison on LFVR
To further validate MFFLNet, experiments were performed on the LFVR dataset.
Table 8 compares the performance of TimeSformer, VideoSwin, I3D [
51], LGAE [
9], and MFFLNet on LFVR. Following the previous study [
9], 32 frames were employed as the input to the neural network. Similar to the experiments conducted on the FSVR dataset, we also present the experimental results of MFFLNet with 16 frames as input.
Generally, the method utilizing 32 frames as training input achieves superior performance compared to the same method employing 16 frames. With 16 frames as input, MFFLNet achieves an accuracy of 95.34% and an F1-score of 95.17%, outperforming other methods that utilize 32 frames. Despite using half the number of input frames compared to other approaches, MFFLNet can still precisely capture the dynamic features of fire via the proposed MFFL block. When 32 frames are employed as input, the performance of MFFLNet on the LFVR dataset is further enhanced. These experimental results demonstrate the superior performance of MFFLNet in the FVR task.
5.7. Discussion and Analysis
To further investigate the effectiveness of MFFLNet, two approaches are employed to visualize and analyze the experimental results.
Figure 4 presents the confusion matrices for the FSVR and LFVR datasets.
As shown in
Figure 4a, MFFLNet achieves the highest accuracy for the Fire category, with a value of 93.4%. This indicates that MFFLNet is capable of stably capturing the discriminative information associated with this category. The accuracies for the Normal and Smoke categories are 67.5% and 78.4%, respectively. Although significantly lower than that of the Fire category, these results remain acceptable given the inherent visual similarities between smoke and normal scenes.
While significantly lower than that of the Fire category, these results remain acceptable. The categories with the highest misclassification rates are Normal and Smoke. This is likely attributed to the visual similarity between smoke and certain normal scenes, which may result in misclassification under low-contrast or complex background conditions. As shown in
Figure 4b, the accuracies for the Non-fire and Fire categories reach 98.1% and 96.1%, respectively, with cross-misclassification rates below 4%. These results demonstrate the outstanding performance of MFFLNet, exhibiting stable and excellent classification across both datasets. Moreover,
Table 10 reports the raw confusion-matrix counts of MFFLNet on FSVR.
The diagonal entries of
Table 10 indicate the number of correctly classified samples, which are 4076 for Normal, 5546 for Fire, and 4725 for Smoke. Regarding misclassifications, 651 and 1309 Normal samples were incorrectly predicted as Fire and Smoke, respectively; 138 and 256 Fire samples were misidentified as Normal and Smoke, respectively; and 517 and 782 Smoke samples were erroneously classified as Normal and Fire, respectively. The row sums show that the total number of test samples per class is 6036, 5940, and 6024, indicating a relatively balanced class distribution. Overall, MFFLNet attains the highest recognition accuracy for the Fire class.
Table 11 presents a comparison of MFFLNet’s performance across three categories on FSVR. Experimental results demonstrate that MFFLNet exhibits the optimal detection performance in the Fire category, achieving a recall of 93.37% and an F1-score of 85.86%, with a corresponding Missed-Detection Rate (MDR) of merely 6.63%. This underscores the model’s robust capability in fire recognition, which effectively ensures a high detection rate for fire. The F1-scores for the Smoke and Normal categories are 76.74% and 75.71%, respectively, indicating that their overall performance lags behind that of the Fire category. The Normal category achieves a specificity of 94.53% and a relatively low False Alarm Rate (FAR) of 5.47%, though its MDR remains comparatively high. The Smoke category, by contrast, has the highest FAR among the three at 13.07%. This discrepancy is closely tied to the properties of smoke, including its variable morphology, ambiguous texture features, and proneness to confusion with fog and other interfering factors. A primary direction for future optimization will be to enhance the smoke feature representation capability, thereby reducing the MDR and improving overall detection performance.
Figure 5 presents typical cases where MFFLNet predicts correctly while the baseline method fails, covering the Normal, Fire, and Smoke categories in the FSVR dataset. For the Normal category, the baseline method is susceptible to interference from smoke-like and flame-like backgrounds, such as mountain clouds, mist, dusk glow, night road scenes, etc. For the Fire category, the baseline method fails to recognize small-area flame and early-stage fire. For the Smoke category, the baseline method exhibits missed detection issues when identifying weak and diffuse smoke, and is prone to confusing smoke with fire. The comparative results demonstrate that MFFLNet exhibits superior FVR capability, achieving robust and accurate classification performance across multiple scenarios.
Figure 6 presents representative False-Positive (FP) and False-Negative (FN) instances of MFFLNet on the FSVR dataset. These samples are primarily attributed to environmental interference, and their visual characteristics resemble those of flames and smoke, such as clouds, fog, sunlight at dawn or dusk, dust from construction sites, and warm light sources at night.
Moreover, FP and FN samples also appear in scenarios involving small-scale, distant, or occluded objects. When smoke or flames occupy a relatively small area in the image, are partially occluded by structures such as buildings, or exist under low-contrast lighting conditions, the model struggles to extract sufficiently prominent and discriminative features. These failure cases highlight the model’s limitations in complex real-world scenarios.
To reduce false alarm rates, future research will focus on two aspects. First, the fine-grained feature decoupling and adversarial discriminative learning algorithms will be developed to distinguish fire from environmental interference at the representation level. Second, the multi-resolution local attention and temporal context modeling algorithms will be developed to enhance the recognition capabilities for fires in scenarios involving small targets, long distances, and occlusions.