ESCFM-YOLO: Lightweight Dual-Stream Architecture for Real-Time Small-Scale Fire Smoke Detection on Edge Devices
Abstract
1. Introduction
- (i)
- Dual-stream fire detector: a lightweight architecture that separates flame/smoke representation learning to mitigate feature interference versus unified models.
- (ii)
- Task-specific attention: ESCFM++ (flames) enhances spatial/channel selectivity for compact, high-frequency structures; ESCFM-RS (smoke) strengthens sensitivity to diffuse, low-contrast textures with depthwise convolutions and residual scaling.
- (iii)
- Edge-validated efficiency and robustness: ONNX/TensorRT deployment on Jetson Xavier NX with real-time throughput, validated through outdoor daytime/nighttime long-range live-stream tests that demonstrate reliable detection of small, early-stage flame cues under challenging illumination.
2. Related Work
2.1. Core Technical Challenges in Fire Detection (Feature/Gradient Interference)
2.2. Single-Stream Approaches (Flame-Centric and Smoke-Centric)
2.2.1. Single-Stream Approaches (Flame-Centric)
2.2.2. Single-Stream Approaches (Smoke-Centric)
2.3. Joint Learning vs. Dual-Stream: Integrated Approaches
- (i)
- Single-stream YOLOv5 baseline (shared backbone): Ahn et al. [43] adopted an off-the-shelf YOLOv5 detector trained as a two-class model (flame and smoke). Their main contribution lies in constructing the dataset and deploying the unmodified YOLOv5 via ONNX/TensorRT for real-time inference, rather than designing separate flame/smoke streams or modifying the network architecture. Pros: simple integration and single-engine deployment. Cons: the shared backbone does not explicitly mitigate interference between flame and smoke representations, which can degrade early-stage sensitivity.
- (ii)
- Detection and classification cascades: Pincott et al. [44] combined YOLOv7/YOLOv8 with EfficientNet-B0 to classify flame, smoke, or background (97% accuracy), while Kim and Ruy [8] used an RGB + IR CNN for shipboard monitoring (94.1% accuracy). Pros: improves reliability via secondary verification or additional modalities. Cons: introduces latency and heavier computation, which can be unfavorable for real-time edge constraints.
- (iii)
- Unified YOLO detectors with generic enhancements: Examples include YOLOv5s with CBAM + GhostConv+BiFPN [45], YOLOv5n with FocalNext+QAHARep-FPN [30], DCGC-YOLO [32] and YOLOGX on YOLOv8 [46], as well as attention/fusion modules such as MLCA [47], AFPN [48], and YOLO-SAD [49]. Pros: strong benchmark gains within a single model. Cons: unified networks may still suffer from feature interference; robustness to dust, steam, and tiny early cues (<30 px) remains limited.
2.4. Attention and Efficiency Mechanisms for Edge Deployment
2.5. Research Gap and Positioning of This Work
3. Proposed Method
3.1. Design Rationale and Contributions
3.2. Overall Pipeline and System Overview
- (i)
- Preprocessing (early-stage emphasis): We load the D-Fire dataset and retain frames where flame or smoke occupies ≤ 15% of the image area to emphasize early-stage cues. The data are then separated into flame and smoke subsets for modality-specific training.
- (ii)
- Dual-stream training (decoupled optimization): Two YOLOv5n detectors are trained independently: the flame stream integrates three ESCFM++ blocks, and the smoke stream integrates three ESCFM-RS blocks. In both cases, blocks are inserted before the P3–P5 detection heads while keeping the backbone and neck unchanged (see Figure 2).
- (iii)
- Model conversion (deployment alignment): Trained models are exported to ONNX in FP32 precision and converted to TensorRT FP32 engines with layer fusion and kernel auto-tuning to preserve accuracy.
- (iv)
- Edge deployment (runtime integration): The TensorRT engines are deployed on NVIDIA Jetson Xavier NX using DeepStream for low-latency streaming inference.
- (v)
- Real-time inference (parallel dual-stream execution): A single camera stream is preprocessed once (resize/normalize) and fed into two parallel engines. Outputs are fused at the system level to support early warning alerts for small flames and thin smoke.

3.3. Dual YOLOv5n Architecture with Specialized Attention Modules
3.4. ESCFM++ Module for Flame Detection
3.5. ESCFM-RS Module for Smoke Detection
3.6. Dual Inference Strategy for Edge Deployment
4. Experimental Results
4.1. Experimental Environment and Dataset Preparation
4.2. Performance Evaluation Metrics
4.3. Comparison with Existing YOLO Models
4.4. Comparison with Related Models
4.5. Ablation Study for Model Optimization
4.6. Edge Computing Performance
- (i)
- Compare parameter sizes before and after TensorRT optimization.
- (ii)
- Benchmark dual stream inference throughput and latency.
- (iii)
- Analyze single-stream performance for flame-only and smoke-only models.
- (iv)
- Perform statistical significance testing to validate mAP@50 improvements.
- (v)
- Visualize qualitative detection results in real-world scenarios.
4.7. Real-World Outdoor Live-Stream Case Study
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Filkov, A.I.; Tihay-Felicelli, V.; Masoudvaziri, N.; Rush, D.; Valencia, A.; Wang, Y.; Blunck, D.L.; Valero, M.M.; Kempna, K.; Smolka, J.; et al. A review of thermal exposure and fire spread mechanisms in large outdoor fires and the built environment. Fire Saf. J. 2023, 140, 103871. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Heilman, W.E.; Potter, B.E.; Clements, C.B.; Jackson, W.A.; French, N.H.; Goodrick, S.L.; Kochanski, A.K.; Larkin, N.K.; Lahm, P.W.; et al. Recent advances in wildland fire smoke dynamics research in the United States. Atmosphere 2025, 16, 1221. [Google Scholar] [CrossRef] [Scilit]
- Bukowski, R.W.; Peacock, R.D.; Averill, J.D.; Cleary, T.G.; Bryner, N.P.; Reneke, P.A. Performance of Home Smoke Alarms: Analysis of the Response of Several Available Technologies in Residential Fire Settings; NIST Technical Note 1455-1; National Institute of Standards and Technology: Gaithersburg, MD, USA, 2008. Available online: https://www.nist.gov/publications/performance-home-smoke-alarms-analysis-response-several-available-technologies-0?pub_id=100900 (accessed on 1 November 2025).
- Agbehadji, I.E.; Mabhaudhi, T.; Botai, J.; Masinde, M. A systematic review of existing early warning systems’ challenges and opportunities in cloud computing early warning systems. Climate 2023, 11, 188. [Google Scholar] [CrossRef] [Scilit]
- Wu, Q.; Ding, Z.; Zhang, W. Research progress on electrochemical gas sensors for fire detection. Int. J. Electrochem. Sci. 2025, 20, 101043. [Google Scholar] [CrossRef] [Scilit]
- Danish, S.; Piran, M.J.; Khan, S.U.; Khan, M.A.; Dang, L.M.; Zweiri, Y.; Song, H.K.; Moon, H. Vision-based fire management system using autonomous unmanned aerial vehicles: A comprehensive survey. Artif. Intell. Rev. 2026, 59, 16. [Google Scholar] [CrossRef] [Scilit]
- Khan, F.; Xu, Z.; Sun, J.; Khan, F.M.; Ahmed, A.; Zhao, Y. Recent advances in sensors for fire detection. Sensors 2022, 22, 3310. [Google Scholar] [CrossRef] [Scilit]
- Kim, D.; Ruy, W. CNN-based fire detection method on autonomous ships using composite channels composed of RGB and IR data. Int. J. Nav. Archit. Ocean. Eng. 2022, 14, 100489. [Google Scholar] [CrossRef] [Scilit]
- Gragnaniello, D.; Greco, A.; Sansone, C.; Vento, B. Fire and smoke detection from videos: A literature review under a novel taxonomy. Expert Syst. Appl. 2024, 255, 124783. [Google Scholar] [CrossRef] [Scilit]
- Lv, C.; Zhou, H.; Chen, Y.; Fan, D.; Di, F. A lightweight fire detection algorithm for small targets based on YOLOv5s. Sci. Rep. 2024, 14, 14104. [Google Scholar] [CrossRef] [Scilit]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar]
- Redmon, J.; Farhadi, A. YOLOv3: An incremental improvement. arXiv 2018, arXiv:1804.02767. [Google Scholar] [CrossRef] [Scilit]
- Tian, Y.; Ye, Q.; Doermann, D. YOLOv12: Attention-centric real-time object detectors. arXiv 2025, arXiv:2502.12524. [Google Scholar]
- Bochkovskiy, A.; Wang, C.-Y.; Liao, H.-Y.M. YOLOv4: Optimal speed and accuracy of object detection. arXiv 2020, arXiv:2004.10934. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.-Y.; Bochkovskiy, A.; Liao, H.-Y.M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 7464–7475. [Google Scholar]
- Wang, C.-Y.; Yeh, I.-H.; Liao, H.-Y.M. YOLOv9: Learning what you want to learn using programmable gradient information. In European Conference on Computer Vision; Springer: Cham, Switzerland, 2024; pp. 1–21. [Google Scholar]
- Khanam, R.; Hussain, M. YOLOv11: An overview of the key architectural enhancements. arXiv 2024, arXiv:2410.17725. [Google Scholar] [CrossRef] [Scilit]
- Jocher, G.; Stoken, A.; Borovec, J.; NanoCode012; Chaurasia, A.; Tao, X.; Liu, C.; Abhiram, V.; Laughing; tkianai; et al. YOLOv5, Version 5.0; [Source Code]; GitHub: Tokyo, Japan, 2020. Available online: https://github.com/ultralytics/yolov5 (accessed on 1 October 2024).
- Han, X.; Wu, Y.; Pu, N.; Feng, Z.; Zhang, Q.; Bei, Y.; Cheng, L. Fire and smoke detection with burning intensity representation. arXiv 2024, arXiv:2410.16642. [Google Scholar] [CrossRef] [Scilit]
- Sultan, T.; Chowdhury, M.S.; Safran, M.; Mridha, M.F.; Dey, N. Deep learning-based multistage fire detection system and emerging direction. Fire 2024, 7, 451. [Google Scholar] [CrossRef] [Scilit]
- Cheng, G.; Chen, X.; Wang, C.; Li, X.; Xian, B.; Yu, H. Visual fire detection using deep learning: A survey. Neurocomputing 2024, 555, 127975. [Google Scholar] [CrossRef] [Scilit]
- Zhao, C.; Zhao, L.; Zhang, K.; Ren, Y.; Chen, H.; Sheng, Y. Smoke and Fire-You Only Look Once: A lightweight deep learning model for video smoke and flame detection in natural scenes. Fire 2025, 8, 104. [Google Scholar] [CrossRef] [Scilit]
- Jangirova, S.; Jankovic, B.; Ullah, W.; Khan, L.U.; Guizani, M. Real-time aerial fire detection on resource-constrained devices using knowledge distillation. arXiv 2025, arXiv:2502.20979. [Google Scholar] [CrossRef] [Scilit]
- Almeida, J.S.; Jagatheesaperumal, S.K.; Nogueira, F.G.; de Albuquerque, V.H.C. EdgeFireSmoke++: A novel lightweight algorithm for real-time forest fire detection and visualization using Internet of Things-human machine interface. Expert Syst. Appl. 2023, 221, 119747. [Google Scholar] [CrossRef] [Scilit]
- Sun, H.; Xu, R.; Luo, J.; Cheng, H. Review of the application of UAV edge computing in fire rescue. Sensors 2025, 25, 3304. [Google Scholar] [CrossRef] [Scilit]
- Maltezos, E.; Petousakis, K.; Dadoukis, A.; Karagiannidis, L.; Ouzounoglou, M.; Krommyda, M.; Hadjipavlis, G.; Amditis, A. A smart building fire and gas leakage alert system with edge computing and NG112 emergency call capabilities. Information 2022, 13, 164. [Google Scholar] [CrossRef] [Scilit]
- Li, C.; Li, L.; Jiang, H.; Weng, K.; Geng, Y.; Li, L.; Ke, Z.; Li, Q.; Cheng, M.; Nie, W.; et al. YOLOv6: A single-stage object detection framework for industrial applications. arXiv 2022, arXiv:2209.02976. [Google Scholar] [CrossRef] [Scilit]
- Reis, D.; Kupec, J.; Hong, J.; Daoudi, A. Real-time flying object detection with YOLOv8. arXiv 2023, arXiv:2305.09973. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Xu, Y.; He, Y.; Cai, Y.; Chen, L.; Li, Y.; Sotelo, M.A.; Li, Z. YOLOv5-Fog: A multiobjective visual detection algorithm for fog driving scenes based on improved YOLOv5. IEEE Trans. Instrum. Meas. 2022, 71, 1–12. [Google Scholar] [CrossRef] [Scilit]
- Geng, X.; Su, Y.; Cao, X.; Li, H.; Liu, L. YOLOFM: An improved fire and smoke object detection algorithm based on YOLOv5n. Sci. Rep. 2024, 14, 4543. [Google Scholar] [CrossRef] [Scilit]
- Peng, B.; Kim, T.K. YOLO-HF: Early detection of home fires using YOLO. IEEE Access 2025, 13, 79451–79466. [Google Scholar] [CrossRef] [Scilit]
- He, Y.; Hu, J.; Zeng, M.; Qian, Y.; Zhang, R. DCGC-YOLO: The efficient dual-channel bottleneck structure YOLO detection algorithm for fire detection. IEEE Access 2024, 12, 65254–65265. [Google Scholar] [CrossRef] [Scilit]
- Pan, W.; Wang, X.; Huan, W. EFA-YOLO: An efficient feature attention model for fire and flame detection. arXiv 2024, arXiv:2409.12635. [Google Scholar]
- Wu, S.; Xia, Y. Enhanced YOLOv7-tiny for small-scale fire detection via multi-scale channel spatial attention and dynamic upsampling. IEEE Access 2025, 13, 126901–126914. [Google Scholar] [CrossRef] [Scilit]
- Khan, S.; Muhammad, K.; Hussain, T.; Del Ser, J.; Cozzolino, F.; Bhattacharyya, S.; Akhtar, Z.; de Albuquerque, V.H.C. DeepSmoke: Deep learning model for smoke detection and segmentation in outdoor environments. Expert Syst. Appl. 2021, 182, 115125. [Google Scholar] [CrossRef] [Scilit]
- Yuan, F.; Li, K.; Wang, C.; Fang, Z. A lightweight network for smoke semantic segmentation. Pattern Recognit. 2023, 137, 109289. [Google Scholar] [CrossRef] [Scilit]
- Celik, T.; Demirel, H.; Ozkaramanli, H.; Uyguroglu, M. Fire detection using statistical color model in video sequences. J. Vis. Commun. Image Represent. 2007, 18, 176–185. [Google Scholar] [CrossRef] [Scilit]
- Töreyin, B.U.; Dedeoğlu, Y.; Güdükbay, U.; Cetin, A.E. Computer vision-based method for real-time fire and flame detection. Pattern Recognit. Lett. 2006, 27, 49–58. [Google Scholar] [CrossRef] [Scilit]
- Li, P.; Zhao, W. Image fire detection algorithms based on convolutional neural networks. Case Stud. Therm. Eng. 2020, 19, 100625. [Google Scholar] [CrossRef] [Scilit]
- Jadon, A.; Omama, M.; Varshney, A.; Ansari, M.S.; Sharma, R. FireNet: A specialized lightweight fire and smoke detection model for real-time IoT applications. arXiv 2019, arXiv:1905.11922. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.-Y.; Kweon, I.S. CBAM: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 3–19. [Google Scholar]
- He, L.; Zhou, Y.; Liu, L.; Zhang, Y.; Ma, J. Research and application of deep learning object detection methods for forest fire smoke recognition. Sci. Rep. 2025, 15, 16328. [Google Scholar] [CrossRef] [Scilit]
- Ahn, Y.; Choi, H.; Kim, B.S. Development of early fire detection model for buildings using computer vision-based CCTV. J. Build. Eng. 2023, 65, 105647. [Google Scholar] [CrossRef] [Scilit]
- Pincott, J.; Tien, P.W.; Wei, S.; Calautit, J.K. Indoor fire detection utilizing computer vision-based strategies. J. Build. Eng. 2022, 61, 105154. [Google Scholar] [CrossRef] [Scilit]
- Yar, H.; Khan, Z.A.; Ullah, F.U.M.; Ullah, W.; Baik, S.W. A modified YOLOv5 architecture for efficient fire detection in smart cities. Expert Syst. Appl. 2023, 231, 120465. [Google Scholar] [CrossRef] [Scilit]
- Li, C.; Du, Y.; Zhang, X.; Wu, P. YOLOGX: An improved forest fire detection algorithm based on YOLOv8. Front. Environ. Sci. 2025, 12, 1486212. [Google Scholar] [CrossRef] [Scilit]
- Fu, J.; Xu, Z.; Yue, Q.; Lin, J.; Zhang, N.; Zhao, Y.; Gu, D. A multi-object detection method for building fire warnings through artificial intelligence generated content. Sci. Rep. 2025, 15, 18434. [Google Scholar] [CrossRef] [Scilit]
- Amjad, A.; Huroon, A.M.; Chang, H.T.; Tai, L.C. Dynamic fire and smoke detection module with enhanced feature integration and attention mechanisms. Pattern Anal. Appl. 2025, 28, 81. [Google Scholar] [CrossRef] [Scilit]
- Yang, R.; Jiang, J.; Liu, F.; Yan, L. YOLO-SAD for fire detection and localization in real-world images. Digit. Signal Process. 2025, 165, 105320. [Google Scholar] [CrossRef] [Scilit]
- Park, J.C.; Kim, M.J.; Kim, G.W. MSA-YOLOv5: An improved lightweight YOLOv5 model for small object detection. In Proceedings of the 15th International Conference on Information and Communication Technology Convergence (ICTC), Jeju Island, Republic of Korea, 16–18 October 2024; pp. 658–663. [Google Scholar]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montreal, Canada, 7–12 December 2015; Volume 28, pp. 91–99. [Google Scholar]
- Tan, M.; Pang, R.; Le, Q.V. EfficientDet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 10781–10790. [Google Scholar]
- Lv, W.; Zhao, Y.; Chang, Q.; Huang, K.; Wang, G.; Liu, Y. RT-DETRv2: Improved baseline with bag-of-freebies for real-time detection transformer. arXiv 2024, arXiv:2407.17140. [Google Scholar]
- Li, X.; Wang, W.; Hu, X.; Yang, J. Selective kernel networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 510–519. [Google Scholar]
- Shaw, P.; Uszkoreit, J.; Vaswani, A. Self-attention with relative position representations. arXiv 2018, arXiv:1803.02155. [Google Scholar] [CrossRef] [Scilit]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 7132–7141. [Google Scholar]
- Liu, Q.; Chen, H.; Lin, D. Research and optimization of a multilevel fire detection framework based on deep learning and classical pattern recognition techniques. Sci. Rep. 2025, 15, 20364. [Google Scholar] [CrossRef] [Scilit]
- Howard, A.; Sandler, M.; Chu, G.; Chen, L.-C.; Chen, B.; Tan, M.; Wang, W.; Zhu, Y.; Pang, R.; Vasudevan, V.; et al. Searching for MobileNetV3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 1314–1324. [Google Scholar]
- Misra, D. Mish: A self-regularized non-monotonic activation function. arXiv 2019, arXiv:1908.08681. [Google Scholar]
- Peng, H.; Li, Z.; Zou, X.; Wang, H.; Xiong, J. Research on litchi image detection in orchard using UAV based on improved YOLOv5. Expert Syst. Appl. 2025, 263, 125828. [Google Scholar] [CrossRef] [Scilit]
- Qiu, S.; Xu, X.; Cai, B. FReLU: Flexible rectified linear units for improving convolutional neural networks. In Proceedings of the 24th International Conference on Pattern Recognition (ICPR), Beijing, China, 20–24 August 2018; pp. 1223–1228. [Google Scholar]
- Ma, N.; Zhang, X.; Liu, M.; Sun, J. Activate or not: Learning customized activation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 8032–8042. [Google Scholar]
- Hendrycks, D.; Gimpel, K. Gaussian error linear units (GELUs). arXiv 2016, arXiv:1606.08415. [Google Scholar]
- Maas, A.L.; Hannun, A.Y.; Ng, A.Y. Rectifier nonlinearities improve neural network acoustic models. In Proceedings of the 30th International Conference on Machine Learning (ICML), Atlanta, GA, USA, 16–21 June 2013; Volume 30, p. 3. [Google Scholar]
- Nair, V.; Hinton, G.E. Rectified linear units improve restricted Boltzmann machines. In Proceedings of the 27th International Conference on Machine Learning (ICML), Haifa, Israel, 21–24 June 2010; pp. 807–814. [Google Scholar]
- Xu, B.; Wang, N.; Chen, T.; Li, M. Empirical evaluation of rectified activations in convolutional network. arXiv 2015, arXiv:1505.00853. [Google Scholar] [CrossRef] [Scilit]
- Godfrey, L.B.; Gashler, M.S. A continuum among logarithmic, linear, and exponential functions, and its potential to improve generalization in neural networks. In Proceedings of the 7th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management (IC3K), Lisbon, Portugal, 12–14 November 2015; Volume 1, pp. 481–486. [Google Scholar]
- YouTube. Fire Scene VR Video. 2018. Available online: https://www.youtube.com/watch?v=R6qJUl9C8HQ&ab_channel=kobecitychannel (accessed on 1 July 2025).
- YouTube. Vent-Enter-Search Rescue Video. 2019. Available online: https://www.youtube.com/watch?v=ECg8bI-ZY0o&ab_channel=StocktonFireHistory (accessed on 1 July 2025).
- YouTube. Working Structure Fire Video. 2016. Available online: https://www.youtube.com/watch?v=S6P2V4I36L4&ab_channel=StocktonFireHistory (accessed on 1 July 2025).








| Item | Details |
|---|---|
| Hardware | GPU: NVIDIA GeForce RTX 4080 (16 GB VRAM), CPU: Intel 9-13900 K, Memory: 64 GB |
| Operating system | Windows 11 |
| Software framework | PyTorch 2.1.0, CUDA 12.1, cuDNN 8.9 |
| YOLOv5 model | YOLOv5n (Nano) |
| Learning hyperparameters | Learning rate: 0.01; weight decay: 0.0005; momentum: 0.937 |
| Max epochs | 3000 |
| Batch size | 8 |
| Warm-up phase | Epochs: 3; warm up momentum: 0.8; warm up bias LR: 0.1 |
| Technique | Description | Probability/Range |
|---|---|---|
| HSV transform | Adjusts hue, saturation, and value (brightness) | Hue: ±1.5%, Saturation.: ±70%, Value.: ±40% |
| Mosaic | Combines four images to generate a new training sample | Probability: 100% |
| Flip | Horizontal flip | Probability: 50% |
| Translation | Random translation | ±10% |
| Rotation | Random rotation | ±10° |
| Scale | Random scaling | ±50% |
| Class | Images | Bounding Boxes |
|---|---|---|
| Only fire | 1164 | 14,692 |
| Only smoke | 5867 | 11,865 |
| Fire and smoke | 4658 | - |
| None | 9838 | - |
| Class (Processed) | Train | Val | Total |
|---|---|---|---|
| Flame only | 4161 | 975 | 5136 |
| Smoke only | 3385 | 798 | 4183 |
| Flame and smoke | 4161 | 975 | 5136 |
| Model | Flame | Smoke | |||||
|---|---|---|---|---|---|---|---|
| Precision (%) | Recall (%) | mAP@50 (%) | Precision (%) | Recall (%) | mAP@50 (%) | Params (M) | |
| YOLOv3 tiny | 57.0 | 52.0 | 49.68 | 83.0 | 46.0 | 72.42 | 8.7 |
| YOLOv4 tiny | 72.0 | 54.0 | 57.54 | 77.0 | 77.0 | 81.39 | 6.1 |
| YOLOv5n | 73.0 | 70.2 | 73.80 | 85.7 | 82.8 | 88.00 | 1.8 |
| MSA YOLOv5n [50] | 73.9 | 68.7 | 73.40 | 85.7 | 83.1 | 88.10 | 1.82 |
| YOLOv6n | 72.3 | 67.0 | 72.30 | 86.4 | 66.5 | 86.50 | 4.7 |
| YOLOv7 tiny | 71.19 | 71.88 | 74.24 | 89.04 | 81.86 | 87.84 | 7.0 |
| YOLOv8n | 73.2 | 66.7 | 73.40 | 86.2 | 83.5 | 88.50 | 3.2 |
| YOLOv9t | 73.0 | 67.8 | 73.10 | 88.8 | 84.6 | 89.00 | 2.0 |
| YOLOv10n | 71.1 | 68.1 | 71.80 | 86.2 | 81.9 | 86.80 | 2.3 |
| YOLOv11n | 74.9 | 64.7 | 72.70 | 85.2 | 80.0 | 86.70 | 2.6 |
| YOLOv12n | 72.1 | 67.2 | 72.80 | 88.9 | 83.2 | 87.70 | 2.6 |
| YOLOv5n ESCFM++ (ours) | 74.1 | 70.4 | 74.50 | - | - | - | 1.89 |
| YOLOv5n ESCFM-RS (ours) | - | - | - | 89.4 | 81.9 | 89.20 | 1.89 |
| Model | Precision (%) | Recall (%) | mAP@50 (%) | Params (M) |
|---|---|---|---|---|
| YOLOv3 tiny | 60.00 | 47.00 | 56.68 | 8.70 |
| YOLOv4 tiny | 72.00 | 57.00 | 66.42 | 6.10 |
| YOLOv5n | 79.60 | 74.50 | 78.80 | 1.80 |
| MSA YOLOv5n [50] | 80.30 | 74.70 | 78.80 | 1.82 |
| YOLOv6n | 78.40 | 61.40 | 78.40 | 4.70 |
| YOLOv7 tiny | 81.08 | 76.44 | 80.16 | 7.00 |
| YOLOv8n | 79.40 | 73.10 | 79.00 | 3.20 |
| YOLOv9t | 81.20 | 73.90 | 79.50 | 2.00 |
| YOLOv10n | 81.10 | 70.90 | 78.70 | 2.30 |
| YOLOv11n | 79.20 | 72.80 | 78.50 | 2.60 |
| YOLOv12n | 80.40 | 73.20 | 79.70 | 2.60 |
| YOLOv5n ESCFM++ (ours) | 80.00 | 72.20 | 78.70 | 1.89 |
| YOLOv5n ESCFM-RS (ours) | 81.36 | 74.62 | 80.20 | 1.89 |
| Model | Class | Precision (%) | Recall (%) | mAP@50 (%) |
|---|---|---|---|---|
| YOLOv5n ESCFM++ (ours) | Flame | 72.30 | 67.00 | 73.10 |
| Smoke | 86.40 | 66.50 | 84.30 | |
| YOLOv5n ESCFM-RS (ours) | Flame | 73.60 | 68.00 | 74.40 |
| Smoke | 87.20 | 67.90 | 85.60 |
| Dataset | Model | Reduction | Precision (%) | Recall (%) | mAP@50 (%) | Params (M) |
|---|---|---|---|---|---|---|
| Flame | Faster R-CNN [51] | - | 71.70 | 49.70 | 71.70 | 41.00 |
| EfficientDet-D0 [52] | - | 30.60 | 49.13 | 64.40 | 3.90 | |
| RT-DETRv2 [53] | - | 40.20 | 58.80 | 74.20 | 20.00 | |
| YOLOFM [30] | - | 71.90 | 69.10 | 72.00 | 3.60 | |
| YOLO-SAD [49] | - | 74.20 | 70.60 | 74.50 | 16.20 | |
| YOLOv5n + SKNet [54] | 24 | 73.80 | 67.90 | 72.40 | 1.98 | |
| YOLOv5n + Self Attention [55] | 24 | 74.60 | 68.50 | 73.40 | 2.05 | |
| YOLOv5n + CBAM [41] | 16 | 72.60 | 70.00 | 74.20 | 1.90 | |
| YOLOv5n + ESCFM++ (ours) | 24 | 74.10 | 70.40 | 74.50 | 1.89 | |
| Smoke | Faster R-CNN | - | 60.70 | 46.80 | 60.70 | 41.00 |
| EfficientDet-D0 | - | 53.50 | 66.70 | 87.70 | 3.90 | |
| RT-DETRv2 | - | 58.90 | 73.20 | 88.70 | 20.00 | |
| YOLOFM | - | 88.20 | 82.10 | 87.30 | 3.60 | |
| YOLO-SAD | - | 87.30 | 84.30 | 88.40 | 16.20 | |
| YOLOv5n + SKNet | 24 | 86.20 | 82.30 | 88.40 | 1.98 | |
| YOLOv5n + Self Attention | 8 | 87.90 | 81.20 | 88.10 | 2.05 | |
| YOLOv5n + CBAM | 24 | 88.00 | 81.90 | 88.60 | 1.90 | |
| YOLOv5n + ESCFM-RS (ours) | 14 | 89.40 | 81.90 | 89.20 | 1.89 |
| Dataset | Model | Post- Processing | IoU | Activation | Precision (%) | Recall (%) | mAP@50 (%) |
|---|---|---|---|---|---|---|---|
| Flame | YOLOv5n | NMS [56] | GloU [50] | SiLU [54] | 72.80 | 69.50 | 73.10 |
| GloU | 73.90 | 68.70 | 73.40 | ||||
| NMS | DluU [55] | SiLU | 74.20 | 67.10 | 72.30 | ||
| MSA-YOLOv5n [57] | CloU [55] | 73.90 | 68.70 | 73.40 | |||
| GloU | 72.40 | 69.60 | 73.50 | ||||
| Soft-NMS [53] | DloU | SiLU | 73.40 | 67.90 | 73.30 | ||
| CloU | 72.40 | 69.60 | 73.50 | ||||
| GloU | 74.20 | 69.10 | 73.80 | ||||
| NMS | DloU | SiLU | 73.00 | 69.00 | 73.40 | ||
| YOLOv5n + ESCFM | CloU | 75.40 | 68.20 | 73.50 | |||
| GloU | 74.70 | 69.00 | 73.20 | ||||
| Soft-NMS | DloU | SiLU | 74.80 | 68.40 | 73.10 | ||
| CloU | 74.70 | 69.00 | 73.20 | ||||
| Smoke | YOLOv5n | NMS | GloU | SiLU | 85.70 | 82.80 | 88.00 |
| GloU | 85.70 | 83.10 | 88.10 | ||||
| NMS | DloU | SiLU | 86.80 | 82.70 | 88.00 | ||
| MSA-YOLOv5n | CloU | 85.70 | 83.10 | 88.10 | |||
| GloU | 86.50 | 82.70 | 87.60 | ||||
| Soft-NMS | Dlou | SiLU | 87.30 | 80.50 | 87.60 | ||
| CloU | 86.50 | 82.70 | 87.60 | ||||
| GloU | 86.80 | 83.00 | 88.60 | ||||
| NMS | DloU | SiLU | 85.80 | 83.40 | 87.00 | ||
| YOLOv5n + ESCFM | CloU | 85.80 | 83.40 | 87.00 | |||
| GloU | 86.20 | 84.30 | 87.80 | ||||
| Soft-NMS | DloU | SiLU | 85.70 | 82.30 | 87.80 | ||
| CloU | 86.20 | 84.30 | 87.70 |
| Dataset | Model | Post-Processing | IoU | Activation | Precision (%) | Recall (%) | mAP@50 (%) |
|---|---|---|---|---|---|---|---|
| YOLOv5n | NMS | GloU | SiLU | 72.8 | 69.5 | 73.1 | |
| SiLU | 74.2 | 69.1 | 73.8 | ||||
| Hardswish [58] | 74.3 | 68.1 | 72.7 | ||||
| Mish [59] | 75.1 | 68.0 | 73.2 | ||||
| EfficientMish [60] | 73.6 | 69.1 | 72.9 | ||||
| FreLU [61] | 70.8 | 69.1 | 71.7 | ||||
| Flame | YOLOv5n + ESCFM | NMS | GloU | AconC [62] | 72.9 | 67.2 | 72.3 |
| MetaAconC [62] | 73.6 | 67.0 | 73.0 | ||||
| GELU [63] | 73.9 | 67.6 | 72.9 | ||||
| Leaky ReLU [64] | 73.0 | 68.6 | 72.3 | ||||
| ReLU [65] | 71.9 | 67.5 | 72.7 | ||||
| ReLUN [66] | 71.0 | 67.9 | 72.1 | ||||
| Soft Exponential [67] | 37.9 | 30.1 | 25.0 | ||||
| YOLOv5n | NMS | GloU | SiLU | 86.3 | 84.4 | 88.5 | |
| SiLU | 86.8 | 83.0 | 88.6 | ||||
| Hardswish | 85.0 | 83.2 | 87.0 | ||||
| Mish | 87.4 | 81.9 | 87.9 | ||||
| EfficientMish | 86.8 | 80.9 | 87.2 | ||||
| FreLU | 86.5 | 80.9 | 85.5 | ||||
| Smoke | YOLOv5n + ESCFM | NMS | GloU | AconC | 86.4 | 80.5 | 86.3 |
| MetaAconC | 84.7 | 79.9 | 86.1 | ||||
| GELU | 85.8 | 81.9 | 87.2 | ||||
| Leaky ReLU | 85.7 | 83.5 | 87.6 | ||||
| ReLU | 87.8 | 80.7 | 87.2 | ||||
| ReLUN | 85.8 | 82.6 | 86.9 | ||||
| Soft Exponential | 63.9 | 54.8 | 53.4 |
| Component | Specification |
|---|---|
| Device | NVIDIA Jetson Xavier NX |
| GPU | 384 core NVIDIA Volta with 48 Tensor Cores |
| CPU | 6 core NVIDIA Carmel ARMv8.2 64 bit @ 1.4 GHz |
| Memory | 8 GB LPDDR4x (51.2 GB/s) |
| Storage | 16 GB eMMC 5.1 |
| Power mode | 20 W (6 core mode) |
| Operating system | Ubuntu 20.04 LTS (JetPack 5.0.2, L4T 35.1.0) |
| Inference resolution | 640 × 640 × 3 |
| Deployment | TensorRT (FP32) with layer fusion and kernel-level optimizations |
| Metrics collected | FPS, GPU utilization, RAM usage, and power consumption (30 run average) |
| Model | Params (M) Before | Params (M) After FP32 | TensorRT Engine Size (MB) |
|---|---|---|---|
| YOLOv5n | 1.80 | 1.76 | 7.04 |
| YOLOv9t | 2.00 | 1.97 | 7.88 |
| YOLOv11n | 2.60 | 2.58 | 10.33 |
| YOLOv12n | 2.60 | 2.56 | 10.23 |
| YOLOFM | 3.63 | 3.63 | 14.52 |
| YOLO-SAD | 7.02 | 7.01 | 28.05 |
| Proposed (Flame) | 1.89 | 1.78 | 7.12 |
| Proposed (Smoke) | 1.89 | 1.79 | 7.15 |
| Model | FPS | GPU Load (%) | RAM (MB) | Power (W) | ||||
|---|---|---|---|---|---|---|---|---|
| Single | Dual | Single | Dual | Single | Dual | Single | Dual | |
| YOLOv5n | 59.97 | 29.66 | 91 | 97 | 4012 | 5721 | 15.75 | 16.70 |
| YOLOv9t | 43.88 | 19.05 | 97 | 97 | 4075 | 5769 | 17.60 | 16.40 |
| YOLOv11n | 50.55 | 21.58 | 97 | 97 | 3846 | 5294 | 17.25 | 17.00 |
| YOLOv12n | 36.77 | 15.27 | 97 | 97 | 3921 | 6854 | 18.00 | 17.80 |
| YOLOFM | 28.31 | 11.66 | 98 | 98 | 4146 | 7195 | 18.10 | 18.30 |
| YOLO-SAD | 31.68 | 13.28 | 98 | 98 | 4049 | 6576 | 18.80 | 18.10 |
| Proposed (Flame) | 59.76 | 28.35 | 90 | 97 | 3952 | 5866 | 15.90 | 17.30 |
| Proposed (Smoke) | 59.96 | 28.34 | 90 | 97 | 3958 | 5866 | 15.70 | 17.00 |
| Videos | Dataset | Model | Precision (%) | Recall (%) | F1 Score (%) |
|---|---|---|---|---|---|
| YOLOv5n | 65.00 | 45.00 | 53.41 | ||
| YOLOv9t | 66.00 | 46.00 | 54.38 | ||
| YOLOv11n | 67.00 | 48.00 | 56.08 | ||
| Flame | YOLOv12n | 67.00 | 49.00 | 56.93 | |
| YOLOFM | 68.00 | 49.00 | 57.55 | ||
| YOLO-SAD | 68.00 | 50.00 | 58.06 | ||
| Japan VR [68] | YOLOv5n + ESCFM++ (ours) | 58.47 | 59.84 | 59.15 | |
| YOLOv5n | 100.00 | 44.86 | 61.93 | ||
| YOLOv9t | 100.00 | 40.46 | 57.61 | ||
| YOLOv11n | 100.00 | 40.00 | 57.14 | ||
| Smoke | YOLOv12n | 95.00 | 47.00 | 62.89 | |
| YOLOFM | 94.00 | 46.00 | 61.77 | ||
| YOLO-SAD | 96.00 | 48.00 | 64.00 | ||
| YOLOv5n + ESCFM-RS (ours) | 100.00 | 50.20 | 66.84 | ||
| YOLOv5n | 73.00 | 69.00 | 71.00 | ||
| YOLOv9t | 84.77 | 86.08 | 85.42 | ||
| YOLOv11n | 87.13 | 88.00 | 87.56 | ||
| YOLOv12n | 86.00 | 85.00 | 85.50 | ||
| YOLOFM | 85.50 | 85.70 | 85.60 | ||
| YOLO-SAD | 86.20 | 86.00 | 86.10 | ||
| USA VES [69] | YOLOv5n + ESCFM++ (ours) | 67.17 | 91.75 | 77.56 | |
| YOLOv5n | 88.57 | 46.27 | 60.78 | ||
| YOLOv9tn | 100.00 | 31.34 | 47.72 | ||
| YOLOv11n | 100.00 | 31.58 | 48.00 | ||
| YOLOv12n | 90.00 | 62.00 | 73.42 | ||
| YOLOFM | 89.00 | 61.00 | 72.39 | ||
| YOLO-SAD | 91.00 | 63.00 | 74.45 | ||
| YOLOv5n + ESCFM-RS (ours) | 93.75 | 67.16 | 78.26 | ||
| YOLOv5n | 93.52 | 70.62 | 80.48 | ||
| YOLOv9tn | 96.47 | 57.34 | 70.10 | ||
| YOLOv11n | 98.15 | 63.10 | 76.81 | ||
| YOLOv12n | 95.00 | 58.20 | 71.80 | ||
| YOLOFM | 95.70 | 59.00 | 72.55 | ||
| YOLO-SAD | 96.20 | 60.10 | 74.01 | ||
| USA WSF [70] | YOLOv5n + ESCFM++ (ours) | 77.01 | 93.71 | 84.54 | |
| YOLOv5n | 87.34 | 67.65 | 76.24 | ||
| YOLOv9tn | 96.61 | 55.88 | 70.80 | ||
| YOLOv11n | 97.44 | 63.33 | 76.77 | ||
| YOLOv12n | 90.00 | 95.00 | 92.43 | ||
| YOLOFM | 89.00 | 94.00 | 91.43 | ||
| YOLO-SAD | 91.00 | 95.00 | 92.96 | ||
| YOLOVv5n + ESCFM-RS (ours) | 93.52 | 99.02 | 96.19 |
| Comparison Pair | T-Statistic | p-Value | Cohen’s d | Significance |
|---|---|---|---|---|
| YOLOv5n vs. Proposed (Flame, ESCFM++) | 4.85 | 0.003 | 1.25 | Yes |
| YOLOv5n vs. Proposed (Smoke, ESCFM-RS) | 4.42 | 0.004 | 1.18 | Yes |
| Model | BBox Area Ratio | Precision (%) | Recall (%) | F1-Score (%) |
|---|---|---|---|---|
| YOLOv5n | >15% of image area | 100.00 | 97.86 | 98.92 |
| YOLOv5n | ≤15% of Image area | 53.36 | 54.92 | 54.13 |
| Proposed | >15% of Image area | 100.00 | 99.15 | 99.57 |
| Proposed | ≤15% of Image area | 100.00 | 37.67 | 54.72 |
| Model | State | Precision (%) | Recall (%) | F1-Score (%) |
|---|---|---|---|---|
| YOLOv5n | Daytime | 59.89 | 62.06 | 60.96 |
| YOLOv5n | Night | 67.17 | 91.75 | 77.56 |
| Proposed | Daytime | 99.31 | 46.41 | 63.26 |
| Proposed | Night | 77.93 | 85.57 | 81.57 |
| Model | State | Precision (%) | Recall (%) | F1-Score (%) |
|---|---|---|---|---|
| YOLOv5n | Daytime | 99.49 | 49.64 | 66.23 |
| YOLOv5n | Night | 99.12 | 56.07 | 71.63 |
| Proposed | Daytime | 99.26 | 48.59 | 65.24 |
| Proposed | Night | 99.03 | 67.72 | 80.43 |
| Condition | YOLOv5n Precision (%) | YOLOv5n Recall (%) | YOLOv5n F1 (%) | Proposed Precision (%) | Proposed Recall (%) | Proposed F1 (%) | ΔF1 (p.p.) |
|---|---|---|---|---|---|---|---|
| Daytime 1 | 95.77 | 69.39 | 80.47 | 90.29 | 94.9 | 92.54 | 12.06 |
| Daytime 2 | 92.31 | 77.42 | 84.21 | 89.55 | 93.75 | 91.6 | 7.39 |
| Nighttime 1 | 100 | 100 | 100 | 100 | 98.8 | 99.39 | −0.61 |
| Nighttime 2 | 96.15 | 97.09 | 96.62 | 94.81 | 97.57 | 96.17 | −0.45 |
| Condition | YOLOv5n Precision (%) | YOLOv5n Recall (%) | YOLOv5n F1 (%) | Proposed Precision (%) | Proposed Recall (%) | Proposed F1 (%) | ΔPrecision (p.p.) | ΔRecall (p.p.) | ΔF1 (p.p.) |
|---|---|---|---|---|---|---|---|---|---|
| Daytime | 94.04 | 73.4 | 82.34 | 89.92 | 94.32 | 92.07 | −4.12 | 20.92 | 9.73 |
| Nighttime | 98.08 | 98.54 | 98.31 | 97.41 | 98.18 | 97.78 | −0.67 | −0.36 | −0.53 |
| Overall | 96.06 | 85.97 | 90.33 | 93.66 | 96.25 | 94.93 | −2.4 | 10.28 | 4.6 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Park, J.-C.; Kim, M.; Choi, S.-M.; Kim, G.-W. ESCFM-YOLO: Lightweight Dual-Stream Architecture for Real-Time Small-Scale Fire Smoke Detection on Edge Devices. Appl. Sci. 2026, 16, 778. https://doi.org/10.3390/app16020778
Park J-C, Kim M, Choi S-M, Kim G-W. ESCFM-YOLO: Lightweight Dual-Stream Architecture for Real-Time Small-Scale Fire Smoke Detection on Edge Devices. Applied Sciences. 2026; 16(2):778. https://doi.org/10.3390/app16020778
Chicago/Turabian StylePark, Jong-Chan, Myeongjun Kim, Sang-Min Choi, and Gun-Woo Kim. 2026. "ESCFM-YOLO: Lightweight Dual-Stream Architecture for Real-Time Small-Scale Fire Smoke Detection on Edge Devices" Applied Sciences 16, no. 2: 778. https://doi.org/10.3390/app16020778
APA StylePark, J.-C., Kim, M., Choi, S.-M., & Kim, G.-W. (2026). ESCFM-YOLO: Lightweight Dual-Stream Architecture for Real-Time Small-Scale Fire Smoke Detection on Edge Devices. Applied Sciences, 16(2), 778. https://doi.org/10.3390/app16020778

