Next Article in Journal
Mechanism Analysis of Monnex Fire Extinguishing Performance and Particular Burning Fragmentation Phenomenon
Next Article in Special Issue
Sensor Layout Optimization and Natural Gas Leakage Source Term Estimation Based on Non-Dominated Sorting Genetic Algorithm
Previous Article in Journal
Simple Spread Models for Understory Surface Fires
Previous Article in Special Issue
Suppression Effects and Mechanisms of Fine Water Mist on Methane Explosions in Large-Scale Roadways via Experimental and CFD Studies
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

SemaFire-YOLO: A Lightweight and Robust Fire-Smoke Detection Model via Semantic Enhancement and Frequency-Aware Perception

1
Department of Emergency Communications and Information Engineering, China Fire and Rescue Institute, Beijing 102202, China
2
School of Safety Science, Tsinghua University, Beijing 100084, China
3
School of Automation Science and Electrical Engineering, Beihang University, Beijing 100191, China
*
Author to whom correspondence should be addressed.
Fire 2026, 9(7), 303; https://doi.org/10.3390/fire9070303
Submission received: 14 May 2026 / Revised: 26 June 2026 / Accepted: 10 July 2026 / Published: 16 July 2026
(This article belongs to the Special Issue Fire and Explosion Safety with Risk Assessment and Early Warning)

Abstract

Accurate detection in the early stages of a fire is a crucial prerequisite for the efficient implementation of fire suppression and emergency rescue operations. Its accuracy and timeliness directly affect the control of disaster loss severity. Traditional fire detection methods mainly include three categories, which are manual inspection, sensor detection, and visual recognition. However, manual inspection is restricted by labor costs and time efficiency, making it difficult to achieve large-scale, high-frequency and real-time fire monitoring. Sensor detection is easily interfered by environmental factors such as temperature, humidity, and dust, leading to frequent false alarms and missed alarms. Visual recognition technology has shortcomings in aspects such as detailed feature perception, dynamic scene modeling, and reasoning robustness in complex environments, making it difficult to meet the requirements of high-precision detection. To address these issues, this study innovatively proposes a lightweight fire and smoke detection model based on semantic enhancement and frequency domain perception modeling, which is named the SemaFire you only look once (SemaFire-YOLO) model. The model constructs a large language and vision assistant (LLaVA) semantic guidance module, which uses a large language model to understand and guide the semantic features of images, thereby enhancing the saliency representation intensity of small and weak target regions. Then, a Haar wavelet-based downsampling module is adopted, which compresses spatial information while preserving high-frequency features such as flame edges and smoke textures, improving the accuracy of target recognition. Next, the convolution modulation mechanism is introduced to replace the traditional attention mechanism, enhancing the overall modeling efficiency and reducing computational overhead. Finally, a Dynamic Tanh normalization module is adopted to replace the batch normalization module in the traditional YOLO algorithm, strengthening the model’s representation stability and reasoning robustness under unstable input distributions. Experimental results show that the SemaFire-YOLO model achieves a mean average precision (mAP@0.5) of 64.30% on the fire image dataset, which is 0.8, 2.0, 0.6, and 3.8 percentage points higher than that of mainstream models such as YOLOv5n, YOLOv8n, YOLOv11n, and YOLOv12n, respectively. It exhibits better boundary detection capability and practical deployment potential. Through visual analysis, the results indicate that the improved SemaFire-YOLO model achieves more accurate detection and higher confidence in actual complex scenarios, further verifying the model’s robustness and accuracy in complex scenarios such as low contrast and dynamic fire conditions.
Keywords: fire detection; smoke recognition; YOLO; large language model; wavelet transform; convolutional modulation; dynamic normalization fire detection; smoke recognition; YOLO; large language model; wavelet transform; convolutional modulation; dynamic normalization

Share and Cite

MDPI and ACS Style

Pei, J.; Zhang, R.; Yan, H.; Hao, Y.; Huang, Y.; Xiao, J. SemaFire-YOLO: A Lightweight and Robust Fire-Smoke Detection Model via Semantic Enhancement and Frequency-Aware Perception. Fire 2026, 9, 303. https://doi.org/10.3390/fire9070303

AMA Style

Pei J, Zhang R, Yan H, Hao Y, Huang Y, Xiao J. SemaFire-YOLO: A Lightweight and Robust Fire-Smoke Detection Model via Semantic Enhancement and Frequency-Aware Perception. Fire. 2026; 9(7):303. https://doi.org/10.3390/fire9070303

Chicago/Turabian Style

Pei, Jiaxu, Ruihuan Zhang, Hualong Yan, Yulu Hao, Yu Huang, and Jin Xiao. 2026. "SemaFire-YOLO: A Lightweight and Robust Fire-Smoke Detection Model via Semantic Enhancement and Frequency-Aware Perception" Fire 9, no. 7: 303. https://doi.org/10.3390/fire9070303

APA Style

Pei, J., Zhang, R., Yan, H., Hao, Y., Huang, Y., & Xiao, J. (2026). SemaFire-YOLO: A Lightweight and Robust Fire-Smoke Detection Model via Semantic Enhancement and Frequency-Aware Perception. Fire, 9(7), 303. https://doi.org/10.3390/fire9070303

Article Metrics

Back to TopTop