Real-Time Detection of Unsafe Worker Behaviors via Adaptive Vision Transformers in Construction Sites
Abstract
1. Introduction
2. Background
2.1. Application of Image Enhancement in Construction Safety Monitoring
2.2. Application of Real-Time Behavior Recognition Technology in Construction Safety Monitoring
3. Methodology
3.1. DAIE
3.2. R-BehaviorNet
3.3. Dual-Stream Processing Architecture and Adaptive Threshold Adjustment
4. Illustrative Examples
4.1. Datasets and Preprocessing
4.2. Baseline Methods and Evaluation Metrics
- Accuracy: Measures the model’s ability to correctly identify positive and negative samples. High accuracy in a safety monitoring system means the model can reliably distinguish between safe and unsafe behaviors, which is critical for the system’s practical usability.
- Recall: Particularly important in safety-related applications, as it measures the model’s ability to identify all actual positive samples (in this context, unsafe behaviors). High recall ensures that nearly all unsafe behaviors are detected, reducing the likelihood of accidents.
- Precision: Refers to the proportion of true positive samples among the samples predicted as positive. In construction safety management, high precision means the system’s alerts are more trustworthy, reducing false alarms and enhancing workers’ trust and response to the safety alert system.
- F1 Score: The harmonic mean of precision and recall, particularly useful in cases where there is an imbalance between positive and negative samples. In construction scenarios, where unsafe events may occur less frequently, the F1 score becomes an important metric as it balances precision and recall, providing a single measure to evaluate the overall performance of the model.
4.3. Experimental Results and Performance Analysis
4.3.1. Performance Evaluation of DAIE Compared to Other Image Processing Methods
4.3.2. Performance Comparison of R-BehaviorNet with Other Machine Learning Models
4.3.3. Visualization Analysis
4.3.4. Category-Specific Recognition Performance
4.3.5. Qualitative Comparison and Visual Analysis of Model Responses
4.3.6. Ablation Study on Module Contributions
4.3.7. Dual-Stream vs. Single-Stream Comparison
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Nguyen, P.H.D.; Tran, D. Exploring the Use of Quality Control Plans for Alternative Contracting Methods in Highway Projects. J. Constr. Eng. Manag. 2024, 150, 04024030. [Google Scholar] [CrossRef]
- Jiang, W.G.; Ding, L.Y. Unsafe hoisting behavior recognition for tower crane based on transfer learning. Autom. Constr. 2024, 160, 105299. [Google Scholar] [CrossRef]
- Golafshani, E.; Khodadadi, N.; Ngo, T.; Nanni, A.; Behnood, A. Modelling the compressive strength of geopolymer recycled aggregate concrete using ensemble machine learning. Adv. Eng. Softw. 2024, 191, 103611. [Google Scholar] [CrossRef]
- Song, H.Q.; Lin, B.C.; Xie, L.C. Dual-Stream Fusion and Multi-scale Analysis: Introducing the Synergistic Dual-Stream Network (SDS-Net) for Image Manipulation Segmentation. Adv. Intell. Syst. 2024, 6, 2300749. [Google Scholar] [CrossRef]
- Zhou, W.Q.; Wu, H.L.; Deng, P. Toward Efficient and Accurate Remote Sensing Image-Text Retrieval with a Coarse-to-Fine Approach. IEEE Geosci. Remote Sens. Lett. 2025, 22, 6000305. [Google Scholar] [CrossRef]
- Pan, X.; Shen, L.X.; Zhong, B.T.; Sheng, D.; Huang, F.; Yang, L.H. Novel blockchain deep learning framework to ensure video security and lightweight storage for construction safety management. Adv. Eng. Inform. 2024, 59, 102334. [Google Scholar] [CrossRef]
- Guo, Z.; She, J.; Li, Z.; Du, J.; Ye, S. Integrating FRAM and BN for enhanced resilience evaluation in construction emergency response: A scaffold collapse case study. Heliyon 2024, 10, e25342. [Google Scholar] [CrossRef]
- Zucca, M.; Tattoni, S.; Di Castri, M.; Simoncelli, M. On the collapse of a post-tensioned reinforced concrete truss bridge during the construction phases. Eng. Fail. Anal. 2024, 158, 107999. [Google Scholar] [CrossRef]
- Kaur, R.; Singh, J.; Sharma, S. Enhanced Helmet Detection in Surveillance Systems with YOLOv6 for Accident Prevention and Safety Compliance. J. Sci. Ind. Res. 2025, 84, 601–613. [Google Scholar] [CrossRef]
- Wu, Z.; Lei, X.; Kumar, M. Advancing construction safety: YOLOv8-CGS helmet detection model. PLoS ONE 2025, 20, e0321713. [Google Scholar] [CrossRef]
- Zhang, Y.; Huang, S.; Qin, J.; Li, X.; Zhang, Z.; Fan, Q.; Tan, Q. Detection of helmet use among construction workers via helmet-head region matching and state tracking. Autom. Constr. 2025, 171, 105987. [Google Scholar] [CrossRef]
- Yuan, H.; Yang, H.; Li, R.Q.; Wang, J.; Tian, L. Personal safety monitoring system of electric power construction site based on AIoT Technology. J. Intell. Fuzzy Syst. 2024, 46, 493–504. [Google Scholar] [CrossRef]
- Xiang, C.C.; Yin, D.F.; Song, F.; Yu, Z.X.; Jian, X.; Gong, H.M. A Fast and Robust Safety Helmet Network Based on a Mutilscale Swin Transformer. Buildings 2024, 14, 688. [Google Scholar] [CrossRef]
- Han, J.; Yoon, S.; Kang, M.; Kim, T. Approach to Enhancing Panoramic Segmentation in Indoor Construction Sites Based on a Perspective Image Segmentation Foundation Model. Appl. Sci. 2025, 15, 4875. [Google Scholar] [CrossRef]
- Kang, S.; Kim, S.; Kim, G.-H. Integration of real-time labor positioning data and 3D laser scan model for dangerous zone access monitoring. J. Build. Eng. 2025, 111, 113445. [Google Scholar] [CrossRef]
- Jiang, Q.; Jia, M.T.; Bi, L.; Zhuang, Z.; Gao, K.X. Development of a core feature identification application based on the Faster R-CNN algorithm. Eng. Appl. Artif. Intell. 2022, 115, 105200. [Google Scholar] [CrossRef]
- Gao, F.Q.; Zhu, Q.Y.; Shao, G.F.; Su, Y.K.; Yang, J.B.; Yu, X.Y. A fast surface-defect detection method based on Dense-YOLO network. Caai Trans. Intell. Technol. 2025, 10, 415–433. [Google Scholar] [CrossRef]
- Nautiyal, D.; Dhir, M.; Singh, T.; Saini, A.; Handa, P. Real-Time, Multi-Task Mobile Application for Automatic Bleeding and Non-Bleeding Frame Analysis in Video Capsule Endoscopy Using an Ensemble of Faster R-CNN and LinkNet. Int. J. Imaging Syst. Technol. 2025, 35, e70171. [Google Scholar] [CrossRef]
- Walia, J.S.; Haridass, K.; Pavithra, L.K. Deep Learning Innovations for Underwater Waste Detection: An In-Depth Analysis. IEEE Access 2025, 13, 88917–88929. [Google Scholar] [CrossRef]
- Xia, C.J.; Ren, M.; Liu, R.Y.; Tian, Z.L.; Song, M.Y.; Dong, M.; Zhang, T.; Miao, J. Tracking moisture contents in the pollution layer on a composite insulator surface using hyperspectral imaging technology. Analyst 2024, 149, 2996–3007. [Google Scholar] [CrossRef]
- Zhang, X.Y.; Yang, S.; Yang, X.; Li, C.; Xu, Y. A Triplet Network Fusing Optical and SAR Images for Colored Steel Building Extraction. Sensors 2024, 24, 89. [Google Scholar] [CrossRef]
- Chen, H.H.; Li, Y.Y.; Wen, H.X.; Hu, X.D. YOLOv5s-gnConv: Detecting personal protective equipment for workers at height. Front. Public Health 2023, 11, 1225478. [Google Scholar] [CrossRef]
- Han, G.J.; Wang, R.X.; Xu, W.Y.; Li, J. Night construction site detection based on ghost-YOLOX. Connect. Sci. 2024, 36, 2316015. [Google Scholar] [CrossRef]
- Rane, M.; Kulkarni, M.; Dalvi, A.; Kulkarni, A.; Singh, A.; Bhave, A.; Arawat, V. AI Algorithm for Image Enhancement. In Proceedings of the 2023 Third International Conference on Advances in Electrical, Computing, Communication and Sustainable Technologies (ICAECT), Bhilai, India, 5–6 January 2023; pp. 1–4. [Google Scholar]
- Brahmaji, R.K.N.; Nagesh, K.K.; Indra, N.M.V.S.S.; Krishna, V.K.C.D.; Premchand, R.M.; Sai, K.P. Image Enhancement of Low Light Image using Deep Learning. Int. J. Adv. Res. Sci. Commun. Technol. 2023, 3, 511–517. [Google Scholar] [CrossRef]
- Blanch, X.; Guinau, M.; Eltner, A.; Abellan, A. Fixed photogrammetric systems for natural hazard monitoring with high spatio-temporal resolution. Nat. Hazards Earth Syst. Sci. 2023, 23, 3285–3303. [Google Scholar] [CrossRef]
- Li, J.Q.; Miao, Q.; Zou, Z.; Gao, H.G.; Zhang, L.X.; Li, Z.B.; Wang, N. A Review of Computer Vision-Based Monitoring Approaches for Construction Workers’ Work-Related Behaviors. IEEE Access 2024, 12, 7134–7155. [Google Scholar] [CrossRef]
- Hu, H.L.; Zheng, X.Y. Augmented and Virtual Reality-Based Cyber Twin Model for Observing Infants in Intensive Care: 6G for Smart Healthcare 4.0 by Machine Learning Techniques. Wirel. Pers. Commun. 2024, 4, 1–17. [Google Scholar] [CrossRef]
- Wang, Z.; Chen, Z.Y.; Ma, L.; Wang, Q.; Wang, H.; Leal, A., Jr.; Li, X.L.; Marques, C.; Min, R. Optical Microfiber Intelligent Sensor: Wearable Cardiorespiratory and Behavior Monitoring with a Flexible Wave-Shaped Polymer Optical Microfiber. ACS Appl. Mater. Interfaces 2024, 16, 8333–8345. [Google Scholar] [CrossRef]
- Wang, Z.; Hua, Z.X.; Wen, Y.C.; Zhang, S.J.; Xu, X.S.; Song, H.B. E-YOLO: Recognition of estrus cow based on improved YOLOv8n model. Expert Syst. Appl. 2024, 238, 122212. [Google Scholar] [CrossRef]
- Li, H.W.; Ni, Y.Q.; Wang, Y.W.; Chen, Z.W.; Rui, E.Z.; Xu, Z.D. Modeling of forced-vibration systems using continuous-time state-space neural network. Eng. Struct. 2024, 302, 117329. [Google Scholar] [CrossRef]
- Bingu, R.; Jothilakshmi, S.; Srinivasu, N. An intelligent multiclass deep classifier-based intrusion detection system for cloud environment. Concurr. Comput. Pract. Exp. 2023, 35, e7840. [Google Scholar] [CrossRef]
- Xiong, C.Y.; Wang, Z.L.; Huang, X.Y. Modelling flame-to-fuel heat transfer by deep learning and fire images. Eng. Appl. Comput. Fluid Mech. 2024, 18, 2331114. [Google Scholar] [CrossRef]
- Walter, T.; Degen, J.; Pfeiffer, K.; Stöckl, A.; Montenegro, S.; Degen, T. A new innovative real-time tracking method for flying insects applicable under natural conditions. BMC Zool. 2021, 6, 35. [Google Scholar] [CrossRef] [PubMed]
- Liu, T.X. Assessing implicit computational thinking in game-based learning: A logical puzzle game study. Br. J. Educ. Technol. 2024, 55, 2357–2382. [Google Scholar] [CrossRef]
- Luo, P.; Niu, Y.P.; Tang, D.X.; Huang, W.Y.; Luo, X.F.; Mu, J. A computer vision solution for behavioral recognition in red pandas. Sci. Rep. 2025, 15, 9201. [Google Scholar] [CrossRef]
- He, J.Y.; Li, C.; Xie, Y.; Luo, H.T.; Zheng, W.; Wang, Y.Q. RMTSE: A Spatial-Channel Dual Attention Network for Driver Distraction Recognition. Sensors 2025, 25, 2821. [Google Scholar] [CrossRef]
- Xie, Y.X.; He, Y.G.; Cheng, A.B.; Zhang, J.W. Study on medical image enhancement based on IFOA improved grayscale image adaptive enhancement. Multimed. Tools Appl. 2016, 75, 14367–14379. [Google Scholar] [CrossRef]
- Liang, X.W.; Chen, X.Y.; Ren, K.Y.; Miao, X.; Chen, Z.H.; Jin, Y.T. Low-light image enhancement via adaptive frequency decomposition network. Sci. Rep. 2023, 13, 14107. [Google Scholar] [CrossRef] [PubMed]
- Tan, H.S.; Zhou, F.Q.; Xiong, Y.; Li, X.K. Adaptive enhancement of image brightness and contrast based on neural networks. J. Optoelectron. Laser 2010, 21, 1881–1884. [Google Scholar]
- Hu, K.; Zhang, Y.W.; Lu, F.Y.; Deng, Z.L.; Liu, Y.P. An Underwater Image Enhancement Algorithm Based on MSR Parameter Optimization. J. Mar. Sci. Eng. 2020, 8, 741. [Google Scholar] [CrossRef]
- Zhang, Y.Y.; Huang, Y.; Lu, B.S.; Ma, Y.M.; Qiu, J.H.; Zhao, Y.N.; Guo, X.H.; Liu, C.X.; Liu, P.; Zhang, Y.G. Real-time sitting behavior tracking and analysis for rectification of sitting habits by strain sensor-based flexible data bands. Meas. Sci. Technol. 2020, 31, 055102. [Google Scholar] [CrossRef]
- Li, S.; Shi, Q. Deep Learning-Based Human Action Recognition in Videos. J. Circuits Syst. Comput. 2025, 34, 2550040. [Google Scholar] [CrossRef]
- Yang, Z.; Zheng, N.; Wang, F. DSSFN: A Dual-Stream Self-Attention Fusion Network for Effective Hyperspectral Image Classification. Remote Sens. 2023, 15, 3701. [Google Scholar] [CrossRef]
- Zhang, W.; Guo, X.; Wang, J.; Wang, N.; Chen, K. Asymmetric Adaptive Fusion in a Two-Stream Network for RGB-D Human Detection. Sensors 2021, 21, 916. [Google Scholar] [CrossRef]
- Erfani, A.; Mansouri, A. Applications of multimodal large language models in construction industry. Adv. Eng. Inform. 2026, 69, 103909. [Google Scholar] [CrossRef]









| Dataset | Category | Training Set Quantity | Validation Set Quantity | Test Set Quantity | Notes |
|---|---|---|---|---|---|
| Construction Action Recognition Dataset | Safety Helmet and Belt Usage | 1398 | 401 | 201 | Various construction site scenarios and safety actions |
| Construction Site Safety Image Dataset | Safety Helmet Usage | 2457 | 702 | 341 | Labeled for correct or incorrect usage of safety helmets |
| Construction Site Safety Image Dataset | Safety Belt Usage | 1748 | 501 | 251 | Labeled for correct or incorrect usage of safety belts |
| Technology | Accuracy | Precision | Recall | F1 Score |
|---|---|---|---|---|
| DAIE | 90.0% | 90.5% | 89.0% | 89.5% |
| GAN | 85.5% | 83.0% | 84.5% | 83.7% |
| ResNet | 83.8% | 81.4% | 82.9% | 82.1% |
| Attention-based CNN | 84.3% | 84.2% | 83.7% | 84.0% |
| Technology | Accuracy | Precision | Recall | F1 Score |
|---|---|---|---|---|
| R-BehaviorNet | 92.0% | 93.0% | 91.0% | 92.0% |
| Faster R-CNN | 89.5% | 90.0% | 89.0% | 89.5% |
| YOLO | 88.0% | 87.5% | 88.5% | 88.0% |
| Mask R-CNN | 90.5% | 91.0% | 90.0% | 90.5% |
| Unsafe Behavior Type | Precision | Recall | F1-Score |
|---|---|---|---|
| Not wearing safety helmet | 94.3% | 91.8% | 93.0% |
| Improper safety-belt use | 92.5% | 89.7% | 91.0% |
| Climbing without protection | 89.4% | 87.6% | 88.5% |
| Standing close to operating machinery | 90.2% | 88.1% | 89.1% |
| Configuration | Precision | Recall | F1-Score | Improvement (ΔF1) |
|---|---|---|---|---|
| Not wearing safety helmet | 88.2% | 85.3% | 86.7% | - |
| Improper safety-belt use | 90.5% | 88.1% | 89.5% | +2.8 |
| Climbing without protection | 92.1% | 89.2% | 90.6% | +3.9 |
| Standing close to operating machinery | 93.2% | 91.4% | 92.3% | +5.6 |
| Model Configuration | Precision | Recall | F1-Score | ΔF1 vs. Raw Stream |
|---|---|---|---|---|
| Raw stream only | 89.1% | 88.3% | 88.7% | - |
| Enhanced stream only | 91.6% | 89.3% | 90.4% | +1.7 |
| Dual stream | 93.2% | 91.4% | 92.3% | +3.6 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Alotaibi, R.T.T.; Ma, S. Real-Time Detection of Unsafe Worker Behaviors via Adaptive Vision Transformers in Construction Sites. Buildings 2025, 15, 4205. https://doi.org/10.3390/buildings15224205
Alotaibi RTT, Ma S. Real-Time Detection of Unsafe Worker Behaviors via Adaptive Vision Transformers in Construction Sites. Buildings. 2025; 15(22):4205. https://doi.org/10.3390/buildings15224205
Chicago/Turabian StyleAlotaibi, Rami Talal T., and Shengbin Ma. 2025. "Real-Time Detection of Unsafe Worker Behaviors via Adaptive Vision Transformers in Construction Sites" Buildings 15, no. 22: 4205. https://doi.org/10.3390/buildings15224205
APA StyleAlotaibi, R. T. T., & Ma, S. (2025). Real-Time Detection of Unsafe Worker Behaviors via Adaptive Vision Transformers in Construction Sites. Buildings, 15(22), 4205. https://doi.org/10.3390/buildings15224205

