Next Article in Journal
Numerical Modelling of Flat Slabs with Different Amounts of Double-Headed Studs as Punching Shear Reinforcement
Next Article in Special Issue
High-Frequency Refined Mamba with Snake Perception Attention for More Accurate Crack Segmentation
Previous Article in Journal
Design of a One-Dimensional Zn3In2S6/NiFe2O4 Composite Material and Its Photocathodic Protection Mechanism Against Corrosion
Previous Article in Special Issue
Evaluating Radiance Field-Inspired Methods for 3D Indoor Reconstruction: A Comparative Analysis
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Enabling High-Level Worker-Centric Semantic Understanding of Onsite Images Using Visual Language Models with Attention Mechanism and Beam Search Strategy

1
School of Civil Engineering and Transportation, South China University of Technology, Guangzhou 510641, China
2
State Key Laboratory of Subtropical Building and Urban Science, Guangzhou 510641, China
3
Department of Civil Engineering, Tsinghua University, Beijing 100084, China
*
Author to whom correspondence should be addressed.
Buildings 2025, 15(6), 959; https://doi.org/10.3390/buildings15060959
Submission received: 7 February 2025 / Revised: 12 March 2025 / Accepted: 14 March 2025 / Published: 18 March 2025
(This article belongs to the Special Issue Intelligence and Automation in Construction Industry)

Abstract

Visual information is becoming increasingly essential in construction management. However, a significant portion of this information remains underutilized by construction managers due to the limitations of existing image processing algorithms. These algorithms primarily rely on low-level visual features and struggle to capture high-order semantic information, leading to a gap between computer-generated image semantics and human interpretation. However, current research lacks a comprehensive justification for the necessity of employing scene understanding algorithms to address this issue. Moreover, the absence of large-scale, high-quality open-source datasets remains a major obstacle, hindering further research progress and algorithmic optimization in this field. To address this issue, this paper proposes a construction scene visual language model based on attention mechanism and encoder–decoder architecture, with the encoder built using ResNet101 and the decoder built using LSTM (long short-term memory). The addition of the attention mechanism and beam search strategy improves the model, making it more accurate and generalizable. To verify the effectiveness of the proposed method, a publicly available construction scene visual-language dataset containing 16 common construction scenes, SODA-ktsh, is built and verified. The experimental results demonstrate that the proposed model achieves a BLEU-4 score of 0.7464, a CIDEr score of 5.0255, and a ROUGE_L score of 0.8106 on the validation set. These results indicate that the model effectively captures and accurately describes the complex semantic information present in construction images. Moreover, the model exhibits strong generalization, perceptual, and recognition capabilities, making it well suited for interpreting and analyzing intricate construction scenes.
Keywords: visual language model; construction scene; image scene understanding; image captioning; attention mechanism visual language model; construction scene; image scene understanding; image captioning; attention mechanism

Share and Cite

MDPI and ACS Style

Deng, H.; Fu, K.; Yu, B.; Li, H.; Duan, R.; Deng, Y.; Lin, J.-r. Enabling High-Level Worker-Centric Semantic Understanding of Onsite Images Using Visual Language Models with Attention Mechanism and Beam Search Strategy. Buildings 2025, 15, 959. https://doi.org/10.3390/buildings15060959

AMA Style

Deng H, Fu K, Yu B, Li H, Duan R, Deng Y, Lin J-r. Enabling High-Level Worker-Centric Semantic Understanding of Onsite Images Using Visual Language Models with Attention Mechanism and Beam Search Strategy. Buildings. 2025; 15(6):959. https://doi.org/10.3390/buildings15060959

Chicago/Turabian Style

Deng, Hui, Kejie Fu, Binglin Yu, Huimin Li, Rui Duan, Yichuan Deng, and Jia-rui Lin. 2025. "Enabling High-Level Worker-Centric Semantic Understanding of Onsite Images Using Visual Language Models with Attention Mechanism and Beam Search Strategy" Buildings 15, no. 6: 959. https://doi.org/10.3390/buildings15060959

APA Style

Deng, H., Fu, K., Yu, B., Li, H., Duan, R., Deng, Y., & Lin, J.-r. (2025). Enabling High-Level Worker-Centric Semantic Understanding of Onsite Images Using Visual Language Models with Attention Mechanism and Beam Search Strategy. Buildings, 15(6), 959. https://doi.org/10.3390/buildings15060959

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop