Next Article in Journal
Post-Socialist Churches and Parish Complexes in Modernist New Towns: Typologies of Spatial Integration in Zagreb
Previous Article in Journal
Children’s Perception of Urban Outdoor Spaces and Playground Design: A Sensory Walk Study in Zagreb, Croatia
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Automated Hazard Identification and Visualisation in Design Using Building Information Modelling and Machine Learning

1
Aston Business School, Aston University, Birmingham B4 7ET, UK
2
School of Built Environment, Engineering and Computing, Leeds Beckett University, Leeds LS2 8AG, UK
3
Hertfordshire Business School, University of Hertfordshire, Hatfield AL10 9EU, UK
*
Authors to whom correspondence should be addressed.
Architecture 2026, 6(2), 93; https://doi.org/10.3390/architecture6020093
Submission received: 30 April 2025 / Revised: 19 July 2025 / Accepted: 20 August 2025 / Published: 9 June 2026

Abstract

The construction industry is recognised globally as one of the most hazardous sectors. Effective hazard management necessitates identifying and communicating these risks early in the project lifecycle. Construction Hazard Prevention through Design (CHPtD) addresses this by incorporating safety information into the design phase that is often cumbersome and heavily reliant on reviewer expertise. The present work enhances hazard recognition and visualisation by automating the process using computational intelligence and building information modelling, aligning with the theoretical framework of CHPtD. The proposed tool provides detailed hazard information, including the nature of the hazard, its causes, and potential resolutions, empowering designers to make informed decisions and mitigate risks proactively. The tool’s performance is evaluated using a confusion matrix, demonstrating promising results with an overall accuracy of 84.77% and a Kappa coefficient of 0.83. While the tool shows strong performance in identifying several hazard classes, further refinement is needed to improve its ability to detect catastrophic events and manage traffic-related hazards.

1. Introduction

The construction industry is globally recognised as one of the most hazardous sectors, a characteristic attributed to the inherent nature of its operations [1]. Within this context, the inability to identify and quantify potential hazards during the transformation of a design model into a completed structure directly correlates with an increased incidence of adverse events, including injuries and fatalities. Insufficient recognition or under-reporting [2] of these hazards during upstream project phases precludes effective risk mitigation. Mitigating these injuries requires management of all possible hazards by identifying and communicating these hazards before they appear during the physical construction phase [3]. Hazard identification is a continuous process during the complete lifecycle of a construction project [4] because new or unidentified hazards can occur with varying environmental parameters. However, researchers have comprehensively attempted to address the mitigation of hazards during design phase by embedding safety information within the design model [5,6]. Construction Hazard Prevention through Design (CHPtD) is considered an eminent process in safety risk reduction that concerns hazard identification and reporting during the design phase of a project [5,7]. The CHPtD process utilises design information either in the form of two-dimensional drawings or three-dimensional building information models (BIMs). The present proposal is aligned with CHPtD and enhances hazard recognition by automating the process using computational intelligence [8].
CHPtD follows the theoretical details of the Prevention through Design (PtD) process [1,7]. PtD categorises hazards as related to work methods, processes, equipment, operations, tools, technologies and industrial products [9]. CHPtD is a review process for recognising hazards from design information related to construction, repairing, maintenance and demolition activities. Along with a review of the design information for hazard recognition, CHPtD also projects the intensity of the hazards and controls it by removing or substituting elements that are considered as source of the hazard. However, this review process is a cumbersome task and is highly dependent on the expertise of the reviewer, which can lead to accidents, as reported by several researchers [2,10]. Research on hazard recognition based on design information is mainly carried out with 2D designs, and very few studies have conducted naïve experiments with 3D designs [11]. However, nearly all these studies manually annotate design models with hazard information [12,13]. Moreover, in this case, hazard identification is highly dependent on the tacit knowledge of the annotator [14]. Involving humans in identification of hazards causes an under-reporting problem, which may lead to a harmful event.
Several researchers have reported a strong relationship between design characteristics and construction accidents and have shown a positive impact of CHPtD on mitigating these incidents by recognizing and managing safety hazards [1,8,14]. Recently, the CHPtD process has been incorporated within BIM solutions; however, a distinctive classification of safety hazards is still unclear, which would make the recognition process more effective and accurate. Moreover, CHPtD lacks consistency between the 2D and 3D design information [15] of the same project in terms of identified hazards. The present proposal enhances the hazard identification process by considering both design characteristics and the formats of design information. It identifies hazards using computational intelligence, particularly natural language processing and machine learning, and subsequently visualises these hazards within the design.
Based on the identified needs and problems with existing solutions, the present work focuses on developing automated methods for identifying and visualising hazards directly from 2D and 3D design information. This approach aims to eliminate the reliance on tacit knowledge and address the under-reporting issues associated with human involvement. The present study aims to achieve the following objectives: (1) to develop and implement a computational intelligence system that leverages natural language processing (NLP) and machine learning (ML) techniques for the automated extraction, identification, and visualisation of potential construction hazards from design documentation; and (2) to validate the accuracy, effectiveness, and usability of the proposal through a system’s output and case studies or expert evaluations, thereby demonstrating its advantages over traditional manual review processes.
For the scope of the present work, hazards that can be identified from the design information are considered situational hazards, which emerge during the construction phase. For example, fall hazards can be identified from the design information, whereas a bucket of oil placed by someone on the construction site is a hazard but is missing in the design information. There is discussion among researchers on the reporting procedure for situational hazards. However, the proposed system can identify material-related hazards because material information is provided within the design, e.g., rolling steel pipes needs careful handling and bonded stacking.

2. Literature Review

The following section will explore the intersection of digital technologies, natural language processing (NLP), predictive modelling, and Building Information Modelling (BIM) in advancing construction safety, particularly through the lens of Prevention through Design (PtD), highlighting how digital tools facilitate critical activities such as knowledge capture, proactive hazard identification and visualisation.

2.1. The CHPtD Framework and Digitalisation

The application of digital technologies holds significant promise for enhancing Prevention through Design (PtD) in construction. Building upon recent research [1], effective PtD tools should address several crucial activities. Firstly, they must excel at knowledge capture [14,16,17], systematically recording construction hazards and corresponding safety measures to facilitate the elimination or mitigation of associated risks. Secondly, these tools should assist designers in proactively identifying and visually [18,19] understanding potential safety hazards within their project designs. Finally, they play a vital role in training designers to develop their own hazard identification skills and effectively apply safety measures in their work. Consequently, the existing body of research on digital technologies for PtD can be broadly categorised [1,5] into four key application areas: knowledge-based systems for storing and retrieving safety information, automatic rule checking to proactively identify design flaws that could lead to hazards, hazard visualisation techniques to provide designers with a clear understanding of potential risks in the built environment, and safety training applications to educate and upskill design professionals. Existing research is mainly carried out in isolation for these categories; however, the present work focuses on hazard visualisation through a knowledge-based system as a PtD tool. The present work presents a review of existing work focusing information extraction (which will become the base for the proposal) from documents to identify hazards specifically in the construction domain.

2.2. Advances in NLP for Safety Information Extraction

Several studies have explored the application of Natural Language Processing (NLP) in the construction industry [8,13], particularly focusing on improving safety management. Feng et al. [14] proposed a natural language data augmentation-based framework for automatic information extraction from accident news reports, aiming to enhance knowledge management in construction safety. The existing methodology involved a cross combination-based text data augmentation algorithm and the use of the BiLSTM-CRF model for character classification. The findings demonstrated the feasibility of establishing reliable information extraction models even with limited training data. In a related study, Zhou et al. [17] concentrated on the automatic extraction of critical information from struck-by accidents using named entity recognition (NER). They compared various deep learning models and found that BERT-LSTM achieved the best performance. While neither Feng et al. [14] nor Zhou et al. [17] directly addressed hazard visualisation, their methods provide valuable techniques for extracting information that could be utilized for visualisation purposes. In contrast, Chung et al. provided a broader perspective by reviewing the applications of NLP in the construction domain and comparing them with state-of-the-art NLP technologies. The analysis by Chung et al. [8] is based on bibliometric analysis and the PRISMA framework, highlighting a decreasing technology gap and identifying future research opportunities in the field. Zhong et al. [13] presented a framework that combines deep learning and text mining to automatically analyse unstructured or semi-structured free-text hazard records. For feature modelling, a CNN model utilises word embedding to transform sequences of words into a vector matrix, and convolution kernels extract text features automatically, avoiding manual processing. This word embedding is formed using the skip-gram (Word2vec) algorithm. Another work [20] focus on deploying machine learning models to predict four specific injury types (upper limbs, lower limbs, head/neck, and back/trunk) from 16,878 construction accident records in Australia. For injury type prediction, they reported Random Forest model have achieved highest accuracy of 79.3%. Similarly, ref. [21] explore the application of deep learning models (CNN, LSTM, and CNN-LSTM) for multi-label classification of Design for Safety (DfS) reports, aiming to overcome the inefficiencies of manual classification. However, for feature modelling, they both [20,21] used traditional pre-processing techniques, whereas Alkaissy et al. [20] combined Principal Component Analysis (PCA) for dimension reduction. These predictive modelling approaches are carried out independently rather than embedded within the design flow.

2.3. BIM-Based Hazard Visualisation

Limited studies have explored different approaches to hazard visualisation in construction. Kulinan et al. [18] integrated Building Information Modelling (BIM) with computer vision to create a real-time system for monitoring workforce safety hazards. This system uses colour coding within a 3D BIM model to represent risk levels and provides detailed information on worker positions and PPE compliance. Similarly, Kulinan et al. [19] developed a 3D BIM-based framework that utilises CCTV data and computer vision to identify proximity hazards. This framework visualizes hazardous zones as 3D heatmaps, displaying hazard indices over time. In contrast, Yuan et al. [14] focused on hazard prevention during the design phase. Their automated inspection tool combines BIM with a Prevention through Design (PtD) knowledge base, providing designers with alerts and risk information within the BIM environment. However, the inspection tool does not address information visualisation. Furthermore, authors [15] have examined the impact of different design information formats (2D and 3D) on hazard recognition. Their findings suggest that the format of design information does not significantly affect hazard recognition performance, focusing instead on the broader context of hazard recognition during design reviews. While Kulinan [19] emphasises real-time, dynamic visualisation of hazards using BIM and computer vision, Yuan et al. [14] prioritise hazard identification and communication during the design phase, and Dylan [15] explores the effectiveness of existing visualisation formats in hazard recognition. Predictive modelling approaches [13,20,21] have provided visualisation of keywords (representing hazards and their associated materials) and their interrelations in a form of graph called a word co-occurrence network (WCN) or word cloud based on a term frequency–inverse document frequency (TF-IDF) measure.

3. Proposed System

A supervised learning in combination with natural language processing (NLP) approach is implemented for the hazard identification or predication. In supervised machine learning, a labelled dataset is provided to the classifier for training purposes [8]. A label is a target class associated with every instance of the training set. These class labels are provided by the domain expert. In the present case, every instance of HSE data [22], tacit knowledge of safety professional, and the design information is labelled with a class or category of the hazards. After extensive review of HSE data and consultation with safety professionals, the hazards are categorized as 10 distinct classes. Every instance of the training set is labelled with one of the ten classes. These hazard classes are as follows: (i) buried services that pose risks from damaging underground utilities during excavation, leading to significant disruption and injury; (ii) catastrophic events, encompassing low-probability but high-impact incidents like structural collapses and fires, demanding proactive avoidance; (iii) driving and (iv) traffic management, highlighting the dangers of vehicle-related activities, emphasizing proper planning and segregation; (v) electricity, presenting severe risks from high-voltage currents, often resulting in burns, fires, and explosions; (vi) excavations, which are hazardous due to potential collapses, falls, and environmental risks; (vii) health, concerning cover exposure to harmful substances, ergonomic issues, and inadequate welfare facilities; (viii) lifting operations, which carry risks from improper equipment use and maintenance; (ix) people-plant interference, describing the dangers of unsafe interactions between workers and machinery; and (x) working at height, which remains a leading cause of fatalities due to falls.
The proposed hazard identification and visualisation application named Auto-BIMHazard predicts and identifies all possible safety hazards in each building design document. The application is implemented as a Revit plug-in using C# .Net framework. Autonomous hazard tagging is performed by utilising a pre-trained machine learning model to predict and identify an associated hazard with an element or the whole building plan. The automated hazard tagging can be visualized in 2D and 3D views. The machine learning model is a one-dimensional (1D) convolutional neural network (CNN) trained on HSE documents and tacit knowledge translated as a feature set. The application carries out multiclass classification for the prediction of possible hazards. Multiple hazards are ranked with probability scores of the prediction made by the model.
The 1D-CNN model was selected due to the specific characteristics of the data and the nature of the hazards, as well as its known advantages, including its ability to identify local patterns (potential hazard are often presented as keywords in design documents for which CNN excels at extracting these local fixed-sized patterns), hierarchical feature learning (learning can be improved by stacking multiple convolutional layers), reduced reliance on long-range dependencies (CNN focuses on presence of specific hazard-related terms or patterns rather than a deep understanding of the entire document’s narrative flow), and efficiency (other models are more computationally expensive and slower to train than 1D CNN) compared to LSTM and Bi-LSTM models. The same has been reported by recent research [13,21], where the CNN model achieved the highest accuracy at 71%, effectively identifying subtle hazard patterns that other models missed.
Training a machine learning model on building design information requires the extraction of the relationship between the design characteristics and the target hazards. For this reason, the proposed system models the design characteristics as a feature vector, i.e., a formal representation for training a machine learning model [23]. This feature vector serves as an input argument for the classifier, and it contains weighting scores for all modelled design characteristics with respect to a possible hazard. For the present case, the modelled feature vector is a supervised weighted vector representation of the design information.
The classification of hazards in the proposed machine learning model was derived from the extracted relationships from the design information such as historical health and safety data, guidance notes from HSE and tacit knowledge captured from construction safety professionals. The text-mining technique is employed for extracting and representing design information into a computer-processable structured model. Using text mining, quantitative analysis can be performed on the textual data. In the present case, the quantitative analysis performed by the machine learning classifier maps the input arguments given in textual format [13] or an element of the design to a hazard class. This mapping function is learnt by the machine learning model through training examples provided as a feature vector. The extraction of meaningful relationship from the unstructured data is carried out using Natural Language Processing (NLP) [8]. NLP is an intelligent process through which data given as text can be represented by a meaningful structure for computers to understand. NLP is a prerequisite for sentiment analysis, text mining, or linguistic analysis. The NLP process starts with pre-processing operations such stop word removal, normalizing text by lowering cases and lemmatization, tokenization for uni- or multi-gram word phrases, and noise removal such as punctuation, special characters and numbers.
For the present case, NLP is applied to analyse the HSE data [22] for the formation of a feature vector. The most important thing in a feature vector is the features themselves, as a feature is an individual measurable characteristic with respect to a target label (hazard for the present case). The machine learning outcome is highly dependent on the quality of identified features. For feature selection during the extraction process using NLP, the present proposal employed WordNet [24] to retrieve semantically related terms with respect to hazard and safety. WordNet is an online lexical repository of English developed by the Cognitive Science Research Group at Princeton University. Wordnet organizes 117,000 synsets that are words appearing in different contexts. For example, closed and shut or car and automobile. WordNet groups together words based on their meanings. Synsets represent an unordered set of words represented under a given context. Word senses are provided by WordNet as words are organized based on close proximity to each other, creating a disambiguated semantic network. Relations among word senses are also modelled, such as hypernym, hyponym, meronym, or IsA relations. Hypernym and hyponym are super-subordinate relations and are most frequently encoded within WordNet. For example, bed is an item of furniture, i.e., {furniture hypernym bed hypernym bedbunk}. From these relations, a conclusion can be made about the relationship between bedbunk and the furniture. Similarly, feature selection in the present case is carried out by computing the relationship between a term and the hazard context. Hypernym and hyponym are transitive relations such that fall is a kind of slip and slip is a kind of mishap; thus, fall is a mishap. WordNet distinguishes between common and proper nouns, verbs and adjectives.
A lexical connected graph of sense ‘fall (a sudden drop from an upright position)’ in WordNet is shown as Figure 1. The proposed feature modelling approach explores the use of semantic relationships and word senses to derive a better set of features to achieve better classification accuracy. The red-coloured nodes represent noun senses where verbs are shown as green-coloured nodes. Straight lines represent the relationship between pairs of nodes, and arrowhead lines are hypernym and hyponym relations. The dotted lines are the derived relations that are not explicitly modelled in WordNet. After pre-processing, the extraction process based on NLP provides the word tokens to the selection process so that each token is evaluated with respect to its relationship with hazard using the semantic relationship of WordNet. WordNet synset similarity scores are used as weights for every individual feature. Computing the similarity score is performed using the NLTK library. The selection procedure performs the same operation for every provided token until the final feature vector is constructed.
A supervised learning approach is implemented for hazard identification or prediction. In supervised machine learning, a labelled dataset is provided to the classifier for training purposes. A label is a target class associated with every instance of the training set. These class labels are provided by the domain expert. In the present case, every instance of HSE data, tacit knowledge of safety professionals, and design information is labelled with a class or category of the hazards. After extensive review of HSE data and consultation with safety professionals, the hazards are categorized in 10 distinct classes. Every instance of the training set is labelled with one of the ten classes.

4. Revit Plug-In

The proposed hazard identification method was implemented as Revit plug-in (Auto-BIMHazard) using C# .Net programming development framework and Revit 2022. The plug-in acts as a frontend to execute machine learning model for a given design document. The Auto-BIMHazard application has three phases, i.e., input, identification and output. The input phase contains the information about hazards, regulations, and construction activities along with underlying design document in Revit file format. The hazard information includes the frequency and severity of any hazard, its category and who will be affected by it. The regulations contain safety, or the protective measures required against all hazards, e.g., “take precautions when working on or near fragile surfaces”. Moreover, common construction activities are enlisted to serve as an item of input to the Auto-BIMHazard. Finally, the application takes the design document as an input to process it for identifying possible hazards.
Figure 2 depicts the hazard visualisations produced by Auto-BIMHazard, an automated Prevention through Design (PtD) tool integrated as a plugin within Autodesk Revit, a widely used Building Information Modelling (BIM) software. The central focus of the interface is a floor plan view of a building design, overlaid with spatially located icons representing identified potential hazards. These icons, varying in appearance, visually distinguish different classes of hazards, allowing designers to readily perceive the distribution and types of risks associated with specific areas and elements within the digital building model. This visual integration directly within the design environment underscores the tool’s aim to embed safety considerations seamlessly into the design workflow.
The overlay of hazard icons on this plan view facilitates a direct correlation between design elements and potential safety risks. The accompanying textual description elaborates on the interactive functionality of these hazard icons, stating that clicking on an icon will trigger the display of detailed hazard information. This information includes the specific nature of the hazard, the underlying reasons or causes contributing to its identification, and suggested possible resolutions or preventative measures. This feature provides designers with immediate access to critical safety-related data, enabling informed decision-making regarding design modifications to mitigate or eliminate identified risks.
Figure 3 presents a modal window overlaid on the Revit environment, indicating its interactive nature within the design workflow. The interface is structured to facilitate consistent hazard identification and visualisation across 2D and 3D formats, as shown in Figure 2 and Figure 3, respectively.
A tabular display as a model window (Figure 3) presents the identified hazards, categorised by drawing specification, category of the risk, and a textual description of the health and safety risk assessment. Each identified hazard is further elaborated with information on “Who is likely to be affected by the risk” and a corresponding “Risk management” strategy. The inclusion of a column for “Action required by” and “PtD applied” indicates a workflow for assigning responsibility and tracking the implementation of preventative measures. The structured format of this tabular information facilitates a systematic review of potential hazards and the associated risk mitigation strategies directly within the BIM environment. Auto-BIMHazard signifies a proactive approach to safety, embedding hazard identification and risk management directly into the design lifecycle, thereby enabling designers to consider safety implications early in the project and design out hazards where feasible. This approach aligns with the core principles of Prevention through Design, aiming to minimise or eliminate risks before they manifest during construction or operation.
Auto-BIMHazard as a Revit plugin leverages the inherent data richness and spatial intelligence of the BIM model for proactive hazard identification. By automating the process of identifying potential hazards through machine learning based on design parameters and embedded information, the tool aims to enhance the efficiency and comprehensiveness of PtD practices. The visual representation of hazards directly within the building plan, coupled with the provision of detailed contextual information and potential solutions upon interaction, empowers designers to address safety concerns iteratively throughout the design process. This approach moves beyond traditional reactive safety measures by integrating hazard identification as an intrinsic component of design, thereby promoting the creation of inherently safer built environments. The use of distinct icons for different hazard classes further enhances the usability and interpretability of the tool, facilitating a more intuitive understanding of the safety landscape of the project.

5. Evaluation and Results

The performance of the Auto-BIMHazard for hazard identification was evaluated through the construction and analysis of a confusion matrix, as presented in Table 1. The matrix provides a detailed breakdown of the Auto-BIMHazard’s classification accuracy across ten distinct hazard classes relevant to construction safety. The diagonal elements, representing true positives, indicate the number of instances where the Auto-BIMHazard correctly predicted the actual hazard class. Conversely, off-diagonal elements signify misclassifications, with false positives (predictions of a class that were actually another) and false negatives (failure to predict a class that was actually present) providing insights into the specific types of errors made by the model.
A class-specific analysis reveals varying levels of predictive capability. For instance, “Class 1—Buried Services” and “Class 5—Excavation” exhibit high precision (90.70% and 8190.24%, respectively) and commendable recall (82.98% and 88.09%, respectively), suggesting a robust ability of the tool to accurately identify these hazard types with a minimal rate of false positives. Similarly, “Class 7—Lifting” and “Class 10—Work at Height” demonstrate strong performance, particularly in recall (88.89% and 91.67%, respectively), indicating a high sensitivity in detecting these critical safety hazards, albeit with slightly lower but still acceptable precision. In contrast, “Class 2—Catastrophic Events” presents a significant challenge, characterized by low recall (30%), indicating a substantial proportion of actual catastrophic events were not identified by the tool. While the precision for this class is relatively high (85.71%), the low recall raises concerns regarding the tool’s capacity to reliably detect these potentially high-consequence hazards. Furthermore, “Class 9—Traffic Management” exhibits a lower precision (73.53%) despite a moderate recall (86.21%), suggesting a tendency for the tool to over-predict this hazard type.
The overall performance metrics provide a consolidated view of the tool’s efficacy. The overall accuracy (OA) of 84.77% indicates that the model correctly classified a substantial majority of the hazard instances. Cohen’s Kappa coefficient of 0.83 further supports the reliability of these results, signifying a substantial level of agreement between the predicted and actual classifications, exceeding what would be expected by chance. This metric is particularly important as it accounts for potential class imbalance, offering a more robust measure of performance than overall accuracy alone.
The high performance observed in several key hazard classes suggests the potential for significant advancements in proactive safety management during the design phase. Early and accurate identification of hazards like buried services, excavation risks, lifting operations, and work at height can facilitate the integration of preventative measures, thereby reducing the likelihood of accidents and improving overall project safety. However, the notable limitations in identifying “Catastrophic Events” underscore a critical area for future development. The rarity and complexity of such events may contribute to the model’s difficulty, necessitating targeted strategies such as the incorporation of more specific training data or the exploration of specialised model architectures capable of handling imbalanced and complex event patterns. The lower precision for “Traffic Management” also warrants attention, as a high rate of false positives could lead to alert fatigue among users, potentially diminishing the overall effectiveness of the tool.
The automated PtD tool for hazard identification demonstrates promising capabilities, achieving a high overall accuracy and substantial agreement with actual hazard classifications. However, the class-specific performance variations, particularly the low recall for catastrophic events and the lower precision for traffic management, highlight critical areas for refinement. Future research should focus on addressing these limitations through enhanced data acquisition strategies, feature engineering, and potentially the adoption of more sophisticated machine learning techniques to improve the robustness and reliability of the tool across all hazard categories, thereby maximizing its potential to contribute to safer design practices in the construction industry.
A semi-structured questionnaire survey was utilized as the primary data collection method to evaluate the effectiveness and efficiency of the developed Auto-BIMHazard tool. The participant pool consisted of research scholars and practitioners whose expertise aligned with the criteria established in the literature by Belton et al. [25].
The survey instrument comprised two distinct sections. The first section provided participants with comprehensive contextual information regarding the research program, specifically detailing the research background, the Auto-BIMHazard tool under investigation, and the pilot project context of its application. This was intended to ensure a consistent and informed understanding among respondents. The second section aimed to gather both the demographic profiles of the expert participants (including job title, age, educational background, and professional experience) and their evaluations of the Auto-BIMHazard’s effectiveness and efficiency. These evaluations were captured through responses to five statements using a five-point Likert scale (1 = strongly disagree to 5 = strongly agree).
A total of 11 completed questionnaires were collected. Demographic analysis, confirmed that all participating research scholars and practitioners met the predefined expert criteria outlined by Belton et al. [25]. Furthermore, the demographic data indicated that all respondents possessed a minimum of a Bachelor’s degree, with over 60% reporting five or more years of relevant professional experience. This profile suggests a high level of domain-specific knowledge among the participants, enhancing the rigor and reliability of the survey data. The survey protocol for this study was reviewed and approved by the Research Ethics Policy and Procedures Board. All procedures involving human participants were conducted in strict accordance with ethical standards.
The expert panel provided their insights on five specific questions designed to assess the effectiveness and efficiency of the Auto-BIMHazard tool. Descriptive statistical analysis of these responses, detailed in Table 2, revealed mean scores above 3 for all five questions. This indicates a statistically positive perception of the Auto-BIMHazard’s effectiveness and efficiency in comparison to conventional safety management practices. The highest level of agreement among the experts observed is improved time efficiency over traditional approaches (4.82 ± 0.40). This strong consensus suggests that the experts overwhelmingly perceive Auto-BIMHazard as a tool that can significantly reduce the time required for hazard identification compared to conventional manual methods. This perceived time efficiency is a crucial advantage in the context of project timelines and resource allocation. Similarly, improvement in safety management also garnered strong agreement (4.73 ± 0.47), indicating a positive outlook among experts regarding the tool’s potential long-term impact on safety practices. The statement “Auto-BIMHazard can enhance the application of safety knowledge” received a high average score of 4.45 with a standard deviation of 0.52. This suggests that experts believe the tool can facilitate a more effective and consistent integration of safety knowledge into the design process. This could be attributed to the tool’s ability to automatically identify hazards, thereby prompting designers to consider safety implications that might be overlooked in traditional workflows. The survey results suggest a consensus that the automated tool can reduce the human effort involved in hazard identification. This reduction in labour can translate to cost savings and the reallocation of resources to other critical project tasks. This finding suggests that the developed PtD tool (Auto-BIMHazard) is perceived by experts as successfully facilitating the integration of safety knowledge into the design phase, a key objective of Prevention through Design and a significant contribution to proactive risk mitigation.

Comparison with Existing Studies

Table 3 presents a comparison of the existing work as compared to Auto-BIMHazard. The existing research, represented by 21, 13 and 20 primarily utilises Convolutional Neural Networks (CNN) or Random Forest models. Kumi et al. [21] leverage a Convolutional Neural Network with basic text pre-processing for feature engineering, achieving an accuracy of 71%. Similarly, Zhong et al. [13] employ a CNN, enhancing feature engineering with term frequency inverse document frequency (TF-IDF) and Word2vec, and report an F1-Score of 0.72 (likely translating to a similar or lower accuracy percentage). Alkaissy et al.’s [20] approach, using a Random Forest model, incorporates text pre-processing and Principal Component Analysis (PCA) for feature engineering, yielding an accuracy of 79.3%. A notable commonality across these existing studies is the absence of explicit hazard visualization techniques within their reported methodologies. Their focus largely remains on the predictive modelling aspect.
In contrast, the Auto-BIMHazard introduces a 1D-Convolutional Neural Network (1D-CNN) that significantly advances both feature modelling and hazard visualisation. The core innovation lies in its knowledge-based feature modelling, specifically utilizing a WordNet-based feature vector. This approach allows for a deeper semantic understanding of the hazard-related text by leveraging WordNet’s lexical database to incorporate relationships such as synonymy, hyponymy, and hypernymy. This enriches the feature representation beyond simple statistical counts (like TF-IDF) or context-agnostic embeddings (like basic Word2vec), capturing more nuanced meanings and conceptual relationships relevant to hazard identification.
The improved accuracy of 84.77% achieved by the Auto-BIMHazard can be directly attributed to this sophisticated feature modelling. By providing the 1D-CNN with semantically rich, WordNet-derived feature vectors, the model can discern more complex patterns and make more accurate classifications of potential hazards. Furthermore, the use of an enhanced 1D-CNN model optimizes the convolutional operations for processing these textual feature vectors, effectively learning intricate local and global dependencies within the data. This combination of knowledge-infused feature engineering and a tailored deep learning architecture leads to a substantial performance gain over the existing methods. Crucially, unlike the other works, the Auto-BIMHazard also proposes hazard visualisation embedded directly within the design workflow, addressing a critical gap in practical application by providing immediate, actionable insights to designers.

6. Conclusions

The research proposes an automated tool, Auto-BIMHazard, to enhance Construction Hazard Prevention through Design (CHPtD) by identifying and visualizing hazards using Building Information Modelling (BIM) and Machine Learning. Auto-BIMHazard is an innovative automated tool designed to enhance safety management within the Architecture, Engineering, and Construction (AEC) industry by leveraging advanced Natural Language Processing (NLP) and deep learning techniques. The core contribution lies in its ability to intelligently identify and classify potential hazards directly from unstructured textual data embedded within Building Information Modelling (BIM) design documents. By employing a 1D Convolutional Neural Network (CNN), the system effectively processes semantically enriched feature vectors, which are meticulously constructed using WordNet-based feature modelling. This novel approach to feature engineering, incorporating lexical and semantic relationships, significantly improved the model’s understanding of hazard contexts, leading to a notable classification accuracy of 84.77%. This performance underscores the efficacy of integrating domain-specific knowledge bases like WordNet into deep learning architectures for complex text classification tasks in specialized domains.
Beyond its robust predictive capabilities, a significant strength of Auto-BIMHazard is its seamless integration with the design workflow through a Revit plug-in. This integration facilitates the visualisation of identified hazards as spatially located icons directly on floor plans and within 2D and 3D BIM views. This intuitive visual feedback mechanism allows designers to immediately perceive the distribution and types of risks associated with specific building elements and areas, supporting a proactive approach to safety-by-design. The provision of interactive detailed hazard information upon clicking these icons, coupled with a comprehensive tabular display of risk assessments, empowers designers with immediate access to critical safety data, enabling informed decision-making and timely implementation of preventative measures. This direct, visual, and interactive feedback loop represents a substantial advancement over traditional, often manual, hazard identification processes, which are prone to human error and are time-consuming.
While Auto-BIMHazard demonstrates significant promise, certain limitations warrant consideration for future research. The model’s performance is inherently dependent on the quality and comprehensiveness of the training data; thus, expanding the dataset to include a wider variety of hazard descriptions and project types could further enhance its robustness and generalizability. Additionally, while WordNet provides a rich semantic foundation, exploring other knowledge graph embeddings or hybrid approaches that combine symbolic AI with deep learning could potentially capture even more nuanced contextual information. Future work will focus on refining the semantic understanding capabilities, potentially through the incorporation of more sophisticated contextual embeddings or domain-specific ontologies. Furthermore, extending the tool’s functionality to include real-time hazard prediction during the operational phase of a building, or integrating it with other project management tools for a more holistic safety management system, presents exciting avenues for continued development. Ultimately, Auto-BIMHazard represents a critical step towards creating more intelligent, automated, and integrated safety management systems in the AEC industry, paving the way for safer and more efficient construction practices.

Author Contributions

Conceptualization, S.A.; Methodology, S.A. and H.A.; Software, M.A.A.; Validation, A.O.; Formal analysis, J.D.; Investigation, A.O.; Resources, A.O.; Writing—original draft, M.A.A.; Writing—review and editing, A.O., J.D. and H.A.; Visualization, M.A.A. and J.D.; Supervision, S.A. and H.A.; Funding acquisition, S.A. All authors have read and agreed to the published version of the manuscript.

Funding

The authors would like to express their sincere gratitude to Innovate UK, the UK’s Innovation agency, for providing the financial support for this study through Grant application number 10081565.

Institutional Review Board Statement

The study was conducted in accordance with the ethical approval of Leeds Beckett University, and approved by the Local Research Ethics Coordinator on 18 April 2023 (Ref: 107274).

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

Some or all data, models or codes that support the findings of this study are available from the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Farghaly, K.; Collinge, W.; Mosleh, M.H.; Manu, P.; Cheung, C.M. Digital information technologies for prevention through design (PtD): A literature review and directions for future research. Constr. Innov. Inf. Process Manag. 2022, 22, 1036–1058. [Google Scholar] [CrossRef]
  2. Behm, M. Linking construction fatalities to the design for construction safety concept. Saf. Sci. 2005, 43, 589–611. [Google Scholar] [CrossRef]
  3. Hardison, D.; Hallowell, M. Construction hazard prevention through design: Review of perspectives, evidence, and future objective research agenda. Saf. Sci. 2019, 120, 517–526. [Google Scholar] [CrossRef]
  4. Perera, U.T.G.; De Zoysa, C.; Abeysinghe, A.A.S.E.; Haigh, R.; Amaratunga, D.; Dissanayake, R. A Study of Urban Planning in Tsunami-Prone Areas of Sri Lanka. Architecture 2022, 2, 562–592. [Google Scholar] [CrossRef]
  5. Hardison, D.; Hallowell, M.; Littlejohn, R. Does the format of design information affect hazard recognition performance in construction hazard prevention through design reviews? Saf. Sci. 2020, 121, 191–200. [Google Scholar] [CrossRef]
  6. Pereira, F.; González García, M.d.l.N.; Poças Martins, J. An Evaluation of the Technologies Used for the Real-Time Monitoring of the Risk of Falling from Height in Construction—Systematic Review. Buildings 2024, 14, 2879. [Google Scholar] [CrossRef]
  7. Chang, S.; Oh, H.J.; Lee, J. Prevention through design (PtD) of integrating accident precursors in BIM. In Proceedings of the International Conference on Construction Engineering and Project Management, Las Vegas, NV, USA, 20–23 June 2022; Korea Institute of Construction Engineering and Management: Seoul, Republic of Korea, 2022; pp. 94–102. [Google Scholar]
  8. Chung, S.; Moon, S.; Kim, J.; Kim, J.; Lim, S.; Chi, S. Comparing natural language processing (NLP) applications in construction and computer science using preferred reporting items for systematic reviews (PRISMA). Autom. Constr. 2023, 154, 105020. [Google Scholar] [CrossRef]
  9. Amber, T.; Memarian, B.; McCleery, T.; Trout, D.; Earnest, G.S. Prevention Through Design to Address Continuing Construction Workplace Deaths and Injuries. Available online: https://blogs.cdc.gov/niosh-science-blog/2024/02/05/ptd-construction-2/ (accessed on 30 January 2025).
  10. Xiahou, X.; Yuan, J.; Li, Q.; Skibniewski, M.J. Validating DFS concept in lifecycle subway projects in China based on incident case analysis and network analysis. J. Civ. Eng. Manag. 2018, 24, 53–66. [Google Scholar] [CrossRef]
  11. Collinge, W.H.; Osorio-Sandoval, C. Deploying a Building Information Modelling (BIM)-based construction safety risk library for industry: Lessons learned and future directions. Buildings 2024, 14, 500. [Google Scholar] [CrossRef]
  12. Hou, H.; Yu, S.; Wang, H.; Huang, Y.; Wu, H.; Xu, Y.; Li, X.; Geng, H. Risk assessment and its visualization of power tower under typhoon disaster based on machine learning algorithms. Energies 2019, 12, 205. [Google Scholar] [CrossRef]
  13. Zhong, B.; Pan, X.; Love, P.E.; Sun, J.; Tao, C. Hazard analysis: A deep learning and text mining framework for accident prevention. Adv. Eng. Inform. 2020, 46, 101152. [Google Scholar] [CrossRef]
  14. Yuan, J.; Li, X.; Xiahou, X.; Tymvios, N.; Zhou, Z.; Li, Q. Accident prevention through design (PtD): Integration of building information modeling and PtD knowledge base. Autom. Constr. 2019, 102, 86–104. [Google Scholar] [CrossRef]
  15. Hardison, D.C. Hazard Recognition in Design: Evaluating the Effects of Design Information on Hazard Recognition Performance. Ph.D. Thesis, University of Colorado at Boulder, Boulder, CO, USA, 2018. [Google Scholar]
  16. Hare, B.; Kumar, B.; Campbell, J. Impact of a multi-media digital tool on identifying construction hazards under the UK construction design and management regulations. J. Inf. Technol. Constr. 2020, 25, 482–499. [Google Scholar] [CrossRef]
  17. Zhou, Z.; Wei, L.; Luan, H. Deep learning for named entity recognition in extracting critical information from struck-by accidents in construction. Autom. Constr. 2025, 173, 106106. [Google Scholar] [CrossRef]
  18. Kulinan, A.S.; Park, M.; Aung, P.P.W.; Cha, G.; Park, S. Advancing construction site workforce safety monitoring through BIM and computer vision integration. Autom. Constr. 2024, 158, 105227. [Google Scholar] [CrossRef]
  19. Kulinan, A.S.; Jeon, Y.; Aung, P.P.W.; Park, M.; Cha, G.; Park, S. BIM-based automated analysis of dynamic hazards for proactive safety measures during the earthwork construction stage using CCTV data. Adv. Eng. Inform. 2025, 65, 103296. [Google Scholar] [CrossRef]
  20. Alkaissy, M.; Arashpour, M.; Golafshani, E.M.; Hosseini, M.R.; Khanmohammadi, S.; Bai, Y.; Feng, H. Enhancing construction safety: Machine learning-based classification of injury types. Saf. Sci. 2023, 162, 106102. [Google Scholar] [CrossRef]
  21. Kumi, L.; Jeong, J.; Jeong, J. Proactive approach to enhancing safety management using deep learning classifiers for construction safety documentation. Eng. Appl. Artif. Intell. 2025, 153, 110889. [Google Scholar] [CrossRef]
  22. HSE Safety Hazards—HSE. Available online: https://www.hse.gov.uk/construction/safetytopics/index.htm (accessed on 12 January 2025).
  23. Feng, D.; Chen, H. A small samples training framework for deep Learning-based automatic information extraction: Case study of construction accident news reports analysis. Adv. Eng. Inform. 2021, 47, 101256. [Google Scholar] [CrossRef]
  24. Fellbaum, C. WordNet. In Theory and Applications of Ontology: Computer Applications; Springer: Dordrecht, The Netherlands, 2010; pp. 231–243. [Google Scholar]
  25. Belton, I.; MacDonald, A.; Wright, G.; Hamlin, I. Improving the practical application of the Delphi method in group-based judgment: A six-step prescription for a well-founded and defensible process. Technol. Forecast. Soc. Change 2019, 147, 72–82. [Google Scholar] [CrossRef]
Figure 1. Lexical graph of sense of fall.
Figure 1. Lexical graph of sense of fall.
Architecture 06 00093 g001
Figure 2. Auto-BIMHazard.
Figure 2. Auto-BIMHazard.
Architecture 06 00093 g002
Figure 3. Hazard identification—Auto-BIMHazard.
Figure 3. Hazard identification—Auto-BIMHazard.
Architecture 06 00093 g003
Table 1. Confusion matrix with Cohen’s Kappa.
Table 1. Confusion matrix with Cohen’s Kappa.
Class 1Class 2Class 3Class 4Class 5Class 6Class 7Class 8Class 9Class 10Classification OverallPrecision
Class 1390010201004390.698
Class 20311100001785.714
Class 3102802100103384.848
Class 4210310030013881.579
Class 5011037011004190.244
Class 6211012601003281.25
Class 7101010320103688.889
Class 8011100014201984.848
Class 9123100012513473.529
Class 10112000010333883.333
Truth Overall47103835422936192936348
Recall82.97977.41973.68488.57188.09589.65588.88984.84886.20789.286
Overall Accuracy (OA)84.77Class 1—Buried ServicesClass 2—Catastrophic eventsClass 3—Driving
Kappa0.83Class 4—ElectricityClass 5—ExcavationClass 6—HealthClass 7—Lifting
Class 8—People-plant interferenceClass 9—Traffic ManagementClass 10—Working at height
Table 2. Descriptive analysis—acceptance testing questionnaire.
Table 2. Descriptive analysis—acceptance testing questionnaire.
No.QuestionResponse AverageStd.Div
1Auto-BIMHazard offers greater accuracy compared to traditional methods.3.730.79
2Auto-BIMHazard provides improved time efficiency over traditional approaches.4.820.40
3Auto-BIMHazard requires less labour than traditional methods.4.090.30
4Auto-BIMHazard can enhance the application of safety knowledge.4.450.52
5Auto-BIMHazard holds promise for improving safety management throughout the project lifecycle.4.730.47
Table 3. Comparison with existing studies.
Table 3. Comparison with existing studies.
ML ModelFeature EngineeringAccuracyHazard Visualisation
[21]Convolutional Neural Network (CNN)Text pre-processing71None
[13]Convolutional Neural Network (CNN)TF-IDF and Word2vec0.72 (F1-Score)None
[20]Random ForestText pre- processing and PCA79.3None
Auto-BIMHazard1D-CNNWordnet based feature vector 84.77Embedded within Design Workflow
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Abbas, M.A.; Ajayi, S.; Oyegoke, A.; Dauda, J.; Alaka, H. Automated Hazard Identification and Visualisation in Design Using Building Information Modelling and Machine Learning. Architecture 2026, 6, 93. https://doi.org/10.3390/architecture6020093

AMA Style

Abbas MA, Ajayi S, Oyegoke A, Dauda J, Alaka H. Automated Hazard Identification and Visualisation in Design Using Building Information Modelling and Machine Learning. Architecture. 2026; 6(2):93. https://doi.org/10.3390/architecture6020093

Chicago/Turabian Style

Abbas, Muhammad Azeem, Saheed Ajayi, Adekunle Oyegoke, Jamiu Dauda, and Hafiz Alaka. 2026. "Automated Hazard Identification and Visualisation in Design Using Building Information Modelling and Machine Learning" Architecture 6, no. 2: 93. https://doi.org/10.3390/architecture6020093

APA Style

Abbas, M. A., Ajayi, S., Oyegoke, A., Dauda, J., & Alaka, H. (2026). Automated Hazard Identification and Visualisation in Design Using Building Information Modelling and Machine Learning. Architecture, 6(2), 93. https://doi.org/10.3390/architecture6020093

Article Metrics

Back to TopTop