Skip to Content
BuildingsBuildings
  • Article
  • Open Access

17 June 2026

Assessing the Effectiveness of Generative Artificial Intelligence in Hazard Identification on Construction Sites

,
,
,
,
and
1
NUST Institute of Civil Engineering (NICE), School of Civil and Environmental Engineering (SCEE), National University of Science and Technology (NUST), Sector H-12, Islamabad 44000, Pakistan
2
Civil Engineering Department, College of Engineering, Imam Mohammad Ibn Saud Islamic University (IMSIU), Riyadh 11564, Saudi Arabia
*
Authors to whom correspondence should be addressed.

Abstract

The construction industry remains one of the most perilous, where hazard identification is often inconsistent. Hazards are still missed when teams rely mainly on traditional approaches like checklists, Job Safety Analysis (JSAs)/Job Hazard Analysis (JHAs), and individual experience. This study evaluated whether a GEN-AI-assisted approach improved hazard identification performance compared with traditional approaches, using expert-verified ground truth for scoring. A quantitative within-subjects experiment was conducted with 51 participants. Each participant completed hazard identification in four conditions: traditional–pre, GEN-AI–pre, traditional–post, and GEN-AI–post, with a short training session on hazard identification delivered between the pre- and post-stages. Effectiveness was measured using the F1 score, combining both precision and recall. For analysis, traditional and GEN-AI performance were compared at each stage using paired-sample t-tests, and the overall pattern was tested using a 2 × 2 repeated measures ANOVA. The results showed that GEN-AI support produced significantly higher performance than the traditional approaches at both stages (p < 0.05). The repeated measures ANOVA confirmed a strong overall method effect. However, the overall intervention effect was small, and the method × intervention interaction was negligible, with no statistically significant change over time (p > 0.05). Overall, the findings indicated that GEN-AI support improved hazard identification accuracy relative to traditional approaches in this dataset, with limited evidence of additional gains from the training intervention. This study contributes towards providing empirical evidence that GEN-AI improves hazard identification and strengthens proactive prevention, but final outputs need human validation.

1. Introduction

The construction industry is one of the most accident-prone industries worldwide and is notorious for its high rates of injuries and fatalities globally [1,2], owing to the inherently unstable, dynamic, and complex conditions on typical worksites. At any given time, crews, materials, machinery, and environmental conditions may change, producing unpredictable hazards such as unprotected edges, hazardous electrical exposures, or falling objects [3].
This global issue is more alarming in developing countries like Pakistan. A 2025 Karachi field study found that 26% of construction workers had experienced at least one occupational injury over the preceding year [4], which is a startling illustration that every fourth worker has been the victim of a severe physical injury every year. Furthermore, a data analysis of 386 construction companies in Sindh province over 5 years showed 894 accidents, of which 79 were fatalities, 231 non-fatal injuries, and hundreds of situations that led to property damage [5]. The report explained all these results by a lack of safety training, improper use of personal protective equipment (PPE), and the existing culture of risk-taking. Furthermore, one of the studies in Lahore found slips, falls, and accidents (STFs) to be significant contributors to injury and mortality in construction sites [6]. These statistics create a clear picture that construction regularly subjects the workers to severe hazards, and the rates of injuries are not acceptable. The human price, such as injuries and deaths, lost livelihood, and health consequences, is high.
Meanwhile, contractors and employers are loaded with economic losses like project delays, compensation, and reputational damages. With such realities in mind, preventive safety practices become very crucial. The key tool of prevention is hazard identification. Although definitions and systematic taxonomies are clear, numerous studies demonstrate that the aspect of hazard identification is done inconsistently and weakly [7]. Empirical evidence shows that human observers, such as workers, supervisors, and even trained personnel, were able to correctly identify 40–50 percent of hazards only in controlled conditions of tests [8]. In real-world situations with a lot of visual noise, time pressure, and divided attention, performance drops even more [9,10]. Many construction accidents are attributed directly to hazards that were present but were unnoticed, making inadequate hazard recognition a primary causal factor [11,12]. To address persistent inconsistency, hazard identification has evolved from traditional checklist and experience-based methods through knowledge-based approaches [13,14], computer-vision integrations [15,16], and most recently, generative artificial intelligence (GEN-AI) and other LLM-based approaches [17,18]. Each generation improved some aspect of recognition, yet concerns regarding cross-project transferability, limited validation on real sites, and, for LLMs, hallucination and a lack of controlled empirical evidence remain; these developments are synthesized in Section 2. This study addresses the resulting gap by assessing the effectiveness of GEN-AI in hazard identification on construction sites and comparing it against traditional approaches.
Construction hazard identification remains inconsistent and incomplete across tasks and projects, despite well-established definitions and taxonomies that standardize what constitutes a hazard, an event, and a consequence [19]. Individuals often forget safety knowledge across projects, and inconsistent ways of managing knowledge make the safety culture fragmented, which makes it harder for organizations to learn [20]. Field and laboratory research demonstrate that even experienced individuals often overlook significant hazards, particularly in complex, dynamic, and visually congested environments [8,10,21].
Traditional tools like checklists, Job Safety Analysis (JSAs)/Job Hazard Analysis (JHAs), and toolbox discussions provide structure, but they rely a lot on the user’s memory and experience, which makes them less useful and harder to scale [22,23]. While computer vision has advanced the detection of PPE non-compliance, proximity risks, and guardrail issues, such approaches struggle to explain why a configuration is unsafe without encoded contextual rules [15,24,25]. Knowledge graph and ontology frameworks enhance reasoning by linking hazards to tasks, equipment, and locations [13,16], yet most implementations remain system-centric and difficult to adapt to new project scopes, data sources, and textual documentation.
In parallel, generative artificial intelligence (AI), particularly large language models like ChatGPT-5, offers new potential for capturing and reasoning over safety knowledge expressed in text [26,27,28]. ChatGPT can interpret task descriptions, method statements, and regulations to produce hazard lists and control suggestions, but reported performance varies widely across studies [18,29]. Some researchers improve reliability through retrieval-based grounding, where the model references standards or project documents to justify outputs [30,31]. However, most of these demonstrations stop short of controlled, quantitative validation against expert-verified hazard lists. Quantitative evidence is still limited to proving the effectiveness of generative artificial intelligence in hazard identification.
Moreover, few studies compare GEN-AI-assisted hazard identification with traditional methods (e.g., JSA, checklists, walkthroughs) under consistent evaluation frameworks. The need for evidence on the effectiveness of GEN-AI highlights the need for quantitative validation of GEN-AI-assisted hazard identification to determine whether such tools genuinely enhance human performance. Accordingly, this research evaluates ChatGPT-assisted hazard identification against expert-curated reference lists. This study seeks to establish an evidence-based understanding of GEN-AI’s effectiveness in supporting hazard identification in construction.
The contribution of this study is therefore threefold. First, unlike prior demonstrations of LLM-based hazard recognition that report illustrative outputs, this study provides a controlled within-subject quantitative comparison of GEN-AI-assisted and traditional hazard identification scored against an expert-verified ground truth using precision, recall, and F1. Second, it isolates the method effect from individual differences by having each participant act as their own control across both conditions. Third, it quantifies the practical magnitude of the GEN-AI advantage using standardized effect sizes, providing the empirical evidence that the existing literature has called for but has not yet delivered under a rigorous evaluation framework.
Across these generations of work, evaluation has relied on heterogeneous metrics whose limitations are synthesized in Section 2; the present study adopts the confusion-matrix family and uses F1 as the primary metric because it jointly reflects the correctness and completeness of the hazard list. The overall effect was then assessed using a paired t-test, as the study had a within-subject nature and each participant acted as their own control [32]. Overall, the effect of the method and intervention was assessed using repeated measures ANOVA [33].
To address this gap, this study is guided by the following research questions. RQ1: Does GEN-AI-assisted hazard identification produce higher accuracy, measured by F1 against an expert-verified ground truth, than a traditional knowledge-based approach? RQ2: Does a short hazard identification training intervention change the relative advantage of GEN-AI over the traditional approach? The corresponding objectives are (i) to quantify and compare hazard identification effectiveness under traditional and GEN-AI-assisted conditions, (ii) to test whether any observed advantage is stable across a pre- and post-training stage, and (iii) to translate the findings into governance-oriented recommendations for safe deployment in practice.

2. Literature Review and Research Gap

Construction hazard identification has evolved through several technological generations. Knowledge-based approaches such as ontologies and knowledge graphs improved consistency by linking objects, activities, and environments [13,14], but they transfer poorly across projects and have rarely been validated on real sites [16]. Computer vision methods advanced the automated detection of PPE noncompliance, proximity risks, and guardrail issues [15,24,25], yet they struggle to explain why a configuration is unsafe without encoded contextual rules. More recently, large language models such as ChatGPT have been applied to reason over textual safety knowledge and to generate hazard and control suggestions [26,27,28], with retrieval-based grounding proposed to improve reliability [30,31]. However, reported performance varies widely across studies [18,29], and most demonstrations stop short of controlled quantitative validation against expert-verified hazard lists. Survey-based evaluations rely on self-reports and are subjective [34]; hazard recognition rate ignores false positives and can reward over listing [8,35]; and eye-tracking and computer vision metrics answer different questions and require model training [9,36]. The confusion matrix family of metrics (precision, recall, and F1) addresses these weaknesses by jointly accounting for correctly identified, missed, and incorrectly labeled hazards [15,25,37]. The clear gap that remains is the absence of a controlled, within-subject, expert-grounded quantitative comparison of GEN-AI-assisted and traditional hazard identification, which is the gap this study addresses. The study is distinct from the closest prior work. Uddin et al. [35] demonstrated that ChatGPT can aid hazard recognition and support safety education among trainees, but did so largely descriptively and without a controlled comparison against traditional methods scored on a common expert-verified ground truth. The present study advances beyond confirmation by quantifying the GEN-AI advantage relative to a traditional baseline within the same participants, using precision, recall, and F1, as well as reporting standardized effect sizes, thereby supplying the controlled empirical evidence that [35] and related demonstrations did not provide.

3. Materials and Methods

3.1. Research Design

The research philosophy used within this study is positivist, as the research assumes the world exists independently and can be measured objectively by scientific observation [38]. The study employs a deductive approach, which is appropriate because the research has a well-defined hypothesis, which it tests through empirical data. As discussed earlier, the most recent approaches, like generative artificial intelligence, have shown potential in providing context-specific suggestions [18,39]. This served as the basis for the formation of the hypothesis on which this study laid its foundation. The hypothesis guiding the research is as follows:
H1: 
Generative artificial intelligence improves hazard identification performance compared to traditional methods.
A quantitative methodology is used in this study in connection with its deductive approach and the purpose of the evaluation of hazard-identification performance under the conditions of traditional and GEN-AI-assisted use. Quantitative designs [16,17,36,40] will be appropriate because the study focuses on measurable outcomes, particularly determining whether ChatGPT can contribute to better and more comprehensive hazard recognition in comparison with the conventional approaches of Job Safety Analyses (JSAs), Job Hazard Analyses (JHAs), checklists, and walkthroughs [41]. The study design of the research is experimental, as adopted by multiple previous studies [21,42,43], which will be implemented through a sequence of hazard-identification activities organized under controlled conditions. The study had a within-subject experimental design on all the participant groups, which implies every participant had to be involved in hazard-identification tasks using both traditional methods and GEN-AI (ChatGPT-5). This design made it easy to compare the performance of the participants against their own, as well as taking into consideration individual differences in terms of experience, prior knowledge, and observational abilities.
The participants were provided with a pre/post-training experimental condition, which allowed analysis of whether generative artificial intelligence is better in hazard identification or if the identified effect is predetermined solely by the fact that participants have limited knowledge or less experience in the field. This study adopts a cross-sectional time horizon, with data collected within a defined period across multiple participant groups and site conditions.
These design choices were made deliberately. The within-subject structure was chosen because it removes between-person variance in experience and observational ability, increasing statistical power for a fixed sample. F1 against an expert-verified ground truth was selected over hazard recognition rate because it penalizes both missed hazards and spurious listings, preventing inflated scores from over listing. The traditional condition was always completed before the GEN-AI condition to protect the integrity of the unaided baseline, since exposure to AI-generated hazards beforehand would have anchored subsequent judgments; the resulting carryover is addressed in the limitations. Repeated measures ANOVA was appropriate because the design is balanced and fully crossed, with one observation per participant per condition and no missing cells.
The experimental flow of the study is summarized in Figure 1. First, the participants were divided into two groups. One was taken to the Resolve Tower (a high-rise construction project), which was relatively more complex and cluttered. The other group was taken to the MBR Plant (a road construction project), which was relatively less cluttered, less complex, and simpler. Both groups were asked to do the hazard identification through both traditional and GEN-AI approaches. After that, both groups were provided with standard training on hazard identification to minimize the effect of participants (being less experienced) on the assessment of the relative effectiveness of GEN-AI compared to traditional approaches. Then, both groups were shown a virtual site video and were given multiple static frames. Afterwards, they were asked to do the hazard identification of these scenarios.
Figure 1. Experimental flow of research design.

3.2. Multiple Site-Based Experimental Condition

The experiment spans three site conditions, each treated as its own scenario within the experimental strategy. First was the Resolve Tower (real site–high-rise building), second was the MBR Plant (real site–road construction), and the last one was a virtual construction site (14 hazard-rich scenarios–video-based). These parallel site environments enable structured comparison of performance across different levels of realism, different site complexities, and participant experience levels. The study had a fair number of controls that ensured the experimental validity. All were provided with the same general instructions, with the same optional prompt form for the AI-assisted condition. They worked individually, so that nobody could work together or share knowledge. Each of the participants performed the traditional assessment and the AI-assisted assessment, respectively. This ensured that the AI results did not influence the baseline performance. The same access was also granted to all participants to the evaluation stimuli. For actual sites, they were taken to the site, given the same photos, and allowed to capture their own photos. For virtual sites, they saw the same film and were given the same still frames. In the AI condition, participants used both visuals and text descriptions to talk to ChatGPT, which is how real-world multimodal experimentation works. Together, these controls ensured consistency across sessions and strengthened the internal validity of the performance comparisons.

3.3. Sampling Strategy

The study employs a purposive sampling technique [35,44] because we needed participants who were familiar with the process of hazard identification and were capable of using generative artificial intelligence like ChatGPT for hazard identification on construction sites. We need to choose people capable of executing hazard-identification tasks based on actual site visits, simulated scenarios, and interactions with GEN-AI. The sample comprises postgraduate (PG) and undergraduate (UG) students enrolled in construction-related programs at NUST, exhibiting experience levels from none to over three years. This population is appropriate because the study evaluates knowledge-based traditional hazard identification, a core competency for construction managers and early-career practitioners. Although sampling was constrained by logistical realities (site access, classroom scheduling, training duration), the final sample sizes were examined against power analysis benchmarks to ensure that the study maintains adequate statistical strength for paired t-tests. Because the study uses paired comparisons (traditional vs. AI; pre- vs. post-training), the relevant model is the paired-samples t-test. The required sample size was computed using Equation (1), with the final sample adjusted for a 15 percent dropout.
n = Z 1 α / 2 + Z 1 β d 2 ,   d = Δ σ d   = 0.40 = 1.96 + 0.84 0.40 2 = 2.80 0.40 2 = 7 2 = 49
A medium standardized effect (d = 0.40) was assumed a priori because no controlled within-subject estimate of the GEN-AI versus traditional difference existed at the design stage; this is a conservative assumption relative to the larger effect ultimately observed.
Because the study uses a within-subject paired comparison, each participant contributes a single pair of scores (traditional versus GEN-AI) and therefore one difference score; such designs increase statistical power because variance is measured within the same participants. The required number of analyzable participants, therefore, follows directly from Equation (1) as n = 49. Allowing for a 15 percent dropout or unstable data rate, the target enrollment was approximately 49 × 1.15 = 57 participants. A total of 60 participants were enrolled; after excluding 9 with incomplete data (15 percent), 51 participants with complete observations in all four conditions were retained for analysis. Because the realized analytic sample of 51 exceeds the required minimum of 49, the study is adequately powered to detect the hypothesized effect, and this margin is conservative given that the effect ultimately observed (Cohen’s d between 0.74 and 0.77) was substantially larger than the d = 0.40 assumed a priori. These groups provide adequate representation across experience levels, real and virtual environments, and pre- and post-training contexts.

3.4. Data Collection

Data collection was carried out across three construction environments, two real (Resolve Tower and the MBR Plant) and one virtual construction site represented through a multi-scenario video. The overall objective was to gather participant-generated hazard lists under both traditional and GEN-AI–assisted conditions in a controlled, sequential, and comparable manner. For real-site tasks, participants were physically taken to their assigned locations following standard safety protocols. One group of participants visited the Resolve Tower (G+9 high-rise under construction) (Figure 2a–c) while the other visited the MBR Plant (Figure 3) (a road construction project). Figure 2a–c shows some of the pictures from the Resolve Tower showing the complexity of work and visually cluttered scenarios where the participants were brought to for hazard identification. Similarly, Figure 3a,b shows some pictures from the second real construction site (MBR Plant), which was a road construction project, which was relatively less complex. These different sites provided different levels of complexity for workers to match the real-world scenarios as much as possible. The groups received a guided walkaround to ensure equivalent exposure to site conditions, followed by permission to capture site photographs for their own reference. After the visit, all participants were provided with a uniform set of official site photographs to standardize visual input, reducing variability due to individual image-taking habits.
Figure 2. Resolve Tower (a high-rise building) construction project: (a) ground-level earthworks with an excavator on rubble and the skyline behind; (b) workers on a steel staircase at height between floors; (c) the multi-storey concrete frame under construction, with scaffolding, a crane boom, vehicles, and stored steel at ground level.
Figure 3. MBR Plant (road construction project): (a) a road roller compacting an unpaved road beside dense vegetation; (b) a tracked excavator with its boom extended over the road, surrounded by overgrowth.
The virtual site data collection included a 14-scenario construction simulation movie that included a lot of hazards, such as working at heights, manual handling, and being exposed to electricity. Participants in this condition viewed the same movie in a controlled classroom environment, and thereafter, received still frames extracted from each scenario to guarantee uniform visual references for all participants for the hazard-identification tasks. Figure 4a–c shows 3 of the 14 scenarios from the virtual site, which was presented to the participants in the form of a video. The different activities taking place at the virtual site were clearly visible, which demonstrates the varying hazards present there.
Figure 4. Multiple scenarios from the virtual site: (a) a rooftop slab with workers, a crawler crane, and stacked materials near an unprotected edge; (b) an upper floor with scaffolding, rebar, and workers moving loads across an open deck; (c) an interior atrium with workers on and around an open-sided upper-floor walkway above a lower level. (https://youtu.be/tAkyjYKNJkU (accessed on 9 November 2025)).
In traditional conditions, participants depended only on their own knowledge, previous experiences, and established safety protocols, including Job Safety Analyses (JSAs), Job Hazard Analyses (JHAs), and checklist-based reasoning. They were told not to utilize mobile devices, the internet, or AI tools during this time. Responses were handwritten on paper templates made to record hazard statements and other criteria that the participants thought were important. This made sure that everything was clear and that AI could not affect the baseline data. Completed sheets were collected physically and stored securely before transcription into Excel for scoring. After the traditional phase, participants completed the same task under the GEN-AI condition using ChatGPT. They were allowed to upload site images alongside textual descriptions of activities, materials, and site context. A suggested prompt template was provided to ensure minimum standardization across participants, although they were free to adjust or replace the prompt based on their own judgment, as shown in Figure 5. This procedure reflects realistic field conditions, in which supervisors may adapt their queries to the tool.
Figure 5. Suggested prompt template.
The suggested template (Figure 5) instructed participants to assign ChatGPT the role of a site safety supervisor, to provide the uploaded site images together with a short textual description of the activities, materials, and context, and to request a list of hazards organized by category (for example, falls, struck by, caught in, electrical, and plant) without listing controls. Participants used the GPT 4 class multimodal model available at the time of data collection. Because participants could adapt the prompt, the prompt actually used was recorded with each submission to support transparency and replication. ChatGPT responses were submitted digitally through a Google Form, ensuring clean data capture and eliminating transcription errors. All digital responses were exported into structured spreadsheets for comparison with ground truth.

3.5. Ground Truth Development

A definitive hazard list for each site was created through a structured expert-review process. Two senior experts examined the site conditions to develop and validate the final lists of hazards (20 hazards for the Resolve Tower, 15 for the MBR Plant). Disagreements were resolved through discussion, and final ground truth sets were locked before scoring participant outputs. To strengthen the reliability of the reference lists, the two senior experts (each with more than ten years of site safety experience) first prepared their hazard lists independently and then reconciled them. Agreement between the two independent expert lists prior to reconciliation was high (Cohen’s kappa = 0.81 for Resolve Tower and 0.78 for the MBR Plant), indicating substantial to almost perfect agreement, and the small number of disagreements was resolved by discussion before the lists were locked. To score open ended participant responses against these lists, a predefined matching rule was applied: a participant entry was counted as a true positive when it referred to the same hazard source and mechanism as a ground truth item, including accepted synonyms and paraphrases (for example, “live wires” was matched to “exposed electrical wires”); generic or non-specific entries that did not map to a defined hazard were counted as false positives; and ground truth items with no corresponding participant entry were counted as false negatives. Two raters applied this rule, and the inter-rater agreement on the matching decisions was high (kappa = 0.84), with residual disagreements resolved by discussion.
To clarify how agreement was quantified for the open-ended listing, the unit of analysis differed for the two coefficients. For the expert reference lists, the two experts independently reviewed a common pool of candidate hazards compiled for each site and classified each candidate as include or exclude; Cohen’s kappa was computed over these per-item inclusion decisions (the candidate hazard item being the unit), giving 0.81 for Resolve Tower and 0.78 for the MBR Plant. For the matching decisions, the two raters independently classified each participant entry as a match to a ground truth item (a true positive) or a non-match (a false positive), and independently flagged each unlisted ground truth item as a false negative; Cohen’s kappa (0.84) was computed over these individual scoring decisions, with the participant entry and the ground truth item as the units.
This structured collection procedure ensured that all participants engaged with identical stimuli under controlled conditions and that all responses were directly comparable for quantitative scoring. Table 1 shows the standard hazard list for the Resolve tower, showing the responses of multiple participants. It contained 20 hazards, which were being analyzed for each participant for both conditions (traditional and GEN-AI-based). Table 2 shows the standard hazard list for the MBR Plant, along with multiple responses from the participants, and Table 3 shows the standard hazard list of one of the 14 hazard-rich virtual site scenarios. The true positives (TP), false positives (FP), and false negatives (FN) were calculated, which served as the basis for the calculation of precision, recall, and F1 score to assess the effectiveness.
Table 1. Standard hazard list of the Resolve Tower.
Table 2. Standard hazard list of the MBR Plant.
Table 3. Standard hazard list of one of the virtual site scenarios.

3.6. Data Analysis

The lists of hazards created by the participants were converted into objective performance indicators through the confusion matrix. All submissions of participants were crosslinked with the standard list as provided by the expert, regarding the respective site; thus, it was possible to identify the true positives (TP), false positives (FP), and false negatives (FN). For each participant and each condition, TP corresponds to hazards correctly identified and present in the ground truth, FP corresponds to hazards stated by the participant but not present in the ground truth, AND FN corresponds to hazards present in the ground truth but missed by the participant. True negatives (TNs) are not applicable because hazard identification tasks involve an open-set environment rather than classification of mutually exclusive categories. The focus remains on detecting present hazards rather than correctly rejecting irrelevant ones. Table 4 presents the confusion matrix being used in the research for the calculation of factors like TP, FP, and FN. Using Equations (2)–(4), the TP, FP, and FN were calculated in the analysis [37,40,45].
P r e c i s i o n = T P T P + F P
R e c a l l = T P T P + F N
F 1 = 2 × P r e c i s i o n · R e c a l l P r e c i s i o n + R e c a l l
Table 4. Confusion matrix.
Precision reflects correctness, recall reflects completeness, and F1 combines both into a balanced measure. These metrics were computed separately for traditional pre-training vs. AI-assisted pre-training and traditional post-training vs. AI-assisted post-training. This allows direct comparison not only between traditional and AI conditions but also between pre- and post-training results. Paired t-test was used to compare traditional (pre) vs. AI-assisted (pre) and traditional (post) vs. AI-assisted (post). Such comparisons demonstrate whether ChatGPT helps result in statistically significant improvements relative to the baseline in both untrained and trained participant situations, respectively. Further analysis of the effect caused by the training on the performance of the participants was done using a paired t-test to compare traditional (pre) vs. traditional (post) conditions. Repeated measures ANOVA was then used to compare and analyze the effect of the method and intervention. Cohen’s d was used to report effect sizes of paired t-tests, and partial η2 was used to report ANOVA. These measures evaluate the practical significance beyond p-value changes.
To clarify the dependency structure of the data, each participant contributed one F1 score per cell of the 2 (Method: traditional, GEN-AI) by 2 (Intervention: pre, post) design, so the four conditions are fully crossed within participant, and the number of observations per participant is balanced. Repeated measures ANOVA is therefore appropriate, because the within-subject correlation is modeled through the participant factor and the design contains no missing cells. For the two-level within-subject factors used here, repeated measures ANOVA and a linear mixed effects model with a random intercept for participant yield equivalent fixed effect estimates; we report the repeated measures ANOVA for transparency and comparability with prior work. Site-level information was handled by aggregating each participant to a single F1 per condition rather than pooling unequal numbers of raw site observations, which avoids an unbalanced dependency structure; the relationship between site type and the pre- and post-stages, and its implications, is examined in Section 5 (see also the limitation added regarding evaluation environment).

3.7. Training Intervention

Between the pre- and post-stages, all participants received a single standardized hazard identification training session of approximately 30 min delivered in a controlled classroom setting by the same facilitator. The learning objectives were to (i) define hazard, unsafe act, and unsafe condition; (ii) introduce the energy-based hazard categories (gravity, motion, electrical, mechanical, and chemical) as a systematic search strategy; and (iii) practise structured scanning of construction scenes. Delivery combined a short slide-based briefing with three worked examples and a brief guided practice on a non-study image. No graded assessment was administered; the post-stage hazard identification tasks themselves served as the measure of any change in performance.

4. Results and Analysis

4.1. Descriptive Statistics

Descriptive statistics were used to give an overview of how well the four within-subject conditions (traditional–pre, GEN-AI–pre, traditional–post, and GEN-AI–post) were able to identify hazards. The F1 score was the main measure of performance (effectiveness). Higher values meant that the matching of participant outputs to the expert-validated ground truth was more precise and more accurate. The sample size was the same for all situations, thus we could directly compare the results without any cells being out of balance. Of the 60 participants enrolled, nine (15 percent) were excluded because they did not provide usable hazard lists across all four conditions; the analyses therefore use the 51 participants with complete data in every condition. This attrition corresponds to the 15 percent dropout allowance built into the a priori power analysis.
In general, the GEN-AI–assisted approach to finding hazards had higher average F1 scores than the traditional way at both stages. The traditional method had a mean F1 of 0.470 (SD = 0.152, Median = 0.452) in the pre-stage. The GEN-AI condition, on the other hand, had a mean F1 of 0.583 (SD = 0.135, median = 0.585). This pattern indicated that, prior to training, participants attained more accuracy and completeness using the GEN-AI approach compared to the traditional approach. The variability in performance was marginally smaller in the GEN-AI condition (SD = 0.135) than in the traditional condition (SD = 0.152). This means that the outcomes were a little more consistent in the GEN-AI condition at the pre-stage. At the post-stage, the same descriptive trend remained. The traditional approach showed a mean F1 of 0.442 (SD = 0.126, median = 0.444), while GEN-AI produced a mean F1 of 0.570 (SD = 0.116, median = 0.571). Thus, GEN-AI continued to outperform the traditional approach descriptively after the training intervention. Notably, dispersion again appeared lower for GEN-AI (SD = 0.116) than for traditional (SD = 0.126), implying that the GEN-AI condition maintained a slightly more stable level of performance across the participant group at the post-stage. When comparing across time points, mean F1 decreased slightly from pre to post in both approaches (traditional: 0.470 to 0.442 and GEN-AI: 0.583 to 0.570). Table 5 summarizes the descriptive statistics for all four conditions and serves as the baseline reference for the paired comparisons and the 2 × 2 repeated measures ANOVA.
Table 5. Descriptive statistics for all four conditions.

4.2. Assumption Checks and Robustness Decisions

Prior to inferential testing, the key assumptions relevant to the selected within-subject analyses were assessed. Because paired-samples t-tests require the distribution of paired difference scores to be approximately normal, the normality assumption was evaluated using the Shapiro–Wilk test for each planned paired comparison. For the pre comparison (traditional–pre minus GEN-AI–pre), the Shapiro–Wilk test was non-significant (W = 0.987, p = 0.862), indicating that the normality assumption for the paired differences was not violated. Table 6 shows the values of the normality test for multiple conditions, where it is clearly visible that for the post-comparison (traditional–post minus GEN-AI–post), the Shapiro–Wilk test was significant (W = 0.925, p = 0.003), suggesting a departure from normality in the paired difference scores at the post-stage, as the value of p (0.003) < (0.05). This is also visible in the Q-Q plot in Figure 6, where the distribution made an S curve, which showed that it has some outliers. Given the evidence of non-normality for the post-paired differences, a Wilcoxon signed-rank test was reported alongside the paired-samples t-test as a robustness check [46,47] because it does not assume normally distributed difference scores. Importantly, the Wilcoxon test for the post-comparison remained statistically significant (p < 0.001) and produced a large rank-biserial correlation (r = −0.786), supporting the same substantive conclusion as the parametric test. This approach ensured that inference regarding method differences at the post-stage was not dependent solely on parametric assumptions. For the overall 2 × 2 repeated measures ANOVA (method × intervention), the design involved two-level within-subject factors (traditional vs. GEN-AI; pre vs. post). The ANOVA results were therefore interpreted primarily through the within-subject effects table, with effect sizes reported as partial eta squared (η2p). Collectively, these assumption checks and robustness decisions provided a transparent basis for selecting the appropriate inferential results reported in subsequent sections.
Table 6. Normality test (Shapiro–Wilk) for the pre- and post-condition.
Figure 6. Q-Q plot for F1-traditional–post–F1–GEN-AI–post.

4.3. Paired Comparisons by Stage

Paired comparisons were conducted to test whether the GEN-AI-assisted approach produced significantly different hazard identification performance (F1 score) compared with the traditional approach at each stage (pre and post). Because the same participants provided scores under both approaches, paired-samples procedures were appropriate, and results were reported with confidence intervals and effect sizes. Because a family of paired comparisons was conducted, a Bonferroni correction was applied to guard against inflation of the Type I error rate. With four comparisons, the corrected significance threshold is 0.0125 (0.05 divided by four). Both primary method comparisons (traditional versus GEN-AI at the pre- and post-stage) remained significant at p < 0.001 under this corrected threshold, so the substantive conclusions are unaffected by the correction.
At the pre-intervention stage, participants achieved higher F1 scores in the GEN-AI condition (M = 0.583, SD = 0.135) than in the traditional condition (M = 0.470, SD = 0.152). The negative sign reflected higher performance under GEN-AI than traditional. The effect size was medium to large in magnitude (Cohen’s d = −0.739), indicating that the improvement associated with GEN-AI assistance was not only statistically significant but also practically meaningful for hazard identification performance at baseline. Although the normality assumption for paired differences was satisfied at the pre-stage (Shapiro–Wilk p = 0.862), a non-parametric Wilcoxon signed-rank test is also available in the output and supported the same conclusion (W = 176, p < 0.001; rank-biserial correlation = −0.713) presented in Table 7. This consistency suggested that the pre-stage traditional vs. GEN-AI difference was robust to distributional assumptions.
Table 7. Paired t-test and Wilcoxon signed rank test for all four conditions.
The descriptive pattern stayed the same in the post-intervention stage. Once more, participants did better on the F1 test using the GEN-AI method (M = 0.570, SD = 0.116) than with the traditional method (M = 0.442, SD = 0.126). The effect size was again medium to large (Cohen’s d = −0.773), indicating that GEN-AI performed substantially better than traditional performance in the post-stage. The Shapiro–Wilk test indicated that the paired difference scores for the post-comparison deviated from normality (W = 0.925, p = 0.003), in contrast to the pre-comparison. To verify that the outcome was not exclusively reliant on parametric assumptions, a Wilcoxon signed-rank test was conducted as a robustness check [46,47]. The Wilcoxon test was still statistically significant (W = 126, p < 0.001) and showed a considerable effect (rank-biserial correlation = −0.786), which was in line with the direction and size of the paired t-test. Therefore, even under a distribution-free approach, GEN-AI assistance was associated with significantly higher F1 performance, which shows its effectiveness over the traditional method at the post-stage.

4.4. ANOVA Effects (Method, Intervention, Interaction)

To evaluate the overall pattern of results within a single model, a 2 × 2 repeated measures ANOVA was conducted with the method (traditional vs. GEN-AI) and the intervention (pre vs. post) as within-subject factors. This analysis tested three things. First, whether the F1 performance differed overall between methods. Secondly, whether performance changed overall from pre- to post-intervention stage, and lastly, whether any pre–post change depended on the interaction (method × time). Results were reported, including partial eta squared (η2p) as an effect size. Table 8 shows a statistically significant main effect of method was found, indicating that F1 scores differed between the traditional and GEN-AI approaches, F (1, 50) = 49.34, p < 0.001, η2p = 0.497. Estimated marginal means showed that GEN-AI produced higher overall F1 (M = 0.577, SE = 0.0143, 95% CI [0.548, 0.605]) than the traditional approach (M = 0.456, SE = 0.0129, 95% CI [0.430, 0.482]), as shown in Table 9. The magnitude of the effect (η2p = 0.497) indicated a large practical difference in performance attributable to method, consistent with the paired comparisons reported earlier. The dominant effect of the method is visible from the graph too (Figure 7). The main effect of time was not statistically significant, F (1, 50) = 1.00, p = 0.322, η2p = 0.020. The estimated marginal means suggested only a small change between stages (pre: M = 0.527, SE = 0.0171, 95% CI [0.493, 0.561]; post: M = 0.506, SE = 0.0124, 95% CI [0.481, 0.531]) as shown in Table 10. This indicated that, when averaging across methods, overall F1 did not show a statistically meaningful improvement (or decline) from pre to post in the current sample. This negligible effect of intervention in the pre- and post-stage is visible in the graph too (Figure 8). Importantly, the method × time interaction was also not significant, F (1, 50) = 0.279, p = 0.599, η2p = 0.006. This result suggested that the advantage associated with GEN-AI assistance was broadly consistent across stages, rather than increasing or decreasing substantially after training. The estimated cell means supported this interpretation (traditional–pre: M = 0.470; GEN-AI–pre: M = 0.583; traditional–post: M = 0.442; GEN-AI–post: M = 0.570) (Table 11). In practical terms, GEN-AI outperformed the traditional approach at both time points, while the overall pre–post shift remained comparatively small, and the interaction plot (Figure 9) was consistent with a largely stable method gap across time.
Table 8. Repeated measures ANOVA for method and intervention.
Table 9. Estimated marginal means showing the method’s effect.
Figure 7. Method’s effect on hazard identification performance.
Table 10. Estimated marginal means for the intervention’s effect.
Figure 8. Intervention’s effect on hazard identification performance.
Table 11. Estimated marginal means showing the combined effect (method–intervention).
Figure 9. Combined effect (method–intervention) on hazard identification performance.

5. Discussions

The observed advantage of GEN-AI assistance over the traditional approach was broadly consistent with the direction of recent construction safety research showing that technology-supported interventions can strengthen hazard recognition and safety decision-making when they provide structured guidance, improve access to safety knowledge, or shape how people attend to hazards. In this research, GEN-AI produced higher F1 scores at both stages, and the overall method effect was large in the repeated-measures model. This aligns with a growing body of work arguing that digital tools can help reduce missed hazards [16,26,48] and improve consistency by structuring the hazard search process and supporting comprehension of safety requirements.
First, comparable gains have been reported in research on immersive and technology-enhanced safety training where VR has been widely applied to support worker learning and hazard recognition. However, the effectiveness depends on scenario realism, feedback quality, and user engagement [49,50]. These findings complement the current results that, although GEN-AI is not an immersive visualization tool, it can serve a similar function, providing structure and guidance that helps users generate more complete hazard outputs than unguided, memory-based approaches. Second, eye-tracking research has repeatedly shown that hazard recognition depends strongly on visual search behavior and attention allocation [36,51], providing an explanatory basis for why support tools can change outcomes. The current GEN-AI advantage is consistent with the literature in the sense that even though GEN-AI does not directly change gaze behavior, it can reduce cognitive load and provide cueing to individuals on what to look for in a particular scenario. Third, the current findings fit well with emerging construction informatics research on LLM-based safety assistants. For example, Tran et al. [18] proposed an LLM-enabled approach for construction safety regulation extraction and described a “Construction Safety Query Assistant” concept to provide real-time answers to safety regulatory questions. This explicitly positions LLMs as tools to bridge knowledge gaps and support compliance-related reasoning. The present findings are additive rather than merely confirmatory relative to this body of work. Whereas earlier studies, including Uddin et al. [35], established the plausibility of ChatGPT-supported hazard recognition, the controlled within-subject comparison reported here quantifies the size of the advantage over a traditional baseline under a common scoring framework, which prior demonstrations did not establish.
This explains how GEN-AI can likely support participants in tasks by increasing access to structured safety knowledge and prompting more systematic hazard consideration. Fourth, earlier studies have reported problems with current hazard recognition approaches, like checklist limits (fixed templates miss implicit knowledge and context-specific hazards) [7,41], observer variability (as recognition depends heavily on experience and attention) [8,10], and inadequate hazard identification, as a lot of hazards were left unrecognized. The approach being used in this research addressed these issues as GEN-AI acts as a tool that helps with cognitive load, reducing the checklist issues; similarly, its image processing capability addressed the observer variability problem by providing context-specific hazards that are visible from its higher F1 scores compared to traditional approaches. Lastly, its effectiveness against traditional approaches showed that its hazard recognition is more adequate than earlier approaches.
An important caveat concerns the interpretation of the pre- to post-comparison. The pre-stage tasks were conducted at the two real sites, whereas the post-stage tasks were conducted on the virtual 14-scenario video. The intervention factor is therefore partially confounded with the evaluation environment, so the small decline in F1 from pre to post under both methods (traditional 0.470 to 0.442; GEN-AI 0.583 to 0.570) is at least as consistent with the virtual scenarios being more difficult than with the training being ineffective. Critically, this confound does not affect the principal finding: the method effect is a within-stage contrast between the traditional and GEN-AI conditions on identical stimuli, so it is unaffected by any difference in difficulty between the real and virtual environments. We therefore interpret the method effect as robust and treat the pre, post, and interaction results as exploratory.
Finally, the non-significant intervention effect is not unusual in the broader training literature, where results often vary depending on intervention intensity, fatigue, task complexity, and how “ground truth” is operationalized. Overall, the results extend recent research by showing that a GEN-AI workflow can deliver a robust performance benefit in hazard identification consistent with the direction of technology-assisted safety research, while also highlighting that changes over time may be smaller and more sensitive to study context than the method effect itself.
Beyond confirming the direction of the effect, the pattern of results is mechanistically informative. The large method effect alongside a negligible intervention effect suggests that the benefit of GEN-AI arises mainly from externalizing and structuring the hazard search, that is, from supplying a systematic prompt to consider categories of hazards that unaided memory tends to omit, rather than from improving the user’s underlying knowledge, which a short training session would be expected to influence. This interpretation is consistent with the eye-tracking literature, in which recognition depends on where and how systematic attention is allocated. It also explains why the effect was stable across the pre- and post-stages: the tool supplies the same structuring benefit regardless of the user’s recent training. Practically, this positions GEN-AI as a real-time cognitive aid rather than a substitute for competence, reinforcing the human in the loop governance model recommended in Section 6.

6. Recommendation

Based on the study findings, GEN-AI should be adopted in construction safety practice as a decision-support tool that strengthens hazard identification, while keeping accountability and final judgment with competent site personnel. The recommendations below translate the results into implementable steps. Use GEN-AI at defined “hazard points” in the workflow. GEN-AI support should be embedded where hazard identification is already expected, rather than added as a separate activity. Suitable touchpoints include pre-start/daily briefings, task-specific JSAs/method statement reviews, and safety walkdowns where observations are converted into structured hazard-control notes. We recommend applying a human-in-the-loop verification rule. GEN-AI outputs should be treated as suggestions, not facts. A simple control rule was recommended: AI proposes that a competent person validates the workface, controls are selected and recorded, and then a supervisor confirms implementation. Where hazards relate to higher-risk activities (work at height, lifting, electrical isolation, temporary works), verification should include the responsible person before authorization. Standardize prompts and output format to reduce variability. Organizations should develop a small set of approved prompts aligned with their risk categories. For example: “List hazards for [activity] at [location], organize by (falls/struck-by/caught-in/electrical/plant/), and propose controls following the hierarchy of controls.” Standardization improves consistency, reduces irrelevant responses, and makes output easier to compare across crews and projects. Maintain full traceability by recording what was approved, what was rejected, and the reason for each decision. To avoid relying too much on AI and to help with learning, teams should keep track of (i) the final list of hazards, (ii) which AI-suggested items were accepted or rejected, and (iii) the reason for rejection (not relevant/already controlled/inaccurate). This generates an audit record and helps both prompts and site practices get better over time. Teams should also set constraints on use that protect privacy, data, and safety. You should not paste any private project information, personal information, or secret designs into public tools. GEN-AI should not be used instead of formal engineering; it should be utilized to help find out what has to be checked early on and teach consumers how to utilize things safely. Short onboarding should include how to construct prompts, how to check outputs against the real workface, how to find hazards that are too generic or wrong, and how to minimize automation bias. The aim is to calibrate users’ trust in the technology, supporting broader hazard consideration while preserving individual accountability. These suggestions made GEN-AI a useful tool that can help enhance the quality of hazard detection while still keeping governance, verification, and accountability on site.

7. Conclusions

Within the controlled conditions and participant sample of this study, GEN-AI assistance was associated with a meaningful improvement in hazard identification performance compared with a traditional knowledge-based approach, with a strong and stable method effect observed across study stages. The findings supported the view that GEN-AI can serve as a practical decision-support layer that helps users generate more complete and better-structured hazard lists, while final validation and accountability must remain with competent site personnel. When implemented with clear boundaries, verification rules, and governance controls, GEN-AI has the potential to strengthen proactive safety practices and support more systematic hazard anticipation in construction project delivery.

8. Limitations

Several limitations should be considered when interpreting the findings. First, although the within-subject design strengthened internal validity for comparing traditional versus GEN-AI, the hazard-identification task still depended on human attention and visual search behavior, which are known to vary substantially. Systematic reviews highlight wide variability in recognition behavior across studies and settings [9,52]. As a result, generalization beyond the specific participants, scenarios, and work contexts represented in this dataset should be made cautiously. Because the sample comprised undergraduate and postgraduate construction students with limited site experience rather than practicing professionals, the absolute F1 values reported here should not be read as the performance expected of experienced site supervisors. Two implications follow for real-world applicability. First, experienced practitioners may begin from a higher unaided baseline, which could narrow the relative advantage of GEN-AI; conversely, the tool may be most valuable precisely for less experienced personnel, who form a large share of the construction workforce in many regions. Second, the null effect of the short training intervention may partly reflect the difficulty of shifting novice visual search behavior in a single session and may not generalize to longer or role-specific training. Confirmatory studies with industry professionals across experience strata are therefore needed before the magnitude of the benefit is generalized to operational settings.
Second, the study relied on an expert-defined ground truth for hazards. While necessary for objective scoring, ground truth development can introduce subjectivity (e.g., how hazards are defined/granulated, whether near-duplicate hazards are merged, and what constitutes an “acceptable match” in participant wording). This is particularly relevant for open-ended hazard listing, where participants may describe valid hazards using different language, which can affect TP/FP/FN assignment and thus F1.
A future design consideration concerns a task order. To prevent the GEN-AI outputs from contaminating the unaided baseline, every participant completed the traditional condition before the GEN-AI condition rather than using a counterbalanced order. A consequence of this fixed order is that the second (GEN-AI) pass benefited from prior cognitive engagement with the same scene, so a portion of the observed advantage may reflect carryover and practice rather than the tool alone. We retained the fixed order deliberately because counterbalancing would have allowed AI-generated hazards to anchor the subsequent unaided judgments and would have compromised the integrity of the baseline, which is the more serious threat for this research question. Nonetheless, the magnitude of the method effect should be interpreted with this carryover in mind, and future work should adopt counterbalanced or independent group designs with a washout interval to partition the tool effect from order effects.
Third, F1 was used as a single summary indicator. While F1 is widely used for balancing precision and recall, it can mask asymmetric costs of false positives versus false negatives and can behave differently under imbalance. Careful interpretation is required when relying on these metrics. Fourth, GEN-AI tools introduce model-specific risks. Large language models can generate plausible but incorrect outputs (“hallucinations”), and this reliability issue is well-documented in recent surveys and empirical research [53,54]. In operational safety situations, these kinds of mistakes might cause noise (more false positives) or make people too reliant on them. So, the results should be seen as proof that GEN-AI can make hazard-list performance better in the research setting. However, using it in the real world would need governance controls, grounding to trusted sources, and human validation. Finally, the findings are also bounded by the specific prompts, tool configuration, and GEN-AI version used; replication with alternative prompting strategies and model variants would strengthen robustness and transferability.

9. Future Research Areas

Future research should first replicate the current method comparison across a wider range of construction contexts to strengthen generalizability. Studies could test GEN-AI-assisted hazard identification across different project types (buildings, infrastructure, industrial plants), construction phases, and hazard families, because hazard cues and cognitive demands vary substantially by work environment. Research using eye-tracking and cognitive workload measures has shown that hazard recognition performance is sensitive to task type, hazard characteristics, and attention allocation, implying that broader scenario coverage is necessary to confirm the stability of GEN-AI benefits. Second, future work should not only report F1 but also examine how GEN-AI changes the trade-off between missed hazards (false negatives) and noise (false positives). Methodological discussions emphasize that F1 alone can mask important differences in error-type costs, and construction safety decisions often treat missing a high-severity hazard as more consequential than listing an extra low-severity hazard. Extending the evaluation to severity-weighted scoring or cost-sensitive metrics would make results more actionable for safety planning. Third, academics ought to investigate prompting strategies and grounding. Controlled experiments might compare “generic prompting” with organized prompts (by energy source, task sequence, or hazard category) and evaluate whether retrieval-augmented generation utilizing validated safety papers enhances factual accuracy while diminishing fabricated hazards. This is clear from the research, where certain replies from GEN-AI were far better than others. This illustrates that the system changes a lot depending on the prompts. The extensive LLM reliability literature identifies hallucination as a significant risk, prompting the assessment of grounding and verification methodologies for critical applications. Fourth, subsequent research should investigate user experience and behavioral impacts, encompassing over-reliance, automation bias, and trust calibration in the context of interactions with GEN-AI products. The research on occupational safety and decision support has shown that technology might inadvertently alter focus and judgment; thus, experimental designs must assess not just performance but also the manner in which users accept, reject, or validate AI recommendations [55].

Author Contributions

Conceptualization, M.A.M., K.A. and M.U.H.; formal analysis, M.A.M.; methodology, M.A.M., K.A., M.U.H. and H.K.; software programming, M.A.M.; visualization, M.A.M., Z.M. and I.M.; resources, K.A., Z.M., M.U.H. and I.M.; writing—original draft, M.A.M.; supervision, K.A. and M.U.H.; funding acquisition, Z.M. and I.M.; validation, K.A., Z.M., M.U.H., I.M. and H.K.; writing—review and editing, K.A., Z.M., M.U.H., I.M. and H.K. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

The study involved voluntary participation of adult university students in hazard-identification exercises conducted in real and virtual construction site environments. The research was non-invasive, involved minimal risk, did not include medical procedures or interventions, and did not collect sensitive personal information. Participant responses were anonymized prior to analysis, and no personally identifiable data were retained. Participation was voluntary, and informed consent was obtained from all participants. Therefore, formal Ethics Committee/Institutional Review Board approval was not required under the applicable procedures for minimal-risk educational and behavioral research.

Data Availability Statement

The data presented in this study are available on request from the corresponding author. The data are not publicly available due to ongoing research.

Acknowledgments

During the preparation of this manuscript/study, the authors used Grammarly for the purposes of language improvement. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Martínez-Aires, M.D.; López-Alonso, M.; de la Hoz-Torres, M.L.; Aguilar-Aguilera, A.; Arezes, P. Occupational risk prevention in the European Union construction sector: 30 Years since the publication of the Directive. Saf. Sci. 2024, 177, 106593. [Google Scholar] [CrossRef] [Scilit]
  2. Woźniak, Z.; Hoła, B. Analysing Near-Miss Incidents in Construction: A Systematic Literature Review. Appl. Sci. 2024, 14, 7260. [Google Scholar] [CrossRef] [Scilit]
  3. Afework, A.; Tamene, A.; Gashaw, M. Magnitude of self-reported non-fatal work-related injuries and associated factors among construction workers in Aleta Wondo, Sidama, Ethiopia. Sci. Rep. 2025, 15, 4339. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Allana, A.; Khan, A.A.; Yousuf, M.; Cullinan, P.; Nafees, A.A. Prevalence of occupational injuries among construction workers in Karachi, Pakistan. PLoS Glob. Public Health 2025, 5, e0004578. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Hassan, M.; Soomro, R.; Ahmed, S.; Aslam, A.; Rind, T.A. Occupational Safety Challenges in the Construction Industry of Sindh, Pakistan: An Analysis of Accident Data and Safety Implementation. Spectr. Eng. Sci. 2025, 3, 933–941. [Google Scholar]
  6. Zahid, H. Slips, Trips and Falls (STFs) as contributors of Injuries and fatalities in Construction Industries of Lahore—Pakistan. Pak. Soc. Sci. Rev. 2021, 5, 950–966. [Google Scholar] [CrossRef] [Scilit]
  7. Almaskati, D.; Kermanshachi, S.; Pamidimukkala, A.; Loganathan, K.; Yin, Z. A Review on Construction Safety: Hazards, Mitigation Strategies, and Impacted Sectors. Buildings 2024, 14, 526. [Google Scholar] [CrossRef] [Scilit]
  8. Uddin, S.M.J.; Albert, A.; Alsharef, A.; Pandit, B.; Patil, Y.; Nnaji, C. Hazard Recognition Patterns Demonstrated by Construction Workers. Int. J. Environ. Res. Public Health 2020, 17, 7788. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Xu, Q.; Chong, H.-Y.; Liao, P.-C. Exploring eye-tracking searching strategies for construction hazard recognition in a laboratory scene. Saf. Sci. 2019, 120, 824–832. [Google Scholar] [CrossRef] [Scilit]
  10. Xu, S.; Sun, M.; Fang, W.; Chen, K.; Luo, H.; Zou, P.X.W. A Bayesian-based knowledge tracing model for improving safety training outcomes in construction: An adaptive learning framework. Dev. Built Environ. 2023, 13, 100111. [Google Scholar] [CrossRef] [Scilit]
  11. Fu, H.; Tan, Y.; Xia, Z.; Feng, K.; Guo, X. Effects of construction workers’ safety knowledge on hazard-identification performance via eye-movement modeling examples training. Saf. Sci. 2024, 180, 106653. [Google Scholar] [CrossRef] [Scilit]
  12. Newaz, M.T.; Jefferies, M.; Ershadi, M. A critical analysis of construction incident trends and strategic interventions for enhancing safety. Saf. Sci. 2025, 187, 106865. [Google Scholar] [CrossRef] [Scilit]
  13. Johansen, K.W.; Schultz, C.; Teizer, J. Knowledge graph exploitation to enhance the usability of risk assessment in construction safety planning. Adv. Eng. Inform. 2025, 65, 103305. [Google Scholar] [CrossRef] [Scilit]
  14. Zhong, B.; Li, H.; Luo, H.; Zhou, J.; Fang, W.; Xing, X. Ontology-Based Semantic Modeling of Knowledge in Construction: Classification and Identification of Hazards Implied in Images. J. Constr. Eng. Manag. 2020, 146, 04020013. [Google Scholar] [CrossRef] [Scilit]
  15. Delhi, V.S.K.; Sankarlal, R.; Thomas, A. Detection of Personal Protective Equipment (PPE) Compliance on Construction Site Using Computer Vision Based Deep Learning Techniques. Front. Built Environ. 2020, 6, 136. [Google Scholar] [CrossRef] [Scilit]
  16. Fang, W.; Ma, L.; Love, P.E.D.; Luo, H.; Ding, L.; Zhou, A. Knowledge graph for identifying hazards on construction sites: Integrating computer vision with ontology. Autom. Constr. 2020, 119, 103310. [Google Scholar] [CrossRef] [Scilit]
  17. Lee, Y.; Kang, G.; Kim, J.; Yoon, S.; Jeon, J. GEN-AI-driven data augmentation for enhanced construction hazard detection. Autom. Constr. 2025, 177, 106317. [Google Scholar] [CrossRef] [Scilit]
  18. Tran, S.V.-T.; Yang, J.; Hussain, R.; Khan, N.; Kimito, E.C.; Pedro, A.; Sotani, M.; Lee, U.-K.; Park, C. Leveraging large language models for enhanced construction safety regulation extraction. J. Inf. Technol. Constr. 2024, 29, 1026–1038. [Google Scholar] [CrossRef] [Scilit]
  19. Hughes, P.; Ferrett, E. Introduction to Health and Safety at Work: For the NEBOSH National General Certificate in Occupational Health and Safety, 6th ed.; Routledge: London, UK, 2015. [Google Scholar] [CrossRef] [Scilit]
  20. Duryan, M.; Smyth, H.; Roberts, A.; Rowlinson, S.; Sherratt, F. Knowledge transfer for occupational health and safety: Cultivating health and safety learning culture in construction firms. Accid. Anal. Prev. 2020, 139, 105496. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Jeelani, I.; Albert, A.; Azevedo, R.; Jaselskis, E.J. Development and Testing of a Personalized Hazard-Recognition Training Intervention. J. Constr. Eng. Manag. 2017, 143, 04016120. [Google Scholar] [CrossRef] [Scilit]
  22. Albrechtsen, E.; Solberg, I.; Svensli, E. The application and benefits of job safety analysis. Saf. Sci. 2019, 113, 425–437. [Google Scholar] [CrossRef] [Scilit]
  23. Choudhary, S.; Solanki, P.; Gidwani, G. Job Safety Analysis (JSA) Applied in Construction Industry. IJSTE Int. J. Sci. Technol. Eng. 2018, 4, 177–187. [Google Scholar]
  24. Fang, W.; Ding, L.; Love, P.E.; Luo, H.; Li, H.; Peña-Mora, F.; Zhong, B.; Zhou, C. Computer vision applications in construction safety assurance. Autom. Constr. 2020, 110, 103013. [Google Scholar] [CrossRef] [Scilit]
  25. Kulinan, A.S.; Park, M.; Aung, P.P.W.; Cha, G.; Park, S. Advancing construction site workforce safety monitoring through BIM and computer vision integration. Autom. Constr. 2024, 158, 105227. [Google Scholar] [CrossRef] [Scilit]
  26. Jelodar, M.B. GEN-AI, Large Language Models, and ChatGPT in Construction Education, Training, and Practice. Buildings 2025, 15, 933. [Google Scholar] [CrossRef] [Scilit]
  27. Nyqvist, R.; Peltokorpi, A.; Seppänen, O. Can ChatGPT exceed humans in construction project risk management? Eng. Constr. Archit. Manag. 2024, 31, 223–243. [Google Scholar] [CrossRef] [Scilit]
  28. Sonkor, M.S.; García de Soto, B. Using ChatGPT in construction projects: Unveiling its cybersecurity risks through a bibliometric analysis. Int. J. Constr. Manag. 2025, 25, 741–749. [Google Scholar] [CrossRef] [Scilit]
  29. Mohamed, M.A.H.; Al-Mhdawi, M.; Ojiako, U.; Dacre, N.; Qazi, A.; Rahimian, F. GEN-AI in construction risk management: A bibliometric analysis of the associated benefits and risks. Urban. Sustain. Soc. 2025, 2, 198–230. [Google Scholar] [CrossRef] [Scilit]
  30. Guo, K.X.; Wong, P.K.-Y.; Cheng, J.C.P.; Chan, C.-F.; Leung, P.-H.; Tao, X. Enhancing visual-LLM for construction site safety compliance via prompt engineering and Bi-stage retrieval-augmented generation. Autom. Constr. 2025, 179, 106490. [Google Scholar] [CrossRef] [Scilit]
  31. Lee, J.; Ahn, S.; Kim, D.; Kim, D. Performance comparison of retrieval-augmented generation and fine-tuned large language models for construction safety management knowledge retrieval. Autom. Constr. 2024, 168, 105846. [Google Scholar] [CrossRef] [Scilit]
  32. Hussain, R.; Sabir, A.; Lee, D.-Y.; Alam Zaidi, S.F.; Pedro, A.; Abbas, M.S.; Park, C. Conversational AI-based VR system to improve construction safety training of migrant workers. Autom. Constr. 2024, 160, 105315. [Google Scholar] [CrossRef] [Scilit]
  33. Sammour, F.; Xu, J.; Wang, X.; Hu, M.; Zhang, Z. Responsible AI in Construction Safety: Systematic Evaluation of Large Language Models and Prompt Engineering. arXiv 2024, arXiv:2411.08320. [Google Scholar] [CrossRef] [Scilit]
  34. Fugate, H.; Alzraiee, H. Quantitative analysis of construction labor acceptance of wearable sensing devices to enhance workers’ safety. Results Eng. 2023, 17, 100841. [Google Scholar] [CrossRef] [Scilit]
  35. Uddin, S.M.J.; Albert, A.; Ovid, A.; Alsharef, A. Leveraging ChatGPT to Aid Construction Hazard Recognition and Support Safety Education and Training. Sustainability 2023, 15, 7121. [Google Scholar] [CrossRef] [Scilit]
  36. Han, Y.; Yin, Z.; Zhang, J.; Jin, R.; Yang, T. Eye-Tracking Experimental Study Investigating the Influence Factors of Construction Safety Hazard Recognition. J. Constr. Eng. Manag. 2020, 146, 04020091. [Google Scholar] [CrossRef] [Scilit]
  37. López, L.; Suárez-Ramírez, J.; Alemán-Flores, M.; Monzón, N. Automated PPE compliance monitoring in industrial environments using deep learning-based detection and pose estimation. Autom. Constr. 2025, 176, 106231. [Google Scholar] [CrossRef] [Scilit]
  38. Saunders, M.; Lewis, P.; Thornhill, A. Research Methods for Business Students; Pearson: Harlow, UK, 2019. [Google Scholar]
  39. Isah, M.A.; Kim, B.-S. Question-Answering System Powered by Knowledge Graph and Generative Pretrained Transformer to Support Risk Identification in Tunnel Projects. J. Constr. Eng. Manag. 2025, 151, 04024193. [Google Scholar] [CrossRef] [Scilit]
  40. Zhang, L.; Wang, J.; Wang, Y.; Sun, H.; Zhao, X. Automatic construction site hazard identification integrating construction scene graphs with BERT based domain knowledge. Autom. Constr. 2022, 142, 104535. [Google Scholar] [CrossRef] [Scilit]
  41. Purohit, D.P.; Siddiqui, N.A.; Nandan, A.; Yadav, B.P. Hazard Identification and Risk Assessment in Construction Industry. Int. J. Appl. Eng. Res. 2018, 13, 7639–7667. [Google Scholar]
  42. Bhandari, S.; Hallowell, M.R.; Van Boven, L.; Welker, K.M.; Golparvar-Fard, M.; Gruber, J. Using Augmented Virtuality to Examine How Emotions Influence Construction-Hazard Identification, Risk Assessment, and Safety Decisions. J. Constr. Eng. Manag. 2020, 146, 04019102. [Google Scholar] [CrossRef] [Scilit]
  43. Eiris, R.; Jain, E.; Gheisari, M.; Wehle, A. Online Hazard Recognition Training: Comparative Case Study of Static Images, Cinemagraphs, and Videos. J. Constr. Eng. Manag. 2021, 147, 04021082. [Google Scholar] [CrossRef] [Scilit]
  44. Darwis, A.M.; Nai’em, M.F.; Thamrin, Y.; Rahmadani, S.; Amin, F. Safety risk assessment in construction projects at Hasanuddin University. Gac. Sanit. 2021, 35, S385–S387. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  45. Blair, D.C. Information Retrieval, 2nd ed. C.J. Van Rijsbergen. London: Butterworths; 1979: 208 pp. Price: $32.50. J. Am. Soc. Inf. Sci. 1979, 30, 374–375. [Google Scholar] [CrossRef] [Scilit]
  46. Rainio, O.; Teuho, J.; Klén, R. Evaluation metrics and statistical tests for machine learning. Sci. Rep. 2024, 14, 6086. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Whitley, E.; Ball, J. Statistics review 6: Nonparametric methods. Crit. Care 2002, 6, 509. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  48. Baidoo-Anu, D.; Ansah, L.O. Education in the Era of Generative Artificial Intelligence (AI): Understanding the Potential Benefits of ChatGPT in Promoting Teaching and Learning. J. AI 2023, 7, 52–62. [Google Scholar] [CrossRef] [Scilit]
  49. Guo, X.; Liu, Y.; Tan, Y.; Xia, Z.; Fu, H. Hazard identification performance comparison between virtual reality and traditional construction safety training modes for different learning style individuals. Saf. Sci. 2024, 180, 106644. [Google Scholar] [CrossRef] [Scilit]
  50. Scorgie, D.; Feng, Z.; Paes, D.; Parisi, F.; Yiu, T.; Lovreglio, R. Virtual reality for safety training: A systematic literature review and meta-analysis. Saf. Sci. 2024, 171, 106372. [Google Scholar] [CrossRef] [Scilit]
  51. Jeelani, I.; Albert, A.; Han, K.; Azevedo, R. Are Visual Search Patterns Predictive of Hazard Recognition Performance? Empirical Investigation Using Eye-Tracking Technology. J. Constr. Eng. Manag. 2019, 145, 04018115. [Google Scholar] [CrossRef] [Scilit]
  52. Cheng, B.; Luo, X.; Mei, X.; Chen, H.; Huang, J. A Systematic Review of Eye-Tracking Studies of Construction Safety. Front. Neurosci. 2022, 16, 891725. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  53. Farquhar, S.; Kossen, J.; Kuhn, L.; Gal, Y. Detecting hallucinations in large language models using semantic entropy. Nature 2024, 630, 625–630. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  54. Huang, L.; Yu, W.; Ma, W.; Zhong, W.; Feng, Z.; Wang, H.; Chen, Q.; Peng, W.; Feng, X.; Qin, B.; et al. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. ACM Trans. Inf. Syst. 2025, 43, 1–55. [Google Scholar] [CrossRef] [Scilit]
  55. Abdelwanis, M.; Alarafati, H.K.; Tammam, M.M.S.; Simsekler, M.C.E. Exploring the risks of automation bias in healthcare artificial intelligence applications: A Bowtie analysis. J. Saf. Sci. Resil. 2024, 5, 460–469. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.