Next Article in Journal
Phase-Resolved Assessment of Helium-Ion-Induced Primary Damage in Nuclear-Facility Concrete Structures
Previous Article in Journal
Complexity and Construction Management: An Integrated Analysis of Logistics, Organizational Processes, and Productivity
Previous Article in Special Issue
Experimental Study on Seismic Performance of Rammed Earth and Rubble Masonry Walls
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

ISC-Perception: A Hybrid Vision Dataset for Robotic Assembly with Novel Intermeshed Steel Connections

1
School of Natural and Built Environment, Queen’s University Belfast, Belfast BT7 1NN, UK
2
School of Civil & Environmental Engineering, and Construction Management, University of Texas at San Antonio, San Antonio, TX 78249, USA
3
School of Electronics, Electrical Engineering, and Computer Science, Queen’s University Belfast, Belfast BT7 1NN, UK
4
Department of Civil and Urban Engineering, Tandon School of Engineering, New York University, Brooklyn, NY 11201, USA
*
Authors to whom correspondence should be addressed.
Buildings 2026, 16(17), 3407; https://doi.org/10.3390/buildings16173407
Submission received: 13 June 2026 / Revised: 11 August 2026 / Accepted: 19 August 2026 / Published: 26 August 2026

Abstract

Smart and sustainable construction increasingly depends on automation, yet robotic steel assembly still lacks task-specific perception data for bespoke connection systems. The Intermeshed Steel Connection (ISC) is a novel steel connection system that can reduce bolting effort and support faster, more reusable assembly, but dependable perception for ISC-aware robotic assembly remains underdeveloped. No public image corpus exists for ISC components, and collecting real site imagery is constrained by access, safety, privacy, and the limited deployment of ISC in practice. This paper introduces ISC-Perception, a hybrid vision dataset for near-field robotic assembly with novel Intermeshed Steel Connections. The dataset combines photorealistic CAD renders from SolidWorks Visualize, automatically annotated synthetic scenes generated in Unity, and a limited curated set of real ISC and human images. For a normalised 10,000-image Unity-based pipeline example, the proposed pipeline reduces estimated human effort to 30.5 h compared with 166.7 h for manual labelling, while the full training and validation set contains 15,928 images. Detectors trained on the hybrid dataset outperform synthetic-only and photorealistic-only alternatives, achieving mAP@0.50 of 0.756 on the complete test set. A near-size-matched comparison indicates that the improved performance is associated with the hybrid composition rather than dataset size alone under the evaluated training conditions. In a 1200-frame multi-view benchtop robotic assembly experiment, the detector achieves mAP@0.50/mAP@[0.50:0.95] of 0.943/0.823. These results show that ISC-Perception provides a practical route to data generation for emerging construction robotics applications where real imagery is scarce and supports the development of perception modules for robotic steel assembly in smart construction.

1. Introduction

1.1. Background and Motivation

Smart and sustainable construction increasingly depends on the ability to automate labour-intensive, safety-critical, and time-sensitive site operations. Robotic manipulators have transformed factory-based manufacturing, but their deployment in construction remains limited because building sites are unstructured, weather-exposed, spatially constrained, and governed by strict safety requirements. Steel frame erection is a strong candidate for automation: cranes often dominate the critical path, structural members are heavy and difficult to manoeuvre, and bolting operations require skilled workers to operate at height [1]. A robotic system capable of recognising, grasping, aligning, and assembling steel members could reduce crane time, improve worker safety, and mitigate skilled labour shortages.
In conventional steel construction, beams and columns are typically fabricated off-site and assembled on-site because complete frame assemblies are difficult to transport. This process involves several key steps: (1) identifying and lifting each structural element from the storage area, (2) transporting it to the installation location; (3) aligning it with the existing structure, and (4) fastening it to the structural frame using bolts or welds [1]. Although both methods are common for connecting steel members, bolts are generally preferred on-site due to their ease of installation, faster connection times, better quality control, and reduced inspection requirements. However, the extensive use of bolts in structural steel connections introduces additional challenges for the deployment of robots in the field.
The recently proposed Intermeshed Steel Connection (ISC) system eliminates most of the bolts required by conventional moment or shear splices. The ISC can be manufactured using cutting-edge technologies such as high-density plasma cutting, water jet cutting, and laser cutting [2]. ISC has two types of components: ISC member and ISC connection plates. The initial design of ISC had 3 connection plates on one side (Figure 1a) but the newer version requires only one connection plate on each side with fewer bolts (Figure 1b). Precision-cut male-female tabs guide members into alignment so that only a handful of set-bolts are needed to secure the joint [2,3,4]. By trimming cycle times and tolerating direct reuse, ISC reduces material waste and greenhouse gas emissions while preserving structural capacity. These benefits align with industry trends towards design-for-manufacture-and-assembly (DfMA) [5] and circular construction. Recent modular-steel-building research has also investigated specialised isolation devices for seismic resilience, illustrating the broader structural engineering context in which modular steel systems are developing [6]. Compared with conventional bolted connections, in which bolt heads, nuts, hole patterns, and plate boundaries provide relatively discrete and regularly arranged visual cues, ISC recognition depends more strongly on the fine geometry and spatial relationships of its precision-cut intermeshing tabs and connection plates. These repeated features can produce visually similar local contours and become progressively occluded as the members are aligned and intermeshed, causing the appearance of the same component to vary with viewpoint and assembly state. In addition, the reflective galvanised surfaces produce angle-dependent specular highlights that can reduce local edge contrast and obscure the geometric features required for detection. These characteristics motivate a data-generation strategy that spans varied viewpoints, assembly configurations, illumination conditions, and complementary synthetic, photorealistic, and real imagery.
Robotic ISC assembly therefore depends on reliable near-field perception. A robot must localise ISC connection plates and member ends in its immediate workcell, provide detections suitable for downstream grasp planning and mating, and monitor human presence for safety. Generic construction datasets, such as MOCS [7] and SODA [8], primarily represent workers, equipment, vehicles, and general site objects. They do not contain ISC geometry, nor do they capture the specific indoor and outdoor assembly cell contexts required for robotic steel assembly. Similarly, general-purpose computer vision datasets, such as COCO [9], lack the object classes, viewpoints, occlusions, and material appearances associated with ISC components.
The economic motivation for this work is substantial. The structural-steel market is substantial and is projected to continue growing, driven by urban development and industrial expansion [10]. Connection design can have a disproportionate influence on structural-steel frame cost because of its detailing, fabrication, and erection labour requirements [11,12]. Perception-enabled robotic assembly is therefore relevant as a technical challenge and as a route towards safer, faster, and more resource-efficient steel construction. A vision-enabled robotic ISC assembly workflow has therefore been proposed, offering the potential for safer construction sites and cost and schedule savings [13]. In the proposed ISC assembly workflow, a robot must (i) localise ISC connection plates and member ends in its immediate workcell, (ii) provide reliable detections for grasp planning and mating, and (iii) monitor human proximity for safety. Conventional datasets, such as Common Objects in Context (COCO) [9], MOCS [7], and SODA [8], contain neither ISC geometry nor the specific indoor/outdoor assembly cell context. ISC-Perception directly addresses this data gap for near-field perception, rather than general stockyard inventory, and is already integrated into a benchtop robotic assembly pipeline (Section 5.5). This paper addresses the perception data gap by introducing ISC-Perception, a hybrid vision dataset for robotic assembly with novel Intermeshed Steel Connections. The dataset combines automatically annotated Unity scenes, photorealistic SOLIDWORKS Visualize 2023 renders, and curated real images of ISC components and humans. It is designed for near-field robotic assembly perception rather than general stockyard inventory or full-site monitoring. ISC-Perception supports object detection of ISC members, ISC connection plates, and humans, and is evaluated through both fixed test set experiments and a benchtop robotic assembly scenario. In this way, ISC-Perception provides a practical route to perception data generation for smart construction applications where the target object is novel, real imagery is scarce, and CAD assets are available.

1.2. Perception Data for Construction Robotics

Progress in deep learning has been strongly shaped by the availability of large, diverse, and accurately labelled image datasets. Construction robotics perception imposes additional demands: detectors must recognise partially occluded objects, operate under variable illumination, and generalise across projects that differ in geometry, material finish, background clutter, and weather conditions [14,15,16,17,18,19]. Image datasets therefore form a central foundation for training and evaluating computer vision models for object detection, classification, and segmentation in construction environments [20].
Generic datasets such as COCO or ImageNet misrepresent site reality: they lack steel members, cranes, PPE, and the dense clutter typical of erection yards. Direct transfer can depress mean Average Precision (mAP) by up to 40% when models are tested on construction imagery [21]. Building an in-domain corpus is equally fraught. Cameras are often barred by safety briefings, union rules, or privacy regulations; outdoor shoots hinge on weather windows; and pixel-accurate annotation of high-resolution frames can consume weeks of person-hours [22]. The hurdle is steeper still for bespoke components such as the ISC, for which no archival photographs yet exist and whose galvanised surfaces frustrate automated labelling.
Data, not algorithms, have thus become the principal bottleneck. An effective remedy must supply (i) scale for deep networks, (ii) fidelity to capture ISC’s subtle tab geometry, and (iii) diversity in backgrounds, lighting, and occlusions, while curbing manual annotation cost. Section 1.3 surveys how synthetic and photorealistic imagery can satisfy those requirements and where current approaches fall short.

1.3. Synthetic and Photorealistic Data: Benefits and Pitfalls

A practical solution to address the challenges of limited access and varying construction site conditions is the creation of annotated synthetic image datasets to supplement real ones [21]. These synthetic datasets can be generated using computer graphics techniques, 3D modelling software or game engines, enabling the simulation of diverse construction environments with different objects and backgrounds.
Computer vision models trained solely on synthetic images often perform worse than those trained on real images. For example, grocery item detection models trained on 400,000 synthetic images performed less effectively than models trained with only 760 real images [23]. Yet, combining just 760 real images with the synthetic images produced superior results compared to both models. Moreover, randomisation techniques (such as lighting conditions, weather conditions, time of day, textures, and camera perspective) are used to generate synthetic images, reduce the sim-to-real gap, and improve dataset diversity [24,25]. Therefore, a hybrid dataset that integrates real and synthetic images could be an effective approach for training computer vision models for construction applications [26]. However, obtaining sufficient real images for many construction scenarios or custom objects, such as the ISC, remains difficult. In such cases, computer-aided design tools can generate and render photorealistic models of custom objects in various settings, reducing the reliance on real images.
Beyond domain randomisation, sim-to-real transfer can also be addressed through domain adaptation. Image-level approaches translate simulated and real observations towards a shared or canonical appearance, whereas feature-level approaches encourage the detector to learn representations that are invariant across source and target domains. More recent teacher–student and self-training methods additionally exploit unlabelled target-domain images through pseudo-labels and consistency objectives [27,28,29,30]. These approaches provide complementary mechanisms for reducing the visual discrepancy between simulated and real observations. Our approach addresses the sim-to-real gap through data design rather than a dedicated domain-adaptation objective, combining controlled randomisation with hybrid training across Unity-generated, SolidWorks Visualize, and limited real imagery.
Additionally, ISC plates pose an additional hurdle: their galvanised coating creates specular highlights that shift with sun angle, and the laser-cut tab patterns differ by millimetres. Capturing these cues demands high-dynamic-range rendering plus fine surface normal maps that are costly to generate at scale. Conversely, photographing ISC plates on active sites remains impractical, because the system is not yet widely deployed. Hence a hybrid strategy [(i) auto-generates large volumes of domain-randomised synthetic frames, (ii) injects photorealistic ray-traced scenes for material fidelity, and (iii) enhances the mix with a small, curated set of real photographs] offers the best trade-off between cost and realism.
Research gap. To date, no public dataset combines these three modalities for steel-connection detection; existing construction corpora (MOCS [7], SODA [8]) neither model bespoke joints nor provide labels. Bridging this gap is therefore prerequisite to closing the perception loop for robotic ISC assembly.

1.4. Research Contribution

Existing vision datasets in construction focus on equipment or personnel safety. None address robotic assembly of structural steel components such as beams, columns, or ISC plates. We fill this void by devising and releasing ISC-Perception, a task-specific, hybrid corpus for object detection in robotic steel erection.
In summary, the main contributions of this paper are
  • A methodology for creating a hybrid dataset for ISC components using real, photorealistic, and synthetic images to tackle the scarcity of real images tailored for robotic assembly tasks, together with a normalised 10,000-image Unity-based pipeline comparison that reduces estimated human effort from 166.7 h for manual annotation to 30.5 h (81.7%).
  • The analysis of training performance of computer vision algorithms for different types of images and validation of the trained computer vision model in small-scale setup. We further demonstrate that detectors trained on ISC-Perception achieve mAP@0.50 of 0.943 and mAP@[0.50:0.95] of 0.823 on a 1200-frame multi-view bench test that mimics a robotic assembly cell.
To contextualise these contributions, Section 2 reviews the current state of the art in real and synthetic computer vision datasets. Subsequently, Section 3 discusses the procedural approach for generating the hybrid dataset. Section 4 provides insight into the ISC-Perception dataset. Section 5 reviews the outcomes of the training and testing phases, followed by a discussion of the results and findings. Finally, Section 6 of the paper summarises the research results and their significant impacts on the construction industry. Our focus is on a task-specific dataset and reproducible data-generation pipeline rather than proposing a new detection architecture. YOLOv8 models are used strictly as reproducible baselines to isolate the effect of dataset composition; we do not claim an algorithmic contribution.

2. Literature Review

This section provides an overview of prominent general-purpose computer vision datasets (see Section 2.1) and datasets specific to the construction industry (see Section 2.2).

2.1. Computer Vision Datasets

As this research focuses on generating an image dataset for ISC, this section provides a comprehensive overview of prominent computer vision datasets. Computer vision datasets can be broadly categorised into two main types: real-world datasets and synthetic datasets.
Real-world datasets consist of images captured from actual environments and are crucial for training and evaluating models across a range of tasks, including object recognition, object detection, segmentation, and scene understanding [31]. Numerous widely used datasets have been developed to support these tasks. Prominent examples include ImageNet [32], COCO [9], Pascal Visual Object Classes [33], Open Images [34], Cityscapes [35], and KITTI [36]. Table 1 provides an overview of these key datasets, highlighting their specific features and contributions to the field.
Synthetic dataset generation in computer vision involves creating artificial images and annotations using tools such as rendering engines (e.g., Blender [37], Unity 3D [38], NVIDIA Omniverse [39], Unreal Engine [40]), physics-based simulation software (e.g., Gazebo [41], Webots [42], CoppeliaSim [43]), and generative AI methods like GANs. These datasets are particularly valuable for generating large-scale, cost-effective, and safe alternatives to real-world data collection. Examples include synthetic datasets derived from video games like Half-Life 2 [44], the SYNTHIA dataset for semantic segmentation [45], Hattori et al.’s 3D pedestrian models using Autodesk 3DS Max [46], and the Virtual-KITTI [47] and Virtual-KITTI 2 [48] datasets, which replicate urban driving scenes with automated annotations via Unity.
Synthetic datasets can be generated using various 3D CAD model rendering and visualisation software, incorporating appropriate lighting and scene generation techniques. For example, Aubry et al. developed a dataset of 86,366 synthesised images by rendering each of 1393 high-quality 3D chair models from 62 distinct viewpoints [49]. Additionally, Peng et al. explored the influence of pose, colour, textures, and background by training a deep convolutional neural network (CNN) using crowd-sourced 3D CAD models, highlighting the potential of synthetic data in improving model performance [50].

2.2. Computer Vision Dataset in Construction Industry

Vision technology has garnered significant interest across multiple sectors, including construction, where its application is transforming how visual data from construction sites is acquired and interpreted. This technology enables the extraction of valuable information such as progress monitoring, object detection, safety condition analysis, and quality control. Through the automated detection and tracking of workers, excavators, cranes, dump trucks, and other equipment, it is possible to efficiently identify unsafe conditions on construction sites [51,52,53,54].
SODA [8], tailored for construction sites, contains 19,846 images of 15 object classes. The Moving Objects in Construction Sites (MOCS) dataset contains 41,668 images depicting 13 types of moving objects, including equipment and workers, commonly found on construction sites [7]. Those images were captured using a camera, UAV, and smartphone from 174 different construction sites involving dams, bridges, buildings, tunnels, and highways [7]. Del et al. created a small dataset of 1046 images comprising eight different object classes for detecting construction equipment and humans [55]. All these datasets were collected from real construction sites, carefully chosen and edited to remove any privacy information, and manually annotated, which is laborious and time-consuming. In contrast, Barrera-Animas and Delgado proposed a method to generate synthetic datasets that closely resemble real-world conditions, using 3D models of construction machinery, workers, site environments, and assets, combined with realistic lighting conditions in different seasons [56]. However, all of the mentioned datasets focus primarily on detecting construction equipment and workers to ensure safe operations, with none designed specifically for robotic assembly tasks.

3. Method of Generating Computer Vision Hybrid Dataset

This research aims to develop a dataset specifically for the robotic assembly of steel structures using ISC. Accordingly, we created a hybrid dataset comprising photorealistic images of ISC components, synthetic images from the simulation environment, and a limited number of real images, and used it to train a YOLOv8 model for robotic assembly.

3.1. Dataset Composition

The creation of a robust computer vision dataset often begins with the selection of target objects for detection or segmentation. In this work, the dataset focuses on three main object classes: (a) ISC member, (b) ISC connection plate, and (c) human; as illustrated in Figure 1, the selection of these classes is driven by the requirements of future robotic assembly tasks. We envisage that the robot must accurately identify ISC components for assembly and detect humans to ensure safety compliance. While typical steel construction sites include equipment such as tower cranes, forklifts, and scaffolds, these are excluded from the dataset since the robot will not interact with them. Our design targets the robot’s near-field assembly workcell and the lightly cluttered yard areas that feed it. Members and plates are placed on the ground or simple supports, giving multi-instance scenes with partial occlusion rather than the dense, opaque stockyard stacks seen in long-term storage; those heavily stacked regimes remain out of scope and are noted as a limitation in Section 6.
Given the novelty of ISC, real images of these components are limited in availability. Hence, to address this, the ISC-Perception dataset integrates three types of images from diverse sources:
1.
Type 1: Photorealistic images from SolidWorks (SW) Visualize (category 1 or C1)
2.
Type 2: Synthetic images from Unity
  • Built-in randomizers (category 2 or C2)
  • Custom randomizers (category 3 or C3)
3.
Type 3: Real images
  • From previous project (category 4 or C4)
  • Human images from Roboflow Universe Public Dataset (category 5 or C5)

3.2. Image Generation

The image generation workflow is shown in Figure 2. Synthetic images in Unity (C2) sometimes suffer from jittering, motion blur, and unrealistic appearances (see Supplementary Figure S1). To overcome these limitations, custom randomizers (C3) were employed to generate images with enhanced variability across indoor and outdoor assembly scenes. Similarly, photorealistic images (C1) were generated with 3D CAD models in SolidWorks Visualize, incorporating diverse lighting and backgrounds. Finally, the dataset includes manually annotated real images of ISC components and humans, augmented through preprocessing techniques to bolster diversity. This hybrid composition ensures the dataset is both diverse and generalisable and provides real-world authenticity to support the development of vision systems capable of detecting ISC components in complex assembly environments.

3.2.1. Photorealistic Images from SolidWorks Visualize

As previously established, real images of ISC components are scarce, hence we use CAD rendering software enabled by SolidWorks Visualize to generate high-quality photorealistic images to supplement the limited availability of real-world ISC data in ISC-Perception.
The first stage involves the creation of several 3D models of the two main ISC components using CAD software, as shown in Figure 2. This is then followed by importing the models into SolidWorks Visualize for scene generation. During this stage, SolidWorks Visualize provides extensive randomisation options to enhance dataset diversity. For randomisation, SolidWorks Visualize selects from 9 total backgrounds and 3 model textures (1 metallic texture and 2 featuring rust), rotates between 0 and 360°, and varies the lighting conditions; see Table 2. The output images from the Scene Generation stage are then annotated in Roboflow to get ground truth bounding boxes of ISC objects. Finally, the photorealistic images are augmented in Roboflow to add more variations to the dataset.

3.2.2. Synthetic Images from Unity

While the use of SolidWorks Visualize produces high-quality photorealistic images, the process of manual annotation is time-consuming. Unity, with its Perception package, offers an efficient alternative for generating large volumes of automatically annotated synthetic images. C3 in Figure 2 shows the synthetic image generation process.
Scene Simulation Generation: To start with, we used the models built using SOLIDWORKS 2023 during the generation of photorealistic images and imported those to Unity. Two steel structure assembly simulation scenes were created; one indoor and one outdoor (see Figure 3 for sample views of robotic steel assembly). The indoor scene included a large workspace with walls displaying custom images to simulate construction environments. Fifty random construction site images were used as wall textures, changing every second to increase variation (Figure 3). The scene was populated with objects such as concrete mixers, dump trucks, scaffolds, and ISC components, placed in various orientations to enhance generalisation. Lighting conditions included directional light mimicking sunlight, dynamically adjusted between −100° and 100°, and multiple indoor light sources for a realistic indoor setup.
The outdoor scene contained environmental elements such as trees and buildings alongside construction equipment, safety barriers, and ISC components placed on pallets or the ground. A directional light simulated sunlight, rotating to create shadows from objects like trees and buildings. ISC components were coloured with solid green, red, and white finishes to further diversify the dataset.
The use of custom randomizers addressed the limitations of Unity’s built-in randomizers, which failed to generate realistic ISC environments. ISC components were rotated incrementally by 5°, and background objects were randomly rotated and translated using custom scripts. These randomisations ensured variability in the dataset. Ground truth labels were assigned using Unity’s Perception package. Annotated images were recorded with a first-person camera capturing the scene through a keyboard and mouse-controlled player.
This approach efficiently produced a diverse dataset of annotated synthetic images, complementing the photorealistic and real images, while addressing the limitations of Unity’s built-in randomizers in simulating real-world ISC environments.
The Unity-generated training/validation and test images were produced from different Unity scenes. The training and validation subsets were formed only from the training-side Unity imagery, whereas the 2748 Unity images in the held-out test set were generated from the separate test scenes and were not used during model training, validation, hyperparameter selection, or checkpoint selection. The same ISC component classes and general randomisation framework were retained across the subsets to preserve a consistent detection task, but the rendered test scenes and image instances were separate from those used for training and validation.
Although the Unity Perception pipeline can generate depth maps, instance masks, and other scene-level labels [38], the present ISC-Perception release standardises on RGB images with 2D bounding boxes. This provides a common annotation space across Unity, SolidWorks Visualize, and the real-image sources, for which equivalent pixel-aligned depth and mask annotations were not available. The bounding boxes support object-level detection and define regions of interest for subsequent modules such as metric localisation, 6-DoF pose estimation, grasp planning, alignment, and human presence monitoring; they do not themselves provide depth, pose, or assembly control outputs. A future multimodal extension could retain Unity-derived depth and instance masks and combine them with calibrated RGB-D or multi-view real capture, enabling cross-modality validation and benchmarks for segmentation, depth-aware localisation, and pose estimation.

3.2.3. Real Images

ISC-Perception includes real images to enhance the dataset’s authenticity and variability. However, with no publicly available ISC dataset, the real images were curated from multiple sources: still frames from ISC assembly videos, manually collected images, and publicly available datasets for humans from Roboflow Universe [57].
For the real ISC component of Dataset 3’s training and validation portion, 29 frames were initially extracted from video footage produced during an earlier ISC research project. Of these frames, 16 were selected for annotation after duplicate and closely similar frames had been removed. The selected frames were annotated using Roboflow’s annotation tool and processed using brightness adjustments ( 15 % to + 15 % ) and rotations ( 10 to + 10 ). Additional augmentations, including 90 rotations, flips, and saturation adjustments, produced 82 real ISC images that were used only for training and validation.
Separately, 207 images were manually photographed using fabricated small-scale ISC members and connection plates with varied appearances and positions and were manually annotated using Label Studio [58]. These images originated from source photographs or frames different from those used to produce the 82-image group and were reserved exclusively as the real ISC component of the held-out test set. They were not used for training, validation, augmentation, hyperparameter tuning, threshold selection, checkpoint selection, or any other model development decision.
The 82 real ISC training and validation images and the 207 real ISC test images were therefore derived from separate source imagery and served different experimental purposes.
Additionally, to address safety considerations, images of humans were included. Since numerous publicly available annotated datasets exist for humans, the Roboflow Universe dataset was used [57]. This dataset contains 235 images of individuals in various standing and sitting poses, with preprocessing effects such as colour, brightness, shear, and stretch. All human images were manually scrutinised to address privacy concerns before inclusion in the dataset.

4. Dataset

The ISC-Perception dataset integrates images from four primary sources as described in Section 3.1. These sources include 13,399 images collected via Unity with custom randomizers (C3), 3599 images via SolidWorks Visualize (C1), 289 images of Real Images from a previous ISC project (C4), and 1728 images of Human Images from Roboflow Universe (C5). The total number of images in the dataset is distributed across training, validation, and testing sets, ensuring diversity and representation of different scenarios. Dataset 3 contains 15,928 training and validation images, and the held-out test set contains 3087 images, giving a complete corpus of 19,015 images. The held-out test set comprises 2748 Unity images, 48 SolidWorks Visualize images, 84 Roboflow human images, and 207 separately collected and manually annotated real ISC images; real ISC images therefore constitute approximately 6.7% of the test set. Test images were manually selected to capture varied scenarios, ensuring robust evaluation. The normalised Unity-based human-time accounting is presented in Table 3. See Table 4 and Table 5 for the summary and distribution of the dataset.

4.1. Time-to-Dataset Accounting (Synthetic vs. Manual)

Table 3 presents a normalised comparison for a 10,000-image Unity-based synthetic pipeline. Under the stated assumptions, manual annotation requires 166.7 h, whereas estimated synthetic pipeline human effort is 30.5 h, an 81.7% reduction; the separate compute wall-clock is 12 h. These values are not measured totals for the complete 15,928-image hybrid training and validation set. The complete hybrid set also includes SolidWorks Visualize, Roboflow, and real-image components that required source-specific manual annotation, selection, checking, curation, or augmentation. Because this additional work was not captured by the Unity timing model, that model cannot be extrapolated to the complete hybrid set without further labour records.

4.2. Statistics of the Datasets

To evaluate the model’s performance, three versions of ISC-Perception datasets were created, with a constant test set across all three versions. Table 5 provides a detailed distribution of images across datasets, while Figure 4 illustrates key statistics.
1.
Image Distribution: Dataset 3 (Hybrid Dataset) contains the largest number of images (15,928) from all sources, while Dataset 2 (SolidWorks Visualize with Roboflow Human) has fewer images (5195). Dataset 1 (Unity custom randomizer) focuses solely on synthetic data, with 10,651 images.
2.
Instances per Class: Figure 4a shows that ISC members dominate with 15,270 instances in the test set, followed by ISC connection plates (14,530) and humans (242). Dataset 1 has the highest average number of instances per image for each class, as shown in Figure 4c. Across the fixed test set, many frames contain several ISC members and plates simultaneously, so detections are typically made in multi-instance, partially occluded scenes rather than isolated single-object views.
3.
Percentage of Sources: In Dataset 3, 67 % of images come from Unity, 22 % from SolidWorks Visualize, 10 % from Roboflow Universe, and approximately 0.5 % from real ISC images (Figure 4d). ISC is a newly developed novel connection that has not yet been commercially adopted in construction, so it is difficult to get real images of ISC. Hence, only 82 real images of ISC could be collected and included in the dataset, as shown in Table 5. These 82 images belong only to Dataset 3’s training and validation portion.

4.3. Example Images

Figure 5 showcases sample images from the dataset:
1.
Unity Synthetic Images: Figure 5a,b highlight images generated using built-in and custom randomizers, with varying lighting, object placements, and occlusion effects.
2.
Photorealistic Images: Figure 5c illustrates high-quality images from SolidWorks Visualize, featuring diverse scenes and objects, including greyscale and construction site settings.
3.
Real Images: Figure 5d shows human images from Roboflow Universe, while Figure 5e presents real images from previous ISC assembly projects. Roboflow Universe aggregates community-contributed images from multiple providers (which may include stock-photography sources); we therefore cite the dataset as Roboflow Universe for panel (d).

5. Performance Analysis

Performance analysis of the datasets was divided into two categories. Initially, three different computer vision models were generated by training the YOLOv8 algorithm with three different datasets (dataset 1, dataset 2 and dataset 3) from Table 5. The training performance was initially analysed using several performance metrics. Finally, the trained model was applied to the test set from Table 5 to predict the desired object. The effects of different image types and datasets were analysed based on prediction performance. We first report full-size training results, then a near-size-matched comparison that evaluates the datasets at comparable training-set sizes while retaining source-composition differences (Section 5.4.2).

5.1. Hardware Configuration

To create the object detection model, the YOLOv8 algorithm was trained with all three versions of the dataset. An Alienware m16 laptop configured with an Intel Core i9 13900HX processor, 32.0GB RAM, and an NVIDIA GeForce RTX 4060 12GB GDDR6 graphics card was used to train the ISC component-detection model. The full-size Dataset 1 run used a 250-epoch cap with patience 30, whereas the full-size Dataset 2 and Dataset 3 runs used a 300-epoch cap with patience 50. The retained best checkpoint from each run was used for the complete test set comparison.

5.2. Training Settings

The full-size Dataset 1, Dataset 2, and Dataset 3 models were trained using COCO-pretrained YOLOv8s weights with Ultralytics 8.0.202 [59], an input size of 640 × 640, batch size 16, and seed 0. Dataset 1 completed its 250-epoch run with patience 30. Dataset 2 and Dataset 3 used a 300-epoch cap with patience 50 and produced 51 and 265 logged epochs, respectively; the retained Dataset 3 checkpoint achieved its best validation result at epoch 215. The later near-size-matched comparison was conducted separately using YOLOv8n with Ultralytics 8.3.198, an input size of 640 × 640, batch size 16, seed 42, deterministic mode, a cosine learning-rate schedule, a 100-epoch budget, and patience 50. All three near-size-matched runs completed the full 100-epoch budget.

5.3. Testing Procedure

The best models trained were obtained from a collective of three distinct dataset configurations: the hybrid dataset (which includes custom randomizer, SolidWorks Visualize, Roboflow, and manually annotated images), a dataset using only the custom randomizer images, and a dataset using SolidWorks Visualize images with Roboflow human images. Tests were then conducted using two groups of test data:
1.
Complete held-out test set: This set contains 3087 images: 2748 Unity images, 48 SolidWorks Visualize images, 84 Roboflow human images, and 207 separately collected and manually annotated real ISC images. The 207 real ISC images were not used for training, validation, augmentation, or model development decisions.
2.
Small scale bench test: This set comprises real-world samples using a multi-view, small scale experimental setup of robotic assembly of ISC, where synchronised cameras provided multiple perspectives of the same scene. This setup allows us to do continuous object detection and tracking of ISC components and humans.
To examine performance at comparable training-set sizes, we also evaluated models trained on near-size-matched subsets; see Section 5.4.2.

5.4. Results and Discussion: Testing with Test Set

5.4.1. Results Overview

The three models trained on the hybrid dataset, custom randomizer dataset, and SolidWorks Visualize and Roboflow dataset were all tested on a held-out test set (Table 5) containing 3087 images, including samples from all data sources: Unity custom randomizer, SolidWorks Visualize, Roboflow human images, and manually annotated real-world images. Four performance metrics (precision, recall, mAP@0.50 and mAP@[0.50:0.95]) were used to assess the trained performance of the model across the test data. We present the results of training in Table 6.
The model trained on the Hybrid dataset achieved the highest overall performance with an mAP@[0.50:0.95] of 0.664 compared to 0.564 for the custom randomizer dataset and 0.321 for the SolidWorks Visualize and Roboflow dataset, indicating that it is able to generalise across a diverse range of image types (see Table 6). The performance was particularly strong for human detection, where it achieved an mAP@[0.50:0.95] of 0.804 and high precision and recall scores. On the other hand, the model’s performance for ISC connection plates was lower with an mAP@[0.50:0.95] of 0.523, likely due to the complexity of detecting these components in varied real-world environments (Table 6).
To assess performance on real imagery separately from the mixed-source test set, the Dataset 3 model was evaluated on the 207 real ISC images within the test set. It achieved a precision of 0.633, recall of 0.355, mAP@0.50 of 0.505, and mAP@[0.50:0.95] of 0.334. The lower performance on this real-image subset relative to the aggregate mixed-source result is consistent with a remaining synthetic-to-real domain gap and limited real-data coverage. These values should therefore be interpreted as a source-specific diagnostic rather than as a replacement for the complete test set results.

5.4.2. Controlled Near-Size-Matched Comparison

We compared the later runs using retained training sets of 5194 images for Dataset 1 and 5195 images each for Dataset 2 and Dataset 3; the held-out test set remained fixed at 3087 images. Dataset 1 and Dataset 3 used sampled subsets of their larger source datasets, whereas Dataset 2 used its full training set. All three models used the common YOLOv8n configuration described in Section 5.2 and completed the full 100-epoch budget. At these near-identical training-set sizes, the hybrid composition achieved mAP@0.50 of 0.675 and mAP@[0.50:0.95] of 0.549, exceeding Dataset 1 (0.546/0.430) and Dataset 2 (0.249/0.206), with higher precision (0.830 versus 0.776/0.649) and recall (0.625 versus 0.505/0.146). This corresponds to +0.129 mAP@0.50 (+23.6%) and +0.119 mAP@[0.50:0.95] (+27.7%) over Dataset 1, and +0.426/+0.343 (+171%/+166%) over Dataset 2. The large gains at near-identical training-set sizes indicate that the improvement is associated with the hybrid composition rather than dataset size alone (Table 7).

5.4.3. Confusion Matrix and Performance Curves

As shown in the confusion matrix in Figure 6, the model trained on the hybrid dataset correctly identified a large proportion of ISC components and human instances compared to the models trained on the custom randomizer dataset and the SolidWorks Visualize and Roboflow human dataset. However, all trained models exhibited some misclassification. For example, the trained model successfully detected 7530 connection plates, while 232 connection-plate instances were misclassified as ISC members and 6768 were missed as background (Figure 6a). The model trained on the hybrid dataset also correctly detected 11,324 instances of ISC members and 195 instances of human (Figure 6a). However, in Figure 6c, the model trained on the SolidWorks Visualize and Roboflow human dataset performed worst, correctly detecting only 1833 instances of ISC connection plates, 2449 instances of ISC members, and 163 instances of human on the test set. In Figure 7a, the F1-confidence curve shows an overall F1 score of 0.74 at a confidence threshold of 0.377. Precision remained high for all classes, as demonstrated in the precision-confidence curve Figure 7c, although ISC connection plate detection showed a noticeable drop-off in recall, confirming that the model sometimes missed these components in challenging scenes.

5.4.4. Discussion on Model Performance Comparison: Hybrid Dataset vs. Custom Randomizer Dataset vs. SolidWorks Visualize and Roboflow Dataset

The evaluation results of the models trained on the three distinct datasets are presented in Table 6, which summarises their performance metrics for each class.
Our experiments were designed primarily to examine how training-data composition affects detection performance while holding the detector architecture fixed within each comparison. In the full-size comparison, Dataset 1, Dataset 2, and Dataset 3 were all trained using YOLOv8s, and Dataset 3 achieved the strongest aggregate performance. The later near-size-matched experiment provides a complementary size-controlled comparison: using YOLOv8n and near-identical training-set sizes, Dataset 3 again achieved the strongest performance. Taken together, these results indicate that the observed improvement is associated with the hybrid composition under the evaluated training conditions rather than dataset size alone. Because both experiments use models from the YOLOv8 family, however, the current evidence does not establish that the magnitude of this advantage will remain unchanged across different detector architectures. Cross-architecture replication with region-based or transformer-based detectors would be required to test whether the same dataset ranking persists under different modelling assumptions.
Hybrid dataset model: The model trained on the Hybrid dataset (Table 6) demonstrated the highest overall performance across all object classes. It achieved a box precision of 0.846 and a recall of 0.666, leading to an overall mAP@0.50 of 0.756 and mAP@[0.50:0.95] of 0.664. The performance in detecting ISC connection plates was slightly lower, with mAP@[0.50:0.95] of 0.523, indicating some difficulty in precise identification. However, the model excelled in ISC member detection, attaining mAP@[0.50:0.95] of 0.664, and performed exceptionally well in detecting humans, with mAP@[0.50:0.95] of 0.804 and a recall of 0.777.
Custom randomizer dataset model: The custom randomizer dataset model (Table 6) displayed a similar trend, but with a lower overall performance compared to the Hybrid dataset model. The box precision of 0.818 and recall of 0.521 resulted in an overall mAP@0.50 of 0.659 and mAP@[0.50:0.95] of 0.564. For ISC connection plates, the performance was comparable to the Hybrid dataset model with mAP@[0.50:0.95] of 0.523. However, the model exhibited reduced accuracy in detecting humans, with mAP@[0.50:0.95] of 0.503, indicating limitations in handling more diverse human instances.
SolidWorks Visualize and Roboflow dataset model: The model trained on the SolidWorks Visualize and Roboflow dataset (Table 6) performed the weakest overall, reflecting its narrow focus on the dataset. It achieved a much lower box precision of 0.404 and recall of 0.321, resulting in an overall mAP@0.50 of 0.386 and mAP@[0.50:0.95] of 0.321. For ISC connection plates, the mAP@[0.50:0.95] was the lowest at 0.176, and ISC member detection also lagged, with mAP@[0.50:0.95] of 0.244. While this model was relatively better at human detection with mAP@[0.50:0.95] of 0.541, its overall ability to generalise to ISC components was clearly limited. From these results, it is evident that the Hybrid Dataset model provides the best performance across all object categories, particularly excelling in human detection and ISC member identification. The custom randomizer model, while decent, struggles with human detection and generalisation to real-world data. Lastly, the SolidWorks Visualize and Roboflow Human model shows clear limitations, particularly for ISC components, due to its narrow training focus.
Within the hybrid training corpus, the 82 real ISC images form a small but deliberately included component. They expose the detector to capture characteristics from the physical assembly environment that are difficult to reproduce fully through rendering alone. However, our experiments evaluate Dataset 3 as a combined training corpus and do not isolate the numerical contribution of those 82 images. We therefore interpret the results as evidence for the hybrid source strategy as a whole, rather than as evidence that the real-image component independently caused the observed improvement. Quantifying its marginal contribution would require a matched source-removal ablation.
The separate evaluation on the 207 real ISC test images also shows that a synthetic-to-real gap remains: performance on this subset is lower than on the complete mixed-source test set. This result places an important boundary on the current findings. The available real imagery represents a limited range of capture and assembly conditions and does not cover the full variation expected across construction environments, including lighting, weather, dust, surface contamination, background clutter, camera motion, component scale, viewpoint, occlusion, and assembly stage. Broader site- or session-separated real-world evaluation is therefore needed before extending the current results to unrestricted construction deployment. The same real-domain gap also motivates subsequent evaluation of domain-adaptation approaches using separately collected real imagery while preserving the current held-out test set for evaluation.
The Roboflow human images serve a narrower role within our experiments: they broaden the variation in human appearance, pose, scale, and background available to the human presence class, but they are not intended to represent the full range of worker behaviour encountered during ISC assembly. Because these images were not collected specifically in ISC assembly settings, they do not capture all construction-specific PPE, task-dependent poses, component-induced occlusions, camera distances, workcell layouts, or human–component interactions. Our results should therefore not be interpreted as validation of a complete industrial safety monitoring system. Extending the dataset with assembly-specific worker imagery collected across representative workcell and construction conditions would provide a stronger basis for that assessment.

5.5. Small Scale Bench Testing

The overall objective of our research is to develop and integrate the object detection model into a broader robotic assembly process. To assess the robustness and performance of the model in real-world conditions, we incorporated its output into the vision module of our assembly framework. The object detection model serves as an intermediary for subsequent vision-based tasks within this framework. Our setup, as shown in Figure 8, consists of a multi-camera system, where synchronised views provide multiple perspectives to minimise occlusion and enhance overall detection robustness. The benchtop layout mirrors a realistic ISC assembly cell: members are fixtured on a worktable, connection plates are presented within the robot’s reachable workspace, and a human operator moves in and out of the scene (Figure 8). From a 2 min, 60 fps bench-test video with two synchronised side views (Figure 8), we temporally subsampled every 10th frame to reduce correlation. Following curation and removal of some sampled frames, the final 1200 frames were manually annotated for ISC components and humans. Using standard detection metrics, the detector achieved mAP@0.50 = 0.943 , mAP@[0.50:0.95] = 0.823 , precision = 0.951 , and recall = 0.930 (see Figure 9 for detailed detection result samples). However, the system was not without its challenges. Failure cases were primarily due to glare in the front-facing camera view (Figure 10), leading to occasional missed detections of connection plates; improving robustness to challenging lighting and appearance shifts is left for future work.
The reflective galvanised surfaces of ISC components remain a challenging source of appearance variation because strong specular highlights can suppress local edge contrast or obscure fine intermeshing geometry. The existing hybrid dataset introduces illumination diversity across rendered and real sources, but further targeted mitigation could strengthen robustness under severe glare. Future extensions should include broader variation in light position, intensity, colour temperature, exposure, and material reflectance during synthetic generation, together with additional real-image capture under deliberately challenging illumination. Practical acquisition measures such as diffuse lighting, polarising filters, exposure bracketing, and high-dynamic-range capture could also be evaluated, alongside glare-aware augmentation or domain-adaptation methods. These measures would complement, rather than replace, the present hybrid source strategy.

6. Conclusions

This research demonstrated a procedural approach to generate a hybrid dataset for detecting ISC components in a steel structure assembly site. As the collection of real images from the steel structure assembly site is an arduous and unsafe task, synthetic and photorealistic images were created to compensate for the need for real images. Synthetic images from Unity 3D provide annotated images, while photorealistic images from SolidWorks Visualize require manual annotation. Multiple datasets were created using various types of images. Dataset 1 was created using only synthetic images generated with Unity’s custom randomizers. Dataset 2 was created using photorealistic images from SolidWorks Visualize and real human images from the Roboflow Universe public dataset. Dataset 3 was created from Unity’s custom randomizer, SolidWorks Visualize, Roboflow, and real-world images. Only three object classes (ISC connection plate, ISC member, and human) were selected for object detection tasks. These were selected because the robot will only manipulate the ISC connection plate and ISC member, while the human was added to ensure safety.
Testing results showed that models trained on the hybrid dataset (Dataset 3) outperformed those trained on either synthetic (Dataset 1) or photorealistic data alone (Dataset 2). The hybrid dataset model demonstrated superior precision and recall across all object classes (ISC connection plate, ISC member, and human) when tested on the complete test set. The custom randomizer (Dataset 1) model achieved reasonable performance in testing but still lagged behind the hybrid model. The model trained on SolidWorks Visualize and Roboflow Human images (Dataset 2) had the lowest performance, especially in detecting ISC connection plates and ISC members, highlighting the difficulty of generalising from such a limited dataset.
To improve the impact of the synthetic images in the hybrid dataset, higher-quality and realistic simulation scenarios will be created in future work. With proper computer-aided design tools and graphics software, accurate colour, shape, and scale will be created for the objects in the steel structure assembly site scene in Unity. Future work will also automate annotation of the SolidWorks Visualize imagery by exploiting information available within the CAD and rendering scene. Component identity, three-dimensional geometry and pose, together with virtual-camera parameters, could be used to project each object into the image plane and generate class labels, bounding boxes, instance masks, or keypoints automatically. Visibility and occlusion checks would be required so that the generated annotations correspond to the observable image region rather than only to the full projected extent of the CAD model. This CAD-to-image workflow would reduce the manual annotation burden while improving the scalability, consistency, and reproducibility of photorealistic dataset generation. More natural lighting and other environmental effects will be introduced for both Unity and the SolidWorks Visualize scene to enhance the photorealism. A key limitation of the current release is that it targets the assembly workcell and lightly cluttered yard scenes; dense stockyard stacks and full-site inventory management remain out of scope and will be addressed in future extensions of ISC-Perception. While this process of generating a hybrid dataset is not completely automatic, this method provides a procedure to generate a hybrid dataset where the collection of real images is very difficult, and the 3D model of the target object is readily available.
Overall, the results of this work reinforce the importance of using datasets for robust object detection in real-world industrial settings and provide a foundation for future research in automating complex tasks in the construction and assembly industries. This method provides a scalable approach that can be further adapted to other robotic applications, particularly in challenging environments where human access is restricted, such as nuclear sites, tunnels, or remote workspaces. In addition, it has been robustly shown that hybrid datasets can provide a rich way to train computer vision models, especially for emerging applications where limited relevant public datasets are available.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/buildings16173407/s1, Figure S1, Sample images generated from Unity with built-in randomisers: (a) sunlight reflection on the floor; (b) random objects with shadows; (c) objects in random positions with camera blur; (d) objects in random positions with hue shift; Figure S2, Sample images generated from Unity with custom randomisers: (a) full scene; (b) randomised objects with mixed illumination; (c) spotlighted poses; (d) dark objects on bright background; Figure S3, Sample images from SolidWorks Visualize: (a) red ISC member and metallic connection plate in a glass building; (b) rusty ISC member and plate at a construction site; (c) single red ISC member in a glass building; (d) greyscale view inside a boiler room; Figure S4, Sample human figures from the Roboflow Universe: (a) on a white background; (b) after rotation; (c) raw capture on site; and Figure S5, Real ISC assembly images from a prior project: (a) attached ISC components; (b) worker at assembly site; (c) lifting an ISC member; (d) fully assembled ISC.

Author Contributions

Conceptualization, M.R., S.A., D.M., K.R. and D.F.L.; methodology, M.R., S.A. and D.H.; software, M.R.; validation, M.R. and S.A.; formal analysis, M.R. and S.A.; investigation, M.R., S.A. and D.A.A.-M.; resources, D.H., D.M., K.R. and D.F.L.; data curation, M.R. and S.A.; writing—original draft preparation, M.R. and S.A.; writing—review and editing, M.R., S.A., D.A.A.-M., D.H., D.M., K.R. and D.F.L.; visualization, M.R. and S.A.; supervision, D.H., D.M., K.R. and D.F.L.; project administration, D.M.; funding acquisition, D.M., K.R. and D.F.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Science Foundation (NSF), grant number 2222815; Science Foundation Ireland, grant number 21/US/3797; and the Department for the Economy (DfE), UK, grant number USI-218.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The ISC-Perception dataset is available from the corresponding author upon reasonable request for research and industrial use.

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

References

  1. Liang, C.J.; Kang, S.C.; Lee, M.H. RAS: A robotic assembly system for steel structure erection and assembly. Int. J. Intell. Robot. Appl. 2017, 1, 459–476. [Google Scholar] [CrossRef] [Scilit]
  2. Shemshadian, M.E.; Labbane, R.; Schultz, A.E.; Le, J.L.; Laefer, D.F.; Al-Sabah, S.; McGetrick, P. Experimental study of intermeshed steel connections manufactured using advanced cutting techniques. J. Constr. Steel Res. 2020, 172, 106169. [Google Scholar] [CrossRef] [Scilit]
  3. Al-Sabah, S.; Laefer, D.F.; Truong Hong, L.; Phuoc Huynh, M.; Le, J.L.; Martin, T.; Matis, P.; McGetrick, P.; Schultz, A.; Shemshadian, M.E.; et al. Introduction of the Intermeshed Steel Connection—A New Universal Steel Connection. Buildings 2020, 10, 37. [Google Scholar] [CrossRef] [Scilit]
  4. Shemshadian, M.E.; Le, J.L.; Schultz, A.E.; McGetrick, P.; Al-Sabah, S.; Laefer, D.F.; Martin, A.; Hong, L.T.; Huynh, M.P. Numerical study of the behavior of intermeshed steel connections under mixed-mode loading. J. Constr. Steel Res. 2019, 160, 89–100. [Google Scholar] [CrossRef] [Scilit]
  5. Montazeri, S.; Lei, Z.; Odo, N. Design for Manufacturing and Assembly (DfMA) in Construction: A Holistic Review of Current Trends and Future Directions. Buildings 2024, 14, 285. [Google Scholar] [CrossRef] [Scilit]
  6. Mo, Z.; Lai, B.; Shu, G.; Yang, T.Y.; Ventura, C.E.; Liew, J.Y.R. Enhancing seismic resilience in modular steel building through three-dimensional isolation. Eng. Struct. 2025, 323, 119269. [Google Scholar] [CrossRef] [Scilit]
  7. Xuehui, A.; Li, Z.; Zuguang, L.; Chengzhi, W.; Pengfei, L.; Zhiwei, L. Dataset and benchmark for detecting moving objects in construction sites. Autom. Constr. 2021, 122, 103482. [Google Scholar] [CrossRef] [Scilit]
  8. Duan, R.; Deng, H.; Tian, M.; Deng, Y.; Lin, J. SODA: A large-scale open site object detection dataset for deep learning in construction. Autom. Constr. 2022, 142, 104499. [Google Scholar] [CrossRef] [Scilit]
  9. Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft COCO: Common Objects in Context. In Proceedings of the European Conference on Computer Vision (ECCV); Springer International Publishing: Cham, Switzerland, 2014; pp. 740–755. [Google Scholar] [CrossRef] [Scilit]
  10. Precedence Research. Structural Steel Market Revenue to Attain USD 177.97 Bn by 2033. 2025. Available online: https://www.precedenceresearch.com/insights/structural-steel-market (accessed on 29 May 2025).
  11. Carter, C.J. Practical Information for Designers: Economy in Steel. In Structures 2004; American Society of Civil Engineers: Reston, VA, USA, 2012; pp. 1–8. [Google Scholar] [CrossRef] [Scilit]
  12. Chambers, R. Connecting the Costs of a Steel Frame. 2022. Available online: https://www.ellandsteel.com/connecting-the-costs-of-a-steel-frame/ (accessed on 29 May 2025).
  13. Rahman, M.; Adebayo, S.; Hester, D.; McPolin, D.; Rafferty, K.; Awolusi, I.; Laefer, D.F. A proposed strategy for automating Intermeshed Steel Connection assembly using robotics. In Proceedings of the 42nd International Symposium on Automation and Robotics in Construction (ISARC 2025), Montreal, QC, Canada, 28–31 July 2025; IAARC Publications: Oulu, Finland, 2025; Volume 42, pp. 477–484. [Google Scholar] [CrossRef] [Scilit]
  14. Jiang, Z.; Messner, J.I. Computer Vision Applications In Construction And Asset Management Phases: A Literature Review. J. Inf. Technol. Constr. (ITcon) 2023, 28, 176–199. [Google Scholar] [CrossRef] [Scilit]
  15. Li, Y.; Zhang, Y. Application Research of Computer Vision Technology in Automation. In Proceedings of the 2020 International Conference on Computer Information and Big Data Applications (CIBDA), Guiyang, China, 17–19 April 2020; pp. 374–377. [Google Scholar] [CrossRef] [Scilit]
  16. Nain, M.; Sharma, S.; Chaurasia, S. Safety and Compliance Management System Using Computer Vision and Deep Learning. IOP Conf. Ser. Mater. Sci. Eng. 2021, 1099, 012013. [Google Scholar] [CrossRef] [Scilit]
  17. Guo, B.H.; Zou, Y.; Fang, Y.; Goh, Y.M.; Zou, P.X. Computer vision technologies for safety science and management in construction: A critical review and future research directions. Saf. Sci. 2021, 135, 105130. [Google Scholar] [CrossRef] [Scilit]
  18. Seo, J.; Han, S.; Lee, S.; Kim, H. Computer vision techniques for construction safety and health monitoring. Adv. Eng. Inform. 2015, 29, 239–251. [Google Scholar] [CrossRef] [Scilit]
  19. Teizer, J. Right-time vs. real-time pro-active construction safety and health system architecture. Constr. Innov. 2016, 16, 253–280. [Google Scholar] [CrossRef] [Scilit]
  20. Zhong, B.; Wu, H.; Ding, L.; Love, P.E.; Li, H.; Luo, H.; Jiao, L. Mapping computer vision research in construction: Developments, knowledge gaps and implications for research. Autom. Constr. 2019, 107, 102919. [Google Scholar] [CrossRef] [Scilit]
  21. Soltani, M.M.; Zhu, Z.; Hammad, A. Automated annotation for visual recognition of construction resources using synthetic images. Autom. Constr. 2016, 62, 14–23. [Google Scholar] [CrossRef] [Scilit]
  22. Mostafa, K.; Hegazy, T. Review of image-based analysis and applications in construction. Autom. Constr. 2021, 122, 103516. [Google Scholar] [CrossRef] [Scilit]
  23. Borkman, S.; Crespi, A.; Dhakad, S.; Ganguly, S.; Hogins, J.; Jhang, Y.C.; Kamalzadeh, M.; Li, B.; Leal, S.; Parisi, P.; et al. Unity Perception: Generate Synthetic Data for Computer Vision. arXiv 2021, arXiv:2107.04259. [Google Scholar] [CrossRef] [Scilit]
  24. Remmas, W.; Lints, M.; Uudmäe, J.J. PCGOD: Enhancing Object Detection with Synthetic Data for Scarce and Sensitive Computer Vision Tasks. IEEE Access 2025, 13, 91325–91333. [Google Scholar] [CrossRef] [Scilit]
  25. Wang, G.; Li, H.; Li, P.; Lang, X.; Feng, Y.; Ding, Z.; Xie, S. M4SFWD: A Multi-Faceted synthetic dataset for remote sensing forest wildfires detection. Expert Syst. Appl. 2024, 248, 123489. [Google Scholar] [CrossRef] [Scilit]
  26. Bayraktar, E.; Yigit, C.B.; Boyraz, P. A hybrid image dataset toward bridging the gap between real and simulation environments for robotics. Mach. Vis. Appl. 2019, 30, 23–40. [Google Scholar] [CrossRef] [Scilit]
  27. Bousmalis, K.; Irpan, A.; Wohlhart, P.; Bai, Y.; Kelcey, M.; Kalakrishnan, M.; Downs, L.; Ibarz, J.; Pastor, P.; Konolige, K.; et al. Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping. In Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia, 21–25 May 2018; pp. 4243–4250. [Google Scholar] [CrossRef] [Scilit]
  28. James, S.; Wohlhart, P.; Kalakrishnan, M.; Kalashnikov, D.; Irpan, A.; Ibarz, J.; Levine, S.; Hadsell, R.; Bousmalis, K. Sim-To-Real via Sim-To-Sim: Data-Efficient Robotic Grasping via Randomized-To-Canonical Adaptation Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 12627–12637. [Google Scholar] [CrossRef] [Scilit]
  29. Li, Y.J.; Dai, X.; Ma, C.Y.; Liu, Y.C.; Chen, K.; Wu, B.; He, Z.; Kitani, K.; Vajda, P. Cross-Domain Adaptive Teacher for Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 7581–7590. [Google Scholar] [CrossRef] [Scilit]
  30. Kennerley, M.; Wang, J.G.; Veeravalli, B.; Tan, R.T. CAT: Exploiting Inter-Class Dynamics for Domain Adaptive Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 16541–16550. [Google Scholar] [CrossRef] [Scilit]
  31. Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. ImageNet Large Scale Visual Recognition Challenge. Int. J. Comput. Vis. 2015, 115, 211–252. [Google Scholar] [CrossRef] [Scilit]
  32. Deng, J.; Dong, W.; Socher, R.; Li, L.J.; Li, K.; Fei-Fei, L. ImageNet: A large-scale hierarchical image database. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, 20–25 June 2009; pp. 248–255. [Google Scholar] [CrossRef] [Scilit]
  33. Everingham, M.; Van Gool, L.; Williams, C.K.I.; Winn, J.; Zisserman, A. The Pascal Visual Object Classes (VOC) Challenge. Int. J. Comput. Vis. 2010, 88, 303–338. [Google Scholar] [CrossRef] [Scilit]
  34. Kuznetsova, A.; Rom, H.; Alldrin, N.; Uijlings, J.; Krasin, I.; Pont-Tuset, J.; Kamali, S.; Popov, S.; Malloci, M.; Kolesnikov, A.; et al. The Open Images Dataset V4. Int. J. Comput. Vis. 2020, 128, 1956–1981. [Google Scholar] [CrossRef] [Scilit]
  35. Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; Schiele, B. The Cityscapes Dataset for Semantic Urban Scene Understanding. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 3213–3223. [Google Scholar] [CrossRef] [Scilit]
  36. Geiger, A.; Lenz, P.; Stiller, C.; Urtasun, R. Vision meets robotics: The KITTI dataset. Int. J. Robot. Res. 2013, 32, 1231–1237. [Google Scholar] [CrossRef] [Scilit]
  37. Basak, S.; Javidnia, H.; Khan, F.; McDonnell, R.; Schukat, M. Methodology for Building Synthetic Datasets with Virtual Humans. In Proceedings of the 2020 31st Irish Signals and Systems Conference (ISSC), Letterkenny, Ireland, 11–12 June 2020; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  38. Unity Technologies. Unity Perception Package. 2020. Available online: https://github.com/Unity-Technologies/com.unity.perception (accessed on 18 August 2026).
  39. Akar, C.A.; Tekli, J.; Jess, D.; Khoury, M.; Kamradt, M.; Guthe, M. Synthetic Object Recognition Dataset for Industries. In Proceedings of the 2022 35th SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI), Natal, Brazil, 24–27 October 2022; pp. 150–155. [Google Scholar] [CrossRef] [Scilit]
  40. Qiu, W.; Yuille, A. UnrealCV: Connecting Computer Vision to Unreal Engine. In Proceedings of the Computer Vision—ECCV 2016 Workshops; Springer: Cham, Switzerland, 2016; pp. 909–916. [Google Scholar]
  41. Koenig, N.; Howard, A. Design and use paradigms for Gazebo, an open-source multi-robot simulator. In Proceedings of the 2004 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE Cat. No.04CH37566), Sendai, Japan, 28 September–2 October 2004; Volume 3, pp. 2149–2154. [Google Scholar] [CrossRef] [Scilit]
  42. Michel, O. Cyberbotics Ltd. Webots™: Professional Mobile Robot Simulation. Int. J. Adv. Robot. Syst. 2004, 1, 5. [Google Scholar] [CrossRef] [Scilit]
  43. Rohmer, E.; Singh, S.P.N.; Freese, M. V-REP: A Versatile and Scalable Robot Simulation Framework. In Proceedings of the 2013 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Tokyo, Japan, 3–7 November 2013; pp. 1321–1326. [Google Scholar] [CrossRef] [Scilit]
  44. Taylor, G.R.; Chosak, A.J.; Brewer, P.C. OVVV: Using Virtual Worlds to Design and Evaluate Surveillance Systems. In Proceedings of the 2007 IEEE Conference on Computer Vision and Pattern Recognition, Minneapolis, MN, USA, 17–22 June 2007; pp. 1–8. [Google Scholar] [CrossRef] [Scilit]
  45. Ros, G.; Sellart, L.; Materzynska, J.; Vazquez, D.; Lopez, A.M. The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 3234–3243. [Google Scholar] [CrossRef] [Scilit]
  46. Hattori, H.; Boddeti, V.N.; Kitani, K.; Kanade, T. Learning scene-specific pedestrian detectors without real data. In Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; pp. 3819–3827. [Google Scholar] [CrossRef] [Scilit]
  47. Gaidon, A.; Wang, Q.; Cabon, Y.; Vig, E. VirtualWorlds as Proxy for Multi-object Tracking Analysis. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 4340–4349. [Google Scholar] [CrossRef] [Scilit]
  48. Cabon, Y.; Murray, N.; Humenberger, M. Virtual KITTI 2. arXiv 2020, arXiv:2001.10773. [Google Scholar] [CrossRef] [Scilit]
  49. Aubry, M.; Maturana, D.; Efros, A.A.; Russell, B.C.; Sivic, J. Seeing 3D Chairs: Exemplar Part-Based 2D-3D Alignment Using a Large Dataset of CAD Models. In Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA, 23–28 June 2014; pp. 3762–3769. [Google Scholar] [CrossRef] [Scilit]
  50. Peng, X.; Sun, B.; Ali, K.; Saenko, K. Learning Deep Object Detectors from 3D Models. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; pp. 1278–1286. [Google Scholar] [CrossRef] [Scilit]
  51. Du, S.; Shehata, M.; Badawy, W. Hard hat detection in video sequences based on face features, motion and color information. In Proceedings of the 2011 3rd International Conference on Computer Research and Development, Shanghai, China, 11–13 March 2011; Volume 4, pp. 25–29. [Google Scholar] [CrossRef] [Scilit]
  52. Rezazadeh Azar, E.; McCabe, B. Automated Visual Recognition of Dump Trucks in Construction Videos. J. Comput. Civ. Eng. 2012, 26, 769–781. [Google Scholar] [CrossRef] [Scilit]
  53. Chi, S.; Caldas, C.H. Automated Object Identification Using Optical Video Cameras on Construction Sites. Comput.-Aided Civ. Infrastruct. Eng. 2011, 26, 368–380. [Google Scholar] [CrossRef] [Scilit]
  54. Park, M.W.; Brilakis, I. Construction worker detection in video frames for initializing vision trackers. Autom. Constr. 2012, 28, 15–25. [Google Scholar] [CrossRef] [Scilit]
  55. Del Savio, A.; Luna, A.; Cárdenas-Salas, D.; Vergara, M.; Urday, G. Dataset of manually classified images obtained from a construction site. Data Brief 2022, 42, 108042. [Google Scholar] [CrossRef] [Scilit]
  56. Barrera-Animas, A.Y.; Davila Delgado, J.M. Generating real-world-like labelled synthetic datasets for construction site applications. Autom. Constr. 2023, 151, 104850. [Google Scholar] [CrossRef] [Scilit]
  57. Tank Detect. Person Dataset. 2025. Available online: https://universe.roboflow.com/tank-detect/person-dataset-kzsop (accessed on 2 September 2025).
  58. Tkachenko, M.; Malyuk, M.; Holmanyuk, A.; Liubimov, N. Label Studio: Data Labeling Software, 2020–2024. Open Source Software. Available online: https://github.com/HumanSignal/label-studio (accessed on 18 August 2026).
  59. Jocher, G.; Chaurasia, A.; Qiu, J. YOLO by Ultralytics. 2023. Available online: https://github.com/ultralytics/ultralytics (accessed on 2 December 2024).
Figure 1. Components of ISC beam-to-beam; (a) earlier version of fabricated ISC [3], (b) CAD drawing of ISC with single connection plate. Colours in panel (b) are used only to visually differentiate the CAD components and do not encode scientific data.
Figure 1. Components of ISC beam-to-beam; (a) earlier version of fabricated ISC [3], (b) CAD drawing of ISC with single connection plate. Colours in panel (b) are used only to visually differentiate the CAD components and do not encode scientific data.
Buildings 16 03407 g001
Figure 2. Source of images and workflow for creating the hybrid dataset combining different types of images.
Figure 2. Source of images and workflow for creating the hybrid dataset combining different types of images.
Buildings 16 03407 g002
Figure 3. View of robotic steel assembly in Unity; (a) indoor scene; (b) outdoor scene. The coloured arrows are Unity transform gizmos and do not encode experimental data.
Figure 3. View of robotic steel assembly in Unity; (a) indoor scene; (b) outdoor scene. The coloured arrows are Unity transform gizmos and do not encode experimental data.
Buildings 16 03407 g003
Figure 4. Dataset statistics; (a) number of instances for each class, (b) percentage of instances in each dataset, (c) number of instances per image for each class, and (d) percentage of images from different sources in Dataset 3.
Figure 4. Dataset statistics; (a) number of instances for each class, (b) percentage of instances in each dataset, (c) number of instances per image for each class, and (d) percentage of images from different sources in Dataset 3.
Buildings 16 03407 g004
Figure 5. Representative samples from ISC-Perception: (a) Unity (built-in randomizers, C2), (b) Unity (custom randomizers, C3), (c) SolidWorks Visualize photorealistic render (C1), (d) Human example from Roboflow Universe (C5), (e) Real ISC frame (C4). Roboflow Universe aggregates contributions from multiple providers (which can include stock libraries); we therefore cite Roboflow Universe as the source for (d).
Figure 5. Representative samples from ISC-Perception: (a) Unity (built-in randomizers, C2), (b) Unity (custom randomizers, C3), (c) SolidWorks Visualize photorealistic render (C1), (d) Human example from Roboflow Universe (C5), (e) Real ISC frame (C4). Roboflow Universe aggregates contributions from multiple providers (which can include stock libraries); we therefore cite Roboflow Universe as the source for (d).
Buildings 16 03407 g005
Figure 6. Confusion matrix plots of trained models on the test set: (a) hybrid dataset, (b) custom randomizer dataset, and (c) SolidWorks Visualize with Roboflow dataset.
Figure 6. Confusion matrix plots of trained models on the test set: (a) hybrid dataset, (b) custom randomizer dataset, and (c) SolidWorks Visualize with Roboflow dataset.
Buildings 16 03407 g006
Figure 7. Performance curves of the model trained on the hybrid dataset and evaluated on the test set: (a) F1–confidence curve, (b) recall–confidence curve, (c) precision-confidence curve, and (d) precision–recall curve.
Figure 7. Performance curves of the model trained on the hybrid dataset and evaluated on the test set: (a) F1–confidence curve, (b) recall–confidence curve, (c) precision-confidence curve, and (d) precision–recall curve.
Buildings 16 03407 g007
Figure 8. Synchronised aerial and frontal camera views of the benchtop ISC assembly experiment, shown before (top row) and after (bottom row) YOLOv8 inference.
Figure 8. Synchronised aerial and frontal camera views of the benchtop ISC assembly experiment, shown before (top row) and after (bottom row) YOLOv8 inference.
Buildings 16 03407 g008
Figure 9. Real-Time object detection tracking performance on ISC objects, connection plates, and human workers.
Figure 9. Real-Time object detection tracking performance on ISC objects, connection plates, and human workers.
Buildings 16 03407 g009
Figure 10. Glaring in the frontal view impacts detection performance.
Figure 10. Glaring in the frontal view impacts detection performance.
Buildings 16 03407 g010
Table 1. Overview of popular datasets for computer vision tasks.
Table 1. Overview of popular datasets for computer vision tasks.
DatasetPurposeYearClassesImagesAnnotationsDomain
ImageNetObject recognition/classification200921,84114,197,122Bounding boxesGeneral
COCOObject detection/segmentation201480328,000Boxes, masksGeneral
Pascal VOCObject detection/classification20052011,530Bounding boxesGeneral
Open ImagesObject detection/classification201619,9589,011,219Bounding boxesGeneral
KITTIAutonomous driving2012974813D/2D boxesUrban driving
CityscapesSemantic segmentation2016305000Segmentation masksUrban
Table 2. Summary of randomisation options used in SolidWorks Visualize.
Table 2. Summary of randomisation options used in SolidWorks Visualize.
ParameterSettings
BackgroundBlack background; empty outdoor parking; Swiss snow; steel building site; black/white background; industrial lot (night); inside glass building; empty indoor garage; boiler room
Model textureCast carbon steel (red); metal rust 1; metal rust 2
Rotation0°–360°
LightingDay; night
Table 3. Normalised human-time accounting for the Unity-based synthetic pipeline at N = 10,000 images. Compute wall-clock (GPU/CPU rendering/export) is listed separately and is not counted as human labour.
Table 3. Normalised human-time accounting for the Unity-based synthetic pipeline at N = 10,000 images. Compute wall-clock (GPU/CPU rendering/export) is listed separately and is not counted as human labour.
StageHuman TimeCompute TimeNotes
Unity setup & assets6.0 hProject, import, and materials
Scene/physics authoring6.0 hColliders and dynamics
Sensor/exporters3.5 hRGB, depth, masks, and ground truth
Domain randomisation3.5 hPoses, lights, and textures
SolidWorks Visualize integration3.0 hMesh overlays and QA hooks
Render/export automation2.0 h12 hBatch scripts; GPU/CPU wall-clock
QA sampling (2% @ 6 s/image)0.33 hVisual checks only
Final end-to-end checks6.2 hSplits, metadata, and hashes
Manual baseline (60 s/image)166.7 hSingle annotator; 5-instance pilot 60 s (max 80 s)
Synthetic total (human)30.5 h12 hEffective 11.0 s/image human time
Note. These values apply only to the normalised N = 10,000 Unity-based pipeline example under the assumptions listed above. The estimated human-effort reduction is 81.7%, from 166.7 h for manual annotation to 30.5 h for the synthetic pipeline; compute wall-clock is 12 h and is reported separately.
Table 4. Summary of the datasets for detecting ISC components.
Table 4. Summary of the datasets for detecting ISC components.
DatasetSource of ImagesImages
Dataset 1 (custom randomizer)Unity custom randomizer indoor/outdoor scenes (C3)10,651
Dataset 2 (SolidWorks Visualize + Roboflow)SolidWorks Visualize images plus Roboflow human annotations (C1 + C5)5195
Dataset 3 (hybrid)Unity + SolidWorks Visualize + Roboflow + real images (C1 + C3 + C4 + C5) 15,928
Table 5. Distribution of images in ISC-Perception across sources.
Table 5. Distribution of images in ISC-Perception across sources.
SourceDataset 1Dataset 2Dataset 3Test Set
Unity (Custom randomizer)10,651010,6512748
SolidWorks Visualize 03551355148
Roboflow Human01644164484
Real Images0082207
Total10,6515195 15,928 3087
Table 6. Performance of trained models on the test set.
Table 6. Performance of trained models on the test set.
DatasetPrecisionRecall mAP@0.50 mAP@[0.50:0.95]
Overall10.820.520.660.56
20.400.320.390.32
30.850.670.760.66
ISC connection plate10.830.470.640.53
20.360.210.210.18
30.810.490.640.52
ISC member10.800.710.750.67
20.520.160.330.24
30.800.730.760.66
Human10.820.380.590.50
20.330.670.610.54
30.920.780.870.80
Table 7. Near-size-matched comparison. Dataset 1 uses a sampled subset of 5194 images; Dataset 2 uses its full set of 5195 images; Dataset 3 uses a sampled subset of 5195 images. The test set remains fixed at 3087 images. * sampled; full.
Table 7. Near-size-matched comparison. Dataset 1 uses a sampled subset of 5194 images; Dataset 2 uses its full set of 5195 images; Dataset 3 uses a sampled subset of 5195 images. The test set remains fixed at 3087 images. * sampled; full.
Train Set (N) mAP@0.50 mAP@[0.50:0.95] Prec.Rec.
Dataset-1 * (5194) 0.5460.4300.7760.505
Dataset-2 (5195)0.2490.2060.6490.146
Dataset-3 * (hybrid, 5195)0.6750.5490.8300.625
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Rahman, M.; Adebayo, S.; Acevedo-Mejia, D.A.; Hester, D.; McPolin, D.; Rafferty, K.; Laefer, D.F. ISC-Perception: A Hybrid Vision Dataset for Robotic Assembly with Novel Intermeshed Steel Connections. Buildings 2026, 16, 3407. https://doi.org/10.3390/buildings16173407

AMA Style

Rahman M, Adebayo S, Acevedo-Mejia DA, Hester D, McPolin D, Rafferty K, Laefer DF. ISC-Perception: A Hybrid Vision Dataset for Robotic Assembly with Novel Intermeshed Steel Connections. Buildings. 2026; 16(17):3407. https://doi.org/10.3390/buildings16173407

Chicago/Turabian Style

Rahman, Miftahur, Samuel Adebayo, Dorian A. Acevedo-Mejia, David Hester, Daniel McPolin, Karen Rafferty, and Debra F. Laefer. 2026. "ISC-Perception: A Hybrid Vision Dataset for Robotic Assembly with Novel Intermeshed Steel Connections" Buildings 16, no. 17: 3407. https://doi.org/10.3390/buildings16173407

APA Style

Rahman, M., Adebayo, S., Acevedo-Mejia, D. A., Hester, D., McPolin, D., Rafferty, K., & Laefer, D. F. (2026). ISC-Perception: A Hybrid Vision Dataset for Robotic Assembly with Novel Intermeshed Steel Connections. Buildings, 16(17), 3407. https://doi.org/10.3390/buildings16173407

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop