1.1. Background and Motivation
Smart and sustainable construction increasingly depends on the ability to automate labour-intensive, safety-critical, and time-sensitive site operations. Robotic manipulators have transformed factory-based manufacturing, but their deployment in construction remains limited because building sites are unstructured, weather-exposed, spatially constrained, and governed by strict safety requirements. Steel frame erection is a strong candidate for automation: cranes often dominate the critical path, structural members are heavy and difficult to manoeuvre, and bolting operations require skilled workers to operate at height [
1]. A robotic system capable of recognising, grasping, aligning, and assembling steel members could reduce crane time, improve worker safety, and mitigate skilled labour shortages.
In conventional steel construction, beams and columns are typically fabricated off-site and assembled on-site because complete frame assemblies are difficult to transport. This process involves several key steps: (1) identifying and lifting each structural element from the storage area, (2) transporting it to the installation location; (3) aligning it with the existing structure, and (4) fastening it to the structural frame using bolts or welds [
1]. Although both methods are common for connecting steel members, bolts are generally preferred on-site due to their ease of installation, faster connection times, better quality control, and reduced inspection requirements. However, the extensive use of bolts in structural steel connections introduces additional challenges for the deployment of robots in the field.
The recently proposed Intermeshed Steel Connection (ISC) system eliminates most of the bolts required by conventional moment or shear splices. The ISC can be manufactured using cutting-edge technologies such as high-density plasma cutting, water jet cutting, and laser cutting [
2]. ISC has two types of components: ISC member and ISC connection plates. The initial design of ISC had 3 connection plates on one side (
Figure 1a) but the newer version requires only one connection plate on each side with fewer bolts (
Figure 1b). Precision-cut male-female tabs guide members into alignment so that only a handful of set-bolts are needed to secure the joint [
2,
3,
4]. By trimming cycle times and tolerating direct reuse, ISC reduces material waste and greenhouse gas emissions while preserving structural capacity. These benefits align with industry trends towards design-for-manufacture-and-assembly (DfMA) [
5] and circular construction. Recent modular-steel-building research has also investigated specialised isolation devices for seismic resilience, illustrating the broader structural engineering context in which modular steel systems are developing [
6]. Compared with conventional bolted connections, in which bolt heads, nuts, hole patterns, and plate boundaries provide relatively discrete and regularly arranged visual cues, ISC recognition depends more strongly on the fine geometry and spatial relationships of its precision-cut intermeshing tabs and connection plates. These repeated features can produce visually similar local contours and become progressively occluded as the members are aligned and intermeshed, causing the appearance of the same component to vary with viewpoint and assembly state. In addition, the reflective galvanised surfaces produce angle-dependent specular highlights that can reduce local edge contrast and obscure the geometric features required for detection. These characteristics motivate a data-generation strategy that spans varied viewpoints, assembly configurations, illumination conditions, and complementary synthetic, photorealistic, and real imagery.
Robotic ISC assembly therefore depends on reliable near-field perception. A robot must localise ISC connection plates and member ends in its immediate workcell, provide detections suitable for downstream grasp planning and mating, and monitor human presence for safety. Generic construction datasets, such as MOCS [
7] and SODA [
8], primarily represent workers, equipment, vehicles, and general site objects. They do not contain ISC geometry, nor do they capture the specific indoor and outdoor assembly cell contexts required for robotic steel assembly. Similarly, general-purpose computer vision datasets, such as COCO [
9], lack the object classes, viewpoints, occlusions, and material appearances associated with ISC components.
The economic motivation for this work is substantial. The structural-steel market is substantial and is projected to continue growing, driven by urban development and industrial expansion [
10]. Connection design can have a disproportionate influence on structural-steel frame cost because of its detailing, fabrication, and erection labour requirements [
11,
12]. Perception-enabled robotic assembly is therefore relevant as a technical challenge and as a route towards safer, faster, and more resource-efficient steel construction. A vision-enabled robotic ISC assembly workflow has therefore been proposed, offering the potential for safer construction sites and cost and schedule savings [
13]. In the proposed ISC assembly workflow, a robot must (i) localise ISC connection plates and member ends in its immediate workcell, (ii) provide reliable detections for grasp planning and mating, and (iii) monitor human proximity for safety. Conventional datasets, such as Common Objects in Context (COCO) [
9], MOCS [
7], and SODA [
8], contain neither ISC geometry nor the specific indoor/outdoor assembly cell context. ISC-Perception directly addresses this data gap for near-field perception, rather than general stockyard inventory, and is already integrated into a benchtop robotic assembly pipeline (
Section 5.5). This paper addresses the perception data gap by introducing ISC-Perception, a hybrid vision dataset for robotic assembly with novel Intermeshed Steel Connections. The dataset combines automatically annotated Unity scenes, photorealistic SOLIDWORKS Visualize 2023 renders, and curated real images of ISC components and humans. It is designed for near-field robotic assembly perception rather than general stockyard inventory or full-site monitoring. ISC-Perception supports object detection of ISC members, ISC connection plates, and humans, and is evaluated through both fixed test set experiments and a benchtop robotic assembly scenario. In this way, ISC-Perception provides a practical route to perception data generation for smart construction applications where the target object is novel, real imagery is scarce, and CAD assets are available.
1.2. Perception Data for Construction Robotics
Progress in deep learning has been strongly shaped by the availability of large, diverse, and accurately labelled image datasets. Construction robotics perception imposes additional demands: detectors must recognise partially occluded objects, operate under variable illumination, and generalise across projects that differ in geometry, material finish, background clutter, and weather conditions [
14,
15,
16,
17,
18,
19]. Image datasets therefore form a central foundation for training and evaluating computer vision models for object detection, classification, and segmentation in construction environments [
20].
Generic datasets such as COCO or ImageNet misrepresent site reality: they lack steel members, cranes, PPE, and the dense clutter typical of erection yards. Direct transfer can depress mean Average Precision (mAP) by up to 40% when models are tested on construction imagery [
21]. Building an in-domain corpus is equally fraught. Cameras are often barred by safety briefings, union rules, or privacy regulations; outdoor shoots hinge on weather windows; and pixel-accurate annotation of high-resolution frames can consume weeks of person-hours [
22]. The hurdle is steeper still for bespoke components such as the ISC, for which no archival photographs yet exist and whose galvanised surfaces frustrate automated labelling.
Data, not algorithms, have thus become the principal bottleneck. An effective remedy must supply (i) scale for deep networks, (ii) fidelity to capture ISC’s subtle tab geometry, and (iii) diversity in backgrounds, lighting, and occlusions, while curbing manual annotation cost.
Section 1.3 surveys how synthetic and photorealistic imagery can satisfy those requirements and where current approaches fall short.
1.3. Synthetic and Photorealistic Data: Benefits and Pitfalls
A practical solution to address the challenges of limited access and varying construction site conditions is the creation of annotated synthetic image datasets to supplement real ones [
21]. These synthetic datasets can be generated using computer graphics techniques, 3D modelling software or game engines, enabling the simulation of diverse construction environments with different objects and backgrounds.
Computer vision models trained solely on synthetic images often perform worse than those trained on real images. For example, grocery item detection models trained on 400,000 synthetic images performed less effectively than models trained with only 760 real images [
23]. Yet, combining just 760 real images with the synthetic images produced superior results compared to both models. Moreover, randomisation techniques (such as lighting conditions, weather conditions, time of day, textures, and camera perspective) are used to generate synthetic images, reduce the sim-to-real gap, and improve dataset diversity [
24,
25]. Therefore, a hybrid dataset that integrates real and synthetic images could be an effective approach for training computer vision models for construction applications [
26]. However, obtaining sufficient real images for many construction scenarios or custom objects, such as the ISC, remains difficult. In such cases, computer-aided design tools can generate and render photorealistic models of custom objects in various settings, reducing the reliance on real images.
Beyond domain randomisation, sim-to-real transfer can also be addressed through domain adaptation. Image-level approaches translate simulated and real observations towards a shared or canonical appearance, whereas feature-level approaches encourage the detector to learn representations that are invariant across source and target domains. More recent teacher–student and self-training methods additionally exploit unlabelled target-domain images through pseudo-labels and consistency objectives [
27,
28,
29,
30]. These approaches provide complementary mechanisms for reducing the visual discrepancy between simulated and real observations. Our approach addresses the sim-to-real gap through data design rather than a dedicated domain-adaptation objective, combining controlled randomisation with hybrid training across Unity-generated, SolidWorks Visualize, and limited real imagery.
Additionally, ISC plates pose an additional hurdle: their galvanised coating creates specular highlights that shift with sun angle, and the laser-cut tab patterns differ by millimetres. Capturing these cues demands high-dynamic-range rendering plus fine surface normal maps that are costly to generate at scale. Conversely, photographing ISC plates on active sites remains impractical, because the system is not yet widely deployed. Hence a hybrid strategy [(i) auto-generates large volumes of domain-randomised synthetic frames, (ii) injects photorealistic ray-traced scenes for material fidelity, and (iii) enhances the mix with a small, curated set of real photographs] offers the best trade-off between cost and realism.
Research gap. To date, no public dataset combines these three modalities for steel-connection detection; existing construction corpora (MOCS [
7], SODA [
8]) neither model bespoke joints nor provide labels. Bridging this gap is therefore prerequisite to closing the perception loop for robotic ISC assembly.
1.4. Research Contribution
Existing vision datasets in construction focus on equipment or personnel safety. None address robotic assembly of structural steel components such as beams, columns, or ISC plates. We fill this void by devising and releasing ISC-Perception, a task-specific, hybrid corpus for object detection in robotic steel erection.
In summary, the main contributions of this paper are
A methodology for creating a hybrid dataset for ISC components using real, photorealistic, and synthetic images to tackle the scarcity of real images tailored for robotic assembly tasks, together with a normalised 10,000-image Unity-based pipeline comparison that reduces estimated human effort from 166.7 h for manual annotation to 30.5 h (81.7%).
The analysis of training performance of computer vision algorithms for different types of images and validation of the trained computer vision model in small-scale setup. We further demonstrate that detectors trained on ISC-Perception achieve mAP@0.50 of 0.943 and mAP@[0.50:0.95] of 0.823 on a 1200-frame multi-view bench test that mimics a robotic assembly cell.
To contextualise these contributions,
Section 2 reviews the current state of the art in real and synthetic computer vision datasets. Subsequently,
Section 3 discusses the procedural approach for generating the hybrid dataset.
Section 4 provides insight into the ISC-Perception dataset.
Section 5 reviews the outcomes of the training and testing phases, followed by a discussion of the results and findings. Finally,
Section 6 of the paper summarises the research results and their significant impacts on the construction industry. Our focus is on a task-specific dataset and reproducible data-generation pipeline rather than proposing a new detection architecture. YOLOv8 models are used strictly as reproducible baselines to isolate the effect of dataset composition; we do not claim an algorithmic contribution.