3.1. Methodological Framework for Rapid Reconfiguration
The proposed approach addresses one of the main barriers preventing the adoption of flexible robotic assembly systems in SMEs: the complexity of reconfiguring both software and hardware when production requirements change. Rather than developing a task-specific solution, the proposed framework combines an open software architecture, AI-based perception, and modular hardware components to simplify system adaptation while limiting engineering effort. As illustrated in
Figure 1, the coordination module orchestrates the interaction between perception, manipulation, gripping, and feeding modules, each encapsulating a specific functionality. This modular decomposition allows individual subsystems to be replaced or extended without affecting the overall control logic.
The framework further promotes rapid deployment through an image-segmentation pipeline based entirely on open-source tools, enabling non-specialist users to create custom datasets, train deep-learning models, and deploy them within the robotic cell using affordable computing resources. At the hardware level, the adoption of commercially available devices combined with rapidly manufacturable 3D-printed tooling minimizes the physical modifications required when introducing new products. Although expert intervention is still required for complex manipulation strategies and system optimization, the proposed methodology significantly reduces the engineering effort needed for reconfiguration. The following sections describe the implementation of this framework through a representative case study involving manual pre-assembly operations for early-stage electrical cabinet wiring.
3.2. Case Study
This section summarizes the case study that inspired this work. The assembly process for electrical switchboards and cabinets, particularly in the case of complex units, remains a largely manual task today due to the need for careful routing of cables and connectors, as shown in
Figure 2.
There are, however, certain preliminary tasks that offer great potential for automation and mechanisation. Often, the cabinet assembly cycles begin with the creation of sub-assemblies and pre-assembled mechanical components that will form the framework and structure upon which the electrical circuit will be built. Although these are mechanical assemblies with a small number of components, they exhibit high variability as they are closely dependent on the final product. Moreover, they are characterised by small size components, as they don’t need to occupy too much space in the cabinets. Traditional automation solutions aim to address each sub-assembly case individually, involving development costs that are difficult for SMEs to sustain. The primary objective of this case study is therefore to develop a robotic cell capable of ensuring high flexibility and reconfigurability to cope with the low-volume high-mix nature, and small-size components. As specific case study, the proposed system is designed to produce a pre-assembled unit consisting of a central primary screw upon which various elements are sequentially stacked and inserted in a predefined order. The manipulated workpieces are categorized into five distinct groups, characterized by varying physical and geometric properties (see
Figure 3):
Figure 3.
Technical illustration of the pre-assembly considered as case study.
Figure 3.
Technical illustration of the pre-assembly considered as case study.
Standard industrial supply conditions dictate that components number 1 through number 4 are delivered in bulk within plastic or cardboard containers, resulting in completely random orientations. While bronze spacers (number 5) are typically supplied resting on their longitudinal axis.
3.3. Hardware Configuration
As shown in
Figure 4, the pivot point of the entire cell is a collaborative robotic platform: the Doosan M0609 (Doosan Robotics Inc., Seongnam-si, Republic of Korea,
https://www.doosanrobotics.com/en/product-solutions/product/m-series/m0609/ accessed on 1 July 2026). It is a 6-axis collaborative robotic platform (cobot) designed for high-speed, high-precision tasks in small workspaces. It features a 6 kg payload capacity and a 900 mm reach. Maximum joints’ speeds are reported in
Table 1.
There are two main reasons that support the choice of this machine. First of all, this robotic platform is equipped with 6 high-performance torque sensors. They allow for hybrid force-position control integration which comes convenient when dealing with assembly operations, where small interferences between components may be required. Additionally, the high-tech torque sensors allow 5 out of 6 axes to rotate a full , and let the robot reach a repeatability of ± mm. Second of all, this robotic platform is ROS2 supported by the manufacturer. The choice of an off-the shelf ROS2 compatible robot is dictated by the principle of open architecture suggested by ROS-I that characterises the aim of this work, facilitating a modular and reconfigurable environment that contrasts with traditional, rigid assembly systems.
Two additional devices were integrated with the Doosan to provide the ability of sensing the surrounding environment and to interact with it. Firstly, a ROS2 compatible tooling was attached to end of the arm. Specifically, the choice was the OnRobot 2FG7 (OnRobot A/S, Odense, Denmark,
https://onrobot.com/en/products/2fg7-finger-gripper accessed on 1 July 2026): a compact, high-performance all-electric parallel gripper designed for tight production environments. This unit eliminates the need for compressed air while delivering precise and programmable control over gripper kinematics and dynamics. The system offers an adjustable gripping force ranging from 20 N to 140 N with a ±5 N tolerance, a total mechanical stroke of 38 mm, and an operational payload capacity of up to 11 kg for form-fit grasping topologies or 7 kg for force-fit applications. Sensor integration within the onboard controller supports closed-loop feedback, enabling real-time grip verification and lost-grip detection with a precision repeatability of ±
mm. Even though this gripper has no OnRobot official ROS2 drivers, community developed drivers to bridge this gap ensuring a flawless integration into the cell, and creating a solid base for future ROS-I applications.
The second device is an industrial CCD camera: a Basler ace acA2040-25gc (Basler AG, Ahrensburg, Germany,
https://www.baslerweb.com/en/shop/aca2040-25gm/ accessed on 1 July 2026) coupled with a Kowa LM6JC (Kowa Optimed Deutschland GmbH, Düsseldorf, Germany,
https://www.kowavision.com/products/lm6jc?srsltid=AfmBOophUdv9DDRq0KCQe-esNttadpJMjIqUF7dHAPivcYw5X0-xldzK accessed on 1 July 2026) high-resolution machine vision lens. This camera is built around a CMV4000 CMOS 1 inch sensor, that yields to a
resolution and (
Megapixels). To facilitate high-speed acquisition, the sensor employs a progressive-scan global shutter mechanism with a pixel size of
, achieving a native throughput of 25 fps at full resolution. Data transmission and power delivery are managed concurrently via Gigabit Ethernet utilizing the GigE Vision standard and Power over Ethernet (PoE IEEE 802.3af), housed within a compact
29 mm enclosure weighing 90 g. This camera ensures high performances in a small footprint. Optically matched via a standard C-mount interface, the Kowa LM6JC lens provides a fixed wide-angle focal length of 6 mm. The lens features completely manual iris and focus rings, operating across an adjustable aperture range of
to
, with a minimum object distance of 100 mm. Pylon camera driver for ROS2 are officially provided as an open-source project by Basler, creating de facto a ROS-I ready device. Moreover, this camera-lens configuration provides a short-exposure vision pipeline suitable for integration into the robotic cell.
The last piece of equipment included in the cell is a feeding system. We implemented the ARS Automation FlexiBowl 500 C (ARS Automation S.r.l., Civitella in Val di Chiana, Arezzo, Italy,
https://www.flexibowl.it/flexibowl#fb-500 accessed on 1 July 2026), a rotating parts-feeding platform designed to handle heterogeneous components with complex geometries and varying material compositions. Operating on a combined mechanism of bidirectional rotary actuation and impulsive flip-vibration, the system prevents part entanglement and optimizes distribution across a total surface load capacity of up to 7 kg. The device is geometrically optimized for components measuring between 5 mm and 50 mm with individual masses under 100 g. Physically, the feeder features an overall diameter of ⌀640 mm, a designated picking height of 270 mm, and a total system mass of 42 kg. To maximize compatibility with computer vision systems, it incorporates an integrated backlit illumination window. It requires both electricity (for the rotation) and pneumatic (for the flip) supplies. Because the manufacturer does not supply an official ROS2 support, a custom, open-source ROS2 hardware driver was developed in-house to integrate the feeder into the cell. Following in the development the ROS-Industrial principles of modularity, and standardized communication, the drivers establish a socket communication interface over standard TCP/IP-UDP protocols, abstracting the device’s native parametrized routines (such as flip, shake, and precise angular disc indexes) into standard ROS2 services and actions. We decided to mount these devices on an Isel USA Inc. (Hicksville, NY, USA) industrial workbench. This bench features all aluminium construction (except for the hardware and mounting feet) and has a T-slot tabletop. The T-slots allow for quick and easy fixturing and adjusting, ensuring an easy reconfiguration of the robotic cell. As shown in the system layout, the M0609 is placed at the back of the table to ensure adequate clearance for the safe loading and unloading of the components. A detailed representation of the cell network is shown in
Figure 5.
Each device has its own controller, which is connected via Ethernet to a D-Link DGS-1100-08P Gigabit Network switch (D-Link (Europe) Ltd., London, United Kingdom). The only exception is the camera, which is directly connected to an external workstation, where also the ROS2 controller runs. We chose a workstation equipped with an Intel Xeon W-2125 CPU, 32 GB of high-speed DDR4 RAM and an NVIDIA GeForce RTX 2080 Ti graphics processing unit (GPU). Built on the Turing (TU102) microarchitecture, this graphics processor features 4352 CUDA cores dedicated to matrix mathematics, and 11 GB of GDDR6 VRAM. This high-performance computational hardware tier acts as the central hub for the custom ROS2 node topology, as well as for the deep-learning-based part localization pipeline. Once finalized, the mechanical and network systems used for our experiments were prototyped, as shown later in the experimental setup (see
Section 4).
3.3.1. Chain of Tolerances
In the design of automated assembly cells, the definition of a rigorous tolerance chain is fundamental to ensure process reliability. The primary objective of this analysis is to determine the feasibility of the assembly task, by evaluating the cumulative effect of geometric variance and mechanical repeatability. This analysis is particularly critical in the context of the suggested case study, where the assembly task requires the precise insertion of washers and spacers through the body of the screw. To determine whether the proposed robotic system can successfully execute these tasks without mechanical interference, a mathematical model can be developed. The analysis begins with the definition of most critical dimension for each part, i.e., the internal diameter, if it’s a receiving element, or the external diameter, if it’s an insertion element. We therefore define:
Internal diameter of the spacer ;
Internal diameter of the washers ;
External diameter of the screw thread .
In a robotic assembly context, the available radial clearance must accommodate the relative radial displacement between the centres of the mating components to avoid mechanical jamming during placement. The relevant clearance
is therefore defined as half the difference between the internal diameter of the receiving component and the external diameter of the inserted component:
Under a conservative worst-case assumption, assembly feasibility requires:
The positioning error of the screw
is conservatively expressed as the sum of the robot’s movement repeatability
, which refers to the precision with which a robot can return to a specific programmed coordinate in the workspace, and the mechanical variance in how the screw is seated within the gripper
.
Similarly, the positioning error of the washer
is a function of
and
By substituting Equations (
3) and (
4) into Equation (
2), we derive the second feasibility criterion, which allows to solve for the maximum allowable positioning error of the receiving part, i.e.,
.
It is important to note that
appears twice in Equation (
5) because the robot’s precision affects both the pick-up pose and the final placement pose.
Figure 6 shows graphically the two feasibility criteria.
Using the nominal dimensions reported in
Table 2, the available radial clearance is
mm. The resulting maximum allowable error for the washer
is
mm. This requirement guided the design of the custom self-centring fingertips described in
Section 3.3.2.
3.3.2. Gripper Design Proposal
Starting from the tolerance constraints derived in
Section 3.3.1, the design of the gripper and in particular of its fingertips was directly driven. The allowable positioning error for the receiving component is limited to
mm, which imposes a very tight request on the consistency with which components are held and centred by the gripper. In this context, the fingertip design was conceived as a geometrically self-aligning system, thereby ensuring compliance with the feasibility criteria defined in Equations (
2) and (
5). As shown in
Figure 7a, the proposed solution adopts the gripper mentioned in
Section 3.3 with custom fingertips specifically designed to closely match the geometries and sizes of the objects to be gripped. The inner curvature increases the contact area and promotes repeatable positioning of screws within the gripper, contributing to lowering
. parts within the gripper, contributing to lowering
and
. In addition to the curved gripping surfaces, a central cavity is introduced between the fingers to address the most critical component in the tolerance chain, namely the washer, whose positioning error directly affects the satisfaction of the feasibility condition. This cavity passively guides ring-shaped components toward the centre of the gripper during the closing phase, effectively implementing a mechanical self-centring mechanism. Furthermore, the external geometry of the fingertips was designed with smoothed and rounded edges, allowing the gripper to interact safely with partially cluttered environments by pushing away interfering elements. The overall structure is compact, reducing lever arms and potential compliance effects that could degrade repeatability during manipulation.
As shown in
Figure 7b, the fingertips represent the only hardware element that must be adapted or replaced to effectively cope with a low-volume, high-mix production scenario, enabling rapid reconfiguration while preserving system accuracy. From a design philosophy perspective, the chosen fingertip solution reflects a preference for geometric coupling and passive compliance rather than relying on active gripping technologies such as suction or electromagnetic systems, which were considered less suitable due to limitations in precision, material compatibility (e.g., non-magnetic bronze spacers), or reliability when dealing with small parts.
To validate whether the designed fingertips effectively reduce gripping-induced variability and satisfy the tolerance requirements, an experimental campaign was carried out focusing on the repeatability and robustness of the grasp. Components with external diameters ranging from 12 mm, corresponding to the screw, to 25 mm, corresponding to the brass spacer, were used. Although the current geometry provides an available opening of approximately 35 mm, grasping reliability beyond the tested dimensional range was not assessed. As shown in
Figure 8, 30 grasping trials were performed for each component type, with the parts deliberately positioned in known and controlled configurations to isolate the performance of the gripper from perception errors. Each trial included not only the closure of the gripper on the component but also the lifting and handling phases, thus evaluating the stability of the grasp throughout the manipulation cycle. The results showed that no grasp failures or part drops occurred across all tested components, indicating that the fingertip geometry provides sufficient constraint and repeatability to maintain consistent positioning. This outcome supports the assumption that the adopted design can effectively limit the additional uncertainty introduced by the gripping process, allowing the system to remain within the error bounds imposed by the tolerance chain analysis. In this sense, the gripper design can be interpreted not merely as a handling device, but as an integral component of the tolerance management strategy, where mechanical design choices are explicitly used to control and minimize uncertainty contributions, ultimately enabling the successful execution of the automated assembly task under the tight geometric constraints identified in the previous section.
3.4. SW Configuration
In this work, a complete automated assembly system was conceived and implemented. The overall system integrates several functional modules, including component recognition, manipulation, robot motion control, and system coordination modules. In the following subsections, the different system modules are explained in more detail. Specific emphasis is given to the vision module and the vision pipeline in general.
3.4.1. Modular Approach
The software architecture adopted in this work follows a modular approach, with the objective of ensuring flexibility, scalability, and ease of reconfiguration in a low-volume, high-mix production context. In contrast to monolithic control systems, where all functionalities are embedded in a single control layer, the proposed solution decomposes the overall application into independent yet interconnected modules. Coordination between the software modules is illustrated in
Figure 9. Overall, the architecture follows the principles of flexibility and modularity consistent with the ROS-Industrial approach, and it is divided into five software modules:
Coordination module;
Robot motion module;
Feeding module;
Gripping module;
Vision module.
Figure 9.
System architecture of the modular robotic cell control framework.
Figure 9.
System architecture of the modular robotic cell control framework.
The
coordination module acts as the central orchestrator of the assembly process, managing the execution flow and coordinating interactions among the perception, feeding, manipulation, and control subsystems. Implemented as a ROS2 node, the module initializes the process parameters, including robot poses, assembly configurations, and the bill of materials (BOM), and supervises the execution of the complete assembly cycle. The coordination logic is based on a finite state machine (FSM), whose states represent the main phases of the process: component distribution, robot positioning, object detection, picking and placement, part repositioning, and task completion. The FSM ensures deterministic execution by activating the appropriate subsystem according to the current process state and the outcome of previous operations. During execution, the module acquires component detections from the vision system, selects the target part according to the BOM, and coordinates the robot and gripper actions required for grasping and placement. If no suitable component is detected, the feeding system is activated to redistribute the parts and a new detection cycle is initiated. The module also manages execution timing, cycle counting, and process monitoring until all components have been assembled. Communication with the other software modules is implemented through ROS2 services and actions, enabling modular, asynchronous, and scalable integration. The logical sequence of the FSM states is illustrated in
Figure 10.
The robot motion module is built upon the ROS2 library provided by Doosan Robotics and is responsible for executing all manipulator movements required during the assembly process. The module supports both joint-space and Cartesian-space motion commands, enabling the robot to perform operations such as homing, object approach, placement, and trajectory retreat. In addition, compliant motion capabilities are available during contact-sensitive tasks, allowing safer and more reliable interactions with the assembly fixture and components. By abstracting low-level robot control, the module provides a standardized interface that enables the coordination module to execute complex manipulation sequences through high-level commands.
The feeding module provides the interface between the robotic cell and the Flexibowl flexible feeding system. Its primary role is to control part distribution and presentation, ensuring that components are continuously rearranged to facilitate reliable detection and grasping. The module exposes feeding operations such as rotation, shaking, and flipping through dedicated ROS2 actions, allowing these functions to be executed asynchronously and coordinated with the other subsystems. By abstracting the device-specific communication layer, the module enables seamless integration of the Flexibowl within the ROS2 architecture and provides a standardized mechanism for triggering feeding actions during the assembly process.
The gripping module provides the interface between the robotic cell and the adaptive gripper used for component manipulation. Based on an extended version of the OnRobot ROS2 driver, the module enables high-level grasping operations while ensuring seamless integration with the overall software architecture. It supports both internal and external gripping strategies and provides real-time feedback on gripper status and opening width. By abstracting device-specific communication and control, the module allows the coordination system to execute grasping tasks through a standardized set of services.
The vision module is responsible for detecting, classifying, and localizing the components required for the assembly process. Image acquisition is performed through an industrial Basler camera, while perception is based on a customized You Only Look Once (YOLO) segmentation model integrated within a ROS2 processing pipeline. The module extracts object attributes relevant to manipulation, including class, position, orientation, and confidence level, and makes this information available to the coordination and manipulation modules. Its modular architecture enables reliable operation under varying conditions and supports the autonomous execution of the assembly workflow. The following sections describe the vision pipeline and the training procedure of the YOLO model in greater detail.
3.4.2. Dataset Creation
The creation of the dataset adopted for training the vision module was carried out using Computer Vision Annotation Tool (CVAT), a widely used platform for image annotation that can be deployed in a self-hosted configuration at no licensing cost. This aspect was particularly advantageous, as it enabled full control over data management and annotation workflows while maintaining compatibility with modern deep learning pipelines. The dataset was composed of images acquired directly from the industrial setup, using the same camera employed in the final system, ensuring consistency between training and deployment conditions. In total, the dataset included 250 labelled samples, distributed across training, validation, and test sets, following a 70-20-10 split.
The class distribution was defined to reflect the components involved in the assembly process, specifically including multiple washer types and orientations, cylindrical spacers, and screws. Each class corresponds to a distinct geometric and functional category, and particular care was taken to ensure that all classes were adequately represented despite inherent differences in visual complexity and occurrence frequency. With reference to
Figure 11, the annotation process was performed manually in CVAT using instance segmentation masks, enabling the extraction of both shape and spatial features required for downstream tasks. Also partially captured objects were labelled as full, to increase the chances of a positive recognition in case of partial cluttering. Special attention was given to classes such as washers, which exhibit low contrast and limited geometric features, making them more challenging to annotate consistently. To improve model robustness and generalization, images were collected under varying illumination conditions, including low, medium, and high lighting scenarios, as well as different component arrangements. From a workflow perspective, CVAT allowed efficient organization of annotation tasks, collaborative labelling if required, and direct inspection of annotation quality. A key advantage of the selected tool is its ability to export annotated datasets directly in formats compatible with YOLO-based frameworks, including segmentation-compatible annotations. Overall, even though well structured, the dataset creation process is the other application-specific element.
3.4.3. Image Segmentation
The image segmentation module is based on the YOLO framework, a state-of-the-art unified architecture that reframes object detection as a single regression problem. Unlike traditional multi-stage pipelines that employ sliding windows or region proposal methods, YOLO evaluates the entire image in a single pass through a convolutional neural network. The system divides the input image into a grid, where each cell is responsible for predicting bounding boxes, confidence scores, and class probabilities simultaneously. This global reasoning approach ensures high speed and significantly fewer background errors compared to architectures like Fast R-CNN. This allows the network to extract detailed shape information, which is essential for downstream tasks such as orientation estimation and grasp point computation [
30]. YOLO was selected for this application primarily due to its real-time processing capabilities and its ability to learn generalizable representations of objects, which is critical for industrial environments with fluctuating lighting. While standard object detection provides bounding boxes, image segmentation is implemented to provide pixel-level masks, identifying the precise contours and centroids of components on the vibratory feeder. This pixel-level precision is fundamental to navigating the strict sub-millimetric tolerance chain required for automated insertion. The trained YOLO model is integrated into the ROS2-based vision module through a dedicated perception node, acting as an action server that returns detection results on request. When triggered by the coordination module, the node receives images from the Basler camera driver, converts them to an OpenCV-compatible format, and performs inference. The resulting segmentation masks are post-processed to extract geometric features, including centroids, bounding polygons, orientation angles, and grasp points according to the detected class. Cylindrical components are analysed using minimum-area bounding rectangles to estimate orientation, while elongated parts such as screws are processed to identify grasp points along their principal axis. The extracted information is converted into metric coordinates in the camera reference frame and returned as structured output, including class identifiers, confidence scores, and pose representations. This enables seamless integration with the manipulation pipeline, while the ROS2 action interface supports asynchronous, non-blocking execution to maintain system responsiveness.