Next Article in Journal
Dynamic Time Warping for System-Level Fault Detection in IoT Devices: An Episode- and Layer-Based, Label-Free Approach
Next Article in Special Issue
MVO: A Magneto-Visual Odometry System for Indoor Positioning
Previous Article in Journal
Sensor-Driven Short-Term Forecasting on the Metropolitan LA Traffic Dataset: A Comparative Study for Multi-Step Prediction
Previous Article in Special Issue
Vision-Based Person-Following Algorithm for Assistive Elderly-Care Quadruped Robots
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Integrating Visual Perception and Control Strategies in Custom Omnidirectional Mobile Robots

by
Radu-Laurențiu Roșca
*,
Andrei-Iulian Iancu
,
Adrian Burlacu
and
Cătălin Dosoftei
Faculty of Automatic Control and Computer Engineering “Gheorghe Asachi” Technical University of Iasi, 700050 Iasi, Romania
*
Author to whom correspondence should be addressed.
Sensors 2026, 26(12), 3918; https://doi.org/10.3390/s26123918
Submission received: 23 April 2026 / Revised: 17 June 2026 / Accepted: 17 June 2026 / Published: 20 June 2026
(This article belongs to the Special Issue Intelligent Sensing for Robotic Control and Visual Perception)

Abstract

Autonomous mobile robots are used in optimizing warehouse logistics, yet achieving precise positioning during docking maneuvers and autonomous planning remains a technical challenge. This study presents a custom vision-based control system designed for an autonomous omnidirectional wheeled robot. The proposed methodology acquires visual feedback using a stereo camera integrated within the Robot Operating System framework. Two visual feedback control laws are formulated and rigorously evaluated: a Classic Position-Based Visual Servoing algorithm, which minimizes pose error using a quaternion-based approach, and a second solution that utilizes Dual Lie Algebra to compute the 3D visual sensor’s velocities, ensuring convergence towards the desired point-feature configuration. Experimental validation reveals that while both methods achieve docking, the dual pose-free approach enables more robust, effortless movement of the robot platform than Classic Position-Based Visual Servoing. Consequently, these findings indicate that integrating depth-based feature recovery with advanced algebraic strategies offers a stable control strategy for automated industrial scenarios.

1. Introduction

Intelligent robotic systems have emerged as indispensable components in navigating dynamically complex environments across diverse domains, including manufacturing, healthcare, and everyday applications. The rapid advancement of mobile robotics technology has been largely driven by the growing demands of the intelligent manufacturing and logistics sectors. In particular, autonomous mobile robots have become integral to modern smart factory operations, where they must execute a broad spectrum of locomotion tasks with precision and reliability. The inherently complex and variable nature of such operational environments requires that mobile robotic platforms exhibit a high degree of flexibility, adaptability, and operational safety. Conventional differential drive mobile robots, however, present significant limitations in this regard due to their non-holonomic kinematic constraints, which fundamentally restrict their maneuverability and spatial adaptability. To address these limitations and enable smooth, flexible material transport within confined industrial environments, an Omnidirectional Mobile Robot (OMR) incorporating four Mecanum wheels was developed. The proposed system integrates a multi-sensor configuration to enhance perceptual awareness and closed-loop control capabilities. Mecanum wheel-based drive mechanisms have gained considerable traction in the development of omnidirectional mobile platforms, which are increasingly being deployed in intelligent warehouse and automated logistics operations [1]. The concept of movements without limits, implemented through these special wheels, offers important advantages over non-holonomic mobile robots, such as Ackerman steering and differential drives, for moving in logistics areas. These environments are collaborative spaces among operators, shelves, conveyors, and forklifts. In these conditions, to improve operational efficiency, the introduction of additional autonomous mobile robots (AMRs) must be accompanied by greater maneuverability in this crowded space. Omnidirectional maneuverability is achieved by precisely controlling each wheel’s speed and direction independently.
Visual servoing constitutes an important pillar of modern robotic control theory. In recent years, the integration of advanced stereo vision systems has offered multiple advantages, including real-time depth estimation. The deployment of advanced, high-frame-rate stereo cameras has catalyzed highly robust implementations across Image-Based Visual Servoing (IBVS), Position-Based Visual Servoing (PBVS), and Hybrid Visual Servoing (HVS) architectures.
IBVS operates on the principle of computing the control error directly within the two-dimensional image plane of the optical sensor. IBVS maps the differential changes in the two-dimensional image features—such as extracted keypoints [2], lines [3] or image moments [4]—directly to the required joint velocities of the robotic platform [5]. The primary advantage of IBVS lies in its robustness to calibration errors and system noise. The controller continues to generate velocity commands until the features mathematically align in the image matrix [6]. However, stereo IBVS is constrained by mathematical and physical disadvantages. The transformation mapping between the two-dimensional image space and the three-dimensional task space is highly non-linear. This non-linearity can frequently lead to control singularities in which the interaction matrix loses full rank, resulting in unbounded velocity commands [7].
PBVS operates by reconstructing the explicit three-dimensional pose of the target object relative to the camera frame [8]. The feedback loop relies heavily on pose estimation algorithms to extract the precise three-dimensional translation vector and rotation matrix, which together represent the target’s pose within the Special Euclidean group S E 3 . The control error function is defined as the geometric distance between the current pose estimate and the desired pose. Once the error is computed, the robotic kinematic model generates the necessary joint torques or velocities to drive the end-effector seamlessly to the target position. The dominant advantage of PBVS is the predictability of the camera and the robotic platform in three-dimensional space. Because the error is minimized directly in Cartesian coordinates, the robot naturally follows an optimal spatial trajectory [9]. Conversely, PBVS relies heavily on precise camera calibration and mathematically perfect three-dimensional feature reconstruction. If the stereo camera drifts over time, the system will converge to an incorrect physical location. Furthermore, pure PBVS provides no mechanism to keep the target within the camera’s optical field of view. As the robot moves, the target features may simply exit the sensor’s field of view [10].
Hybrid Visual Servoing (2.5D Visual Servoing) integrates the advantages of IBVS and PBVS while discarding their vulnerabilities. In 2.5D partitioned visual servoing, the fundamental control law is mathematically decoupled. The camera’s rotational velocity is estimated by evaluating the homography matrix to determine the partial camera displacement in three-dimensional space, closely resembling PBVS. Simultaneously, the translational velocity of the system is controlled directly in the two-dimensional image plane, mirroring IBVS [11]. Alternatively, switching-based Hybrid Visual Servoing architectures constantly monitor specific geometric or state variables, such as target depth and the condition of the Jacobian matrix. Based on these observations, a logic controller dynamically transitions the control command between a pure IBVS controller and a pure PBVS controller, depending on the spatial proximity to the target or the calculated risk of field-of-view loss [12].
This paper presents the design of an omnidirectional mobile platform and evaluates two control strategies within the Position-Based Visual Servoing (PBVS) framework. The experimental results demonstrate that the performance of Classical PBVS is significantly enhanced by using a dual-number-based formulation.

2. Custom OMR System Architecture

Custom Design for an Autonomous Omnidirectional Wheeled Robot

Modern warehouse logistics demand precision and adaptability, particularly when handling heavy payloads in highly constrained environments. In this section, the focus will be on the design and implementation of a custom-built Omnidirectional Wheeled Mobile Robot engineered specifically for these challenges. By integrating a dual-stereo camera perception system, this project solves critical bottlenecks in autonomous docking, precise spatial positioning, and reliable package delivery. Supplementary video footage validating the docking and payload tests can be viewed at https://youtu.be/e9q2MC5HzEs (accessed on 1 April 2026).
From a hardware perspective, the Omnidirectional Mobile Autonomous Robot is a mechatronic assembly that focuses on the mechanical elements and associated drive structures, which, based on the energy provided by the power system (usually of an electrical nature), act in the workspace to perform various tasks with the ability to understand the environment through the perception system. The decision-making component is usually implemented on a multi-layered control architecture. The software component that integrates various working algorithms must meet several important objectives, as defined by [13]: programmability, reactivity, autonomy and adaptability, robustness, coherent behavior, evolution, and observability.
Managing the complexity associated with high-performance omnidirectional robotic platforms requires the adoption of a structured, layered control framework. Within this framework, authority flows exclusively from higher-level layers to lower-level ones, establishing a well-defined order of precedence among functional components. Beyond simplifying system management, this organization provides considerable design freedom, since changes to a single module have negligible effects on the rest of the architecture. Two principal paradigms exist for formulating the control problem of a mobile robot: one rooted in kinematics and one rooted in dynamics. The kinematic formulation structurally separates the problem into two nested control loops operating at distinct timescales, whereas the dynamic formulation consolidates all control objectives into a single unified loop that captures the full system dynamics. Despite the theoretical appeal of the latter, its practical implementation is unattractive due to the substantial computational burden it imposes, making real-time execution demanding and the associated mathematical complexity significant. The kinematic formulation, being analytically more tractable, permits formal guarantees of closed-loop stability [14]. From a practical standpoint, the dynamic behavior of the robot need not be explicitly modeled, provided that the drive actuators have sufficient response time to commanded inputs at rates far exceeding those required by the task. Under this assumption, actuator torque can be treated as instantaneously available on the timescale of the higher-level control loops, thereby justifying a control design founded exclusively on the kinematic description of the platform [15].
As depicted in Figure 1, the proposed control structure is organized as a three-tier cascade, encompassing a trajectory generation loop at the outermost level, an intermediate kinematic loop, and an innermost dynamic loop. The stability of the composite system is ensured by maintaining a sufficient timescale across the three tiers, so that each inner loop operates on significantly faster dynamics than its enclosing outer loop [16]. The role of the vehicle controller is to aggregate motion commands from various sources, resolve conflicts using a defined priority scheme, and enforce kinematic feasibility by saturating both velocity and acceleration commands to keep the platform within its physical limits before dispatching the resulting signals to the actuators. Actuator velocity references are derived through the application of the inverse kinematic mapping of the OMR, which relates the local velocity vector comprising the longitudinal component v x , the lateral component v y , and the yaw rate ω z to the individual angular velocities ω i ( i = 1 , 4 ¯ ) of the four wheels. This mapping accounts for the geometric parameters of the platform as introduced in [16], namely the wheel radius R, the lateral half-track l x , and the longitudinal half-wheelbase l y . The resulting inverse kinematic relationship, in matrix form, is given by Equation (1):
ω 4 ω 3 ω 2 ω 1 = 1 R 1 1 ( l x + l y ) 1 1 ( l x + l y ) 1 1 ( l x + l y ) 1 1 ( l x + l y ) v x v y ω z .
The vehicle controller uses the direct kinematic model (2) to determine the actual OMR’s relative speed to ground:
v x v y ω z = R 4 1 1 1 1 1 1 1 1 1 l x + l y 1 l x + l y 1 l x + l y 1 l x + l y ω 4 ω 3 ω 2 ω 1 .
Supplementary, as illustrated in Figure 1, the proposed system architecture relies on a highly distributed hardware topology. The perception cycle originates at the mobile platform, where the ZED 2 Stereo Camera acquires raw visual data. The image is transmitted to the Application Controller, where a dedicated ArUco detection node extracts the 3D point features ( P i ) of the target. A fundamental difference between this proposed architecture and Classic PBVS lies in the way visual feedback is handled. Classic PBVS requires a fragile, computationally expensive intermediate step to extract the target’s exact rigid-body pose (translation and rotation) before computing a control error. The dual pose-free architecture fundamentally eliminates this intermediate bottleneck. Using Dual Lie Algebra, the controller parametrizes the raw 3D point measurements directly into dual vectors ( a ^ i ) and calculates the velocity field ( ω c ^ ). The control algorithm generates the platform velocity vector [ v x , v y , ω z ] T . This high-level command is then given to the vehicle controller, who applies the inverse kinematics matrix to distribute specific angular velocities ( [ ω 1 , ω 2 , ω 3 , ω 4 ] T ) to the integrated motor controllers.
The hardware architecture can be described by the simplified schematic in Figure 1, which illustrates the connectivity among low-level components required to drive the robot and to provide sensory feedback, as well as the network connections and communication protocols between high-level components. The system shown in this paper is composed of three computational units:
  • Vehicle Controller: Implemented by an industrial Programmable Logic Controller (PLC) used for implementing the control logic for the movement of the robot based on the kinematic equations and the constraints imposed by the safety controller, and also managing the execution of the actuators. The reference velocity inputs can come either manually from a user using a joystick or using a numerical unit (in this case, the NVIDIA Jetson AGX Xavier) that communicates via a network connection with the PLC.
  • Safety Controller: Its duty is to secure the safety of both the robot and the objects or people in its proximity, based on data acquired from an LiDAR system. The data represents distances to nearby obstacles, and the system uses this data to determine whether an object is in any of the user-defined regions around the robot, triggering control signals that tell the Schneider PLC to either slow down or stop the robot to avoid a collision.
  • Application Controller: The most powerful processing unit in the system, which integrates and executes all of the software functionalities needed for local decision-making and data acquisition that require computationally heavy algorithms, which work with a big volume of data that needs to be acquired from network communication or the peripherals of the unit. The response of such algorithms needs to be fast to compensate for the network latency and to ensure the stability of the system and communicate anomalies, whether external, in the outside environment, or internal, in the system, so it can be as dynamic as possible, hence the need for computational power.
The software architecture is designed to prioritize modularity, scalability, and maintainability. Furthermore, it requires robust interfaces for efficient data acquisition, analysis, and debugging. To meet these functional requirements, the Robot Operating System (ROS) Noetic [17] framework was selected as the core development environment.
ROS operates in a simple, structured way and offers a highly versatile set of tools for robotics across most programming languages, for both development and visualization. The development tools help break complex functionality into nodes in a directed graph, making it easier to follow the system’s workflow and understand each individual subsystem. Any node can be a publisher, subscriber, or both (Figure 2). Publishers transmit data to the subscribers to listen through a topic. Abstractly speaking, a topic represents an edge between two nodes in the graph, where its purpose is to define the type of message that needs to be sent between two nodes, and which publishers communicate to which subscribers, where data is transmitted asynchronously.
The management of these functionalities is handled by the ROS Master, which adds nodes, ensures connectivity between them, whether they are on the same local network or not, and orders messages sent by publishers and in requests/responses.

3. Visual Perception and Control Strategy

The proposed control architecture for our custom Omnidirectional Mobile Robot prioritizes reliable visual perception for warehouse maneuvers. Using ArUco markers [18] and a dual-stereo camera integrated into the ROS-based Application Controller, the robot achieved robust docking and orientation. The proposed methods are Classic Pose-Based Visual Servoing and Dual Pose-Free Visual Servoing. In testing, the Dual Pose-Free Visual Servoing law outperformed the Classic Visual Servoing method, delivering a more stable and more accurate control response. In the YouTube video from Section Custom Design for an Autonomous Omnidirectional Wheeled Robot (Project Functionality Tests, at 1:00 min), the occupancy grid map of the operational environment is shown, which serves as the human–machine interface for the fleet management system. The map was constructed offline using simultaneous localization and mapping (SLAM) and is used at runtime for robot localization. The two orange rectangles represent the current poses of the two omnidirectional platforms (Rosy 1 and Rosy 2), with the green square marker indicating each robot’s heading. The current navigation goal vector is assigned by the task manager. To dispatch the leader to a docking location, an operator or high-level cloud task manager issues a target pose command specified as a 2D position and azimuth angle, which is forwarded to the leader’s navigation stack as a global goal. The robot then autonomously navigates to the vicinity of the target using odometric control mode. When the robot is close to the objective, the visual servoing controller is activated to execute the final precision docking maneuver. The follower robot receives no explicit goal from the task manager and remains in continuous visual servoing mode, maintaining its relative formation with the leader throughout the mission. This setup was used in order to validate the robustness and better performance of the Dual Pose-Free Visual Servoing controller within Section 4.

3.1. Visual Perception

The architecture design can be divided into two categories:
  • Connections and communication between the hardware components.
  • Frameworks and topology used for software functionalities.
Given the inherent complexity of the robotic platform, a distributed hardware architecture is adopted. This design ensures safe and reliable operation by distributing tasks across specialized components to prevent computational bottlenecks.
Stereo vision has proven to be effective in applications such as indoor navigation, obstacle detection and avoidance, or mapping buildings. In the field of autonomous robots, environmental dynamics are important factors in operational mode, response time, and safety measures. In this sense, a wide range of sensors is used to obtain information about the space in which the robot operates. Artificial vision systems are used in robot systems because they provide valuable information, such as depth, color, shape, and recognition of persons or objects, as well as, together with other sensors (barometer, magnetometer), positioning and orientation information.
Given these aspects, the stereo camera is a viable solution for this custom OMR project due to fast dynamics, the possibility of calculating the distance in real time, the ability to detect objects, shapes, and colors, along with the conditions of operation ensured by the working environment (optimum brightness, low presence of disturbances and exogenous elements). For this project, a ZED 2 camera manufactured by Stereolabs was used. It is a stereo camera used in other projects, such as quality control [19], visual SLAM [20,21], and underwater applications [22]. The camera can capture images at up to 2K resolution and is integrated with ROS, enabling the interconnection of the robot’s subsystems. This camera, along with the incorporated sensors (accelerometer, gyroscope, magnetometer), was used in project development for marker detection, positioning, and orientation.
The operational stability of the omnidirectional platform is highly dependent on the synchronized management of data across the distributed hardware architecture. To ensure safe and continuous trajectory execution, the data flow is strictly shared between the perception sensors, the Application Controller (NVIDIA Jetson AGX Xavier), and the vehicle controller (Schneider PLC). The perception cycle begins with the ZED 2 stereo camera, which is configured to capture high-definition RGB-D frames at 60 Hz. The raw optical and depth data are published into the ROS environment. A dedicated computer vision algorithm subscribes to this stream to detect fiducial markers. Upon successful detection, the node isolates the 3D point features ( P i ) using the ArUco detect package, which implements the core ArUco algorithms explained in [23] and publishes the relative pose, fiducial area, and detection confidence within the ROS network.
While the camera hardware operates at 60 Hz, the visual feedback control node is implemented with a 10 Hz callback to avoid computational bottlenecks and network congestion. Within this 100 ms period, the Application Controller subscribes to the latest available asynchronous messages, parametrizes the 3D coordinates for the control law, and generates the final platform velocity command vector u = [ v x , v y , ω z ] T .
Once computed, the command vector is published to the network via the topic. The Schneider PLC (vehicle controller) acts as the subscriber for this high-level command. Operating on a significantly faster internal dynamic loop, the PLC receives the velocity command and instantly applies the inverse kinematics matrix to distribute the necessary angular velocities ( [ ω 1 , ω 2 , ω 3 , ω 4 ] T ) to the four Mecanum wheel drives.

3.2. Visual Feedback Control

Visual feedback control integrates concepts from computer vision and control theory to guide robotic systems in performing various tasks. The goal of the control algorithm is to compute velocities to minimize the error between the desired position and the actual positions of the image features. Therefore, researchers have increasingly focused on visual servoing, developing innovative solutions for a range of sensor types and applications. To handle highly complex and unstructured environments, researchers are increasingly combining visual servoing with Deep Reinforcement Learning methods, as in [24,25]. Furthermore, cutting-edge predictive frameworks now incorporate visual servoing methods and visual feedback Control Barrier Functions (CBFs) to control robots in novel spatial environments and to guarantee real-time safety constraints for mobile robots navigating highly dynamic spaces [26,27].

3.2.1. Classic PBVS

The Classic PBVS is designed to minimize the pose deviation between the current feature state and a defined reference state. For the rotational components, we utilize a quaternion-based representation, where a quaternion is expressed as Q = [ q 0 q 1 q 2 q 3 ] T , where q 0 is the scalar part and the units q 1 to q 3 form the vector part. Calculating the orientation error requires comparing the target quaternion q r e f against the measured quaternion of the omnidirectional robot q m . Adopting the methodology outlined by [28], the orientation error q e r r is obtained by multiplying the reference quaternion by the conjugate of the measured quaternion:
q e r r = q r e f q m * ,
where q r e f is the reference quaternion and q m * is the conjugate of the estimated quaternion. The operator ∗ denotes the standard quaternion product (Hamilton product). For any two given quaternions p = [ p 0 , p v ] T and q = [ q 0 , q v ] T , where p 0 and q 0 represent the scalar parts and p v , q v R 3 represent the vector parts, the quaternion product is mathematically defined as
p q = p 0 q 0 p v · q v p 0 q v + q 0 p v + p v × q v .
The command that is transmitted to the motors is
u = k p · [ δ X δ y q 1 e r r ] ,
where k p is the gain and δ X δ y are the position biases.

3.2.2. Dual Pose-Free Visual Servoing

Pose-free visual servoing is a control framework that drives a robotic system directly from 3D point features, bypassing any intermediate recovery of the rigid object’s exact pose. The motion is parameterized through dual numbers and dual vectors, which together generate the orthogonal dual tensor group—a structure isomorphic to the special Euclidean group S E 3 . A central contribution of this work is to adapt the pose-free scheme of [29] to the OMR architecture and to benchmark it against the classical PBVS.
The set of real dual numbers R ^ = R + ε R is written as
R ^ = { a ^ = a + ε a 0 | a , a 0 R , ε 2 = 0 , ε 0 } ,
in which a = R e ( a ^ ) denotes the real part of a ^ and a 0 = D u ( a ^ ) its dual part. Correspondingly, the set of dual vectors V ^ 3 = V 3 + ε V 3 reads
V ^ 3 = { a ^ = a + ε a 0 ; a , a 0 V 3 , ε 2 = 0 , ε 0 } ,
where V 3 is the three-dimensional linear space of free Euclidean vectors, and a = R e ( a ^ ) and a 0 = D u ( a ^ ) are, respectively, the real and dual parts of a ^ . Three operations are available on dual vectors: the scalar product (written a ^ · b ^ ), the cross-product (written a ^ × b ^ ), and the triple-scalar product (written < a ^ , b ^ , c ^ > ).
Suppose a depth-capable visual sensor observes a rigid body in two configurations, a current and a desired one. The body is described by n point features, given in the current image as p i = ( x i , y i ) , i { 1 , n } and in the desired image as p i * = ( x i * , y i * ) , i { 1 , n } . From the sensor’s intrinsic parameters, together with the per-feature depth, the corresponding 3D positions P i , P i * , i { 1 , n } are reconstructed. The objective is to determine the 3D sensor velocities that steer the features toward the desired configuration. Denote the centroids of the 3D point sets { P i } i = 1 , n ¯ and { P i * } i = 1 , n ¯ by
G ( t ) = i = 1 n P i ( t ) n ; G * = i = 1 n P i * n .
The image error p i p i * is back-projected into 3D using the intrinsic parameters and the depth measurements:
E ( t ) = P i ( t ) P i * , i = 1 , n ¯ .
Differentiating with respect to time yields
E ˙ i ( t ) = P ˙ i ( t ) , i = 1 , n ¯ .
Imposing an exponential decay of the error,
E ˙ i ( t ) = λ E i ( t ) , i = 1 , n ¯ ,
and combining Equation (9) with v i ( t ) = P ˙ i ( t ) , it follows that the desired behavior is obtained whenever the velocity field of the rigid-body features satisfies
v i ( t ) = λ P E i ( t ) , i = 1 , n ¯ .
With these quantities, the following dual vectors are formed:
a ^ i = P i G + ε G × P i ,
a ^ ˙ i = v i v G + ε ( v G × P i + G × v i ) ,
for i = 1 , n ¯ , with v G = λ G ( G G * ) .
The pose-free formulation of the visual feedback problem [10,29] collects the desired linear and angular sensor velocities [ v c ω c ] = [ v x v y v z ω x ω y ω z ] into a single dual vector ω ^ c = ω c + ε v c , which encodes the prescribed velocity field:
ω ^ c = 1 2 i = 1 n ( A ^ 1 a ^ i ) × a ˙ ^ i ,
with A ^ = a ^ 1 a ^ 1 + . . . + a ^ n a ^ n . Equation (14) is straightforward to implement and to test. For the omnidirectional robot, the applied command reduces to
u = [ v x v y ω z ] .
The experimental results obtained on the omnidirectional platform are presented next.

4. Experimental Results

The developed omnidirectional platform is illustrated in Figure 3. The same picture illustrates several markers considered for the project (16-bit binary matrix).
The camera captures HD images (720p resolution) at 60 Hz. In every frame, a computer vision algorithm detects ArUco markers. If one marker is detected, a ROS topic (e.g.,/fiducial_transforms) sends data through the system, including the relative pose of the marker with respect to the camera, fiducial area, and a confidence value for the detected object. Then, knowing the camera’s position relative to the robot’s center, the robot’s position is determined relative to the camera and the marker.
To experimentally validate and compare the proposed control strategies, a dual-robot framework operating in a leader–follower configuration was established. It is important to note that the two platforms operate under fundamentally different control regimes throughout the mission. The leader robot is not continuously engaged in visual servoing; rather, it is dispatched by a task manager to a designated target location via a cloud-based command (Task and Trajectory Tracking Control from Figure 1). During the transit phase, the leader operates under waypoint-following control, and the visual servoing controller is activated only upon approaching the target, at which point the docking procedure is initiated. This behavior is reflected in the command profiles of Figure 4, where the leader’s visual servoing signals are constant during the initial transit and become active only after approximately 11 s. The follower robot operates in continuous closed-loop visual servoing mode throughout the entire duration, persistently tracking the leader as its dynamic visual reference. Since only the leader is assigned a physical destination by the task manager—as evidenced in the accompanying video—the follower’s sole objective is formation maintenance, making it permanently dependent on the quality and stability of its visual control law. In this way, the follower is exposed to a longer, continuous period of active visual servoing, making the robustness of the control law particularly critical to its smooth operation. To assess the performance of the visual servoing algorithms under different perceptual conditions, the experimental design incorporated two distinct fiducial marker arrangements: (i) a mono-marker configuration, featuring a single tag attached to both the docking station and the leader robot; and (ii) a multi-marker configuration, employing a group of four tags at each target location to describe the rigid body. The docking action starts by detecting the fiducial marker and extracting the translation vector and rotation quaternion. Using these data, the visual feedback algorithm minimizes the error between the robot’s desired docking position and its actual position. We have implemented a ROS node with a 10 Hz callback that processes data from the detection algorithm and computes a motor command, which is transmitted to the system via a dedicated topic (e.g., /cmd_vel).
These experiments aim to evaluate the differences between the two solutions described above: Classic and Dual Pose-Free VS. The goal is to compare linear and angular velocities and analyze the commands generated for the four mechanum wheels. The gains for the control laws were set to k p = 0.3 and λ P = λ G = 0.5 .
At HD resolution, the camera accurately detects a marker from 1.5–2 m up to 25 cm. At this distance, marker orientation is not precise, which is why the Dual Pose-Free VS method, which uses only marker positions, is preferred, thereby increasing the algorithm’s robustness.
The experimental results demonstrate a significant improvement in control signal integrity when utilizing the Dual Pose-Free VS formulation compared to the classic approach. Figure 4 presents the velocity command profiles and chattering for the dual-robot setup under both Classic PBVS and the proposed Dual Pose-Free Visual Servoing controller.
As depicted in Figure 4a, the Dual Pose-Free VS demonstrates a superior convergence trajectory. After 11 s, the dual method initiates a smooth, robust adjustment across all axes. The velocity commands approach zero asymptotically, indicating stable exponential error decay. In contrast, the Classic PBVS controller exhibits a delayed response followed by abrupt and aggressive step changes. This is most evident in the rotational Z-axis command, where the classic method forces the system into large-amplitude chattering throughout the task duration.
To rigorously quantify the impact of these algorithms, the derivatives of the control commands ( Δ u x , Δ u y , Δ ω z ) were plotted to reveal high-frequency signals, indicating aggressive changes in acceleration that induce substantial mechanical stress on the actuators and gears. This phenomenon leads to increased energy consumption and potential hardware degradation. Figure 4b shows the advantage of the dual method, which mitigates the severe high-frequency oscillations seen within Classic PBVS. This chattering is caused by the intermediate pose estimation method, which is altered at close range. If the stereo camera is near the fiducial marker, minor depth ambiguities cause the pose recovery block to fluctuate rapidly, forcing the proportional controller to issue spiky commands. The Dual Pose-Free VS entirely solves this disadvantage. While there is a brief, bounded transient response as the maneuver initiates (t = 11 s), the derivative of the dual command quickly converges and remains near zero throughout the critical final docking phase.
The follower data corroborates the findings from the leader data, demonstrating the superiority of the Dual Pose-Free approach. As observed in the velocity command from Figure 4c, the Dual Pose-Free VS exhibits a much smoother and more anticipatory tracking behavior. On the X-axis translation, the dual method initiates a gradual, exponential deceleration phase significantly earlier. The Classic PBVS remains at −0.2 m/s for a longer duration, resulting in a late and aggressive kinematic correction. The Z-axis rotational command further highlights the dual method’s stability. The dual controller computes a low-amplitude, smooth trajectory. In comparison, the Classic PBVS initiates with an excessively high angular velocity (0.5 rad/s) and exhibits high-amplitude oscillations throughout the entire 20-s maneuver.
Figure 4d provides a better quantification of the high-frequency noise. The Classic PBVS controller exhibits severe chattering, particularly on the Y- and Z-axes. This phenomenon appears from the intermediate pose estimation algorithm, which induces noise into the controller. Classic PBVS struggles to maintain a consistent orientation while tracking the leader robot.
The Dual Pose-Free VS algorithm successfully filters out this instability. Because it directly parametrizes 3D feature points into a dual-vector velocity field, it is immune to pose estimation errors. While the dual method registers a single, brief transient spike on the X-axis and Y-axis at t = 11s, it immediately dampens the signal. Most importantly, the rotational chattering ( Δ ω z ) is completely mitigated, remaining at a near-perfect zero for the entirety of the operation.
The seamless integration of robust visual perception with the continuous control strategies of a custom Omnidirectional Mobile platform is crucial for executing high-precision logistical tasks. Implementing a Pose-Free Visual Servoing architecture parameterized by dual vectors, the system achieves a high level of operational flexibility. Because the dual control law relies purely on the mathematical relationships and velocity fields of 3D point sets, the method can use a variety of point features, such as infrared point features, Artificial Fiducial Markers (e.g., ArUco, AprilTags) and Deep Learning-Based Semantic Keypoints (e.g., SuperPoint). This represents an advantage over Classic Position-Based Visual Servoing (PBVS). Classic PBVS strictly requires an a priori 3D geometric model of the target to explicitly compute a rigid-body pose estimation before it can generate any control velocities. This intermediate estimation step is computationally rigid; if the robot rotates and a feature leaves the field of view, traditional PBVS often fails. Conversely, the pose-free dual method entirely addresses this vulnerability by directly computing the required velocities from the visible 3D measurements. If occlusions occur during a maneuver, the dual controller simply reduces the contribution of the occluded feature while maintaining trajectory stability. In this way, the Dual Pose-Free VS represents a robust method for pure visual tracking, enabling the omnidirectional robot to dock seamlessly in dynamic, highly unstructured warehouses.
Consequently, the Dual Pose-Free VS method exhibits superior trajectory generation capabilities, characterized by more intuitive, kinematically consistent robot behavior. The dual approach initiates a gradual deceleration phase much earlier in the trajectory; this way, the robot converges with a fluid motion that is both safer and more predictable.
The real behavior of the visually controlled omnidirectional robot can be observed in these two videos: https://youtu.be/scYwB4Gz-58 (accessed on 1 April 2026) and https://youtu.be/4BAgJl4RAQA (accessed on 1 April 2026).

5. Conclusions

This research aims to address the precision-positioning bottlenecks inherent to autonomous, high-payload warehouse logistics. This paper presents a custom OMR architecture and contributes to the design of an integrating protocol of visual perception and control. This platform was explicitly chosen because traditional differential-drive systems lack the lateral maneuverability required for docking and heavy-payload positioning in constrained warehouses.
Robust visual perception underpins the autonomous navigation and maneuvering capabilities of our custom Omnidirectional Mobile Robot. The primary feedback loop is driven by a stereo camera system paired with an ArUco fiducial marker detection pipeline, fully integrated into the Robot Operating System framework on the main application controller. By leveraging dense depth data from the stereo camera, the system achieves continuous marker extraction and tracking, demonstrating high resilience against variable illumination and dynamic warehouse environments.
Experimental validation of the proposed control strategies—Classic Pose-Based Visual Servoing and Dual Pose-Free Visual Servoing—revealed significant performance differences. The Dual Pose-Free controller outperformed the classic baseline, demonstrating superior steady-state positioning accuracy and smoother trajectories. Consequently, the integration of this dual method markedly enhanced the system’s overall robustness and minimized operational failure rates under dynamic conditions.
Another contribution highlights how visual feedback control can be integrated into the OMR architecture while achieving autonomous behavior. The evaluation of Classical Position-Based Visual Servoing and an innovative pose-free approach to visual servoing provides a solid foundation for integrating visual control. While Classical PBVS is computationally rigid—relying heavily on explicit a priori 3D models and risking control stability if features leave the field of view—the Dual Pose-Free VS method proved fundamentally more robust. The dual approach bypassed fragile intermediate pose estimations entirely, offering a highly flexible and robust alternative. Optimal kinematic behavior can be achieved by systematically tuning the two exponential-decrease parameters ( λ P and λ G ). By allowing the robot to maintain operational continuity even if markers were lost, provided at least three markers were visible, the system demonstrated high resilience. This fault-tolerant capability is particularly advantageous in warehouse environments, where fluctuating lighting conditions and physical obstructions often interfere with visual tracking.
To validate these theoretical advantages within a practical logistical paradigm, the control architecture was deployed on a custom-engineered Omnidirectional Mobile Robot. Empirical results substantiated the operational superiority of the dual method, demonstrating sustained trajectory stability and precise autonomous docking execution. Notably, the pose-free algorithm dynamically discarded unobserved visual features, thereby maintaining control-loop continuity without compromising asymptotic convergence. By decoupling and scaling the translational and rotational error-decay rates, the parameter-tuning framework effectively mitigated mechanical overshoot, thereby ensuring smooth, precise, and bounded velocity profiles essential for transporting heavy logistical payloads. Future work will focus on extending this framework to multi-robot cooperative logistics and on integrating dynamic obstacle-avoidance capabilities.

Author Contributions

Conceptualization, A.B., R.-L.R. and A.-I.I.; methodology, A.B. and R.-L.R.; software, R.-L.R. and A.-I.I.; validation, R.-L.R., A.-I.I.; writing, R.-L.R., A.B., A.-I.I. and C.D. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The datasets presented in this article are not readily available because these data are part of a broader ongoing research project focused on long-term fleet management. Requests to access the datasets should be directed to [the corresponding author].

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Qian, J.; Zi, B.; Wang, D.; Ma, Y.; Zhang, D. The Design and Development of an Omni-Directional Mobile Robot Oriented to an Intelligent Manufacturing System. Sensors 2017, 17, 2073. [Google Scholar] [CrossRef] [PubMed]
  2. Chaumette, F.; Hutchinson, S. Visual servo control Part I: Basic approaches. IEEE Robot. Autom. Mag. 2006, 13, 82–90. [Google Scholar] [CrossRef]
  3. Andreff, N.; Espiau, B.; Horaud, R. Visual Servoing from Lines. Int. J. Robot. Res. 2002, 21, 679–699. [Google Scholar] [CrossRef]
  4. Chaumette, F. Image moments: A general and useful set of features for visual servoing. IEEE Trans. Robot. 2004, 20, 713–723. [Google Scholar] [CrossRef]
  5. Mäkinen, P.; Mustalahti, P.; Launis, S.; Mattila, J. Vision-aided precise positioning for long-reach robotic manipulators using local calibration. Adv. Robot. 2024, 38, 82–94. [Google Scholar] [CrossRef]
  6. Xu, F.; Wang, H.; Gao, H. What Matters in Constructing a Visual Servoing Scheme: A Review of Key Issues and Solutions. IEEE/ASME Trans. Mechatron. 2026, 31, 1509–1523. [Google Scholar] [CrossRef]
  7. Enyedy, A.; Aswale, A.; Calli, B.; Gennert, M. Stereo Image-based Visual Servoing Towards Feature-based Grasping. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; pp. 7325–7331. [Google Scholar] [CrossRef]
  8. Chaumette, F.; Hutchinson, S.; Corke, P. Visual Servoing. In Handbook of Robotics; Siciliano, B., Khatib, O., Eds.; Springer: Berlin/Heidelberg, Germany, 2016; pp. 841–867. [Google Scholar]
  9. Samadikhoshkho, Z.; Lipsett, M.G. Visual Servoing for Aerial Vegetation Sampling Systems. Drones 2024, 8, 605. [Google Scholar] [CrossRef]
  10. Burlacu, A.; Condurache, D. A different approach to solving the PBVS control problem. In Proceedings of the IEEE 29th International Symposium on Industrial Electronics (ISIE), Delft, The Netherlands, 17–19 June 2020; pp. 1359–1364. [Google Scholar]
  11. Rahmatillah, A.; Trilaksono, B.R.; Hindersah, H. 2.5-D visual servoing experiment using Adept Viper s850. In Proceedings of the 2013 3rd International Conference on Instrumentation, Communications, Information Technology and Biomedical Engineering (ICICI-BME), Bandung, Indonesia, 7–8 November 2013; pp. 296–301. [Google Scholar] [CrossRef]
  12. Kan, J.; Wu, Y.; Dong, R.; Yao, S.; Zhao, X.; Zou, T.; Kang, B.; Li, J. A Progressive Hybrid Automatic Switching Visual Servoing Method for Apple-Picking Robots. Agriculture 2026, 16, 620. [Google Scholar] [CrossRef]
  13. Fleury, S.; Herrb, M.; Chatila, R. Design of a modular architecture for autonomous robot. In Proceedings of the 1994 IEEE International Conference on Robotics and Automation, San Diego, CA, USA, 8–13 May 1994; Volume 4, pp. 3508–3513. [Google Scholar] [CrossRef]
  14. Indiveri, G.; Nuchter, A.; Lingemann, K. High Speed Differential Drive Mobile Robot Path Following Control With Bounded Wheel Speed Commands. In Proceedings of the 2007 IEEE International Conference on Robotics and Automation, Rome, Italy, 10–14 April 2007; pp. 2202–2207. [Google Scholar] [CrossRef]
  15. Gracia, L.; Tornero, J. Kinematic control of wheeled mobile robots. Lat. Am. Appl. Res. 2008, 38, 7–16. [Google Scholar]
  16. Dosoftei, C.C.; Popovici, A.T.; Sacaleanu, P.R.; Gherghel, P.M.; Budaciu, C. Hardware in the Loop Topology for an Omnidirectional Mobile Robot Using Matlab in a Robot Operating System Environment. Symmetry 2021, 13, 969. [Google Scholar] [CrossRef]
  17. Quigley, M. ROS: An open-source Robot Operating System. In Proceedings of the IEEE International Conference on Robotics and Automation, Kobe, Japan, 12–17 May 2009. [Google Scholar]
  18. Garrido-Jurado, S.; Muñoz-Salinas, R.; Madrid-Cuevas, F.; Marín-Jiménez, M. Automatic generation and detection of highly reliable fiducial markers under occlusion. Pattern Recognit. 2014, 47, 2280–2292. [Google Scholar] [CrossRef]
  19. Tran, M.T.; Kim, D.H.; Kim, C.K.; Kim, H.K.; Kim, S.B. Determination of Injury Rate on Fish Surface Based on Fuzzy C-means Clustering Algorithm and L*a*b* Color Space Using ZED Stereo Camera. In Proceedings of the 2018 15th International Conference on Ubiquitous Robots (UR), Honolulu, HI, USA, 26–30 June 2018; pp. 466–471. [Google Scholar]
  20. Afanasyev, I.; Ibragimov, I. Comparison of ROS-based Visual SLAM methods in homogeneous indoor environment. In Proceedings of the 2017 14th Workshop on Positioning, Navigation and Communications (WPNC), Bremen, Germany, 25–26 October 2017. [Google Scholar] [CrossRef]
  21. Gupta, T.; Li, H. Indoor mapping for smart cities—An affordable approach: Using Kinect Sensor and ZED stereo camera. In Proceedings of the 2017 International Conference on Indoor Positioning and Indoor Navigation (IPIN), Sapporo, Japan, 18–21 September 2017; pp. 1–8. [Google Scholar] [CrossRef]
  22. Wang, C.; Zhang, Q.; Lin, S.; Li, W.; Wang, X.; Bai, Y.; Tian, Q. Research and Experiment of an Underwater Stereo Vision System. In Proceedings of the OCEANS 2019, Marseille, France, 17–20 June 2019; pp. 1–5. [Google Scholar] [CrossRef]
  23. Salinas, R.M. ArUco: A Minimal Library for Augmented Reality Applications Based on OpenCv; Technical Report; Universidad de Cordoba: C’ordoba, Spain, 2011. [Google Scholar]
  24. Shin, C.; Ferguson, P.W.; Pedram, S.A.; Ma, J.; Dutson, E.P.; Rosen, J. Autonomous Tissue Manipulation via Surgical Robot Using Learning Based Model Predictive Control. In Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada, 20–24 May 2019; pp. 3875–3881. [Google Scholar] [CrossRef]
  25. Karnan, H.; Warnell, G.; Xiao, X.; Stone, P. VOILA: Visual-Observation-Only Imitation Learning for Autonomous Navigation. In Proceedings of the 2022 International Conference on Robotics and Automation (ICRA), Philadelphia, PA, USA, 23–27 May 2022; pp. 2497–2503. [Google Scholar] [CrossRef]
  26. Tang, J.; Deng, Z.; Zhang, H.; Jiang, T.; Wang, C.; Ke, Z. Control Barrier Function-Based Quadrotor Interception Control Using Visual Servoing. In Proceedings of the 2024 IEEE International Conference on Unmanned Systems (ICUS), Nanjing, China, 18–20 October 2024; pp. 319–324. [Google Scholar] [CrossRef]
  27. Jiang, J.; Wang, Y.; Jiang, Y.; Miao, Z. Adaptive NN based Visual Servoing Control for Robot Manipulator with Field of View Constraints and Dynamic Uncertainties. In Proceedings of the 2021 IEEE International Conference on Robotics and Biomimetics (ROBIO), Sanya, China, 27–31 December 2021; pp. 1694–1699. [Google Scholar] [CrossRef]
  28. Fresk, E.; Nikolakopoulos, G. Full quaternion based attitude control for a quadrotor. In Proceedings of the 2013 European Control Conference (ECC), Zurich, Switzerland, 17–19 July 2013; pp. 3864–3869. [Google Scholar]
  29. Burlacu, A.; Rosca, R.L.; Iancu, A.I.; Cervera, E. Pose-free visual servoing from 3D measurements. In Proceedings of the 2023 European Control Conference (ECC), Bucharest, Romania, 13–16 June 2023; pp. 1–6. [Google Scholar] [CrossRef]
Figure 1. Architectural comparison between (a) Classic Position-Based Visual Servoing and (b) the proposed Dual Pose−Free Visual Servoing. The classic framework requires a computationally rigid intermediate step to estimate the full rigid-body pose, making the control loop highly vulnerable to visual occlusions. The proposed dual approach directly maps extracted 3D point features ( P i ) into a dual−vector parameterization block, bypassing traditional intermediate rigid−body pose estimations.
Figure 1. Architectural comparison between (a) Classic Position-Based Visual Servoing and (b) the proposed Dual Pose−Free Visual Servoing. The classic framework requires a computationally rigid intermediate step to estimate the full rigid-body pose, making the control loop highly vulnerable to visual occlusions. The proposed dual approach directly maps extracted 3D point features ( P i ) into a dual−vector parameterization block, bypassing traditional intermediate rigid−body pose estimations.
Sensors 26 03918 g001
Figure 2. ROS architecture.
Figure 2. ROS architecture.
Sensors 26 03918 g002
Figure 3. Marker-based docking.
Figure 3. Marker-based docking.
Sensors 26 03918 g003
Figure 4. Comparison of Classic PBVS and Dual Pose−Free Visual Servoing velocity commands and control signal volatility for the leader (top row) and follower (bottom row) omnidirectional platforms.
Figure 4. Comparison of Classic PBVS and Dual Pose−Free Visual Servoing velocity commands and control signal volatility for the leader (top row) and follower (bottom row) omnidirectional platforms.
Sensors 26 03918 g004
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Roșca, R.-L.; Iancu, A.-I.; Burlacu, A.; Dosoftei, C. Integrating Visual Perception and Control Strategies in Custom Omnidirectional Mobile Robots. Sensors 2026, 26, 3918. https://doi.org/10.3390/s26123918

AMA Style

Roșca R-L, Iancu A-I, Burlacu A, Dosoftei C. Integrating Visual Perception and Control Strategies in Custom Omnidirectional Mobile Robots. Sensors. 2026; 26(12):3918. https://doi.org/10.3390/s26123918

Chicago/Turabian Style

Roșca, Radu-Laurențiu, Andrei-Iulian Iancu, Adrian Burlacu, and Cătălin Dosoftei. 2026. "Integrating Visual Perception and Control Strategies in Custom Omnidirectional Mobile Robots" Sensors 26, no. 12: 3918. https://doi.org/10.3390/s26123918

APA Style

Roșca, R.-L., Iancu, A.-I., Burlacu, A., & Dosoftei, C. (2026). Integrating Visual Perception and Control Strategies in Custom Omnidirectional Mobile Robots. Sensors, 26(12), 3918. https://doi.org/10.3390/s26123918

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop