Next Article in Journal
Horse Herd Leadership Optimization: A Trust-Aware Metaheuristic for Resource Allocation and Secure Wireless Sensor Networks
Previous Article in Journal
Evaluating the Potential of Decision Tree Modeling to Augment Return-to-Duty Decisions Following Major Limb Injury
Previous Article in Special Issue
A Hierarchical Distributed Control System Design for Lower Limb Rehabilitation Robot
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Anomaly Detection Using Machine Learning for Robotics Environments on 5G Networks

by
Mikel Dean Oses
1,
Aitor Domec Paz
1,*,
Santiago Figueroa-Lorenzo
2,
Saioa Arrizabalaga
1,2,3,* and
Ricardo Rodriguez-Jorge
1,2
1
Ceit-Basque Research and Technology Alliance (BRTA), Manuel de Lardizábal 15, 20018 Donostia/San Sebastián, Spain
2
Universidad de Navarra, Tecnun, Manuel de Lardizábal 13, 20018 Donostia/San Sebastián, Spain
3
Institute of Data Science and Artificial Intelligence (DATAI), Universidad de Navarra, 31009 Pamplona, Spain
*
Authors to whom correspondence should be addressed.
Technologies 2026, 14(2), 108; https://doi.org/10.3390/technologies14020108
Submission received: 12 December 2025 / Revised: 23 January 2026 / Accepted: 5 February 2026 / Published: 9 February 2026
(This article belongs to the Special Issue AI Robotics Technologies and Their Applications)

Abstract

This work underscores the importance of developing and refining machine learning (ML) methods to meet the specific demands of anomaly detection in 5G-powered environments. It addresses key challenges, including the deployment of robotics within industrial settings that require robust low-latency communication and high data throughput. The proposed architecture thus delves into innovative ML-driven approaches that not only optimize anomaly detection but also maintain high performance under the constraints and requirements imposed by 5G-enabled industrial applications. Our experiments demonstrate the effectiveness of these techniques in accurately identifying anomalies while minimizing false positives. The practical implications of integrating anomaly detection into robotics processes are discussed, with potential applications in autonomous driving, warehouse automation, and remote inspection. Finally, this research contributes to the development of robust robotic systems in real-world environments.

1. Introduction

The emergence of advanced wireless communication technologies signals a significant shift in connectivity, ushering in a new era with transformative potential across various sectors. These technologies bring about improvements in data transfer speeds, reduced latency, and expanded network capacity. This evolution in communication infrastructure paves the way for the seamless integration of cutting-edge technologies such as the Internet of Things (IoT) [1]. Rapid data transmission speeds enable the nearly instantaneous exchange of extensive information, promoting real-time connectivity, and unlocking previously impractical applications. The ability to support numerous simultaneous connections not only enhances user experiences, but also drives the development of interconnected smart devices, such as robots, fostering a period of unparalleled connectivity and technological innovation [2].
In the domain of robotic processes, the introduction of wireless technologies, such as 5G, marks a significant leap forward in terms of performance and capabilities. The key advantages of 5G, including high data transfer rates, low latency, and increased device connectivity, synergize to empower robots in diverse applications. With enhanced data transfer speeds, these robots can efficiently process and act on information in real time, allowing quicker decision-making and more dynamic responses to their environments [3]. In addition, the increased connectivity capacity enables seamless communication between robotic devices, facilitating collaborative efforts and coordination in tasks such as warehouse automation or search and rescue missions. As 5G becomes more widely used, robotics processes can achieve new levels of autonomy, efficiency, and adaptability, making significant contributions to fields such as logistics, manufacturing, and beyond [4].
Despite the advantages mentioned above, in a 5G-based environment for robotic processes, security is crucial to prevent errors that could potentially have serious consequences. The integration of advanced technologies such as 5G, 6G, IoT, and legacy OT systems has greatly expanded the capabilities of robotics [1]. However, this technological convergence also introduces new vulnerabilities. A low latency secure setup is essential to ensure real-time responses, enabling systems to avoid critical issues such as entering restricted areas or encountering hazardous conditions, such as high voltage levels, that could cause severe damage to both robots and their surrounding environment.
The increasing complexity of the stack supporting robotics has inadvertently widened the attack surface [5]. Each new technology layer, while offering innovative functionalities, becomes a potential entry point for malicious actors [6]. This expanded surface amplifies the risk of vulnerabilities being exploited throughout the interconnected architecture [7], particularly in environments where data flows between legacy OT systems and modern IoT or 5G networks [8]. Such risks require robust cybersecurity measures that are capable of dynamically addressing potential threats before they escalate into operational disruptions or safety hazards.
Machine learning significantly improves the security of robotic processes by empowering them to detect and respond autonomously to cyber threats; [9] explains the importance of security and [10] presents a solution using advanced machine learning techniques. Existing approaches, such as AI-driven AIDS, have shown promise, particularly in their theoretical capability to identify zero-day vulnerabilities that traditional signature-based systems often miss. These approaches can implement countermeasures such as the integration of information systems and cybersecurity strategies to mitigate risk exposure [11], detailed frameworks that address specific vulnerabilities and attacks specific to robotics [12], and state-of-the-art protection schemes that combine multiple layers of defense [13]. However, these methods face notable challenges. For example, they are prone to false positives, which can cause unnecessary system interruptions, and often lack the adaptability required to navigate the complexities of multilayered architectures involving diverse protocols, legacy technologies, and cutting-edge frameworks [14]. The main contributions of this paper are summarized as follows.
  • The design and implementation of a standalone anomaly detection system for robotics environments on 5G networks, incorporating supervised and unsupervised algorithms.
  • The design of adversary emulation as security tests for both availability and integrity attacks, particularized to Modbus over the TCP communication protocol.
  • Investigate the effectiveness of each model at capturing attacks in terms of recall, accuracy, precision, and other relevant metrics.
The subsequent sections of this paper are organized as follows. Section 2 presents the background, while Section 3 details the design and implementation. Section 4 presents the results of the experimentation, followed by the final Section 5 that summarizes the conclusions and outlines plans for future work.

2. Background

The purpose of this section is to provide an overview of the cybersecurity challenges in robotic processes and the relevant standards that can help address them. Robots are increasingly integrated into IoT ecosystems, making their security a critical concern. However, currently there is no specific standard that directly addresses the unique cybersecurity challenges of robotic environment in 5G networks [15].
To address this gap, this work references the Internet of Things (IoT) reference architecture defined by the ISO/IEC 30141 standard [16]. This standard provides a structured framework for designing IoT systems with reliability, scalability, and interoperability, offering a foundation that can be adapted to enhance the security of robotic systems.
Due to the high use of robots and the rapid development of the IoT, cybersecurity of these robotic processes has become an essential aspect of their design and deployment, as highlighted in [8].

2.1. URSIM

URSim (version 3.2) [17] is a simulation software that is used for offline programming, robot program simulation, and manual robot movements. URSim allows simulation of robot kinematics through digital inputs (I/O) and the definition of different metrics and parameters. It is also able to communicate to a robot controller using TCP sockets via Modbus, as is done in this project.
Since URSim is a simulation software, and not a physical robot, it has some limitations and some functions are not available, e.g., the emergency stop or collisions with itself or surrounding objects [18]. Following the ISO/IEC 30141:2024 standard approach, URSim is located in the device layer, as it simulates robotic hardware such as sensors and actuators.

2.2. Data Transportation Layer

The network where the architecture is deployed is a 5G network. The ultra-low latency and high bandwidth required by this architecture makes it necessary to use a wireless network such as 5G. The network is a private 5G network, which allows one to make changes in bandwidth, making it possible to check the minimum requirements of our architecture. This private network also allows one to avoid interferences of other devices, as other external devices are not permitted to connect to the network. It represents a fundamental element of the network interface capabilities in an IoT system, adhering to the framework defined by the ISO/IEC 30141:2024 standard.

2.2.1. Modbus Protocol

Modbus is a request-response Industrial IoT protocol based on a client–server model, where one device initiates a request and awaits a response. Typically, the master is an HMI (Human–Machine-Interface) or SCADA system, while the slave is a sensor, PLC, or PAC. The protocol’s most common versions are Modbus RTU, Modbus ASCII, and Modbus TCP/IP, with Modbus TCP/IP being used in this project. It is part of the data transfer capability of an IoT system, based on the ISO/IEC 30141:2024 standard approach.
From a performance perspective, Modbus IIoT environments can be considered restricted environments, as latencies cannot exceed 100 ms [19]. It is also known that the major limitations of Modbus are in the area of security, as it cannot be secured directly, but security layers can be added over it. For this reason, references such as [20] or [21] present an extensive taxonomy of attacks for Modbus TCP.

2.2.2. Apache Kafka

Apache Kafka is a widely used open source scalable distributed streaming platform, originally developed by LinkedIn for log processing [22].
It operates on a publisher/subscriber mechanism for data streams and uses a topic-based model for communication. The Kafka design supports data partitioning between multiple servers, enabling a scalable, fault-tolerant system. In addition, it offers security features such as authentication and encryption to secure data transmission.
Apache Kafka makes possible the creation of a new generation of distributed applications capable of handling billions of streamed events per minute. It constitutes a key aspect of the data exchange functionality within an IoT system, following the framework established by the ISO/IEC 30141:2024 standard.

2.2.3. gRPC Protocol

gRPC is a modern open-source high performance remote procedure call (RPC) framework that can run in any environment. It can efficiently connect services in and across data centers with pluggable support for load balancing, tracing, health checking, and authentication. It provides machine-to-machine (M2M) communication and due to its simplicity and transparency is a very popular protocol in distributed applications [23]. It forms part of the data transmission capabilities of an IoT system, aligned with the approach described in the ISO/IEC 30141:2024 standard.
The gRPC protocol provides four types of machine-to-machine communication: unary, server streaming, client streaming, and bidirectional streaming.
Another huge advantage is that the gRPC framework provides a special data type, called “Any”. This data type allows the client and server to freely exchange messages regardless of the data form, which means that the data type can be changed during the stream process without the need to adjust any schema. The “Any” type is supported by all the communication methods explained before.

3. Materials and Methods

This section establishes the design and implementation elements defined for the general architecture of the system, composed of a field device, a cyber-attack simulation tool, the gateway for data transmission, the MEC and cloud for monitoring, and machine learning (ML) models. The system is designed for scalability, reliability, and adaptability, especially within complex real-time environments where secure data handling and analysis are essential.

3.1. Infrastructure Details

The architecture consists of four main components: the Field Device, Gateway, Multi-access Edge Computing (MEC), and Cloud. All components are connected across three network areas: the OT network area where the Field Device resides, the 5G network area for Gateway and MEC and the IT network area for the Cloud. The proposed architecture is presented in Figure 1.
Each component and network area play a vital role in data flow, prediction, and monitoring. Microservice-based architecture is used, allowing the infrastructure to leverage benefits such as rapid deployment, simplified maintenance, and improved scalability. All the parts are connected in a linear way, but the attacker can bypass the connection between the field device and the gateway to alter the transmission of data.

3.1.1. Field Device

The robotic arm, the Universal Robots UR3e model, is moved in a specific pattern that simulates common behavior. This normal operation is defined by the repeatable, deterministic execution of a pre-programmed trajectory. This baseline behavior is characterized by the smooth, controlled movement of the six-degrees-of-freedom robotic arm as it simulates its designated task sequence. During this mode, the system maintains a highly predictable profile across its 37 variables. The URSim tool has been used as a custom synthetic data generator to produce realistic, timestamped sensor streams and actuator logs. The generator can emulate sensor noise, drift, latency, support multiple sensor channels, and inject configurable anomalies (e.g., an unexpected-movement event was injected into the data stream at a regular interval of 1000 data records).
For the baseline scenario, the kinematic state, including the six joint angles, the six joint velocities, and the revolution counts, exhibits periodic and temporally consistent signatures. The dynamic variables (six joint currents, robot current, I/O current, and tool current) operate within tight ranges, reflecting minimal and expected loading. Finally, the thermal profile of the system, comprising six joint temperatures and the tool temperature, remains stable and within known operational bounds. Table 1 shows all the captured process variables.
The device’s data are extracted and relayed through the Modbus TCP protocol, establishing a robust connection for high-frequency data flow. This continuous data collection results in substantial data throughput, supporting high-resolution real-time analytics.

3.1.2. Gateway

Acting as the system’s data entry and exit point, the Gateway, entirely implemented in Python (version 3.12.7) and containerized with Docker, is responsible for receiving and dispatching streaming data from the Field Device to the MEC. Here, a Modbus container establishes the initial connection to the Field Device, retrieving sensor data via Modbus TCP protocol. Once received, the data undergoes preprocessing to meet the specific input requirements of the ML models, a crucial step to minimize latency, since the data processing speed in the Gateway must match the faster publishing rate of the producer in the MEC. After preprocessing, a Kafka Publisher connects via WebSocket to the Modbus container to continuously receive and publish data to a Kafka Broker topic. This step transitions data from the Gateway to the MEC. The ML prediction generated in the MEC is sent back to the Kafka Consumer, which then forwards it to the Modbus container in the Gateway. Finally, a decision-making algorithm, also housed within the Modbus container, determines actions for the Field Device based on predictions, implementing commands (e.g., stop, pause, start, change registers) via Modbus TCP.

3.1.3. MEC

Serving as the system’s computational core, the MEC hosts the Kafka broker and ML models for data classification. The MEC handles large volumes of data efficiently to ensure that predictions and monitoring can occur in near real-time. The MEC handles two ML models with both supervised and unsupervised approaches. Both models are trained exclusively with process data and are described in Section 3.2.

3.1.4. Cloud

The proposed architecture is implemented using the Elastic Stack and is connected directly to the Kafka Broker. The cloud infrastructure ingests data from one or more Kafka topics for real-time monitoring and historical data analysis. This setup enables continuous observation of the health, trends, and performance of the system, thus giving administrators the tools to create custom dashboards, query historical data, and set alerts for any significant trends or anomalies. With Elastic Stack’s visualization capabilities, the cloud environment serves as a comprehensive monitoring and analysis layer, offering insights for continuous improvement.

3.2. Machine Learning Model Selection and Working Flow

Machine learning models are taking over as a powerful way to identify patterns that could indicate malicious activity in an environment where cyber threats are becoming more complex and varied. There are two principal approaches to ML models, supervised and unsupervised.

3.2.1. Supervised Learning

In the case where data labels are available, where it is known when normal and abnormal behavior is happening, a dataset containing 33 min and 32 s of behavior for a total of 10,000 data-points collected at a sampling rate of 5 Hz (one sample every 200 ms) is created. A cyber-attack is simulated every 3.5 min, during these attacks, the movement of the robotic arm is altered for the same amount of time. The dataset was divided into three parts, training (80%), independent testing (20%), and validation (20% of the training dataset).
To ensure a robust detection mechanism, four different models were tested, LightGBM, XGBoost, Random Forest, and a Dense Neural Network (DNN). The first three models are ensemble-based decision-tree models, while DNN is a deep learning model. While ensemble methods like XGBoost and LightGBM are renowned for their efficiency and high performance, DNNs offer the capacity to model complex, non-linear relationships.
To optimize these architectures, the Tree-structured Parzen Estimator (TPE) was used to perform an automated hyperparameter search. The hyperparameters we optimized for 500 trials. The distribution of these 500 trials was non-uniform (265 for LightGBM, 125 for XGBoost, 53 for Random Forest and 57 for DNN) due to how the sampler works. As TPE favors the models that consistently yield the best results as the study progresses, the first 200 trials were carried out at random to establish a more diverse baseline distribution. The DNN uses a sigmoid activation function to convert predictions into probability scores, where any score greater than 0.5 signals an anomaly. The DNN models were trained for 100 epochs with an early stopping mechanism where the training stopped if the validation loss did not improve in 10 epochs.
The data was normalized by removing the mean and scaling to unit variance. For input size reduction, PCA (Principal Component Analysis) was applied to reduce the initial 37 variables to 10 (keeping 95% of variance explained), improving computational efficiency without sacrificing prediction accuracy.

3.2.2. Unsupervised Learning

In the case where data labels for anomalies are unavailable, a dataset with the same characteristics as the supervised one was created, consisting of 33 min and 32 s of normal behavior for a total of 10,000 data-points with a sampling rate of 5 Hz. This dataset was not divided, it was fully used for training and the supervised-one for testing.
To address the challenge of detecting unforeseen deviations, four different unsupervised models were tested, Local Outlier Factor (LOF), Autoencoders, One-Class Support Vector Machine (OC-SVM) and Isolation Forest (iForest). These models were trained exclusively on benign data. LOF identifies outliers by measuring the local density deviation of a data point relative to its neighbors. The Autoencoder uses a symmetric neural architecture to learn to deconstruct and reconstruct data, the reconstruction error is used as a metric to see the deviation of the data from normal behavior. OC-SVM projects the normal data into a high-dimensional feature space and then constructs an optimal separating hyperplane to differentiate normal from abnormal behavior. Lastly, iForest utilizes an ensemble of trees to isolate anomalies based on proximity.
The optimization of these models was similar to the supervised models. A TPE sampler was used to perform an automated hyperparameter search. The first 200 of the 500 trials were random. The total distribution of trials were 338 for LOF, 61 for the Autoencoder, 53 for OC-SVM and 48 for iForest. The anomaly detection threshold for the autoencoder was dynamically determined as a hyperparameter during training. The NN was trained for 100 epochs each time and the threshold parameter was defined as the nth percentile used as the boundary for normalcy.
The data was normalized as with the supervised dataset and its dimensionality reduced equally, with PCA, reducing the initial 37 variables to 10 (keeping 95% of variance explained). PCA was not necessary for the Autoencoder, as it already reduces the dimensionality of the data.

3.3. Adversary Emulation

This section describes the use of adversary emulation as a form of security testing within the OT environment, using the MITRE ATT&CK framework adapted specifically for Modbus over TCP, which is the communication protocol employed by the Industrial Control Systems (ICS) in our setup. We prioritize availability and integrity over confidentiality because interruptions to service can produce substantial economic losses in ICS contexts, and integrity breaches may create safety hazards when humans and robots interact. A breakdown of the relevant MITRE ATT&CK tactics is provided in Table 2.
These tactics have been conducted within the Modbus over TCP-based environment through an open source tool called ModTester. Table 3, based on our previous work [24], presents adversary emulation techniques against availability and integrity principles using the streamlined MITRE ATT&CK tactics for an ICS environment shown in Table 2. In the tool column, the attack tool is highlighted in bold, while the rest describes the attack vector. The tactic column is used in conjunction with the Table 2 ID column.
Availability attacks aimed at a mobile robot seek to make sensor signals inaccessible or prevent control signals from reaching the physical system [24]. To achieve this, denial-of-service techniques will be used, specifically the TCP SYN flood and a Modbus DoS attack that writes to multiple registers.
An integrity attack involves tampering with sensor measurements transmitted over the network or introducing unauthorized devices that impersonate legitimate ones. Its goal is to cause the robot to make decisions based on false information [24]. In this work, the integrity attacks implemented are data manipulation and an Adversary-in-the-Middle, both capable of injecting erroneous data into the robot’s data stream.

4. Results

This section provides a comprehensive analysis of the performance of our anomaly detection system. Based on the system architecture and infrastructure detailed in Section 3.1, adversary emulation has been carried out for both availability and integrity attacks, as described in Section 3.3. This section describes the results for both types of attacks and shows the performance of both supervised (Section 3.2.1) and unsupervised machine learning algorithms (Section 3.2.2) deployed in the MEC for the integrity attacks.

4.1. Availability Attacks

This section evaluates the resilience and detection ability of the system in a series of availability-oriented cyber-attacks. These attacks were selected to assess how network-level disruptions and resource exhaustion conditions affect both the operational stability of the robotic process and the performance of the anomaly detection model. All attacks were executed using the ModTester framework, following the configurations outlined in Table 3, which lists the specific tactics and techniques applied during experimentation.
The availability attacks included T0806—ModTester DoS: Multiple Register Write Attack, T0814—ModTester DoS: TCP SYN Flood Attack, T0846—ModTester Network Sniffing, and T0885—ModTester Connection to a Universally Known Port.
  • The first two, T0806 and T0814, represent active Denial-of-Service (DoS) scenarios targeting the control and telemetry channels. In the multiple-register-write attack, ModTester continuously issued unauthorized write commands to several Modbus registers, progressively saturating the communication channel. Similarly, the TCP SYN flood attack generated a large number of half-open TCP connections, exhausting the network and processing resources. As these attacks persisted, the robotic system exhibited gradual performance degradation. The command execution latency increased, control loops became unstable, and system resources such as CPU utilization, memory, and network buffers experienced significant strain. Ultimately, both attacks led to complete system unresponsiveness and service collapse.
  • In contrast, the other two attacks—T0846 and T0885—were passive or semi-passive reconnaissance operations. The network sniffing attack collected communication packets and telemetry data without actively disrupting the ongoing process, while the connection-to-a-universally-known-port technique probed network services to identify accessible endpoints and potential entry points. These activities provided valuable information on the infrastructure and network configuration of the system, but did not produce observable effects on the dynamics of the process or the availability of the system during execution.
None of the availability attacks were expected to be detected by the deployed anomaly detection model, as they are entirely based on process-level data—such as sensor readings, actuator positions, and control signals—while omitting any telemetry or communication metrics from the network-layer. Consequently, the network-based anomalies introduced by ModTester produced a minimal deviation within the process dataset, rendering them indistinguishable from normal operational fluctuations. This limitation underscores a critical shortcoming of process-only anomaly detection approaches: attacks that compromise system availability or target communication channels can severely impact performance and stability without generating detectable signatures in process-domain features. Future work should therefore incorporate network-level and protocol-aware features (e.g., connection statistics, packet rates, and Modbus register access patterns) to improve detection capabilities against availability and reconnaissance attacks.

4.2. Integrity Attacks

This section details the experiments that targeted data integrity in the robotic system and evaluates how the process-level anomaly-detection model responded. These integrity-focused tests complement the availability experiments described in the previous section (see Table 3) by exercising attack vectors that directly alter the observables of the process rather than primarily disrupting communication channels. All integrity scenarios were executed using the ModTester framework and configured according to the attack identifiers listed in Table 3.
Specifically, two integrity attacks were executed: T1565—ModTester Data Corruption through False Data Injection and T0830—ModTester Man-in-the-Middle (MiTM).
  • In the false-data-injection scenario (T1565), ModTester inserted spurious sensor readings and forged actuator feedback into the telemetry stream that the robotic arm sent to the Gateway; the attacker’s manipulations were concentrated at the Gateway entry point so that the data recorded downstream (at the Gateway, MEC, and Cloud logging services) no longer reflected the true physical state of the robot.
  • The MiTM scenario (T0830) involved intercepting and altering messages in transit specifically at or immediately before the Gateway: control commands and sensor reports were modified before being forwarded to downstream components. In both cases, the controller continued to issue the intended trajectory commands, while the recorded process logs and monitoring feeds diverged from the true system behavior as observed in the robotic arm.
The observable consequences of these integrity attacks were immediate and distinct from normal process variability. False data injection produced abrupt, implausible jumps in sensor channels (position, velocity, or force) and inconsistencies between correlated signals (for example, actuator command values that did not correspond to reported joint positions). The MiTM attack yielded similarly anomalous signatures: altered timestamps, mismatched command–response pairs, and temporally inconsistent sensor histories across the Gateway, MEC, and Cloud telemetry stores. Depending on the magnitude and timing of the injected values, these manipulations could produce simulated unexpected movements, transient control errors, or temporary loss of synchronization between the commanded and reported states.
Unlike with the availability experiments, the anomaly-detection models successfully identified integrity attacks. Because the detector was trained on process-domain data—sensor streams, actuator logs, and derived process features collected from the robotic arm’s telemetry as received at the Gateway and propagated through the MEC and Cloud logging pipeline—deliberate manipulations that produced out-of-distribution process signals generated clear, detectable deviations from learned normal behavior. The model raised alarms when falsified readings produced statistical inconsistencies (aberrant magnitudes, unrealistic dynamics, broken correlations across channels) or when the temporal patterns of the signals no longer matched those of the training set. In other words, integrity attacks that directly altered the robot’s process data at the Gateway produced detectable signatures within the process-only feature space and were therefore flagged by the deployed detection pipeline.
Nevertheless, detection of integrity attacks revealed important nuances and limitations. Gross or sustained data corruptions were detected reliably, but very small-magnitude manipulations or carefully crafted, stealthy injections that preserved local statistical properties over short windows could be more difficult to distinguish from legitimate noise or drift. Moreover, while the model flagged anomalies in the recorded data, detection alone does not guarantee safe remediation: some integrity manipulations could still induce unsafe actuator behavior before automated mitigation or human intervention was enacted. Finally, because Gateway-focused MiTM attacks can simultaneously affect both control and monitoring channels, correlating process-domain alarms with network and protocol-level indicators would improve confidence and enable faster attribution.
In summary, Gateway-targeted integrity manipulations implemented via ModTester (T1565 and T0830) produced clear deviations in process observables and were successfully detected by a detector trained in process-level features. These results underscore that models trained on process-domain data are effective at capturing attacks that corrupt sensor and actuator data at the data-collection ingress (the Gateway), while also highlighting the value of complementary defenses—such as sensor redundancy, signal validation, and the inclusion of network/protocol telemetry—to increase robustness against stealthy or multi-vector integrity threats.
Building upon these results, the subsequent analysis focuses on assessing the detection performance of different machine learning models when exposed to the described data integrity attacks. As detailed in Section 3.2.1 and Section 3.2.2, eight different models for two distinct ML architectures were deployed to evaluate their efficacy in detecting the simulated data integrity attacks. The comparative results of the best models of each type are summarized in Table 4.
The models were trained and optimized in an Ubuntu 22.04 server with two intel xeon gold 5420+ CPUs, two RTX A4500 graphic cards, and 256 GB of DDR5 RAM.
The results demonstrate that the supervised models always outperformed their counterparts in all precision metrics, as expected. For the supervised models, the LightGBM model outperforms the rest. For the unsupervised models, the Local Outlier Factor and the Autoencoder yield very similar results, but since the training time of LOF is significantly shorter, it can be considered the better option. The training time of both, LOF and OC-SVM are very low, but if resources are extremely limited in the system, it could be argued that OC-SVM is the better model, since its metrics are not that far behind LOF, but its training time is one quarter of it.
All models demonstrated high precision (≳0.95), indicating a near-absence of false positives, meaning that when the system classifies data as an attack, it most likely is. However, the lower recall rates (around 0.8 for supervised and 0.68 for unsupervised) suggest that some attacks are being misclassified as benign, especially in the unsupervised models.
It can be concluded that while supervised ensemble models, specifically LightGBM and XGBoost, achieve the highest overall performance with F1-scores of ≈0.89, they require labeled data that may not always be available in evolving 5G environments. Conversely, unsupervised models provide a robust alternative for anomaly detection, with the Autoencoder achieving a near-perfect precision of 0.991. This indicates an exceptional ability to minimize false positives, which is considered the most important task.

5. Conclusions

This study has presented a practical implementation of a scalable and robust robotic system in a 5G hybrid environment that prioritizes a low latency communication for critical industrial processes. The proposed architecture successfully uses ML-based anomaly detection algorithms based in process data to enhance the cybersecurity of the system, while keeping the required performance speed of the communications. Adversary emulation has been carried out within this realistic scenario based on the well-known MITRE ATT&CK framework, and the performance of the entire system in terms of resilience and detection ability has been evaluated. Both types of models (supervised and unsupervised) have been evaluated with acceptable results for integrity attack detection, as presented in Table 4.
Two primary limitations were identified. Firstly, the reliance on process-level data restricts the types of attacks that the system is able to identify to only those who physically alter the movement of the robotic arm. This highlights the necessity of integrating network-based data to extend the types of cyber-attacks that the system can detect. Secondly, using only the classification of the models seems insufficient for real classification of attacks. The models are not perfect and the frequency of the data is relatively high. Only one anomalous data point in multiple seconds of activity should not be enough to determine whether there is an attack or not, some kind of decision-making engine should be implemented so that the system can make more accurate decisions.
Further research will be focused on enhancing the system architecture for decreasing the latency of the communications within the components and automating the deployment and operation of the software components in hybrid 5G environments. Additionally, work will be done in the development of a proactive decision-making engine on top of the real-time anomaly detection system that will be extended by network-based traffic intrusion detection systems.

Author Contributions

Conceptualization, M.D.O. and A.D.P.; methodology, M.D.O.; software, M.D.O. and A.D.P.; validation, M.D.O. and A.D.P.; formal analysis, S.A.; investigation, M.D.O. and A.D.P.; data curation, R.R.-J.; writing—original draft preparation, M.D.O. and A.D.P.; writing—review and editing, R.R.-J., S.F.-L. and S.A.; visualization, M.D.O. and A.D.P.; supervision, S.F.-L.; project administration, S.A. and S.F.-L.; funding acquisition, S.A. All authors have read and agreed to the published version of the manuscript.

Funding

This research was partially funded by the AGENCIA ESTATAL DE INVESTIGACION through the PICRAH4.0 project—Plataforma inteligente y cibersegura para optimización adaptativa en la operación simultánea de robots autónomos heterogéneos with grant number PLEC2023-010353. The opinions, findings, conclusions, and suggestions articulated in this article are exclusive to the authors and do not reflect the perspectives of the sponsors.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are available on request from the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Lessi, C.C.; Gavrielides, A.; Solina, V.; Qiu, R.; Nicoletti, L.; Li, D. 5G and Beyond 5G Technologies Enabling Industry 5.0: Network Applications for Robotic. Procedia Comput. Sci. 2024, 232, 675–687. [Google Scholar] [CrossRef] [Scilit]
  2. Asavasirikulkij, C.; Mathong, C.; Sinthumongkolchai, T.; Chancharoen, R.; Asdomwised, W. Low Latency Peer to Peer Robot Wireless Communication with Edge Computing. In 2021 IEEE 11th International Conference on System Engineering and Technology, ICSET 2021-Proceedings; IEEE: Piscataway, NJ, USA, 2021; pp. 100–105. [Google Scholar] [CrossRef] [Scilit]
  3. Kadena, E.; Dai Nguyen, H.P.; Ruiz, L. Mobile Robots: An Overview of Data and Security. In Proceedings of the 7th International Conference on Information Systems Security and Privacy ICISSP; SciTePress: Setúbal, Portugal, 2021; pp. 291–299. [Google Scholar] [CrossRef] [Scilit]
  4. Lacava, G.; Marotta, A.; Martinelli, F.; Saracino, A.; La Marra, A.; Gil-Uriarte, E.; Mayoral-Vilches, V. Cybsersecurity issues in robotics. J. Wirel. Mob. Netw. Ubiquitous Comput. Dependable Appl. (JoWUA) 2021, 12, 1–28. [Google Scholar] [CrossRef]
  5. Jiang, X.; Lora, M.; Chattopadhyay, S. An Experimental Analysis of Security Vulnerabilities in Industrial IoT Devices. ACM Trans. Internet Technol. 2020, 20, 16. [Google Scholar] [CrossRef] [Scilit]
  6. Mittal, M.; Kumar, K.; Behal, S. Deep learning approaches for detecting DDoS attacks: A systematic review. Soft Comput. 2023, 27, 13039–13075. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Ziegler, V.; Schneider, P.; Viswanathan, H.; Montag, M.; Kanugovi, S.; Rezaki, A. Security and Trust in the 6G Era. IEEE Access 2021, 9, 142314–142327. [Google Scholar] [CrossRef] [Scilit]
  8. Shafique, K.; Khawaja, B.A.; Sabir, F.; Qazi, S.; Mustaqim, M. Internet of Things (IoT) for Next-Generation Smart Systems: A Review of Current Challenges, Future Trends and Prospects for Emerging 5G-IoT Scenarios. IEEE Access 2020, 8, 23022–23040. [Google Scholar] [CrossRef] [Scilit]
  9. Botta, A.; Rotbei, S.; Zinno, S.; Ventre, G. Cyber security of robots: A comprehensive survey. Intell. Syst. Appl. 2023, 18, 200237. [Google Scholar] [CrossRef] [Scilit]
  10. Ness, S.; Eswarakrishnan, V.; Sridharan, H.; Shinde, V.; Venkata Prasad Janapareddy, N.; Dhanawat, V. Anomaly Detection in Network Traffic Using Advanced Machine Learning Techniques. IEEE Access 2025, 13, 16133–16149. [Google Scholar] [CrossRef] [Scilit]
  11. Riggs, H.; Tufail, S.; Parvez, I.; Tariq, M.; Khan, M.A.; Amir, A.; Vuda, K.V.; Sarwat, A.I. Impact, Vulnerabilities, and Mitigation Strategies for Cyber-Secure Critical Infrastructure. Sensors 2023, 23, 4060. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Yaacoub, J.P.A.; Noura, H.N.; Salman, O.; Chehab, A. Robotics cyber security: Vulnerabilities, attacks, countermeasures, and recommendations. Int. J. Inf. Secur. 2022, 21, 115–158. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Sangoleye, F.; Johnson, J.; Eleni Tsiropoulou, E. Intrusion Detection in Industrial Control Systems Based on Deep Reinforcement Learning. IEEE Access 2024, 12, 151444–151459. [Google Scholar] [CrossRef] [Scilit]
  14. Apruzzese, G.; Laskov, P.; Montes de Oca, E.; Mallouli, W.; Brdalo Rapa, L.; Grammatopoulos, A.V.; Di Franco, F. The Role of Machine Learning in Cybersecurity. Digit. Threat. 2023, 4, 8. [Google Scholar] [CrossRef] [Scilit]
  15. Adejimi, A.; Sodiya, A.; Ojesanmi, O.; Falana, O.; Tinubu, C. A Dynamic Intrusion Detection System for Critical Information Infrastructure. Sci. Afr. 2023, 21, e01817. [Google Scholar] [CrossRef] [Scilit]
  16. ISO/IEC 30141:2024; Internet of Things (IoT)—Reference Architecture. ISO/IEC: Geneva, Switzerland, 2024.
  17. Collaborative Robotic Automation|Universal Robots Cobots. Available online: https://www.universal-robots.com/ (accessed on 4 February 2026).
  18. Universal Robots Offline Simulator Limits 2023. Available online: https://www.universal-robots.com/download/software-e-series/simulator-non-linux/offline-simulator-e-series-ur-sim-for-non-linux-5126-lts/ (accessed on 4 February 2026).
  19. Figueroa-Lorenzo, S.; Añorga, J.; Arrizabalaga, S. A survey of IIoT protocols: A measure of vulnerability risk analysis based on CVSS. ACM Comput. Surv. (CSUR) 2020, 53, 1–53. [Google Scholar] [CrossRef] [Scilit]
  20. Chen, B.; Pattanaik, N.; Goulart, A.; Butler-Purry, K.L.; Kundur, D. Implementing attacks for modbus/TCP protocol in a real-time cyber physical system test bed. In Proceedings-CQR 2015: 2015 IEEE International Workshop Technical Committee on Communications Quality and Reliability; Curran Associates, Inc.: Red Hook, NY, USA, 2015. [Google Scholar] [CrossRef] [Scilit]
  21. Bhatia, S.; Kush, N.S.; Djamaludin, C.; Akande, A.J.; Foo, E. Practical modbus flooding attack and detection. In Proceedings of the Twelfth Australasian Information Security Conference (AISC 2014); Conferences in Research and Practice in Information Technology; Australian Computer Society: Sydney, Australia, 2014; Volume 149, pp. 57–65. [Google Scholar]
  22. Kreps, J.; Narkhede, N.; Rao, J. Kafka: A distributed messaging system for log processing. In Proceedings of the NetDB, Athens, Greece, 12 June 2011; Volume 11, pp. 1–7. [Google Scholar]
  23. Wang, X.; Zhao, H.; Zhu, J. GRPC: A communication cooperation mechanism in distributed systems. SIGOPS Oper. Syst. Rev. 1993, 27, 75–86. [Google Scholar] [CrossRef] [Scilit]
  24. Machaka, V.; Figueroa-Lorenzo, S.; Arrizabalaga, S.; Hernantes, J. Comparative analysis of the standalone and Hybrid SDN solutions for early detection of network channel attacks in Industrial Control Systems: A WWTP case study. Internet Things 2024, 28, 101413. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Deployed architecture overview. It consists of four main components: the Field Device, Gateway, Multi-access Edge Computing (MEC), and Cloud.
Figure 1. Deployed architecture overview. It consists of four main components: the Field Device, Gateway, Multi-access Edge Computing (MEC), and Cloud.
Technologies 14 00108 g001
Table 1. Categorization of the 37 Sensor Variables.
Table 1. Categorization of the 37 Sensor Variables.
Sensor CategorySub-CategoryVariables (Units)Count
Kinematic StateJoint AngleBase, Shoulder, Elbow,
Wrist1, Wrist2, Wrist3 (mrad)
6
Joint VelocityBase, Shoulder, Elbow,
Wrist1, Wrist2, Wrist3 (mrad/s)
6
Joint RevolutionBase, Shoulder, Elbow,
Wrist1, Wrist2, Wrist3 (Count)
6
Positionx, y (Position Units)2
Subtotal (Kinematic)20
Dynamic ProfileJoint CurrentBase, Shoulder, Elbow,
Wrist1, Wrist2, Wrist3 (mA)
6
System and Tool StatusRobot Current, I/O Current,
Tool Current (mA), Tool State (Status)
4
Subtotal (Dynamic)10
Thermal ProfileJoint TemperatureBase, Shoulder, Elbow,
Wrist1, Wrist2, Wrist3 (°C)
6
Tool TemperatureTool Temperature (°C)1
Subtotal (Thermal)7
Total sensor variables37
Table 2. MITRE ATT&CK Tactics for ICS.
Table 2. MITRE ATT&CK Tactics for ICS.
IDNameDescription
T0806Brute Force I/OAdversaries may repetitively or successively change I/O point values to perform an action. Brute Force I/O may be achieved by changing either a range of I/O point values or a single point value repeatedly to manipulate a process function.
T0846Remote System DiscoveryAdversaries may attempt to get a listing of other systems by IP address, hostname, or other logical identifier on a network.
T0814Denial of ServiceAdversaries may perform Denial-of-Service (DoS) attacks to disrupt expected device functionality.
T0885Commonly Used PortAdversaries may communicate over a commonly used port to bypass firewalls or network detection systems and to blend in with normal network activity, to avoid more detailed inspection.
T0830Adversary-in-the-MiddleAdversaries with privileged network access may seek to modify network traffic in real time using adversary-in-the-middle (AiTM) attacks.
T1565Data ManipulationAdversaries may insert, delete, or manipulate data in order to influence external outcomes or hide activity, thus threatening the integrity of the data.
Table 3. Design of adversary emulation (adapted from [24]).
Table 3. Design of adversary emulation (adapted from [24]).
PrincipleTacticTool
AvailabilityT0806ModTester DOS Multiple Register Write Attack
T0814ModTester DOS TCP SYN flood attack
T0846ModTester Network sniffing
T0885ModTester Connection to a universally known port
IntegrityT1565ModTester Data corruption via False Data Injection
T0830ModTester MiTM
Table 4. Performance of the models.
Table 4. Performance of the models.
ML TypeML ModelAccuracyPrecisionRecallF1Train Time (s)
SupervisedLightGBM0.8670.9930.8060.8900.658
SupervisedXGBoost0.8650.9940.8030.8880.457
SupervisedR.F.0.8640.9940.8010.8870.872
SupervisedDense NN0.8640.9940.8010.88740.857
UnsupervisedLOF0.8270.9440.6940.8000.239
UnsupervisedAutoencoder0.8320.9910.6710.80098.821
UnsupervisedOC-SVM0.8230.9330.6970.7980.064
UnsupervisediForest0.8230.9330.6960.7973.103
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Dean Oses, M.; Domec Paz, A.; Figueroa-Lorenzo, S.; Arrizabalaga, S.; Rodriguez-Jorge, R. Anomaly Detection Using Machine Learning for Robotics Environments on 5G Networks. Technologies 2026, 14, 108. https://doi.org/10.3390/technologies14020108

AMA Style

Dean Oses M, Domec Paz A, Figueroa-Lorenzo S, Arrizabalaga S, Rodriguez-Jorge R. Anomaly Detection Using Machine Learning for Robotics Environments on 5G Networks. Technologies. 2026; 14(2):108. https://doi.org/10.3390/technologies14020108

Chicago/Turabian Style

Dean Oses, Mikel, Aitor Domec Paz, Santiago Figueroa-Lorenzo, Saioa Arrizabalaga, and Ricardo Rodriguez-Jorge. 2026. "Anomaly Detection Using Machine Learning for Robotics Environments on 5G Networks" Technologies 14, no. 2: 108. https://doi.org/10.3390/technologies14020108

APA Style

Dean Oses, M., Domec Paz, A., Figueroa-Lorenzo, S., Arrizabalaga, S., & Rodriguez-Jorge, R. (2026). Anomaly Detection Using Machine Learning for Robotics Environments on 5G Networks. Technologies, 14(2), 108. https://doi.org/10.3390/technologies14020108

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop