Next Article in Journal
Observer Pre-Synchronization for Smooth Transition in Wide-Speed-Range Sensorless Control of IPMSM Drives
Previous Article in Journal
Synchronization of Chaotic Buck Converters via Control-Signal Injection
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Integrating Real Tool Interaction and Multimodal Operator Monitoring in Immersive Simulation for Human-Centred Assessment

Dipartimento di Ingegneria Industriale e dell’Informazione, Università degli Studi di Pavia, Via A. Ferrata 5, 27100 Pavia, Italy
*
Author to whom correspondence should be addressed.
Electronics 2026, 15(16), 3525; https://doi.org/10.3390/electronics15163525
Submission received: 31 March 2026 / Revised: 4 August 2026 / Accepted: 5 August 2026 / Published: 8 August 2026

Abstract

The transition from Industry 4.0 to Industry 5.0 is increasing the need for design approaches that place human centrality, safety, and ergonomics at the core of system development. In this context, immersive simulation is evolving from a training-oriented technology into a controlled environment for the observation of human behaviour and task execution. However, many Virtual Reality applications still rely on generic interaction devices that are not sufficiently representative of real tool-mediated operations, limiting the reliability of ergonomic and behavioural assessment. This paper proposes an anthropocentric framework for human-centred assessment in immersive simulation. The framework integrates four main components: a task-oriented Scenario Digital Twin, a Physical-Tool-in-the-Loop module based on a real instrument synchronised with its virtual counterpart, an Interaction Engine for state-dependent action management, and an Operator-in-the-Loop module coupled with a Human-Centred Assessment Layer. The framework is instantiated through an immersive hazelnut pruning simulator, selected as a representative case study because pruning involves non-neutral postures, irreversible actions, and strong dependence on tool handling. The reference implementation combines reality-based reconstruction of hazelnut trees, interactive branch-cutting logic, an instrumented electric pruning shear, full upper-body embodiment, and multimodal data acquisition. In particular, the proposed architecture supports the extraction of body motion through markerless multi-camera pose estimation and the acquisition of eye-related variables through the head-mounted display. A preliminary experimental session demonstrates the technical feasibility of the proposed architecture by showing that multimodal data, including reconstructed 3D body kinematics and eye-tracking signals, can be successfully acquired during immersive task execution, providing a basis for subsequent ergonomic analysis. The results show that a real instrumented tool, a task-oriented digital twin, and a continuous monitoring pipeline can be integrated within a single immersive platform, as well as that the assessment pipeline is sensitive to differing postural configurations during simulated task execution. They do not establish equivalence between behaviour in the simulator and behaviour during real-world pruning; dedicated cross-modal validation is therefore required. The framework is intended to support the evolution of immersive simulation toward assessment-oriented applications, contributing to a broader Safety-by-Design perspective in which human behaviour and interaction quality are considered as early-stage design inputs.

1. Introduction

The transition from Industry 4.0 to Industry 5.0 is redefining how industrial systems are designed and validated. Beyond automation and productivity, current approaches must account for human centrality, resilience, and social sustainability [1]. In this perspective, Virtual Commissioning (VC) is no longer only a tool for validating the system dynamic behaviour or the control logic but an environment in which operator interaction, usability, and procedure-related criticalities can be assessed before physical deployment [2].
This shift is closely connected to Safety-by-Design: in conventional workflows, many safety and ergonomic issues emerge only after prototyping or pilot commissioning, when corrective actions are costly and constrained. Immersive simulation, instead, enables early exploration of critical scenarios and turns safety into an anticipatory design variable [3]. Accordingly, the digital twin becomes a space where operator behavior, technical-system response, and workspace constraints can be analysed jointly, supporting proactive risk identification [4].
The need for this approach is amplified by increasingly dynamic work environments shared with automation. In mixed human–robot settings, such as industrial applications involving autonomous mobile robots, human behavior is harder to predict and therefore more relevant for safety-oriented design [5]. Similar requirements also emerge in high-risk sectors and in agriculture, where semi-autonomous machinery requires early evaluation of tool-mediated human actions in complex outdoor contexts [6]. In this context, immersive simulation offers a controlled way to reproduce hazardous or cognitively demanding situations before real-world deployment.
Over the last decade, XR technologies have shown strong potential for training and procedural rehearsal [7]. In particular, VR has helped overcome several structural limits of field-based training, including cost, repeatability constraints, and exposure to unsafe conditions [8,9]. However, most VR applications remain primarily oriented to skill acquisition and familiarization. The present work addresses a complementary objective: the design of an immersive platform that integrates high-fidelity tool interaction and continuous operator monitoring, so that the simulation can also serve as an assessment environment in which human behaviour and task execution are observable and measurable.
A major limitation of existing systems concerns interaction fidelity. Standard controllers, gesture-based interfaces, and generic haptic devices are often insufficient to reproduce task-specific physical interaction conditions, particularly in terms of contact realism, force rendering, compliance, and object-specific haptic cues [10,11]. When physical interaction differs substantially from that required by the target task, posture, muscular effort, and motion strategy may also differ, limiting the interpretation of ergonomic and safety-related analyses. For this reason, integrating dedicated real tools into the virtual loop becomes a key requirement for assessment-oriented simulation.
In a Physical-Tool-in-the-Loop (PTL) configuration, the operator manipulates an instrumented real tool synchronised with its virtual counterpart, preserving task-congruent constraints and action dynamics [12,13]. In parallel, advances in AI-based Human Pose Estimation (HPE) [14,15] make it possible to move from static observational ergonomics toward continuous postural monitoring and dynamic indicator extraction during task execution [16]. Recent studies also suggest that immersive systems can provide behavioral and physiological cues related to cognitive load and stress, further supporting human-centred assessment [17,18].
Despite these advances, physical interaction fidelity, operator monitoring, and ergonomic assessment are still frequently addressed as separate dimensions in the literature. This paper proposes an integrated framework that combines these three components within a single immersive platform and demonstrates its feasibility through a reference implementation based on a hazelnut pruning simulator. The contribution is methodological: the framework defines an architecture and a set of design principles, instantiates them in a functional system, and shows that multimodal data relevant to ergonomic analysis can be acquired continuously during immersive task execution. The present study evaluates technical feasibility and within-simulation sensitivity to different postural configurations; it does not establish ecological validity or equivalence with real task execution. Cross-modal comparison with real-world pruning is therefore identified as future work.
The remainder of the paper is organised as follows: Section 2 reviews the relevant literature. Section 3 presents the proposed framework. Section 4 describes the reference implementation. Section 5 reports the preliminary experimental results. Section 6 concludes the paper and outlines future directions.

2. Background and Literature Analysis

The recent diffusion of immersive technologies in industrial and agro-industrial contexts has progressively expanded the role of Virtual Reality (VR) from a purely demonstrative medium to an operational environment for training, rehearsal, and task-oriented interaction. In agriculture, this trend is especially relevant because workforce preparation is still affected by structural limitations such as seasonality, scarcity of skilled labor, safety concerns, and the economic cost of repeated field-based instruction. In this sense, immersive simulation is increasingly regarded as a means to decouple learning activities from the physical availability of crops, tools, and favorable environmental conditions, while also enabling more repeatable and controllable training sessions [19].
Within this broader context, the literature available in the documents considered here shows that agriculture is no longer marginal in the XR domain, even if it still lags behind more consolidated sectors such as manufacturing and healthcare. A review reports that a large share of XR applications in agro-industry has been developed for educational and training purposes, with Virtual Reality being the dominant technological choice thanks to its immersive capabilities [20]. In parallel, the more general literature on immersive virtual training indicates that high-fidelity environments tend to improve user engagement, sense of presence, and learning effectiveness more than lower-fidelity or desktop-based solutions [21].

2.1. From Generic Haptics to Task-Specific Physical Interfaces

A recurrent limitation of immersive systems concerns the fidelity of the interaction layer. In many XR applications, interaction is still mediated by standard controllers, gesture-based commands, or generic haptic devices that are effective for basic object manipulation, yet only partially representative of the physical conditions of real manual tasks. Recent studies on virtual grasping have shown that interaction modalities entail a structural trade-off between precision, naturalness, workload, and presence. Controller-based solutions often provide greater reliability and efficiency, whereas hand tracking tends to enhance intuitive use and perceived naturalness, sometimes at the cost of reduced stability or increased effort. Interaction fidelity, therefore, should not be regarded as a purely technological property of the interface but evaluated against the physical and functional logic of the simulated action [11].
This issue becomes even more evident in the haptic domain. Although contemporary immersive systems can provide convincing visual and auditory immersion, their ability to reproduce the physical qualities of contact remains limited. Recent reviews have highlighted that insufficient haptic fidelity and reduced biomechanical realism are still major barriers to the use of VR/AR for ergonomic assessment and workspace-oriented simulation [22]. When the interface fails to reproduce relevant task constraints, such as contact configuration, resistance, compliance, or object-specific manipulation cues, user behavior may deviate substantially from real execution conditions. Accordingly, comparative studies have emphasised the need for richer haptic rendering and more physically coherent interaction design when the virtual environment is intended not only for immersion but also for analysis and evaluation [10].
From a methodological perspective, the literature is progressively moving away from symbolic or controller-triggered interaction toward more structured forms of task-specific embodiment. A relevant example is offered by recent work on reconfigurable glove-based systems for physical and virtual grasp reconstruction. In that line of research, conventional symbolic grasping is explicitly criticised for attaching the virtual object to the hand once a grasp event is triggered, thus producing interactions with limited realism and weak correspondence to actual contact conditions. To address this limitation, the proposed architecture reconstructs fine-grained hand motion, detects stable grasp configurations through collision geometry, and provides haptic feedback during interaction. It is further extended to capture contact points, force information, and tool-related action effects, moving toward a richer representation of manipulation events [23].
A similar tendency can be observed in controller-free procedural simulators based on hand tracking. For example, a VR training system for dental local anesthesia used Leap Motion to support direct syringe manipulation through gesture-based interaction, allowing the user to control tool orientation, grip state, and injection activation without standard VR controllers. Although this approach does not reproduce the full mechanical behavior of a real instrument, it shows that the interaction layer can be shaped around the operational logic of a specific tool-mediated procedure rather than around generic input metaphors [24].
For design-oriented immersive simulation, this transition is particularly significant. Generic interfaces may still be acceptable for familiarization or procedural rehearsal, but when the objective shifts toward ergonomic assessment, human-centred validation, or the comparison of alternative design solutions, the physical coherence of the interaction becomes a methodological requirement. In such cases, task-specific physical interfaces, dedicated props, instrumented tools, or hybrid physical–virtual interaction architectures should be regarded not merely as devices for increasing realism but as components intended to reduce interaction mismatch and support subsequent investigation of ecological validity, posture, movement strategy, contact behavior, and human–system compatibility [2,4,22].
Taken together, these studies suggest that the transition from generic haptics to task-specific physical interfaces marks a key step in the evolution of immersive simulation. The relevant question is no longer only whether users can interact within a virtual environment but whether the simulated constraints are sufficiently aligned with the target task and subsequently validated to support ergonomic, behavioral, and design-oriented interpretations.

2.2. Toward Ergonomic Assessment and Design-Oriented Simulation

A further step in the evolution of immersive systems concerns their use not only for training or procedural rehearsal but also for ergonomic assessment and design-oriented validation. In this perspective, VR and XR environments are increasingly regarded as digital workspaces in which posture, task execution, and human–system interaction can be analysed before physical deployment. Recent reviews have emphasised this transition, showing that immersive technologies are progressively being adopted for workspace design, ergonomic analysis, and human-centred evaluation, while also highlighting that their validity depends on the realism of the interaction and on the quality of the measurement pipeline [22]. This shift is particularly relevant with respect to traditional observational methods such as RULA (Rapid Upper Limb Assessment) and REBA (Rapid Entire Body Assessment), which, although still widely used, remain inherently static, evaluator-dependent, and often unable to capture the temporal dynamics of risk exposure during task execution.
Recent work based on computer vision and markerless tracking has started to address this limitation by transforming ergonomic assessment into a continuous and automated process. A notable example is provided by a bio-inspired monitoring system that combines 3D pose estimation, depth sensing, and real-time RULA scoring to evaluate posture continuously during dynamic load-handling tasks, showing that critical postural risk emerges more clearly during movement than in static conditions [25]. At the same time, the literature shows that the validity of virtual ergonomic assessment cannot be taken for granted. Comparative studies on manual assembly tasks performed in both real and virtual settings suggest that immersive VR can support early-stage ergonomic analysis but may underestimate postural risk when tasks involve high physical demands or precise manual alignment. In such cases, the absence of physical feedback affects execution strategy and reduces the representativeness of the measured posture, confirming that ergonomic assessment in immersive environments is closely intertwined with interaction fidelity and with the extent to which the simulated task preserves real execution constraints [26].
Another relevant development concerns the integration of AI-based sensing pipelines that extend beyond purely postural metrics. Recent studies suggest that immersive systems can increasingly support the extraction of behavioral and physiological indicators related to user state during task execution. Although developed in a clinical domain, validation studies on VR-based eye-tracking platforms have shown that headset-integrated eye and pupil measurements can achieve good reproducibility and user acceptability, confirming the feasibility of embedding objective monitoring functions directly into immersive systems [27]. In parallel, VR-based stress monitoring has shown that behavioral cues observed in controlled immersive environments can be exploited for real-time stress detection, especially when complemented by lightweight physiological sensing [17]. Taken together, these studies indicate that immersive simulation is progressively moving toward a more comprehensive role in human-centred assessment, in which ergonomic scoring, markerless HPE, and behavioral or physiological monitoring can provide complementary evidence for design-oriented analysis and, following appropriate real-world validation, future predictive safety assessment [2,22].

2.3. XR-Based Training in Agriculture: From Early Applications to Immersive Simulators

The first wave of agricultural VR applications was mainly oriented toward awareness, basic education, or simplified procedural familiarization. Early examples such as [28,29] introduced virtual agricultural scenarios through screen-based interfaces and conventional controllers, allowing users to explore farming-related tasks and concepts without requiring direct access to the field. These systems were important as proof-of-concept platforms, but their interaction model remained limited in terms of embodiment, realism, and motor transfer.
Subsequent developments moved toward more explicit training objectives. VR environments have been designed for agricultural surveyors and greenhouse operations, where users could practice observation, environmental control, and operational procedures in simulated yet more structured scenarios [30,31,32]. These works demonstrated that VR can support the rehearsal of practical skills in a protected environment, while reducing the risk of crop damage and improving accessibility to complex or resource-intensive learning experiences.
Pruning represents a particularly meaningful case within this landscape because it combines procedural knowledge, biomechanical effort, spatial reasoning, and irreversible action. Once a cut is performed in the real world, it cannot be undone, and ineffective training may damage the plant itself. For this reason, pruning has been identified as a task that can strongly benefit from simulation-based learning. The literature already includes early virtual pruning systems, such as the orchard simulator [33] and the vine-pruning solution based on Leap Motion proposed by [34]. More recent commercial solutions, including [35,36], have further increased the level of immersion by introducing head-mounted displays and more structured task guidance. However, even these systems still appear mainly conceived as training tools, rather than as platforms for predictive assessment of interaction quality, ergonomic risk, or design validation.
A recurring issue emerging from both uploaded documents is that realism in agricultural VR cannot be reduced to visual quality alone. Systems based on monitors, handheld controllers, or simplified gesture interaction may support procedural familiarization, but they often remain insufficient when the goal is to reproduce the physical constraints associated with manipulating vegetation and tools. This distinction is particularly important for the present work, which is not limited to the question of whether VR can train operators but asks which conditions and validation steps are required before information about operator behaviour and task execution in simulation can be interpreted in relation to the real task.

3. Anthropocentric Framework for Human-Centred Assessment in Immersive Simulation

This work proposes a general framework for anthropocentric human-centered assessment in immersive simulation where the simulator is no longer interpreted merely as a training environment but as an assessment-oriented design platform in which operator behaviour becomes an observable component of the commissioning process. Predictive use would require subsequent validation against corresponding real-world tasks.
The framework is intentionally general and transferable across different domains, including manufacturing, infrastructure maintenance, and agriculture. Its purpose is to support the early validation of scenarios in which effectiveness and safety depend not only on technical performance but also on the way operators move, interact, and respond to the surrounding environment and to the tools involved in the task.

3.1. Design Principles

The proposed framework is grounded in four complementary principles that define how immersive simulation can provide assessment data for safety-oriented design and future predictive validation:
  • Environmental fidelity: The simulator must reproduce the operative context with sufficient realism to induce credible operator behaviour. This includes not only visual appearance but also spatial structure, task-relevant objects, operative constraints, and the affordances that guide action.
  • Interaction fidelity: Realism also concerns biomechanical and procedural coherence. When the task involves tools or machine interfaces, the interaction model must preserve the constraints that shape movement strategy, body positioning, and action timing.
  • Human observability: The operator is treated as an observable component of the system rather than a simple user immersed in the scene. Posture, motion, and task execution must therefore be monitored continuously so that ergonomic and safety-related variables can be extracted together with task outcomes.
  • Design transferability: The evidence gathered during simulation must be translated into concrete engineering decisions. The outputs of immersive execution should therefore support the redesign of tools, layouts, interfaces, and procedures, embedding safety and human factors directly into the design process.

3.2. Framework Architecture

Based on these principles, the framework is structured as an integrated architecture that connects design inputs, an anthropocentric immersive simulation core, a human-centred assessment layer, and design outputs within an iterative loop, as shown in Figure 1. In this way, simulation is not treated as an isolated virtual environment but as a closed design cycle in which task requirements, operator behaviour, and assessment results jointly support system refinement before physical implementation.
The framework starts from the design inputs, which include the task or procedure to be simulated, the tool or machine involved, the spatial organization of the workspace, the relevant safety requirements, and the ergonomic targets motivating the analysis. Their role is to define the validation problem not only in functional terms but also with respect to the environmental, procedural, and human-centred constraints under which the system must be evaluated.
At the centre of the architecture lies the anthropocentric immersive simulation core, which integrates four tightly coupled modules. The first is the Scenario Digital Twin, representing the operative context through the environment, the relevant objects or machines, the task constraints, and the surrounding conditions. Its purpose is not limited to visual reconstruction but to generate a task-oriented virtual context capable of supporting interaction, event generation, and analysis. Depending on the application domain, this module may derive from CAD assets, structured scene modelling, or reality-based acquisition workflows.
The second module is the Physical-Tool-in-the-Loop, which introduces the real or sensorised tool as an active component of the simulation. In this configuration, the operator interacts with a physical interface synchronised with its virtual counterpart, so that the action is performed under constraints closer to those of the target task than when generic controllers are used. More specifically, rather than triggering a virtual action through an abstract controller-based command, the operator manipulates the actual task-specific tool, whose relevant operational state is acquired and synchronised in real time with its digital counterpart. This configuration preserves physical and biomechanical characteristics that are typically absent in generic controller-based interaction, including grip geometry, mass distribution, inertia, trigger mechanics, tool orientation, and the hand–tool relationship. Since these characteristics can affect reaching strategy, upper-limb configuration, body positioning, and action timing, the PTL paradigm is intended to reduce interaction mismatch and limit artefacts introduced by unrelated input devices. For this reason, the framework gives a central role to dedicated physical tools when the objective is assessment-oriented safety and ergonomic analysis [2,22]. However, whether this increased physical coherence produces behaviour corresponding to real-world task execution remains to be established through dedicated comparative validation.
The third module is the Interaction Engine, which governs the dynamic coupling between the digital twin, the physical tool, and the operator. As shown in Figure 1, it handles events and collisions, updates task states, and generates action outcomes. It therefore constitutes the computational layer through which the environment responds to operator behaviour and through which the consequences of each action are propagated consistently across the virtual scene.
The fourth module is the Operator-in-the-Loop, which concerns the embodiment and monitoring of the human participant within the simulation. It includes avatar embodiment, hand and body tracking, posture and motion reconstruction, and task execution as an observable process. In the proposed framework, the operator is not treated as an external user of the simulator but as an integral element of the commissioned system, whose movement patterns, positioning, and interaction strategies directly affect the simulated outcome.
The data generated by the core are then processed by the Human-Centred Assessment Layer, which transforms immersive execution into measurable evidence. This layer supports the extraction of performance metrics, the identification of unsafe states or critical events, the computation of ergonomic indicators such as RULA- or REBA-like scores, and the estimation of postural quality or workload-related proxies.
Finally, the assessment results are translated into design outputs, including the redesign of tools, layouts, interfaces, and procedures. This passage closes the iterative Safety-by-Design loop shown in Figure 1, through which the evidence gathered during immersive simulation feeds back into the initial design assumptions.

4. An Immersive Pruning Simulator as Reference Implementation

To demonstrate the practical applicability of the framework introduced in the previous section, this paper considers a reference implementation developed for immersive hazelnut pruning simulation [8,9]. Although the demonstrator originates from an agricultural training use case, its architecture was conceived to support a broader class of operations in which task-specific tool interaction, operator embodiment, and detailed environmental representation are required. In this sense, the simulator should not be interpreted merely as a standalone VR application but rather as a concrete instantiation of the proposed anthropocentric human-centered assessment in immersive simulation paradigm, in which the digital twin of the environment, the physical interaction tool, and the human monitoring layer are integrated within a single experimental platform.
Pruning is strongly tool-dependent, requires interaction with targets located at different heights and orientations, and often involves non-neutral postures. Moreover, each cut is irreversible, so execution quality depends directly on the operator’s movement strategy and local body positioning. Traditional training is also constrained by seasonality, resource consumption, and the risk of damaging real plants during the learning phase. For these reasons, the pruning simulator provides a meaningful case study for testing an assessment-oriented architecture and its continuous monitoring capabilities. The application allows users to interact with detailed digital reconstructions of hazelnut trees, receive guidance from a virtual trainer, and manipulate a dedicated physical pruning tool synchronised with its virtual counterpart. The current implementation was developed in Unreal Engine 5.3 and designed for SteamVR-compatible head-mounted displays, with the Varjo XR-4 adopted in the most recent configuration.
As shown in Figure 2, the simulator can therefore be interpreted as a reference instantiation of the proposed framework, in which the general layers of the architecture are translated into application-specific components for immersive execution and human-centred assessment before real deployment. The Scenario Digital Twin is instantiated through a photogrammetric and LiDAR-based reconstruction of real hazelnut trees integrated into a virtual agricultural environment. The Interaction Engine corresponds to the Unreal Engine 5.3 runtime, which manages collision detection, branch-cutting logic, and task state updates. The Physical-Tool-in-the-Loop module is realised through the instrumented C3X electric pruning shear by Pellenc, whose blade opening signal is wirelessly synchronised with its virtual counterpart in real time. The Operator-in-the-Loop module provides full upper-body avatar embodiment driven by head and hand tracking, embedding the operator as an active and observable element of the simulation. Finally, the Human-Centred Assessment Layer integrates three concurrent data streams: ocular variables acquired through the head-mounted display, including pupil diameter and gaze direction; three-dimensional body kinematics reconstructed through a markerless multi-camera pose estimation pipeline; and the resulting ergonomic indicators derived from the reconstructed joint angles.

4.1. Scenario Digital Twin: Reconstruction Pipeline and Virtual Environment

In the present case study, the scenario digital twin module was implemented through a pipeline that combined reality-based acquisition, tree model parametrization, and integration into an immersive virtual environment for guided pruning execution. As shown in Figure 3, the workflow started from an acquisition campaign on real hazelnut trees and then progressed through photogrammetric and laser-data alignment, scan-to-model parametrization, foliage generation, collision setup, and final insertion into the virtual agricultural scene.
The main requirement of this module was to preserve the geometric identity of the real trees while transforming the reconstruction into a controllable asset suitable for simulation. For this reason, the surveyed hazelnut specimens were acquired during the winter season at the Università Cattolica del Sacro Cuore campus in Piacenza through LiDAR scanning and image-based photogrammetry, using an RTC360 laser scanner (Leica Geosystems AG, Heerbrugg, Switzerland) and a D5600 digital camera (Nikon Corporation, Tokyo, Japan) [37]. The acquired datasets were processed to obtain a dense and metrically coherent digital representation of the trees, followed by cleaning and refinement steps aimed at preserving morphological fidelity while removing reconstruction artifacts. Previous validation of the workflow showed millimetric agreement between laser and photogrammetric data over most of the structure, confirming that the resulting model could reliably reproduce the size and geometry of the real specimen for simulator use [9,38].
Since a static high-resolution mesh was not sufficient for branch-level interaction and cutting events, the surveyed geometry was subsequently converted into a parametric model through a scan-to-model workflow implemented in SpeedTree. This step enabled the generation of a controllable tree structure that preserved the appearance of the original specimen while supporting live manipulation in the simulator [39]. The same logic was extended to foliage generation and collision placement, so that the digital twin could be used not only as a realistic visual asset but also as an interactive object within the pruning task.
The resulting tree models were then integrated into an Unreal Engine scene representing the agricultural environment. The virtual setting was conceived not as a generic background but as an operative context supporting perception, orientation, and task execution. In this way, the Scenario Digital Twin module combines realistic replication and task-oriented scene construction, providing the environmental basis for the subsequent modules of physical interaction, operator embodiment, and human-centred assessment.

4.2. Physical-Tool-in-the-Loop Module: Instrumented Pruning Shear and Virtual Counterpart

Within the proposed framework, the Physical-Tool-in-the-Loop module introduces a task-specific physical interface into the immersive simulation through an instrumented electric pruning shear synchronised with its virtual counterpart. In the present case study, this solution preserved more tool-specific handling characteristics than the earlier glove-centred approach adopted in the first prototype. In that initial version, pruning interaction was mediated by a haptic glove, whose normalised finger flexion values were used to reconstruct hand motion and control tool actuation. Although effective for testing tactile feedback and hand animation, that solution did not reproduce the handling characteristics of an actual pruning instrument, motivating the transition toward a physical-tool interface [8,9].
The selected device is a modified C3X electric pruning shear (PELLENC SAS, Pertuis, France), instrumented for real-time communication with the simulator. The tool was adapted to extract the blade opening state continuously, generating a signal ranging from fully open to fully closed (0–1). This signal is acquired locally and transmitted wirelessly to the workstation running the immersive application. In the current implementation, the communication chain includes serial signal conversion, local acquisition through a Raspberry Pi 5 (Raspberry Pi Ltd., Cambridge, United Kingdom), and UDP transmission to the host PC, allowing the tool to operate as an active component of the simulation loop without cumbersome wired connections around the user. This design choice allows the operator to manipulate the task-specific tool rather than an abstract handheld controller, thereby retaining its grip geometry, trigger mechanics, mass distribution, and actuation pattern. For safe use in the simulator, the blades were removed; consequently, the setup preserves selected handling characteristics but does not reproduce physical branch contact.
An additional point concerns haptic perception during branch cutting. The selected pruning shear uses a servo-assisted cutting mechanism, which is expected to reduce the branch reaction force transmitted to the operator’s hand during real operation. However, the present study did not measure hand-level forces during real pruning or compare them with the simulated condition; the physical tool preserves selected properties, including inertia, grip geometry, trigger mechanics, and actuation pattern, and thus reduces the mismatch associated with a generic controller. Nevertheless, the absence of the branch and blades means that contact-force correspondence remains unverified. Dedicated force and torque measurements, together with comparative trials in the corresponding real task, are required to quantify this gap. Applications involving larger manual or contact forces may additionally require an active force-feedback mechanism.
The real pruning shear is associated with a virtual counterpart implemented in Unreal Engine as an articulated actor. The digital tool is modeled as a multi-part object including the main body, trigger, and blade, connected through constrained kinematic relationships so that only the intended rotational degrees of freedom are allowed. Angular limits were defined to reproduce the mechanical range of the real instrument, ensuring consistency between the transmitted tool state and its graphical representation in the virtual scene. As shown in Figure 4, the module therefore combines the instrumented physical shear with a coherent digital replica, in which the main moving parts and their reference rotation axes are explicitly modeled to reproduce the kinematics of the real device.

4.3. Interaction Engine: Runtime Cutting Logic and Action Outcomes

Within the proposed framework, the Interaction Engine module is responsible for translating tool–environment contact into task states and action outcomes. In the present case study, its main role is to manage the runtime cutting logic that transforms the hazelnut tree digital twin from a static visual asset into an interactive object whose geometry evolves coherently with the operator’s action.
The cutting procedure is based on the interaction between the collision box positioned between the blades of the virtual pruning shear (Figure 4, right) and the hierarchical representation of the tree. The tree actor is organised as a parent–child structure of static branch meshes, each associated with a corresponding null procedural mesh. When the collision box overlaps a branch, the original static mesh is temporarily hidden and its information is copied to the corresponding procedural mesh, which becomes the active representation of the candidate branch for cutting.
If the overlap ends without blade closure beyond the cutting threshold, the procedural mesh is hidden and the original static branch is restored. If, on the contrary, the cutting command is validated, the engine checks the child branches connected above the cutting location and aggregates the relevant geometry into the procedural representation to be processed. The cut is then executed by defining the slicing point at the center of the collision box and the secant plane according to the normal vector of the blades. In this way, the resulting cutting surface preserves both the inclination imposed by the tool and the appearance of the original branch texture.
After slicing, the detached portion is subjected to physics simulation and falls under gravity with terrain collision enabled, while the remaining structure is preserved as the updated tree configuration. As illustrated in Figure 5, the sequence managed by the Interaction Engine therefore includes branch detection, activation of the procedural counterpart, secant-plane definition, runtime slicing, physical fall of the detached geometry, and restoration or update of the remaining tree. This logic is important because it moves the simulator beyond static visualization: the operator does not simply observe the digital twin but alters it through a stateful and geometrically coherent action, making the simulation suitable for realistic task execution and subsequent design-oriented assessment.

4.4. Operator-in-the-Loop Module

Within the proposed framework, the Operator-in-the-Loop module represents the embodied presence of the user inside the simulation. In the present case study, head and hand tracking are used to animate a full upper-body avatar in real time, so that the operator is represented not as an external observer but as an active element of the task. The avatar was generated through the MetaHuman pipeline [40] and driven through an inverse-kinematics solution based on tracked end-effectors [41]. The free hand is tracked with Leap Motion Controller 2, while the working hand holding the pruning shear is tracked through an HTC VIVE Tracker 3.0 mounted on the back of the hand. The operator’s first-person view during the immersive pruning task is shown in Figure 6.
This module is essential because the operator does not simply observe the virtual environment but acts within it. Guided by the virtual instructor, the user selects branches, moves around the tree, and performs pruning actions whose consequences modify the state of the digital twin during execution. In this way, the operator becomes part of the simulation loop, linking embodiment, task execution, and environmental change.

4.5. Human-Centred Assessment Layer

Within the proposed framework, the Human-Centred Assessment Layer enables the collection of multimodal data during task execution, transforming the immersive simulator into an assessment-ready platform. In the present implementation, three main sources of information can be acquired during the operator’s activity: ocular data from the head-mounted display, body motion reconstructed through a markerless pose-estimation pipeline, and task-related signals extracted from the virtual environment.
From the perspective of human monitoring, the most relevant component concerns body tracking. The integration of a markerless AI-based Human Pose Estimation module, described in details in previous works [14,15], makes it possible to reconstruct posture continuously during the pruning task and to derive joint-angle information over time. This provides a basis for the computation of ergonomic metrics and for the identification of non-neutral configurations, repeated movement patterns, or potentially critical postural strategies during execution. Compared to marker-based optoelectronic systems or wearable IMU suits, the adopted markerless pipeline requires only a small number of standard RGB cameras and a mid-range workstation, without imposing any instrumentation on the operator. This represents a significant practical advantage for deployment in industrial or agricultural settings, where cost, setup complexity, and operator intrusiveness are relevant constraints.
At the same time, the immersive setup can also provide eye-related variables through the headset, supporting the observation of visual behaviour during the task.
A further contribution comes from the pruning tool itself: the blade opening signal is continuously transmitted and can be used not only to animate the virtual counterpart but also to extract task-related indicators such as the number of cuts performed, the degree of blade closure, and the timing of each actuation. The following section presents representative outputs from this assessment layer in order to illustrate the type of data that can be extracted from the proposed architecture.

5. Experimental Validation

5.1. Pilot Study Design and Objectives

The experimental session presented in this chapter addresses two complementary validation objectives: first, to verify that the full multimodal acquisition pipeline, covering body kinematics, ergonomic indicators, and eye-tracking variables, can be executed correctly and completely within a single immersive session; second, to assess whether the framework operates consistently across different participants, confirming its robustness beyond a single-subject demonstration. To this end, six participants aged 18–30 years were recruited from within the research group and the department. All participants were right-handed, none had prior pruning experience, and three had previous experience with virtual reality headsets. Informed consent was obtained from all participants before the experimental session. This choice reflects the nature of the pilot study: the goal is to expose the framework to a range of body morphologies and interaction styles, not to constitute a representative sample of agricultural operators. Each participant completed the same standardised session under identical conditions, ensuring that any variability in the acquired data reflects differences between participants rather than differences in the experimental protocol. The results reported in the following sections are presented through representative outputs from a single participant, selected to illustrate the type and quality of data that the framework can produce. The consistent acquisition across all six sessions is reported separately as evidence of pipeline reliability.

5.2. Experimental Setup

5.2.1. Hardware Configuration

The experimental setup integrates four hardware components operating simultaneously during task execution: the head-mounted display, the instrumented pruning shear, the motion capture system, and the host workstation. The participant was immersed in the virtual environment through a Varjo XR-4 head-mounted display, which provided both visual rendering and continuous acquisition of eye-tracking variables. Two SteamVR base stations were positioned in the experimental area to enable six-degrees-of-freedom tracking of the headset and of the HTC VIVE Tracker 3.0 mounted on the back of the participant’s right hand, which held the pruning shear throughout the session. The free left hand was tracked through a Leap Motion Controller 2 mounted on the headset. Body motion was reconstructed through a markerless pose estimation pipeline based on four RGB monocular cameras, mounted on a dedicated fixed structure and arranged frontally with respect to the operational area, as shown in Figure 7. This configuration was selected to maximise coverage of the upper body during the pruning task, where the most ergonomically relevant actions occur. Each camera operated at 30 fps with a resolution of 1984 × 1264 pixels. The simulation was executed on a high-performance workstation (Alienware Aurora R15, NVIDIA RTX 4090 GPU), which simultaneously managed real-time rendering, communication with the instrumented tool, and synchronisation with the external acquisition system.

5.2.2. Experimental Protocol

Each experimental session comprised three sequential phases: an acclimatisation phase, a standardised pruning task, and a post-task quieting period.
During the first phase, which lasted approximately 60 s, the participant remained immersed in the virtual agricultural environment without being assigned a pruning task. The participant was allowed to visually explore the scene and become familiar with the immersive setup, while remaining within the designated operational area. This phase served both as a familiarisation period and to compute the reference condition for the pupillometric analysis as explained in Section 5.3.1.
During the second phase, the participant performed a standardised pruning sequence under the guidance of the virtual instructor. Ten target branches were presented sequentially in a fixed order that was identical for all participants. Each target branch was highlighted within the virtual scene, and the participant was instructed to cut it at its base, at the junction with the parent branch. The sequence and spatial distribution of the targets were designed to require different reaching configurations and body postures, so that the acquired data captured the postural variability associated with the simulated pruning operation.
After completion of the tenth and final cut, the session entered a post-task quieting period lasting approximately 30 s. Participants were instructed to stop performing the pruning task, remain within the operational area, keep the head-mounted display on, and refrain from further interaction with the tree or the pruning shear while continuing to look naturally at the virtual scene. Eye-tracking acquisition continued throughout this period to characterise the post-task evolution of the pupillometric response. The start of the quieting period was defined operationally as the timestamp associated with the completion of the final pruning action. In the recorded data, this instant was identified from the last transition of the branch identifier logged by the Unreal Engine application, corresponding to the final branch-cut completion marker. The quieting period extended from this marker to the end of the recording.

5.3. Data Extraction and Processing

5.3.1. Eye-Tracking Data Processing

The eye-tracking data were acquired at 200 Hz through the Varjo XR-4 head-mounted display. Table 1 lists the raw signals recorded during each session and the metrics derived from the processing pipeline, which constitute the outputs used for analysis and visualisation.
Since gaze direction is measured relative to the HMD, head movements would otherwise appear as gaze shifts even when the eyes remain fixed on the same point in the scene. To correct for this effect, the raw gaze vector was transformed from HMD-local coordinates into world-space coordinates using the head-orientation quaternion logged synchronously by the Unreal Engine application. This transformation decoupled eye movement from head movement and allowed subsequent event detection to operate on the direction of visual attention within the virtual scene. The coordinate-frame mapping between the Varjo gaze convention and the Unreal Engine reference frame was determined through a dedicated calibration procedure and stored in a per-session JSON file provided as input to the processing pipeline.
The pupillometric signal was processed through a multi-stage cleaning procedure. Blink detection was performed using an adaptive participant-specific threshold on eye openness, computed as median k · 1.4826 · MAD over samples with valid gaze tracking, where MAD denotes the Median Absolute Deviation and the factor 1.4826 provides a normal-consistent estimate of scale. Pre- and post-blink exclusion windows were then applied to remove pupillometric artefacts associated with eyelid closure and reopening [42]. Samples outside the physiological pupil-diameter range were rejected, and dilation-speed outliers were removed following the method proposed by Kret and Sjak-Shie [43]. Residual outliers were further filtered using a MAD-based rejection criterion [44]. Short gaps not associated with blink-contaminated intervals were filled by linear interpolation, and the resulting signal was smoothed using a second-order Butterworth low-pass filter with a 4 Hz cutoff frequency.
The acclimatisation phase lasted 60 s; however, only its final 10 s, corresponding to the interval 50 < t 60 s, were used as the participant-specific pupillometric baseline reference window. The preceding 50 s were treated as a settling period and were excluded from the calculation of the baseline statistics. After preprocessing, all valid cleaned pupil-diameter samples within the final 10 s window were used to calculate the baseline median, p ˜ B , and the corresponding scaled Median Absolute Deviation, 1.4826 · MAD B . Samples classified as blink-contaminated, physiologically implausible, or residual outliers, as well as samples remaining unavailable after preprocessing, were excluded from these calculations.
The cleaned pupil-diameter signal was normalised using the robust z-score
z ( t ) = p ( t ) p ˜ B 1.4826 · MAD B ,
where p ( t ) is the cleaned pupil diameter at time t, p ˜ B is the median of the valid samples within the final 10 s of the acclimatisation phase, and MAD B is the corresponding Median Absolute Deviation. A minimum scale value of 0.01 mm was imposed to prevent numerical instability when baseline variability was very small. Percentage change was calculated relative to the same baseline median as
Δ p ( t ) = 100 · p ( t ) p ˜ B p ˜ B .
For the 1 s window analysis, the mean cleaned pupil diameter within each window was compared with the same participant-specific baseline median. The baseline median and scaled MAD were calculated once for each participant and were then kept fixed for the sample-level z-score, percentage-change analysis, 1 s window summaries, and phase-level results reported in the manuscript.
Oculomotor event detection was performed on the world-space gaze signal using REMoDNaV [45], which classified the continuous gaze stream into fixations, saccades, and smooth-pursuit episodes. The resulting events were characterised by their duration, spatial amplitude, and mean angular velocity. Summary metrics were computed separately for the acclimatisation and task phases, enabling a direct comparison of oculomotor behaviour between the resting condition and active task execution. Branch-cut markers, derived from transitions in the branch identifier logged by Unreal Engine, were overlaid on the output plots to contextualise metric variations with respect to individual task events.

5.3.2. Markerless Pose Estimation and Ergonomic Indicators

The markerless pose estimation pipeline was applied to reconstruct the three-dimensional body posture of each participant continuously throughout the immersive pruning session. For each camera frame, the YOLOv8-Pose model detected 17 body keypoints per view, and the corresponding 3D skeleton was reconstructed through multi-view triangulation. The resulting keypoint trajectories were then used to compute time-series joint angles [46] for the main upper- and lower-body segments, providing the kinematic input required for subsequent ergonomic scoring.
The reconstructed joint angles were processed through a semi-automatic pipeline for the computation of RULA [47] and REBA [48] scores, applied independently to both the left and right body sides for each frame of the session. This continuous scoring approach differs from conventional observational ergonomic assessment, which relies on the evaluator selecting a small number of representative postures. In the proposed pipeline, scores are computed for all frames [49], and the frame captured at each branch-cut event is automatically extracted and saved alongside the corresponding kinematic data, providing a set of task-anchored reference postures for targeted ergonomic review.
Two limitations of the YOLOv8-Pose keypoint model required dedicated solutions to enable complete RULA and REBA scoring. First, the model does not provide hand keypoints, making it impossible to estimate the wrist joint angle directly from the markerless reconstruction. This limitation was addressed through a manual assessment of the wrist configuration at the identified critical postures, with the corresponding wrist angle values entered directly into the scoring pipeline. This aspect represents a direction for future development, to be addressed through new body-tracking models or the training of dedicated models capable of hand recognition.
Second, the facial keypoints associated with eye, ear, and nose landmarks, which are used to infer the neck flexion angle, were systematically occluded by the head-mounted display worn by participants during the session, resulting in unreliable or missing detections. The neck angle estimation was then resolved by exploiting the head orientation data available from the Varjo XR-4 head-mounted display, synchronised with the markerless acquisition stream. Following a rigid alignment of the HMD coordinate frame with the markerless skeleton reference frame, the head orientation quaternion was expressed in the local coordinate system of the reconstructed trunk segment. This transformation allowed the neck flexion, lateral bending, and rotation angles to be derived from the relative orientation between the head and the torso, independently of the availability of facial keypoints. The resulting neck angles were integrated into the RULA and REBA scoring pipeline, replacing the occluded landmark estimates with a sensor-based measurement of equivalent physical meaning.
The overall pipeline is therefore characterised as semi-automatic. The joint angle computation and the frame-level RULA and REBA score calculation are performed automatically for the full duration of the session in order to identify the most critical moments during task execution.
A subset of scoring parameters, including the assessment of load or force category, the distinction between static and dynamic muscle use, and the classification of posture as balanced or unbalanced, require contextual judgement that cannot be reliably inferred from skeletal data alone. They are therefore assigned manually by the evaluator, preliminarly or at the level of the selected critical frames. This approach is consistent with the intended use of RULA and REBA as structured evaluation tools that combine objective kinematic measurements with informed ergonomic interpretation, and it reflects the current state of practice in automated ergonomic assessment based on markerless pose estimation [16,25].
To assess the reliability of the proposed semi-automatic ergonomic pipeline, a preliminary static validation was performed by comparing the automatically generated RULA and REBA scores with manual assessments. The validation protocol was derived from nine of the static postures proposed by Manghisi et al. [49], and the obtained results were qualitatively compared with the performance reported by Manghisi et al. [49] and Battini et al. [50]. The nine selected postures were performed by two participants, resulting in 18 static acquisitions. For each acquisition, the right and left body sides were evaluated separately, both manually and through the proposed system. Therefore, the agreement analysis comprised 36 side-specific comparisons for RULA and 36 side-specific comparisons for REBA. Exact agreement was obtained in 31 out of 36 comparisons for RULA, corresponding to 86.1%, and in 29 out of 36 comparisons for REBA, corresponding to 80.6%. In all 36 comparisons, the automatically generated scores remained within ± 1 point of the corresponding manual assessment. Most discrepancies were associated with forward-bent postures, where small inaccuracies in trunk orientation arose from the viewing geometry of the camera setup. These preliminary results are consistent with the agreement levels reported in the literature for markerless and wearable-based ergonomic assessment systems and support the ability of the proposed pipeline to identify critical postures during immersive task execution.

6. Results

6.1. Pupillometric Analysis

Figure 8 reports the pupillometric outputs obtained for a representative participant (P3) during the immersive pruning session. Subplot (a) shows the raw pupil diameter signal over time, together with the baseline median (grey dash-dot horizontal line, 2.39 mm) and the corresponding ± 1.4826 · MAD band (light blue shading, ±0.10 mm). The yellow shading identifies the baseline reference window used for normalisation (last 10 s of the acclimatisation phase), while the black dashed vertical line marks the end of the baseline phase and the orange dashed line marks the start of the quieting period. Green dotted vertical lines indicate the ten branch-cut events throughout the session. Subplot (b) shows the mean pupil diameter aggregated over consecutive 1 s windows, colour-coded by direction of change relative to the baseline: orange bars correspond to the baseline reference window, red bars indicate dilation exceeding 5%, blue bars indicate constriction exceeding 5%, and grey bars indicate stable windows within ±5% of baseline. Subplot (c) shows the normalised pupil z-score over time, with blink-contaminated samples masked. The dash-dot horizontal line at z = 0 represents the baseline reference, and the light red and blue shaded regions indicate periods of positive and negative z-score, respectively.
In Subplot (a), the raw diameter during the task phase is visibly and persistently elevated above the baseline median of 2.39 mm, with the signal remaining largely above the upper bound of the ± MAD band for most of the task duration. Subplot (b) confirms this trend: the overwhelming majority of 1 s windows in the task phase appear in red, indicating dilation exceeding 5% relative to baseline, with a single annotated peak of +33.2%. Subplot (c) provides the clearest characterisation of this response: the z-score rises sharply at the onset of the task phase, marked by the black dashed line, and remains predominantly positive throughout, with values frequently exceeding + 4 and reaching up to + 8 during the most demanding pruning actions.
The aggregated results across all six participants are reported in Table 2 and illustrated in Figure 9. Subplot (a) of Figure 9 shows the pupil z-score over time for each participant as a coloured line, with the group mean overlaid as a thick black line. The black dashed and orange dashed vertical lines mark the baseline end and quieting start, respectively, consistent with the single-subject representation. Subplot (b) summarises the per-subject comparison between the task phase and the quieting period: coloured bars represent the mean pupil z-score during the task phase for each participant, while black diamond markers connected by a dashed line indicate the corresponding mean z-score during the quieting period. The dash-dot horizontal line at z = 0 represents the baseline reference. All six participants showed a positive mean pupil z-score during the task phase, with values ranging from 0.37 for P4 to 5.32 for P5. This indicates that mean pupil diameter during task execution was above the participant-specific baseline in every participant, although the magnitude of the response varied substantially across subjects. The percentage of task time spent in dilation ranged from 65.5% (P4) to 100.0% (P5), with a mean of 86.7% across the group. The mean percentage change in pupil diameter relative to baseline ranged from +4.2% (P4) to +17.3% (P6), with a group mean of +12.5%.
The quieting period analysis, visible in subplot (b) of Figure 9 through the position of the diamond markers relative to the z = 0 line, reveals that the pupillary response did not fully recover in most participants. Three out of six participants maintained a positive mean z-score during the quieting period, indicating that pupil diameter remained above the individual baseline even after task completion. P2, P4, and P6 (quieting z-scores of −0.27, −0.17, and −1.90, respectively) showed a return below baseline, suggesting a more rapid physiological recovery in these cases. Taken together, these results show that the proposed framework can acquire pupillometric signals continuously throughout immersive task execution across participants with different baseline pupil diameters, ranging from 1.90 mm for P6 to 3.12 mm for P4.
To investigate whether pupillary activation differed in the temporal neighbourhood of the pruning actions compared with the periods between consecutive cuts, an exploratory event-based analysis was performed using the branch-cut completion markers recorded by the virtual environment. For each marker, an event neighbourhood extending from 1 s before to 1 s after the recorded completion of the cut was considered. This interval was selected to include both the preparation of the cutting action and the delayed pupillary response following its completion.
The event neighbourhoods were compared with the intermediate inter-cut periods. To prevent temporal overlap, each inter-cut interval started 1.25 s after the preceding marker and ended 1.25 s before the subsequent marker, thus excluding the 1 s event neighbourhood and an additional 0.25 s guard interval on each side. An inter-cut interval was retained only when at least 1 s remained after these exclusions. The pupil z-score was averaged separately within the complete event neighbourhood, the pre-cut interval from 1 to 0 s, the post-cut interval from 0 to + 1 s, and the valid inter-cut periods. Repeated events were first aggregated within each participant, so that the participant, rather than the individual cut, represented the statistical unit. A total of 58 valid cutting events and 46 valid inter-cut intervals were included in the analysis.
As shown in Figure 10, five of the six participants exhibited higher mean pupil z-scores in the neighbourhood of the cutting events than during the inter-cut periods. At the group level, the mean pupil z-score was 3.319 within the complete event neighbourhood and 3.027 during the inter-cut periods. Positive mean differences were also observed in both the pre-cut and post-cut intervals. These findings provide descriptive evidence of task-related variations in physiological activation around the pruning actions. The observed pattern is also compatible with a possible increase in cognitive engagement during action preparation and execution; however, this interpretation remains exploratory because pupil diameter can be influenced by multiple physiological, perceptual, and environmental factors.

6.2. Body Kinematics and Ergonomic Assessment

Figure 11 illustrates the output of the semi-automatic ergonomic assessment pipeline for a representative participant. The lower panel reports the continuous RULA final score computed frame by frame throughout the session, with dashed vertical lines indicating the branch-cut events. This continuous analysis allows postures with differing RULA scores to be identified within the simulated session. The two starred markers identify a pair of postures selected for detailed inspection: a low-risk upright configuration (left marker, approximately frame 1800) and a higher-risk forward-bent configuration (right marker, approximately frame 3400). The corresponding reconstructed 3D skeletons are shown in the upper panels, where the coloured axes attached to the head segment represent the head orientation derived from the HMD tracking data and expressed in the trunk local reference frame, used for neck angle estimation.
The RULA Score B parameters extracted at the two selected frames are reported in Table 3. In the upright configuration, the trunk flexion angle was negligible (−4.6°), with no trunk twisting, side bending, or leaning detected. The neck exhibited a flexion of 54.3°, classified in the highest neck score category, while the leg support was assessed as unsupported and unbalanced, contributing a score of 2. The resulting Group B table lookup (neck 3/trunk 1/legs 2) yielded a Score B of 3, and the final RULA score was 3, corresponding to the risk level “investigate”.
In the forward-bent configuration, the postural load increased substantially. The trunk flexion reached 41.3°, accompanied by lateral side bending and a leaning condition, resulting in a trunk score of 3 and a trunk lookup input of 3 + 1 = 4 . The neck flexion was 20.0°, but the simultaneous presence of neck twisting and lateral bending raised the neck lookup input to 3 + 1 + 1 = 5 . The leg support remained unsupported and unbalanced (score 2). The Group B table lookup (neck 3/trunk 3/legs 2) returned a Score B output of 8, yielding a final Score B of 7 and a RULA final score of 6, corresponding to the risk level "act soon" This difference between the two postures—RULA scores of 3 and 6, respectively—illustrates the capacity of the proposed pipeline to discriminate between postural configurations of different ergonomic severity within the same session. This result demonstrates within-simulation sensitivity to differing postures without requiring wearable instrumentation on the operator; it does not show that the measured postures or scores are equivalent to those of real-world pruning. Cross-validation with the corresponding real task is required to assess such correspondence.

7. Conclusions

This paper presented an anthropocentric framework for human-centred assessment in immersive simulation, addressing the need to integrate operator behaviour, task-specific interaction, and safety-oriented analysis within early-stage system design. Starting from the limitations of conventional Virtual Reality applications and from the evolving perspective of Industry 5.0, the proposed approach extends immersive simulation beyond purely training-oriented applications by demonstrating the technical feasibility of integrating interaction, operator monitoring, and assessment functions within a single platform.
The framework combines a task-oriented Scenario Digital Twin, a Physical-Tool-in-the-Loop module, an Interaction Engine for the management of state-dependent actions, and an Operator-in-the-Loop configuration coupled with a Human-Centred Assessment Layer. The architecture was instantiated through an immersive hazelnut-pruning simulator in which a real instrumented pruning shear was synchronised with its virtual counterpart and used to interact with a dynamically updated digital twin of the tree. Although pruning was selected because of its postural variability, tool dependence, and irreversible actions, the proposed modules are not specific to the agricultural domain and may be adapted to other tool-mediated applications. Their transfer to manufacturing, infrastructure maintenance, and related contexts will nevertheless require domain-specific implementation and validation.
A pilot study involving six participants showed that the platform can acquire and temporally integrate multimodal data during immersive task execution. The markerless pose-estimation pipeline provided continuous three-dimensional body kinematics that supported frame-level RULA and REBA assessment through a semi-automatic procedure. The extracted indicators distinguished postural configurations associated with different levels of ergonomic severity within the same immersive session. In parallel, all six participants exhibited a positive mean pupil z-score during the task phase, with a group mean of 2.55 and a mean pupil-diameter increase of 12.5% relative to baseline. These results show that the framework can capture task-related postural and pupillometric variations continuously. They do not, however, establish equivalence with real-world pruning or validate pupil diameter as a specific measure of cognitive workload.
The present findings therefore support the use of the proposed architecture as an assessment-ready experimental platform. The acquired data may contribute to early-stage Safety-by-Design activities by enabling the observation of task execution, posture, and physiological activation under controlled immersive conditions. The use of these observations to predict real-world critical conditions or to support validated design decisions will require comparison with the corresponding physical task.
Future work will focus on a dedicated cross-modal validation study comparing operator kinematics, interaction strategies, and task execution between the immersive simulation and the real environment. Future experiments will also include independent workload and performance measures to further investigate the relationship between pupillometric variations, physiological activation, and task-related cognitive engagement. Finally, the framework will be evaluated across different application domains and user populations to assess its generalisability beyond the agricultural case study.

Author Contributions

Conceptualization, D.F., M.C. and H.G.; Methodology, D.F. and M.C.; Software, D.F.; Validation, H.G.; Formal analysis, M.C.; Investigation, D.F. and M.C.; Resources, M.C.; Data curation, D.F.; Writing—original draft, D.F.; Writing—review & editing, M.C.; Supervision, M.C. and H.G.; Project administration, M.C. and H.G.; Funding acquisition, M.C. All authors have read and agreed to the published version of the manuscript.

Funding

This article is produced within the NODES project, financed by the MUR on M4C2 funds—Investment 1.5 Notice “Innovation Ecosystems”, within the PNRR financed by the European Union—NextGenerationEU (Grant agreement Cod. n.ECS00000036). The NODES—Digital and Sustainable North West ecosystem offers innovation opportunities for the territory through tenders aimed at institutions and businesses.

Data Availability Statement

Data are contained within the article.

Acknowledgments

The authors gratefully acknowledge Pellenc Italia S.R.L. for providing the 3D model of the pruning shears, for enabling the presentation of the application at EIMA 2024, and for supplying the sensorised C3X pruning shear integrated into the system. During the preparation of this manuscript, the authors used the Prism corrector instrument for the purposes of grammar and spelling checks. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Rosário, A.T.; Raimundo, R.J.G. AI, Optimization, and Human Values: Mapping the Intellectual Landscape of Industry 4.0 to 5.0. Appl. Sci. 2025, 15, 7264. [Google Scholar] [CrossRef]
  2. Valentini, L.; Weistroffer, V.; Grandi, F.; Peruzzini, M. Digital toolkit for human-centered machine design: Development and testing of an innovative system integrating virtual reality and HMI digital prototypes. Int. J. Adv. Manuf. Technol. 2025, 144, 7717–7735. [Google Scholar] [CrossRef]
  3. Nunes, M.; Silva, E.; Sousa, N.; Sousa, E.; Nunes, E.; Margolis, I. VR Virtual Prototyping Application for Airplane Cockpit: A Human-centred Design Validation. In Proceedings of the 18th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications (VISIGRAPP 2023)—HUCAPP; INSTICC: Lisboa, Portugal; SciTePress: Setúbal, Portugal, 2023; pp. 177–184. [Google Scholar] [CrossRef]
  4. Röhrich, M.; Abramuszkinová Pavlíková, E.; Ulrich, R. From Digital Motion Capture to Human-Friendly Forestry Machines: A Digital Human Modeling Framework—Case Study in Design and Prototyping of Forestry Machines. Forests 2026, 17, 235. [Google Scholar] [CrossRef]
  5. Haney, J.M.; Ammons, D.; Choi, H. Perceived safety during human-robot interaction with an autonomous mobile robot. Appl. Ergon. 2026, 135, 104747. [Google Scholar] [CrossRef] [PubMed]
  6. Di Gironimo, G.; Lanzotti, A.; Melemez, K.; Renno, F. A top-down approach for virtual redesign and ergonomic optimization of an agricultural tractor’s driver cab. In Engineering Systems Design and Analysis; American Society of Mechanical Engineers: New York, NY, USA, 2012; Volume 3, pp. 801–811. [Google Scholar] [CrossRef]
  7. Meşe, E.A.; Basmaci, F.; Bulut, A.C.; Topallı, D.; Şenel, F.; Cagiltay, N.E. The role of extended reality technologies in hygiene education and training: A systematic review of applications, benefits, and challenges. BMC Infect. Dis. 2026, 26, 272. [Google Scholar] [CrossRef] [PubMed]
  8. Fabiocchi, D.; Sergenti, C.; Placa, S.L.; Carnevale, M.; Giberti, H. An Immersive Hazelnut Tree Pruning Simulator for Agricultural Operator Training. In Proceedings of the 2025 IEEE International Conference on Metrology for eXtended Reality, Artificial Intelligence and Neural Engineering (MetroXRAINE); IEEE: New York, NY, USA, 2025; pp. 483–488. [Google Scholar] [CrossRef]
  9. Martinelli, A.; Fabiocchi, D.; Picchio, F.; Giberti, H.; Carnevale, M. Design of an Environment for Virtual Training Based on Digital Reconstruction: From Real Vegetation to Its Tactile Simulation. Designs 2025, 9, 32. [Google Scholar] [CrossRef]
  10. DeVrio, N.; Harrison, C. Reel Feel: Rich Haptic XR Experiences Using an Active, Worn, Multi-String Device. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems; ACM: New York, NY, USA, 2025. [Google Scholar] [CrossRef]
  11. Lee, J.; Kim, J.; Kang, H.; Ryu, H.; Kim, Y. Trade-offs in Virtual Grasping: The Interplay of Interaction Fidelity and Object Affordance. In Proceedings of the VRST ’25: 31st ACM Symposium on Virtual Reality Software and Technology; VRST ’25; ACM: New York, NY, USA, 2025. [Google Scholar] [CrossRef]
  12. Rückert, P.; Sievers, T.S.; Ewers, J.; Heine, J.; Kuschel, N.; Marhenke, L.; Wassermann, N.; Wedler, G.; Tracht, K. Adaptive Tool Replicas with Haptic Feedback for Increased Presence Perception in Virtual Reality. Procedia CIRP 2024, 130, 797–801. [Google Scholar] [CrossRef]
  13. Dewberry, N.K.; AlHmoud, I.; Benton, K.; Suarez, D.; Chen, Y.P.; Karkaria, V.; Tsai, Y.K.; Brock, M.; Alazzawi, N.; Chowdhury, S.; et al. A real-time VR-enabled digital twin framework for multi-user interaction in Industry 4.0. Manuf. Lett. 2025, 44, 1486–1497. [Google Scholar] [CrossRef]
  14. Giulietti, N.; Todesca, D.; Carnevale, M.; Giberti, H. A Real-Time Human Pose Measurement System for Human-In-The-Loop Dynamic Simulators. IEEE Access 2025, 13, 24954–24969. [Google Scholar] [CrossRef]
  15. Giulietti, N.; Fabiocchi, D.; Todesca, D.; Carnevale, M.; Giberti, H. A Vision-Based marker-less measurement system for Real-Time Estimation of Human Body Inertial properties. Measurement 2026, 260, 119843. [Google Scholar] [CrossRef]
  16. Benharkat, N.E.H.; Bentaalla Kaced, S.; Chakhrit, A.; Chergui, A. Automatic real-time ergonomic posture assessment using digital models: A case study of manual handling tasks. Int. J. Adv. Manuf. Technol. 2026, 142, 963–980. [Google Scholar] [CrossRef]
  17. Rah, A.; Chen, Y. Virtual Reality as a Stress Measurement Platform: Real-Time Behavioral Analysis with Minimal Hardware. Sensors 2025, 25, 5323. [Google Scholar] [CrossRef] [PubMed]
  18. Poupard, M.; Larrue, F.; Bertrand, M.; Liguoro, D.; Sauzéon, H.; Tricot, A. From movement to learning: Leveraging VR behavioral metrics to evaluate cognitive load and curiosity. Int. J.-Hum.-Comput. Stud. 2026, 209, 103751. [Google Scholar] [CrossRef]
  19. Radhakrishnan, U.; Konstantinos, K.; Chinello, F. A systematic review of immersive virtual reality for industrial skills training. Behav. Inf. Technol. 2021, 40, 1310–1339. [Google Scholar] [CrossRef]
  20. Anastasiou, E.; Balafoutis, A.; Fountas, S. Applications of Extended Reality (XR) in Agriculture, Livestock Farming, and Aquaculture: A Review. Smart Agric. Technol. 2022, 3, 100105. [Google Scholar] [CrossRef]
  21. Buttussi, F.; Chittaro, L. Effects of Different Types of Virtual Reality Display on Presence and Learning in a Safety Training Scenario. IEEE Trans. Vis. Comput. Graph. 2017, 24, 1063–1076. [Google Scholar] [CrossRef] [PubMed]
  22. Singh, A.K. VR/AR in ergonomics and workspace design: A dual-perspective analysis of applications and implications. Appl. Ergon. 2025, 129, 104612. [Google Scholar] [CrossRef] [PubMed]
  23. Liu, H.; Zhang, Z.; Jiao, Z.; Zhang, Z.; Li, M.; Jiang, C.; Zhu, Y.; Zhu, S.C. A Reconfigurable Data Glove for Reconstructing Physical and Virtual Grasps. Engineering 2024, 32, 202–216. [Google Scholar] [CrossRef]
  24. Kim, J.S.; Kim, K.W.; Kim, H.J.; Moon, S.Y. Development and Usability Evaluation of a Leap Motion-Based Controller-Free VR Training System for Inferior Alveolar Nerve Block. Appl. Sci. 2026, 16, 1325. [Google Scholar] [CrossRef]
  25. Zamorano Núñez, G.A.; Norambuena, N.; Quezada, I.C.; Valín Rivera, J.L.; Olmos, J.N.; Ketterer, C.G. Real-Time Automated Ergonomic Monitoring: A Bio-Inspired System Using 3D Computer Vision. Biomimetics 2026, 11, 88. [Google Scholar] [CrossRef] [PubMed]
  26. Ciccarelli, M.; Papetti, A.; Germani, M. Investigating the Reliability of Ergonomic Assessment in Immersive Virtual Environments. In Proceedings of the Design Tools and Methods in Industrial Engineering V; Berselli, G., Andrisano, A.O., Di Stefano, P., Rizzi, C., Gherardini, F., Eds.; Springer: Cham, Switzerland, 2026; pp. 154–165. [Google Scholar]
  27. Karaer, I.; Hunt, C.; Yoon, H.J.; Ma, R.; Savant, R.; Rodwell, V.; Shenoy, R.; Tu, Z.; Arshad, Q.; Mukaetova-Ladinska, E.B.; et al. Technical validation of a virtual reality-based eye tracker for neuro-ophthalmic assessment: A reliability and reproducibility study. Sci. Rep. 2026, 16, 1134. [Google Scholar] [CrossRef] [PubMed]
  28. Cheung, K.K.F.; Jong, M.S.Y.; Lee, F.L.; Lee, J.H.M.; Luk, E.T.H.; Shang, J.; Wong, M.K.H. FARMTASIA: An online game-based learning environment based on the VISOLE pedagogy. Virtual Real. 2008, 12, 17–25. [Google Scholar] [CrossRef]
  29. Yongyuth, P.; Prada, R.; Nakasone, A.; Prendinger, H. AgriVillage: 3D multi-language Internet game for fostering agriculture environmental awareness. In Proceedings of the International Conference on Management of Emergent Digital EcoSystems; ACM: New York, NY, USA, 2010; pp. 145–152. [Google Scholar] [CrossRef]
  30. Rosmansyah, Y.; Achiruzaman, M.; Hardi, A.B. A 3D Multiuser Virtual Learning Environment for Online Training of Agriculture Surveyors. J. Inf. Technol. Educ. Res. 2019, 18, 481–507. [Google Scholar] [CrossRef] [PubMed]
  31. Carruth, D.; Hudson, C.; Fox, A.; Deb, S. User Interface for an Immersive Virtual Reality Greenhouse for Training Precision Agriculture; Springer: Cham, Switzerland, 2020; pp. 35–46. [Google Scholar] [CrossRef]
  32. Kim, R.; Jungyu, K.; Lee, I.B.; Yeo, U.H.; Lee, S.Y.; Valentin, C. Development of three-dimensional visualisation technology of the aerodynamic environment in a greenhouse using CFD and VR technology, Part 2: Development of an educational VR simulator. Biosyst. Eng. 2021, 207, 12–32. [Google Scholar] [CrossRef]
  33. Lv, M.; Lu, S.; Guo, X. Interactive Virtual Fruit Tree Pruning Simulation; Atlantis Press: Paris, France, 2015. [Google Scholar] [CrossRef]
  34. Zhao, P.F.; Chen, T.E.; Wang, W.; Chen, F.Y. Research on the Agricultural Skills Training Based on the Motion-Sensing Technology of the Leap Motion. In Proceedings of the International Conference on Computer and Computing Technologies in Agriculture; Springer: Cham, Switzerland, 2016; Volume 479, pp. 277–286. [Google Scholar] [CrossRef]
  35. VINUM. AMPÉLOS|Simulateur de Formation à la Taille de la Vigne; Programme VINUM—Programmevinum.fr. Available online: https://programmevinum.fr/les-services-vinum/ampelos/ (accessed on 27 May 2025).
  36. SkillsVR. Kiwi Fruit Pruning by SkillsVR. Available online: https://skillsvr.com/modules/kiwifruit-pruning (accessed on 27 May 2025).
  37. La Placa, S.; Doria, E. Digital Documentation and Fast Census for Monitoring the University’s Built Heritage. Int. Arch. Photogramm. Remote. Sens. Spat. Inf. Sci. 2024, 48, 271–278. [Google Scholar] [CrossRef]
  38. Digumarti, S.T.; Nieto, J.; Cadena, C.; Siegwart, R.; Beardsley, P. Automatic Segmentation of Tree Structure From Point Cloud Data. IEEE Robot. Autom. Lett. 2018, 3, 3043–3050. [Google Scholar] [CrossRef]
  39. SpeedTree—3D Vegetation Modeling and Middleware. Available online: https://store.speedtree.com/ (accessed on 17 July 2025).
  40. Metahuman. MetaHuman—Realistic Person Creator. Available online: https://www.unrealengine.com/en-US/metahuman/ (accessed on 27 May 2025).
  41. MimicPRO, VR IK Body System. Available online: https://jakeplayable.gumroad.com/l/MimicPro (accessed on 4 August 2026).
  42. Mathôt, S. Pupillometry: Psychology, Physiology, and Function. J. Cogn. 2018, 1, 16. [Google Scholar] [CrossRef] [PubMed]
  43. Kret, M.E.; Sjak-Shie, E.E. Preprocessing pupil size data: Guidelines and code. Behav. Res. Methods 2019, 51, 1336–1342. [Google Scholar] [CrossRef] [PubMed]
  44. Mathôt, S.; Vilotijević, A. Methods in cognitive pupillometry: Design, preprocessing, and statistical analysis. Behav. Res. Methods 2023, 55, 3055–3077. [Google Scholar] [CrossRef] [PubMed]
  45. Dar, A.H.; Wagner, A.S.; Hanke, M. REMoDNaV: Robust eye-movement classification for dynamic stimulation. Behav. Res. Methods 2021, 53, 399–414. [Google Scholar] [CrossRef] [PubMed]
  46. Todesca, D.; Fabiocchi, D.; Fuentes, J.F.; Giulietti, N.; Carnevale, M.; Giberti, H. Comparison of Real-Time Marker-Less and Optoelectronic 3D Human Pose Estimation Systems for Cyclist Pose Analysis; IEEE: New York, NY, USA, 2025; pp. 417–422. [Google Scholar] [CrossRef]
  47. McAtamney, L.; Corlett, E.N. RULA: A survey method for the investigation of work-related upper limb disorders. Appl. Ergon. 1993, 24, 91–99. [Google Scholar] [CrossRef] [PubMed]
  48. Hignett, S.; McAtamney, L. Rapid Entire Body Assessment (REBA). Appl. Ergon. 2000, 31, 201–205. [Google Scholar] [CrossRef] [PubMed]
  49. Manghisi, V.M.; Uva, A.E.; Fiorentino, M.; Bevilacqua, V.; Trotta, G.F.; Monno, G. Real time RULA assessment using Kinect v2 sensor. Appl. Ergon. 2017, 65, 481–491. [Google Scholar] [CrossRef] [PubMed]
  50. Battini, D.; Berti, N.; Finco, S.; Guidolin, M.; Reggiani, M.; Tagliapietra, L. WEM-Platform: A real-time platform for full-body ergonomic assessment and feedback in manufacturing and logistics systems. Comput. Ind. Eng. 2022, 164, 107881. [Google Scholar] [CrossRef]
Figure 1. General framework for human-centred assessment in immersive simulation. The framework integrates task-oriented digital twins, physical-tool-in-the-loop interaction, operator monitoring, and human-centred assessment to support the redesign of tools, layouts, interfaces, and procedures. Blue, green, and orange boxes denote inputs and simulation modules, assessment, and design outputs, respectively; solid arrows show the process flow and the dashed arrow the iterative feedback loop.
Figure 1. General framework for human-centred assessment in immersive simulation. The framework integrates task-oriented digital twins, physical-tool-in-the-loop interaction, operator monitoring, and human-centred assessment to support the redesign of tools, layouts, interfaces, and procedures. Blue, green, and orange boxes denote inputs and simulation modules, assessment, and design outputs, respectively; solid arrows show the process flow and the dashed arrow the iterative feedback loop.
Electronics 15 03525 g001
Figure 2. Framework instantiation for the immersive hazelnut pruning simulator. The figure shows how the general architecture is translated into the specific case study through the integration of the tree digital twin, the Physical-Tool-in-the-Loop pruning shear, the operator-in-the-loop module, and the human-centred assessment layer. Blue, green, and orange boxes denote inputs and simulation modules, assessment, and design outputs, respectively; solid arrows show the process flow and the dashed arrow the iterative feedback loop.
Figure 2. Framework instantiation for the immersive hazelnut pruning simulator. The figure shows how the general architecture is translated into the specific case study through the integration of the tree digital twin, the Physical-Tool-in-the-Loop pruning shear, the operator-in-the-loop module, and the human-centred assessment layer. Blue, green, and orange boxes denote inputs and simulation modules, assessment, and design outputs, respectively; solid arrows show the process flow and the dashed arrow the iterative feedback loop.
Electronics 15 03525 g002
Figure 3. Pipeline adopted for the implementation of the Scenario Digital Twin module in the immersive pruning simulator. The workflow starts from the acquisition campaign, combines photogrammetric and laser-data alignment, and then proceeds through tree model parametrization, leaf generation and collision setup, up to the integration of the resulting asset into the agricultural virtual environment.
Figure 3. Pipeline adopted for the implementation of the Scenario Digital Twin module in the immersive pruning simulator. The workflow starts from the acquisition campaign, combines photogrammetric and laser-data alignment, and then proceeds through tree model parametrization, leaf generation and collision setup, up to the integration of the resulting asset into the agricultural virtual environment.
Electronics 15 03525 g003
Figure 4. Implementation of the Physical-Tool-in-the-Loop module. (Left): instrumented electric pruning shear connected to the local acquisition and transmission unit. (Right): virtual counterpart of the tool in Unreal Engine, showing the articulated structure and the reference rotation axes used to reproduce the kinematics of the real device.
Figure 4. Implementation of the Physical-Tool-in-the-Loop module. (Left): instrumented electric pruning shear connected to the local acquisition and transmission unit. (Right): virtual counterpart of the tool in Unreal Engine, showing the articulated structure and the reference rotation axes used to reproduce the kinematics of the real device.
Electronics 15 03525 g004
Figure 5. Illustrative sequence of the cutting logic managed by the Interaction Engine. The images are schematic views generated to highlight the main processing steps and do not correspond to the full immersive simulation, in which the pruning shear is manipulated by the operator within the agricultural virtual environment. From (left) to (right) and from (top) to (bottom), the sequence shows branch detection through blade-region overlap, conversion of the candidate branch from a static to a procedural mesh (schematically highlighted in pink), definition of the secant plane (green), runtime slicing of the selected geometry, physical fall of the detached portion, and restoration/update of the remaining tree configuration.
Figure 5. Illustrative sequence of the cutting logic managed by the Interaction Engine. The images are schematic views generated to highlight the main processing steps and do not correspond to the full immersive simulation, in which the pruning shear is manipulated by the operator within the agricultural virtual environment. From (left) to (right) and from (top) to (bottom), the sequence shows branch detection through blade-region overlap, conversion of the candidate branch from a static to a procedural mesh (schematically highlighted in pink), definition of the secant plane (green), runtime slicing of the selected geometry, physical fall of the detached portion, and restoration/update of the remaining tree configuration.
Electronics 15 03525 g005
Figure 6. Operator’s first-person view during the immersive pruning task. The image illustrates the interaction with the virtual tree through the synchronised pruning shear within the reconstructed agricultural environment.
Figure 6. Operator’s first-person view during the immersive pruning task. The image illustrates the interaction with the virtual tree through the synchronised pruning shear within the reconstructed agricultural environment.
Electronics 15 03525 g006
Figure 7. Experimental laboratory setup for the pilot study. The image shows the frontal camera array mounted on the fixed overhead structure, providing coverage of the operational area, and the two HTC base Station for headset tracking. The cross marked on the floor indicates the reference position of the participant during task execution.
Figure 7. Experimental laboratory setup for the pilot study. The image shows the frontal camera array mounted on the fixed overhead structure, providing coverage of the operational area, and the two HTC base Station for headset tracking. The cross marked on the floor indicates the reference position of the participant during task execution.
Electronics 15 03525 g007
Figure 8. Pupillometric signals for the representative participant (P3). (a) Raw pupil diameter with baseline reference (median = 2.39 mm, ±1.4826·MAD = 0.10 mm). (b) Mean pupil diameter per 1 s window, colour-coded by direction of change relative to baseline: orange = baseline reference window, red = dilation >5%, blue = constriction >5%, and grey = stable (±5%). The annotated value indicates the peak dilation observed during the task phase. (c) Normalised pupil z-score over time (blink-masked). In all subplots, the black dashed vertical line marks the end of the baseline phase, the orange dashed line marks the start of the quieting period, and the green dotted lines indicate branch-cut events.
Figure 8. Pupillometric signals for the representative participant (P3). (a) Raw pupil diameter with baseline reference (median = 2.39 mm, ±1.4826·MAD = 0.10 mm). (b) Mean pupil diameter per 1 s window, colour-coded by direction of change relative to baseline: orange = baseline reference window, red = dilation >5%, blue = constriction >5%, and grey = stable (±5%). The annotated value indicates the peak dilation observed during the task phase. (c) Normalised pupil z-score over time (blink-masked). In all subplots, the black dashed vertical line marks the end of the baseline phase, the orange dashed line marks the start of the quieting period, and the green dotted lines indicate branch-cut events.
Electronics 15 03525 g008
Figure 9. Aggregated pupillometric results across all six participants. (a) Pupil z-score over time for each participant (coloured lines) and group mean (black line). The black dashed vertical line marks the end of the baseline phase; the orange dashed line marks the start of the quieting period. (b) Mean pupil z-score during the task phase (bars) and during the quieting period (diamond markers) for each participant. The dash-dot line indicates the baseline reference (z = 0).
Figure 9. Aggregated pupillometric results across all six participants. (a) Pupil z-score over time for each participant (coloured lines) and group mean (black line). The black dashed vertical line marks the end of the baseline phase; the orange dashed line marks the start of the quieting period. (b) Mean pupil z-score during the task phase (bars) and during the quieting period (diamond markers) for each participant. The dash-dot line indicates the baseline reference (z = 0).
Electronics 15 03525 g009
Figure 10. Participant-level comparison of the pupil z-score during the intermediate periods between consecutive cuts (Inter-cut), during the 1 s interval preceding each cut marker (Pre-cut), during the 1 s interval following the marker (Post-cut), and across the complete ± 1 -s event neighbourhood (Full event). Each participant is shown in a different colour, with thin lines connecting their values, while the thicker pink line with diamond markers represents the group mean.
Figure 10. Participant-level comparison of the pupil z-score during the intermediate periods between consecutive cuts (Inter-cut), during the 1 s interval preceding each cut marker (Pre-cut), during the 1 s interval following the marker (Post-cut), and across the complete ± 1 -s event neighbourhood (Full event). Each participant is shown in a different colour, with thin lines connecting their values, while the thicker pink line with diamond markers represents the group mean.
Electronics 15 03525 g010
Figure 11. Continuous RULA final score and reconstructed 3D skeletons for a representative participant. Lower panel: frame-level RULA score throughout the session; dashed vertical lines indicate branch-cut events; starred markers identify the two postures selected for detailed analysis. Upper panels: reconstructed 3D skeleton at the selected frames—upright posture (left, RULA C = 3) and forward-bent posture (right, RULA C = 6). The coloured triad attached to the head segment represents the head orientation derived from the Varjo XR-4 HMD and used for neck angle estimation.
Figure 11. Continuous RULA final score and reconstructed 3D skeletons for a representative participant. Lower panel: frame-level RULA score throughout the session; dashed vertical lines indicate branch-cut events; starred markers identify the two postures selected for detailed analysis. Upper panels: reconstructed 3D skeleton at the selected frames—upright posture (left, RULA C = 3) and forward-bent posture (right, RULA C = 6). The coloured triad attached to the head segment represents the head orientation derived from the Varjo XR-4 HMD and used for neck angle estimation.
Electronics 15 03525 g011
Table 1. Raw signals recorded during each session (left) and derived metrics extracted by the processing pipeline (right).
Table 1. Raw signals recorded during each session (left) and derived metrics extracted by the processing pipeline (right).
Recorded SignalRateDerived MetricCategory
Gaze direction ( x , y , z ) 200 HzPupil diameter (mm)Pupillometry
Gaze status200 HzPupil z-scorePupillometry
Left/right pupil diameter (mm)200 HzPupil % change/1 s windowPupillometry
Left/right eye openness200 HzBlink rate (count/min)Blink
HMD orientation quaternion200 HzBlink duration (ms)Blink
HMD position ( x , y , z ) 200 HzMean fixation duration (ms)Oculomotor
Branch identifier200 HzFixation rate (count/min)Oculomotor
Mean saccade amplitude (°)Oculomotor
Saccade rate (count/min)Oculomotor
Table 2. Summary of pupillometric metrics for all six participants. Task z ¯ : mean pupil z-score during the task phase. Task dil.: percentage of task time with pupil diameter above baseline. Task Δ : mean percentage change in pupil diameter relative to baseline. Quiet. z ¯ and Quiet. Δ : same metrics computed during the quieting period. Return: whether mean pupil z-score returned below baseline during the quieting period.
Table 2. Summary of pupillometric metrics for all six participants. Task z ¯ : mean pupil z-score during the task phase. Task dil.: percentage of task time with pupil diameter above baseline. Task Δ : mean percentage change in pupil diameter relative to baseline. Quiet. z ¯ and Quiet. Δ : same metrics computed during the quieting period. Return: whether mean pupil z-score returned below baseline during the quieting period.
SubjectBaseline (mm)Task z ¯ Task dil. (%)Task Δ (%)Quiet. z ¯ Quiet. Δ (%)Return
P12.052.1199.1+11.81.46+6.6No
P22.711.2183.8+9.1−0.27+0.7Yes
P32.444.2397.1+15.83.23+11.5No
P43.120.3765.5+4.2−0.17+2.3Yes
P52.595.32100.0+16.71.93+6.1No
P61.902.0374.7+17.3−1.90+3.1Yes
Mean2.472.5586.7+12.50.71+5.13/6
Table 3. RULA Score B parameters for the two representative postures identified during the immersive pruning session. Values refer to the right body side. The upright posture corresponds to approximately frame 1800; the forward-bent posture corresponds to approximately frame 3400.
Table 3. RULA Score B parameters for the two representative postures identified during the immersive pruning session. Values refer to the right body side. The upright posture corresponds to approximately frame 1800; the forward-bent posture corresponds to approximately frame 3400.
ParameterUpright PostureForward-Bent Posture
Neck
Neck flexion (°)54.320.0
Neck twistedNoYes
Neck side bendingNoYes
Neck lookup input3 + 0 = 33 + 1 + 1 = 5
Trunk
Trunk flexion (°)−4.641.3
Trunk twistedNoNo
Trunk side bendingNoYes
Person leaningNoYes
Trunk lookup input1 + 0 = 13 + 1 = 4
Legs & Summary
Legs & feetUnsupported/unbalanced
Group B table (N/T/L)3/1/25/4/2
Group B output38
Muscle useNoNo
Force/loadUnder 2 kg
Score A33
Score B final37
Score C (RULA final)36
Risk levelInvestigateAct soon
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Fabiocchi, D.; Carnevale, M.; Giberti, H. Integrating Real Tool Interaction and Multimodal Operator Monitoring in Immersive Simulation for Human-Centred Assessment. Electronics 2026, 15, 3525. https://doi.org/10.3390/electronics15163525

AMA Style

Fabiocchi D, Carnevale M, Giberti H. Integrating Real Tool Interaction and Multimodal Operator Monitoring in Immersive Simulation for Human-Centred Assessment. Electronics. 2026; 15(16):3525. https://doi.org/10.3390/electronics15163525

Chicago/Turabian Style

Fabiocchi, Davide, Marco Carnevale, and Hermes Giberti. 2026. "Integrating Real Tool Interaction and Multimodal Operator Monitoring in Immersive Simulation for Human-Centred Assessment" Electronics 15, no. 16: 3525. https://doi.org/10.3390/electronics15163525

APA Style

Fabiocchi, D., Carnevale, M., & Giberti, H. (2026). Integrating Real Tool Interaction and Multimodal Operator Monitoring in Immersive Simulation for Human-Centred Assessment. Electronics, 15(16), 3525. https://doi.org/10.3390/electronics15163525

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop