Next Article in Journal
A Grid-Forming Control Strategy Based on a Hybrid Approach Combining a Physical Model and LSTM for Photovoltaic and Energy Storage Systems
Previous Article in Journal
Green Vehicle Routing Model and Optimization Algorithm with Soft Time Window and Dynamic Demand
Previous Article in Special Issue
An Efficient Job Insertion Algorithm for Hybrid Human–Machine Collaborative Flexible Job Shop Scheduling with Random Job Arrivals
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Digital-Twin-Enabled Human–Machine Collaboration Systems in Sustainable Smart Manufacturing: System Architecture, Development Methods, Applications, and Future Trends

1
College of Mechanical Engineering, Nanjing University of Industry Technology, Nanjing 210023, China
2
College of Mechanical and Electrical Engineering, Nanjing University of Aeronautics and Astronautics, Nanjing 210016, China
*
Author to whom correspondence should be addressed.
These authors contributed equally to this work.
Electronics 2026, 15(17), 3781; https://doi.org/10.3390/electronics15173781
Submission received: 30 June 2026 / Revised: 9 August 2026 / Accepted: 20 August 2026 / Published: 24 August 2026
(This article belongs to the Special Issue Human–Robot Interaction and Communication Towards Industry 5.0)

Abstract

Digital-twin-enabled human–machine collaboration (HMC) has increasingly been proposed as a system-level approach for connecting human operators, robots, sensors, artificial intelligence modules, and manufacturing resources. However, the literature varies substantially in what is called a digital twin, how physical and virtual models are coupled, whether models are updated from physical data, and how far systems have progressed beyond simulation or controlled laboratory demonstrations. This structured integrative review examines the conditions under which a digital twin can function as an integration layer for HMC in sustainable smart manufacturing, rather than assuming that such integration is already established industrial practice. The literature corpus was assembled through searches of the Web of Science Core Collection, Scopus, and IEEE Xplore, complemented by Google Scholar-based citation tracking and backward and forward citation tracing. The core search focused on studies published from 1 January 2020 to 5 August 2026, while earlier seminal studies were retained to support definitions and historical context. Studies were screened using explicit criteria for manufacturing relevance, physical–virtual coupling, state synchronization or model updating, feedback capability, and validation setting, and were critically coded by model type, integration mechanism, deployment maturity, and sustainability evidence. The review compares multimodal perception and human-state modeling, intention understanding and augmented interaction, task allocation and shared planning, digital-twin architectures, adaptive control and safety verification, and human–AI decision-making. The evidence indicates that digital twins are promising as coordination and verification layers, but many reported systems remain conceptual, simulation-based, or limited to controlled physical prototypes. Key barriers include model fidelity, online model updating, real-time synchronization, cross-platform interoperability, safety assurance, human-data governance, and the limited availability of directly measured sustainability outcomes. Future work should prioritize validated hybrid models, traceable model-update mechanisms, staged virtual-to-physical deployment, interoperable data contracts, and longitudinal evaluation of technical, human, economic, and environmental performance.

1. Introduction

Manufacturing systems are being redesigned around shorter product life cycles, higher product variety, increased quality requirements, carbon reduction pressure, and the need for resilient supply chains. The shift toward Industry 5.0 provides a human-centric and sustainability-oriented framing for this redesign [1]. Conventional smart manufacturing has supported mass production by isolating robots and machines from human operators, stabilizing process parameters, and repeating predefined trajectories; comparative work on smart and intelligent manufacturing clarifies the data-driven foundations of this model [2]. This model remains effective for high-volume and low-variation production, yet it becomes less efficient when product configuration changes frequently, workpieces are uncertain, or local disturbances require immediate interpretation. Under these conditions, purely automated systems often need expensive reprogramming, additional fixtures, or long commissioning cycles. Human-centered production offers a complementary route because operators possess contextual knowledge, tactile judgment, abnormality recognition, and flexible problem-solving capabilities that are difficult to encode completely in fixed control logic. Human–machine collaboration therefore seeks to combine the strengths of human workers and machine agents rather than treating them as mutually exclusive resources, which is consistent with proactive human–robot collaboration research [3]. Interaction technologies such as AR further support this cooperation by making robot actions and human instructions more transparent on the shop floor [4].
The term human–machine collaboration in this review is used in a broad manufacturing sense. It includes physical cooperation between human workers and robotic systems, cognitive cooperation between human experts and AI-based decision-support systems, and production-level coordination among operators, machines, sensors, and management resources. A digital twin is not treated as an independent collaborating agent by default. Instead, it may provide a virtual representation and integration layer that supports state synchronization, model updating, simulation, verification, and coordination across these interactions. This distinction is important because the presence of a simulation, dashboard, AR interface, or AI module alone does not demonstrate a digital twin.
Human–machine collaboration is closely aligned with the transition from Industry 4.0 to Industry 5.0. Industry 4.0 emphasizes cyber–physical connectivity, data-driven control, and flexible automation, while Industry 5.0 places stronger emphasis on human-centricity, sustainability, and resilience. HMC can operationalize these principles by keeping workers in the decision loop, reducing physically demanding and repetitive operations, enabling data-supported training, and allowing production systems to adapt to disturbances without sacrificing human control. In assembly lines, the robot may hold, position, or fasten components while the operator performs alignment, verification, and exception handling. In quality inspection, AI may identify candidate defects while experts make final judgments on ambiguous cases. In maintenance and remanufacturing, digital guidance and robotic assistance can help workers extend product life and recover value from used components. These examples show that HMC has both technical and sustainability implications.
The sustainability value of HMC should be interpreted carefully. Collaboration can reduce ergonomic risk, minimize rework, shorten travel distance, optimize resource utilization, and avoid unnecessary energy consumption when task characteristics and agent capabilities are matched appropriately [5]. Multimodal perception can further support these benefits by improving human and robot awareness of collaborative operations [6]. At the same time, HMC may introduce additional sensing devices, computing infrastructure, data storage, and integration complexity. Augmented-reality interaction and digital-twin infrastructure may improve guidance and integration, but they also add interface complexity and modeling cost [7,8]. A collaborative system that improves task time but increases failure recovery time or cognitive workload may not be sustainable at the system level. Therefore, this review treats sustainability as a multi-dimensional objective rather than as a rhetorical label. Environmental performance, economic efficiency, operator well-being, safety, flexibility, and long-term maintainability must be considered together. This approach is consistent with the growing emphasis on life-cycle thinking and human-centered design in smart manufacturing.
Existing research on human–robot collaboration, multimodal perception, task allocation, digital twins, adaptive control, and human–AI decision-making has grown rapidly, but these topics are often studied in isolation and reported at very different levels of maturity. Some studies evaluate only module-level recognition or planning performance; others provide virtual models without continuous physical data; and only a limited subset demonstrates bidirectional, sustained, and industrially validated physical–virtual coupling. Consequently, what remains insufficiently understood is not whether these technologies can be drawn within a common architecture, but the extent to which they have actually been integrated through digital twins, how the physical and virtual models are constructed and updated, and whether the reported systems represent conceptual architectures, simulations, controlled prototypes, pilot systems, or deployable industrial solutions.
This review therefore does not assume that the digital twin has already become a universal integration backbone for HMC. Instead, it critically examines the technical conditions, evidence levels, and deployment maturity required for a digital twin to serve as an effective integration layer. The contribution is fourfold. First, the review provides an operational boundary that distinguishes a digital model, a one-way digital shadow, and a bidirectionally coupled digital twin, and separates core digital-twin evidence from adjacent enabling technologies. Second, it compares model-based, data-driven, and hybrid approaches, with particular attention to the distinction between state synchronization and model updating from physical data. Third, it analyzes how perception, human modeling, planning, interaction, safety control, and human–AI decision-making are connected—or remain disconnected—within physical–virtual loops. Fourth, it evaluates industrial applications using deployment maturity and sustainability-evidence categories rather than treating laboratory feasibility or operational improvement as equivalent to industrial sustainability. The remainder of the review is organized into nine sections: Section 1 introduces the background and scope; Section 2 explains the review methodology and analytical framework; Section 3 defines the conceptual boundary and HMC taxonomy; Section 4 reviews enabling technologies and their actual connection to digital twins; Section 5 examines digital-twin architecture, model construction, model updating, synchronization, and interoperability; Section 6 analyzes adaptive control, safety assurance, and human–AI decision-making; Section 7 reviews applications, deployment maturity, and sustainability evidence; Section 8 identifies challenges and future research; and Section 9 concludes the review. The analytical structure used for this synthesis is summarized in Figure 1.

2. Review Methodology and Analytical Framework

2.1. Review Design and Research Questions

This study adopts a structured integrative review design. The objective is not to calculate a pooled effect size, because the reviewed studies differ substantially in tasks, hardware, human participants, digital-twin definitions, and evaluation metrics. Instead, the protocol is designed to make literature discovery, eligibility decisions, critical comparison, and evidence classification transparent. The reporting structure follows the transparency principles of PRISMA 2020 and PRISMA-S, adapted to an interdisciplinary integrative review rather than a clinical meta-analysis [9,10]. Five research questions guide the synthesis: (RQ1) How are digital twins defined and implemented in manufacturing HMC? (RQ2) What physics-based, data-driven, and hybrid models are used to represent physical humans, machines, products, and processes? (RQ3) How are twin states and model parameters updated using physical data? (RQ4) How are perception, interaction, planning, control, safety verification, and human–AI decision-making connected to the twin? (RQ5) What levels of validation, industrial deployment, and sustainability evidence have been reported?

2.2. Information Sources and Search Strategy

The primary literature search used the Web of Science Core Collection, Scopus, and IEEE Xplore. Google Scholar was used only for supplementary citation tracking, recently published-item checking, and locating full bibliographic records. Backward reference checking and forward citation tracing were applied to highly relevant studies. The core search window covered 1 January 2020 to 5 August 2026; earlier studies were retained only when they provided foundational definitions, widely used methods, or historical context necessary for interpreting the recent literature. The search combined three concept groups: (i) digital-twin terms (“digital twin”, “digital shadow”, “physical–virtual synchronization”, and “human digital twin”); (ii) collaboration terms (“human–machine collaboration”, “human–robot collaboration”, “human–AI collaboration”, “human-in-the-loop”, and “cobot”); and (iii) manufacturing terms (manufacturing, production, assembly, inspection, maintenance, disassembly, remanufacturing, logistics, smart factory, and Industry 5.0). The Boolean structure was DIGITAL-TWIN TERMS AND COLLABORATION TERMS AND MANUFACTURING TERMS, with database-specific field and syntax adaptations.

2.3. Eligibility and Screening

Studies were considered core digital-twin evidence when they met all of the following conditions: (1) a manufacturing or production-related HMC context; (2) an identifiable physical entity or process and a corresponding virtual representation; (3) an explicit physical-to-virtual data connection; and (4) a described mechanism for state synchronization, model updating, virtual verification, decision support, or physical feedback. Studies on multimodal perception, AR/MR, task allocation, LLMs, adaptive control, or human-state modeling that did not satisfy these criteria were retained only as adjacent enabling evidence when they clarified a module that could be connected to a twin. Studies were excluded when they only mentioned digital twins rhetorically, used stand-alone simulation without a physical counterpart, addressed non-manufacturing domains, lacked sufficient technical information, or duplicated another publication. Titles and abstracts were first screened for domain relevance; full texts were then assessed against the physical–virtual coupling and evidence criteria.

2.4. Critical Coding and Evidence Classification

Each included study was coded along four analytical dimensions. First, the virtual representation was classified as physics-based, data-driven, hybrid, geometric/kinematic, process/logical, or human-state oriented. Second, physical–virtual coupling was coded by data source, synchronization mode, update frequency, state-update mechanism, model-parameter or structure update, and virtual-to-physical feedback. Third, validation maturity was classified as Level 1 (conceptual architecture), Level 2 (simulation or offline validation), Level 3 (controlled physical prototype), or Level 4 (pilot or sustained industrial deployment). Fourth, sustainability evidence was classified as directly measured outcomes, indirect operational outcomes with a plausible sustainability pathway, or aspirational claims without corresponding measurements. These categories are used throughout the review to prevent conceptual architectures, laboratory prototypes, and industrial deployments from being treated as equivalent evidence.

2.5. Methodological Limitations

The review has two methodological limitations. First, the original corpus was developed iteratively across several interdisciplinary searches rather than through a prospectively registered single-query protocol; exact database-specific yield counts for every exploratory query were therefore not retained. To avoid retrospective pseudo-precision, this revision reports the information sources, search window, search logic, eligibility criteria, and coding framework, rather than inventing record counts. The final reference base contains 104 sources, including methodological and foundational references, while the chapter tables identify representative core and adjacent studies used in the critical synthesis. Second, heterogeneity in terminology, task design, hardware, and reporting quality limits formal quantitative comparison. The synthesis therefore emphasizes traceable evidence categories, explicit deployment maturity, and study-level limitations rather than meta-analysis.

3. Conceptual Foundations, System Boundary, and Taxonomy of Digital-Twin-Enabled HMC Systems

Before discussing the conceptual architecture, the system boundary must distinguish HMC functions from the digital-twin mechanism used to connect them. In this review, HMC refers to coordinated human and machine perception, decision, and execution through structured information exchange and authority sharing. Digital-twin-enabled HMC is a narrower subset: the physical work system must have an identifiable virtual representation and an explicit data connection that supports synchronization, model updating, virtual verification, decision support, or feedback. Multimodal perception, AR/MR, task allocation, adaptive control, or LLM-based planning are therefore not automatically classified as digital-twin components; they become core evidence only when their inputs, models, or outputs are explicitly coupled to the twin.

3.1. Operational Definition and Inclusion Boundary

A first distinction concerns deployment maturity. A controlled laboratory cell can demonstrate feasibility under stable lighting, known communication networks, limited product variation, supervised operation, and short trials. It is not equivalent to an industrially deployable architecture, which also requires sustained synchronization, interface maintenance, cybersecurity, failure recovery, operator training, safety certification, and evidence under production disturbances. The maturity levels defined in Section 2.4 are therefore applied when interpreting reported results.
A second distinction concerns state updating and model updating. Updating a robot joint angle, operator position, workpiece status, or task-progress variable changes the current state represented by the twin. Updating model parameters, model structure, uncertainty bounds, or learned representations from newly observed physical behavior changes the model itself. The latter may involve parameter identification, data assimilation, online calibration, incremental learning, or hybrid physics–data adaptation. Many published systems demonstrate state synchronization but provide little evidence of model updating; this difference is explicitly considered in the critical synthesis.
A third distinction concerns the degree of physical–virtual coupling. For the purposes of this review, a digital model is a virtual representation that is updated manually or offline and has no automatic physical data connection. A digital shadow receives automatic or event-driven information from the physical system, but the information flow is predominantly physical-to-virtual. A digital twin requires a maintained physical–virtual association and uses synchronized data for analysis, prediction, verification, or decision support; where virtual outputs influence physical execution, the system additionally exhibits a closed-loop or bidirectional capability. This operational distinction follows the manufacturing-oriented classification of digital models, digital shadows, and digital twins [11] and the characterization of physical-to-virtual and virtual-to-physical twinning processes [12]. It is used here to avoid labeling every simulation, visualization, or AI-enabled interface as a digital twin. The resulting taxonomy and system boundary are summarized in Figure 2.
HMC can be understood as a socio-technical control system with four interacting layers. The first layer is the physical execution layer, where humans, robots, tools, fixtures, products, and workstations interact. The second layer is the perception and communication layer, where sensors, wearable devices, human interfaces, and data networks collect and transmit information. The third layer is the decision and planning layer, where task models, optimization algorithms, AI agents, and human supervisors generate plans and interventions. The fourth layer is the evaluation and governance layer, where safety rules, sustainability targets, authority policies, and performance indicators constrain and assess the system. Treating HMC as a layered system prevents the review from reducing collaboration to one device or one algorithm.
The difference between automation and collaboration lies in the distribution of uncertainty and authority. Automation aims to execute predefined rules with high stability. Collaboration aims to coordinate complementary capabilities under uncertainty. A robot may be fast and precise, but it may not understand implicit production context or abnormal workpiece conditions. A human operator may understand context and priorities, but may be slower, more variable, and exposed to physical fatigue. HMC attempts to allocate tasks and authority according to the capability profile of each agent. This allocation may change over time as the environment, workload, skill level, and process state evolve. Therefore, HMC requires dynamic coordination rather than static division of labor.
A useful taxonomy should include collaboration object, collaboration level, task coupling, autonomy distribution, interaction channel, and sustainability contribution. Collaboration object distinguishes whether the human collaborates with a robot, an AI decision system, a digital twin, or a production resource network. Collaboration level distinguishes coexistence, sequential cooperation, synchronized collaboration, and shared autonomy. Task coupling describes whether tasks are independent, weakly coupled through timing, strongly coupled through shared workpieces, or physically coupled through direct contact. Autonomy distribution describes whether the human leads, the machine assists, both negotiate, or the machine leads under human supervision. Interaction channel includes physical contact, visual guidance, speech, gesture, haptic feedback, augmented reality, and digital dashboards. Sustainability contribution includes ergonomic improvement, waste reduction, energy saving, quality improvement, and life-cycle extension.
The architecture of HMC should also include feedback loops. The first feedback loop links machine execution and human observation: the operator monitors robot behavior, detects abnormality, and intervenes if needed. The second loop links human behavior and machine adaptation: the system observes human posture, workload, and intent, then adjusts speed, trajectory, interface content, or task sequence. The third loop links physical production and virtual models: digital twins update the state of products, tools, and resources, allowing simulation and verification before execution. The fourth loop links performance evaluation and system redesign: production data, safety incidents, quality metrics, and sustainability indicators are used to redesign the collaborative process. Without these loops, HMC remains a static automation cell with superficial human presence.
The most important design principle is complementarity. A collaborative system should not assign tasks to humans merely because machines cannot yet perform them, nor should it assign tasks to machines simply because automation is available. Instead, allocation should be based on task requirements, human skill, ergonomic risk, machine capability, safety constraints, quality sensitivity, and sustainability goals. In dynamic workstations, this allocation also depends on predicting human motion and updating digital representations of the collaborative process [13,14]. For example, heavy lifting, repetitive fastening, or high-frequency inspection may be machine-suitable, while flexible fitting, final judgment, and abnormal recovery may remain human-centered. This principle is consistent with review work on cobot deployment in manufacturing [15]. In high-mix production, the optimal allocation may change across product variants and operator skill levels. Therefore, HMC requires formal task models and dynamic allocation rules [16], human-activity prediction for anticipating operator behavior [17], and task-classification mechanisms for updating the division of labor [18].

3.2. Human-Centered Design Logic

Human-centered design is the foundation of HMC. It requires the system to be designed around the capabilities, limitations, and responsibilities of the human worker rather than around the robot alone. In early industrial robotics, safety was achieved mainly through separation. Humans and robots worked in different spaces, and the robot was programmed for repeatable execution. HMC changes this arrangement by allowing proximity, shared tasks, and dynamic adaptation. However, proximity should not be confused with human-centeredness. A robot placed near a human worker is not necessarily collaborative. Collaboration requires mutual awareness, understandable interaction, safe behavior, and an appropriate distribution of control authority. This human-centered design logic is summarized in Figure 3.
Human-centered HMC should reduce physical and cognitive burden rather than transferring hidden workload to the operator. A poorly designed collaborative cell may require workers to monitor unpredictable robot behavior, respond to confusing interface messages, or correct machine errors continuously. This increases cognitive load and may reduce trust. Good design should provide clear status information, predictable robot motion, intuitive intervention mechanisms, and training support. The operator should understand what the machine is doing, why it is doing it, when it may fail, and how to stop or adjust it safely. Explainability is therefore not only an AI requirement; it is a manufacturing requirement.
Another human-centered issue is skill development. HMC may change the role of workers from manual executors to supervisors, trainers, quality judges, or maintenance collaborators. This role shift can increase job quality if workers receive adequate training and authority, but it can also create skill gaps if deployment focuses only on equipment installation. Sustainable HMC should include training workflows, feedback mechanisms, and adaptive interfaces that match operator experience. Novice operators may need step-by-step augmented instructions, while experts may prefer compact status panels and direct control modes. A flexible interface design supports both productivity and worker acceptance.

3.3. Collaboration Mechanisms and Sustainability Linkages

HMC contributes to sustainability through several mechanisms. At the workstation level, robots can reduce repetitive strain, heavy lifting, awkward postures, and exposure to hazardous environments. At the process level, collaborative perception and decision-making can reduce defects, rework, and scrap by detecting abnormal conditions earlier. At the system level, digital twins and task planning can shorten commissioning time, reduce unnecessary trial-and-error, and improve resource utilization. At the product life-cycle level, collaborative disassembly and remanufacturing can recover materials and components more efficiently. These mechanisms show that sustainability is not an add-on; it is embedded in task design, system architecture, and evaluation. The main sustainability linkages are summarized in Figure 4.
Nevertheless, the sustainability effect of HMC must be measured rather than assumed. A robot that reduces operator fatigue may consume additional energy, require maintenance, and increase system complexity. A vision system that improves quality may generate data management overhead. A digital twin that supports optimization may require modeling effort and continuous data synchronization. Therefore, the sustainability assessment of HMC should include life-cycle considerations, such as equipment utilization, energy consumption, material saving, defect reduction, training cost, maintenance effort, and worker well-being. Only a multi-dimensional assessment can determine whether a collaborative system provides net value.
The link between collaboration mechanisms and sustainability also depends on deployment scale. A single collaborative workstation may show limited environmental benefit, but if the same architecture reduces rework or improves recovery across a production line, the system-level benefit can be substantial. Conversely, a high-cost collaborative cell may not be justified if the product mix is stable and task complexity is low. HMC is most valuable where uncertainty, variation, ergonomic risk, and quality sensitivity are high. This explains why collaborative assembly, inspection, maintenance, logistics, disassembly, and remanufacturing have become major application areas.
Leng et al. [1] clarified the human-centric and sustainability-oriented direction of Industry 5.0 for future manufacturing. In addition, Wang et al. [2] compared smart and intelligent manufacturing, clarifying the data-driven basis of intelligent production. Li et al. [3] framed proactive human–robot collaboration as a cognitive manufacturing paradigm enabling anticipatory cooperation. Further, Hietanen et al. [4] developed an augmented-reality interaction scheme that makes robot actions and human instructions more transparent on the shop floor. Liau et al. [5] proposed a task-allocation method for human–robot collaboration based on task characteristics and agent capability. In a related effort, Duan et al. [6] built a dual-robot intelligent assembly system that fuses multimodal perception for collaborative operation.
Blankemeyer et al. [7] introduced a hand-interaction model that enhances augmented-reality-based human–robot collaboration. In addition, Malik et al. [19] proposed a complexity-based task-allocation method for human–robot collaborative assembly. Kong et al. [20] investigated human–robot joint task assignment under varying task complexity. Further, Sleeman et al. [21] surveyed multimodal classification, summarizing its landscape, taxonomy, and future directions. Xue et al. [22] reviewed human–machine augmented intelligence and its industrial applications.
Critical synthesis. Existing HMC taxonomies are useful for describing collaboration objects, autonomy, task coupling, and interaction channels, but most do not specify the physical–virtual coupling required for a digital twin. Conceptual frameworks offer broad coverage and support system design, yet they can overstate integration if each listed module has only been validated independently. Human-centered taxonomies are strongest in authority, ergonomics, and interaction design, whereas digital-twin taxonomies are stronger in state representation, synchronization, and lifecycle traceability. A deployable architecture must combine both perspectives and explicitly state which links are implemented, which are inferred from adjacent studies, and which remain aspirational. Representative studies are summarized in Table 1.

4. Enabling Technologies and Their Connection to Digital Twins

The enabling technologies reviewed in this section provide sensing, interpretation, interaction, planning, and decision functions that may be connected to a digital twin. They should not be treated as proof of digital-twin integration by themselves. A perception model becomes part of a twin when its outputs update an identified virtual state or model; a planning or LLM module becomes part of the loop when its proposals are verified against synchronized twin data; and an AR/MR interface becomes twin-enabled when it visualizes or modifies traceable physical–virtual information. This boundary separates core digital-twin evidence from adjacent HMC technologies while still allowing the latter to be compared as potential integration modules.
The technology stack can be divided into three categories. The first category is sensing and interpretation. It includes cameras, depth sensors, LiDAR, force/torque sensors, tactile sensors, microphones, wearable devices, eye tracking, electromyography, and human skeleton estimation. These signals are used to identify objects, recognize actions, infer intent, estimate fatigue, and detect safety risks. The second category is interaction and coordination. It includes speech interfaces, gesture control, teach-by-demonstration, augmented reality, mixed reality, haptic feedback, task allocation, and shared planning. The third category is integration and assurance. It includes digital twins, edge computing, safety controllers, adaptive motion control, explainable AI, standards, and performance evaluation. A mature HMC system requires all three categories to operate together.
A key trend is the movement from rule-based systems to data-driven and model-based hybrid systems. Early collaborative systems relied heavily on predefined work zones, fixed logic, and manually programmed task sequences. Recent systems use deep reinforcement learning to improve task allocation under changing manufacturing states [23]. CNN-based and RNN-based methods support action recognition and intention prediction in collaborative assembly [24,25]. Vision-based motion recognition and skeleton-based action recognition further improve context awareness for proactive assistance [26,27]. Hybrid recurrent architectures can reduce intent-prediction errors [28], while task-complexity metrics keep allocation decisions aligned with human capability and ergonomic constraints [19]. However, data-driven intelligence does not eliminate the need for engineering constraints. Industrial systems require predictable timing, verified safety, maintainable code, traceable decisions, and robust fallback modes. Therefore, a practical HMC architecture should combine learning-based perception and reasoning with deterministic safety logic and formal process constraints.

4.1. Multimodal Perception and Workspace Understanding

Multimodal perception is the entry point of HMC because collaboration requires machines to understand the state of humans, objects, tools, and the environment. Vision sensors provide object recognition, pose estimation, workpiece localization, human skeleton tracking, and spatial mapping. Depth cameras and LiDAR can reconstruct three-dimensional relationships and support collision avoidance. Force and torque sensors capture contact states, applied loads, and assembly conditions. Microphones and natural language processing modules capture spoken commands and operator feedback. Wearable sensors and physiological signals can provide information about posture, fatigue, workload, and stress. Each modality has limitations, but their combination can improve robustness in dynamic industrial environments. The corresponding sensor-fusion pipeline is illustrated in Figure 5.
Vision-based perception has received extensive attention because it provides rich spatial information. RGB images can identify tools, parts, and visual states, while depth maps can estimate distance and support safety monitoring. Skeleton tracking enables action recognition, gesture interpretation, and prediction of future human motion. In collaborative assembly, a robot may use human pose and hand trajectory to infer whether the operator is reaching for a part, placing a component, or waiting for robotic assistance. In inspection, cameras can detect surface defects and guide human attention to suspected areas. However, industrial vision is sensitive to illumination, occlusion, reflective surfaces, clutter, and domain shift. A model trained on a laboratory dataset may perform poorly on the shop floor if the camera angle, workpiece appearance, or background changes.
Force and tactile sensing are essential when collaboration involves contact. In co-manipulation, the robot must infer human intention from interaction forces and adjust impedance accordingly. In assembly, force profiles can indicate insertion success, misalignment, jamming, or excessive contact. Tactile arrays can detect local contact distribution, while torque observers can estimate unexpected collisions without additional external sensors. These signals complement vision because they remain informative when the contact state cannot be visually observed. Nevertheless, force-based interpretation is often task-specific. The same force magnitude may indicate normal operation in one process and abnormal contact in another. Therefore, force perception should be combined with task context and process knowledge.
Speech and language interfaces are attractive because they allow operators to communicate naturally with machines. Spoken instructions can specify task goals, part attributes, sequence changes, or abnormal conditions. Language also supports explanation: a machine can report why a task is delayed, why a trajectory is modified, or what safety condition has been triggered. The challenge is that industrial language is often incomplete, context-dependent, noisy, and mixed with gestures or visual references. An instruction such as “hold this side first” cannot be interpreted without knowing the current workpiece and human hand position. Multimodal grounding is therefore required to connect language with perception and task models.
Human-state perception extends workspace understanding from external actions to internal conditions. Posture, joint angles, muscle activation, fatigue, attention, and stress may affect task performance and safety. Ergonomic assessment methods can be combined with skeleton tracking to estimate awkward postures or accumulated load. Physiological signals may indicate fatigue or high workload, although they require careful calibration and privacy protection. In sustainable HMC, human-state perception should be used to improve work design and reduce risk, not to impose excessive surveillance. The design should clarify what data are collected, how they are used, who can access them, and how operator consent is handled.
Data fusion is the technical core of multimodal perception. Early fusion combines raw signals or low-level features, intermediate fusion integrates learned representations, and late fusion combines decisions or confidence scores. Attention mechanisms and transformer architectures have improved cross-modal alignment, allowing models to learn relationships among vision, language, force, and temporal context. Pre-trained foundation models provide strong representation capabilities, but they must be adapted to industrial domains where data are limited, classes are specialized, and safety consequences are high. The most reliable approach is often a hybrid architecture: deep models provide flexible recognition, while rule-based constraints and process models verify feasibility. The relationship between observed modalities, human-state modeling, and applications is summarized in Figure 6.

4.2. Intention Understanding and Human–Machine Interaction

Perception becomes useful for collaboration only when it is converted into intention understanding. Intention refers to the human operator’s likely goal, next action, or desired assistance. It can be inferred from hand trajectory, gaze direction, object state, task sequence, speech command, interaction force, and historical behavior. For example, if the operator reaches toward a specific part while the assembly sequence requires that part next, the robot can prepare the corresponding tool or adjust its position. If the operator slows down, pauses, or repeatedly corrects a component, the system may infer uncertainty or abnormality. Intention understanding allows the machine to act proactively rather than waiting for explicit commands. The resulting bidirectional interaction loop is illustrated in Figure 7.
Several modeling approaches are used for intention inference. Probabilistic models represent uncertainty and update beliefs as new observations arrive. Sequence models such as recurrent neural networks and temporal convolutional networks capture action evolution over time. Graph models represent relationships among humans, tools, parts, and tasks. Knowledge graphs can encode process rules, part relationships, and safety constraints. Large language models and vision-language models can interpret textual or visual task descriptions, but their outputs must be grounded in verified process data. In industrial applications, intention inference should be conservative. A wrong proactive action may cause safety risks or process errors, so the system should maintain confidence estimates and request human confirmation when uncertainty is high.
Interaction technology provides the communication channel between human intention and machine behavior. Traditional interfaces include buttons, teach pendants, foot switches, touch panels, and emergency stops. These remain important because they are predictable and safety-certified. Newer interfaces include speech, gesture, gaze, haptic feedback, wearable displays, augmented reality, virtual reality, and mixed reality. Augmented interaction can overlay robot trajectory, assembly sequence, torque limit, quality alert, or safe zone information onto the physical scene. This reduces the need for operators to switch attention between the workpiece and a separate screen. It also supports training and remote assistance.
AR, VR, and MR interfaces have different roles. AR is useful for in-situ guidance because it superimposes information onto the real workstation. VR is useful for training, simulation, and remote process design because it can reproduce the workcell without interrupting production. MR combines physical and virtual objects more tightly, allowing interaction with both real and virtual resources. In HMC, these technologies can visualize hidden machine states, planned paths, danger zones, part placement positions, and decision explanations. Their limitations include device weight, field of view, calibration drift, user discomfort, and integration with industrial safety systems. Therefore, augmented interaction should be designed around task value rather than visual novelty.
Human–machine interaction should support bidirectional adaptation. Humans give instructions, corrections, and approvals; machines provide status, warnings, predictions, and explanations. A collaborative system should avoid both under-communication and over-communication. If the system gives too little information, the operator may not trust or understand it. If it gives too much information, the operator may become distracted. Adaptive interfaces can adjust the level of detail according to task complexity, operator skill, risk level, and current workload. For high-risk operations, the system may require explicit confirmation. For routine operations, it may provide compact status information and allow quick override.
Trust is a central interaction issue. Trust depends on reliability, predictability, transparency, and the operator’s experience with the system. Excessive trust may lead to automation complacency, while insufficient trust may cause operators to ignore useful assistance. Explanations can improve trust when they are concise, task-specific, and linked to observable evidence. For example, a robot path modification can be explained by a detected human hand entering a safety zone. A task allocation change can be explained by predicted operator fatigue or machine availability. Such explanations are more useful than generic statements about optimization. HMC systems should therefore include explanation design as part of interface design.

4.3. Task Allocation and Shared Planning

Task allocation determines which agent performs which subtask, when it is performed, and under what constraints. In HMC, allocation cannot be based only on machine cycle time. It must consider human skill, ergonomic load, safety risk, task complexity, part availability, machine capability, quality requirements, and production objectives. A task may be assigned to a robot because it is repetitive and heavy, to a human because it requires flexible judgment, or to both because it requires simultaneous holding and fastening. Allocation may be static during process design or dynamic during execution. Dynamic allocation is increasingly important for high-mix production and disturbance recovery. A hierarchical allocation and shared-planning structure is shown in Figure 8.
Task allocation usually begins with task decomposition. A manufacturing operation is divided into subtasks such as fetching parts, positioning, aligning, fastening, inspecting, and transferring. Each subtask is described by attributes, including required force, precision, workspace, tool, duration, sequence dependency, safety risk, and knowledge requirement. Human and machine resources are described by capability models, including reachability, payload, speed, accuracy, skill, availability, and current workload. Allocation then becomes a constrained optimization or decision-making problem. The objective may minimize cycle time, balance workload, reduce ergonomic risk, or improve energy efficiency.
Optimization-based allocation methods include integer programming, constraint programming, heuristic search, genetic algorithms, and scheduling algorithms. These methods are useful when task structure and constraints are explicit. Learning-based methods, including reinforcement learning and imitation learning, can adapt allocation decisions from data or simulation. Deep learning can support action recognition, duration prediction, and human-state estimation for allocation. However, learning-based allocation must be interpretable and safe. A production supervisor needs to know why a subtask was reassigned and whether the new assignment violates safety or quality constraints. Therefore, allocation algorithms should be connected with process knowledge and human approval mechanisms.
Shared planning extends task allocation by coordinating the temporal and spatial actions of agents. In collaborative assembly, the robot may need to wait for the operator to finish alignment, approach only when the operator’s hand has moved away, or adjust trajectory based on the operator’s position. Shared planning must handle precedence constraints, mutual exclusion zones, synchronization points, and recovery actions. It should also consider communication delays and perception uncertainty. A plan that is optimal in simulation may fail if the human action time varies or if sensor detection is delayed. Robust planning therefore includes buffers, alternative actions, and replanning triggers.
Large language models have introduced new possibilities for task planning because they can parse natural language instructions, decompose tasks, and generate high-level plans. When connected with knowledge bases and digital twins, they can help transform human descriptions into executable workflows. However, LLM-based planning has risks. Models may generate plausible but infeasible plans, ignore safety constraints, or rely on incomplete context. For industrial HMC, LLMs should not directly control robots without verification. A safer architecture uses LLMs for semantic interpretation and plan proposal, while deterministic planners, digital twins, and safety controllers verify feasibility before execution. Human approval remains necessary for critical operations.
The quality of task allocation and planning should be evaluated beyond task time. Important indicators include human idle time, robot idle time, workload balance, ergonomic score, intervention frequency, plan stability, recovery time, quality defects, and energy consumption. A system that minimizes cycle time by forcing the operator to monitor the robot continuously may not be acceptable. Similarly, a plan that reduces human workload but increases robot travel distance and energy use may not be sustainable. Multi-objective evaluation is therefore essential. The most useful allocation methods are those that make trade-offs explicit and allow engineers to tune priorities according to production goals.
Ding et al. [13] developed a scenario-enhanced network for diverse human motion prediction in proactive collaboration. In addition, Cao et al. [29] examined AI-driven design of soft robots for adaptive interaction. Hussain et al. [30] developed a human-centric attention model with deep multiscale feature fusion for activity recognition. Further, Zhang et al. [31] proposed a skeleton–RGB integrated method for predicting highly similar human actions in collaboration. Nadeem et al. [32] combined vision-enabled large language and deep learning models for image-based emotion recognition. In a related effort, Hazmoune et al. [33] reviewed transformer-based approaches to multimodal emotion recognition.
Liu et al. [34] designed a transformer-encoder multimodal fusion network for sentiment analysis. In addition, Sun et al. [35] explored AI-enabled flexible sensing systems for human-state perception. Li et al. [36] introduced a self-supervised multi-level correlation framework for multimodal learning. Further, Wang et al. [37] developed a data-efficient multimodal scheme for human action recognition in proactive collaboration. Yang et al. [38] outlined technical enablers for natural-language-based human–robot collaboration. In a related effort, Balamurugan et al. [39] demonstrated wearable, sensor-enabled augmented interaction for human–machine systems.
Wang et al. [40] proposed IndVisSGG, a VLM-based scene-graph-generation method for industrial spatial intelligence.
Critical synthesis. Vision-based methods provide rich spatial information and integrate naturally with geometric representations, but their performance is sensitive to illumination, occlusion, reflective surfaces, camera relocation, and domain shift. Force and tactile sensing are more informative for contact-intensive operations, although their interpretation is highly task-dependent and often requires process context. Wearable and physiological sensing can enrich human-state models, but calibration burden, privacy, worker acceptance, and inter-individual variability limit deployment. For task allocation, optimization-based approaches offer explicit constraints and traceable decisions, whereas learning-based approaches adapt more readily to uncertainty but require data, uncertainty estimation, and independent safety verification. LLM-based planning improves semantic flexibility but remains unsuitable for direct physical control without deterministic feasibility and safety checks. Across these categories, many studies report module-level accuracy or task time without demonstrating that the result updates a digital-twin model or closes a physical feedback loop. Representative enabling-technology studies are summarized in Table 2.

5. Digital-Twin-Enabled System Architecture and Integration

This section examines the digital twin as a potential system-level integration and verification layer for HMC. The term integration backbone is treated as a target architectural role rather than an established capability of every reviewed system. Evidence is therefore evaluated according to the maintained physical–virtual association, model construction, state and model updating, synchronization latency, interoperability, feedback pathway, and validation maturity.

5.1. Digital-Twin Definition, Model Construction, and Updating

Digital-twin models in HMC can be grouped into physics-based, data-driven, and hybrid approaches. Physics-based and kinematic models offer interpretability, enforce geometric or process constraints, and support safety or feasibility verification, but they are expensive to construct and may omit human variability or unmodeled disturbances. Data-driven models can capture complex correlations in perception, human behavior, quality, and equipment condition, but they depend on representative data and may extrapolate poorly. Hybrid models combine mechanistic constraints with learned residuals, parameter estimation, or adaptive components and are particularly promising when industrial traceability and adaptability are both required.
The physical-to-virtual update mechanism determines whether the virtual representation remains decision-relevant. Periodic sampling is suitable for slow process and resource states; event-driven updates can reduce unnecessary communication; and continuous or near-real-time streaming is required for dynamic safety and motion functions. However, a high update frequency does not guarantee fidelity. Timestamp alignment, coordinate-frame consistency, missing data, sensor drift, communication delay, and uncertainty propagation must be managed explicitly.
Model updating is more demanding than ordinary state evolution. State synchronization changes variables such as position, force, workload, or task progress. Model updating revises parameters, uncertainty, or structure in response to observed physical behavior. Parameter identification and online calibration are interpretable but require informative excitation and stable measurements; online learning can adapt to complex changes but raises stability, validation, and certification concerns. The reviewed literature contains substantially more evidence of state synchronization than of traceable online model updating, which remains a major research gap.
The role of digital twins in HMC can be divided into three levels. At the workstation level, a twin models the layout, robot, tools, fixtures, workpiece, and human position. It can detect reachability, collision risk, and process deviation. At the line level, a twin models task flows, resource availability, buffers, and scheduling constraints. It can evaluate dynamic allocation and production resilience. At the life-cycle level, a twin links product design, process planning, operation, maintenance, disassembly, and remanufacturing. It can support sustainability assessment by tracing energy use, defects, rework, and material recovery. These levels may be integrated gradually according to deployment maturity.
A digital-twin-enabled HMC architecture usually contains physical entities, data acquisition, communication, virtual models, decision modules, and feedback control. Physical entities include humans, robots, products, tools, and sensors. Data acquisition collects geometry, motion, force, status, quality, and human-state information. Communication protocols transmit data with known latency and reliability. Virtual models reconstruct the scene and process state. Decision modules use simulation, optimization, AI, or rule-based logic to recommend actions. Feedback control sends verified commands to robots, interfaces, or management systems. The key requirement is synchronization: the virtual model must remain sufficiently consistent with the physical process for decisions to be useful.
Digital twins also provide a bridge between engineering design and production operation. During design, the twin can simulate collaborative layouts, test safety zones, evaluate ergonomics, and compare task allocations before equipment is installed. During commissioning, it can verify robot paths, detect collisions, and reduce trial-and-error on physical hardware. During operation, it can monitor progress, update state variables, and trigger replanning. During maintenance, it can diagnose abnormal patterns and guide technicians. This full-cycle role is important for sustainability because many environmental and economic costs are determined before production starts, while many human factors emerge during operation.
Another important function is virtual validation of AI-generated decisions. As LLMs and learning-based planners enter manufacturing, digital twins can act as a verification layer. A proposed task sequence can be checked against assembly constraints, robot reachability, tool availability, and safety zones. A proposed robot path can be simulated before execution. A proposed schedule can be evaluated for resource conflicts and human workload. This reduces the risk of directly applying unverified AI outputs to physical systems. In high-risk HMC, digital twins should be integrated with safety controllers and human approval mechanisms rather than treated as visualization tools only.
The main barriers to digital-twin-enabled HMC are modeling cost, data consistency, interface interoperability, and real-time performance. Building accurate models of robots, fixtures, tools, products, and human behavior requires engineering effort, especially when optimization and line-balancing constraints must be represented [41]. Maintaining synchronization is difficult when sensors fail, communication delays occur, or manual operations are not fully observed; digital-twin-driven assembly studies show both the value and the difficulty of keeping virtual and physical states aligned [42]. Transfer learning between virtual and physical systems can reduce trial cost, but it also increases model validation and data consistency requirements [43]. Interoperability is limited because different devices and software platforms use different data formats and protocols. Generative AI and LLM-based task modeling can automate parts of planning, but their outputs still require validation before execution [44,45]. MLLM/VLM and machine-vision systems extend perception–decision–execution capabilities [46,47]. Large-model decision-making highlights the need to connect reasoning with shop-floor authority [48]. Task-complexity-based allocation emphasizes the link between planning logic and production constraints [20], while digital-twin-driven assembly shows how simulation must be tied to executable shop-floor workflows [49]. Real-time performance becomes challenging when high-fidelity simulation, AI inference, and control must occur within short cycle times. These barriers explain why many digital twin studies remain at demonstration level rather than full industrial deployment. A reference architecture and its critical integration points are summarized in Figure 9.

5.2. Human Digital Twins and Human-State Representation

A human digital twin represents the operator’s state, capability, and interaction with the production system. It may include anthropometric data, posture, motion, workload, fatigue, skill level, task history, and interaction preferences. In HMC, the human digital twin can support ergonomic assessment, task allocation, training, safety monitoring, and adaptive interface design. For example, if the system detects that a worker has maintained an awkward posture for too long, it can reassign a lifting task to a robot or adjust workstation height. If the operator is a novice, the interface can provide more detailed instructions. If the operator is experienced, it can reduce unnecessary prompts. The corresponding human digital twin modeling and feedback structure is illustrated in Figure 10.
Human digital twins raise technical and ethical challenges. Human behavior is variable, context-dependent, and difficult to model with the same precision as machine kinematics. Physiological signals may be noisy and sensitive to individual differences. Continuous monitoring can create privacy concerns and resistance if workers feel surveilled rather than supported. Therefore, human digital twins should be designed with data minimization, transparency, and consent. The model should collect only the data required for safety, ergonomics, and process improvement. Its outputs should be used to support workers and improve work design, not to penalize normal human variability.
The accuracy of a human digital twin should be assessed according to its purpose. A model used for ergonomic risk screening does not need millimeter-level motion accuracy if it can reliably identify high-risk postures. A model used for collision avoidance requires low latency and conservative safety margins. A model used for fatigue estimation requires longitudinal calibration and uncertainty reporting. This purpose-driven modeling approach avoids excessive modeling cost and improves deployability. It also clarifies how human-state information connects to action: not every detected state should trigger automatic adaptation; some states should trigger recommendations, warnings, or requests for confirmation.

5.3. Data Architecture, Interoperability, and Traceability

HMC depends on data flows across heterogeneous devices and software systems. Sensors collect raw signals, controllers execute commands, planning modules generate task decisions, and management systems store production records. If these data remain isolated, collaboration cannot scale. A data architecture should define common identifiers for workpieces, tasks, tools, robots, operators, and events. It should also define timestamps, coordinate frames, data quality indicators, and access rules. Traceability is particularly important because decisions in HMC may involve both human and machine contributions. When a defect occurs or a safety event is triggered, the system should reconstruct what information was available, what decision was made, and who had authority.
Interoperability is difficult because industrial environments contain legacy equipment, proprietary controllers, and customized databases. Digital twins can reduce this fragmentation if they provide a common information model. However, a twin can also become another isolated platform if it is not connected through standard interfaces. Practical HMC deployment should therefore prioritize interface definition, data governance, and maintainability. Engineers need to know how a new sensor, robot, or AI module can be added without rewriting the whole system. A modular architecture with clear communication contracts improves both flexibility and long-term sustainability.
Edge computing is increasingly important for HMC because many functions require low latency. Human detection, collision avoidance, force control, and interface updates cannot always rely on cloud computing. Edge devices can process sensor data near the workstation, reduce communication delay, and preserve sensitive data locally. Cloud or server-based systems can still support heavy model training, historical analytics, and cross-line optimization. A hierarchical computing architecture is therefore suitable: safety-critical perception and control run locally, while long-term optimization and model improvement use larger computing resources. This hierarchy should be reflected in the digital twin design.
Baratta et al. [8] reviewed how digital twins enhance human–robot collaboration in manufacturing systems. In addition, Liu et al. [14] integrated a vision–language model with embodied intelligence for digital-twin-assisted human–robot collaboration. Wang et al. [50] presented a digital-twin-based approach to the design and operation of human–robot collaborative assembly. Further, Piardi et al. [51] analyzed the role of digital technologies in enhancing human integration within industrial cyber–physical systems. Ling et al. [52] developed a real-time data-driven scheme for human–machine synchronization and proactive ergonomics. In a related effort, Havard et al. [53] built a digital-twin and virtual-reality co-simulation environment for collaborative design and assessment.
Krupas et al. [54] reviewed the enablers of human-centric digital twins for human–machine collaboration. In addition, Zafar et al. [55] explored the synergies among collaborative robotics, digital twins, and augmentation technologies. Piardi et al. [56] discussed digital technologies that empower human activities in cyber–physical systems. Further, Choi et al. [57] developed an integrated mixed-reality system for safety-aware human–robot collaboration using a digital twin. Liu et al. [58] integrated a vision–language model into an embodied multi-agent system for digital-twin-assisted collaborative assembly.
Critical synthesis. Existing studies demonstrate several valuable but uneven integration patterns. Geometry- and process-oriented twins are effective for layout design, offline verification, collision checking, and commissioning; data-driven twins support prediction and adaptation but often provide weaker traceability; and human digital twins add ergonomic and cognitive context but face calibration and privacy constraints. Real-time synchronization, model fidelity, and interoperability form a three-way trade-off: higher-fidelity models increase computation and maintenance effort, while simplified models may meet latency requirements but miss safety-relevant behavior. Cross-platform integration is further constrained by proprietary controllers, heterogeneous timestamps and coordinate frames, and inconsistent information models. Consequently, most studies support the claim that digital twins can serve as coordination or verification layers, whereas evidence for a continuously updated, cross-platform, industrial integration backbone remains limited. Representative digital-twin integration studies are summarized in Table 3.

6. Adaptive Control, Safety Assurance, and Human–AI Decision-Making in Digital-Twin-Enabled HMC Systems

Adaptive control, safety assurance, and human–AI decision-making can be supported by digital twins, but the literature does not show that every such algorithm is simulated, verified, and closed through a twin. In this section, a method is treated as digital-twin-enabled only when synchronized virtual information is used for feasibility checking, risk assessment, prediction, adaptation, or feedback. Other methods are discussed as adjacent evidence that may be integrated into a twin-based architecture.
Safety should be treated as a layered system. The first layer is safe layout design, including separation distance, reachable workspace, visibility, and emergency access. The second layer is sensing-based monitoring, including human detection, speed and separation monitoring, force limitation, and protected zones. The third layer is control adaptation, including speed reduction, path replanning, impedance adjustment, and collision response. The fourth layer is governance, including standards, risk assessment, operator training, maintenance procedures, and incident reporting. These layers must work together. A strong perception model cannot compensate for poor workspace design, and a good safety standard cannot compensate for unreliable control implementation.
Adaptive motion control in HMC includes path planning, trajectory modification, impedance control, admittance control, force control, and collision handling. Path planning and dynamic task allocation can generate feasible robot motion under geometric and workload constraints [59]. Vision-guided robotic assembly supports workpiece localization and path planning in precision operations [60], while digital-twin-based design and operation can verify collaborative assembly before execution [50]. Trajectory modification adapts motion online when humans or obstacles move. Impedance and admittance control regulate the dynamic relationship between force and motion, enabling safer contact and co-manipulation. Collision handling detects unexpected contact and reduces impact. These methods may be combined with multimodal classification [21], AI-enabled sensing and robotic design [29], activity recognition [30], and skeleton-RGB action prediction [31]. However, learning-based prediction must be conservative because human motion can be abrupt and difficult to forecast. Safety margins and fallback states remain necessary.
Human–AI collaborative decision-making operates at a higher cognitive level. AI systems may recommend task allocations, predict quality risk, diagnose faults, or generate process plans. Humans provide domain knowledge, ethical judgment, and final authority for ambiguous or high-impact decisions. The goal is not to maximize machine autonomy at all costs, but to distribute decision work according to reliability, accountability, and context. For example, an AI model may identify a defect candidate and provide evidence, while a human inspector confirms the decision. A planning agent may propose a revised sequence after a disturbance, while the supervisor approves the change. This shared decision mode is essential when production consequences are significant.
Explainability is central to human–AI collaboration. Operators and engineers need to understand why the system recommends an action, what evidence supports it, and what uncertainty remains. Explanations should be linked to task context: a safety warning should identify the detected risk zone; a quality recommendation should point to the relevant defect feature; a scheduling decision should explain the bottleneck or resource constraint. Generic explanations are less useful than operational explanations. At the same time, explanations should not overload the operator. The system should provide layered information: concise alerts during execution, detailed records for analysis, and full traceability for audit or improvement.
Authority management is another critical issue. HMC systems must define who can approve task changes, who can override machine decisions, when automatic stopping is required, and how responsibility is assigned. Authority may vary by risk level. Low-risk adjustments, such as changing interface prompts, can be automatic. Medium-risk adjustments, such as changing task sequence, may require operator confirmation. High-risk actions, such as entering a shared workspace during robot motion, require strict safety logic and possibly supervisor authorization. Clear authority rules prevent ambiguity and support trust. They also make HMC more acceptable in regulated industrial environments.

6.1. Safety Standards and Engineering Verification

Safety engineering in HMC should begin before algorithm development. Risk assessment identifies hazards, estimates severity and probability, and defines risk reduction measures. Collaborative operation requires careful consideration of robot speed, payload, tool shape, contact force, workspace layout, and human access patterns. Standards and technical specifications provide structured guidance, but they must be translated into concrete design requirements. For example, speed reduction must be linked to detection range and stopping distance. Power and force limitation must consider tool geometry and body contact. Emergency stops and protective stops must be accessible and tested. Verification should include normal operation, abnormal operation, and recovery scenarios.
Algorithmic safety must be verified at several levels. Perception models should be tested under lighting variation, occlusion, clothing variation, and background clutter. Motion planning should be tested for collision-free behavior under dynamic obstacles. Control systems should be tested for stability, force limitation, and response time. Interface systems should be tested for correct warning delivery and operator understanding. Digital twins should be validated against physical measurements. AI-generated plans should be checked against task constraints and safety rules. This multi-level verification is essential because failures can occur at the interface between modules even if each module performs well in isolation.
Simulation is useful but insufficient by itself. Digital twins and virtual environments can generate scenarios, test trajectories, and reduce physical trial cost. However, real workstations include sensor noise, human variability, mechanical tolerances, communication delay, and unmodeled disturbances. Therefore, virtual validation should be followed by staged physical testing. A practical deployment path may begin with offline simulation, then supervised dry runs, then limited-speed operation, then monitored production trials. Each stage should collect evidence on safety, task performance, human workload, and recovery behavior. This evidence-based deployment approach reduces risk and supports continuous improvement.

6.2. Learning-Based Adaptation and Reliability Constraints

Learning-based methods can improve HMC by adapting to human behavior, product variation, and environmental changes. Reinforcement learning can learn policies for task allocation, motion assistance, or shared control. Imitation learning can transfer human demonstrations to robot behavior. Online learning can adjust models as new production data arrive. Foundation models can improve semantic understanding and support flexible planning. These capabilities are valuable because industrial environments contain uncertainty that is difficult to describe fully with manual rules. However, learning-based adaptation also raises reliability concerns. A policy that performs well in training may fail under distribution shift or rare safety-critical events.
Reliability constraints should be embedded in learning-based HMC. First, safety constraints should be enforced independently of learned policies whenever possible. A learned planner may propose a path, but a safety controller should still check separation distance and force limits. Second, uncertainty estimates should be used to trigger conservative behavior. If the perception system is uncertain about a human hand position, the robot should slow down or request confirmation. Third, learned models should be monitored after deployment. Changes in product appearance, operator behavior, or sensor configuration can degrade performance over time. Model monitoring and update procedures should be part of the production system.
The combination of learning and engineering constraints is more promising than either approach alone. Pure rule-based systems may be too rigid for high-mix production, while pure learning systems may be difficult to certify and maintain. Hybrid architectures can use learning for perception, prediction, and recommendation, while using process rules, digital twins, and safety logic for verification and execution. This architecture is also easier to explain to operators and engineers. It supports adaptation without sacrificing accountability. For sustainable manufacturing, such reliability is essential because frequent downtime, rework, or unsafe interventions can eliminate the benefits of collaboration.
Joo et al. [23] applied deep reinforcement learning to task allocation in human–machine manufacturing systems. In addition, Lim et al. [45] used large language models to enhance human–robot collaborative assembly. Chen et al. [46] proposed a perception–decision–execution coordination mechanism for dynamic autonomous collaboration. Further, Chen et al. [47] developed an LLM-based method for human–robot autonomous collaboration in smart manufacturing. Liu et al. [48] introduced large-model-based interactive collaborative decision-making for human–machine systems. In a related effort, Liu et al. [61] proposed a real-time hierarchical control method for safe human–robot coexistence.
Xia et al. [62] leveraged error-assisted fine-tuning of large language models for manufacturing tasks. In addition, Laplaza et al. [63] enhanced robotic collaborative tasks through contextual human motion prediction. Wang et al. [64] applied reinforcement learning with imitative behaviors to humanoid robot navigation. Further, Ma et al. [65] proposed a hybrid human–machine decision-making strategy for collaborative control. Masehian et al. [66] investigated assembly sequence and path planning for complex assemblies. In a related effort, Zhang et al. [67] developed a mixed-reality remote collaboration system with adaptive instruction generation.
Liu et al. [68] built an AR-assisted collaborative assembly system integrating a vision–language model and deep reinforcement learning for task planning. In addition, Liu et al. [69] proposed an LLM-enhanced embodied multi-agent manufacturing system for self-organizing production. Zhu et al. [70] reviewed large language models for industrial embodied intelligence.
Critical synthesis. Rule-based and optimization-based safety mechanisms offer predictability and are easier to validate, but they can be conservative and difficult to adapt to high-mix production. Learning-based control and task allocation improve flexibility but introduce distribution-shift, explainability, and certification risks. Digital-twin simulation can reduce physical trial cost and detect infeasible actions, yet simulation validity depends on model fidelity and cannot replace staged physical testing. LLM and multimodal foundation models are most defensible as semantic interpretation and plan-proposal modules; deterministic planners, safety controllers, and human approval should remain responsible for executable decisions. The strongest evidence therefore comes from layered architectures that separate learned recommendations from independently enforced physical safety constraints. Representative control, safety, and human–AI decision-making studies are summarized in Table 4.

7. Industrial Applications, Deployment Maturity, and Sustainability Evidence

The industrial value of digital-twin-enabled HMC should be judged by deployment maturity and measured outcomes rather than by architectural completeness alone. The reviewed applications span assembly, inspection, logistics, maintenance, disassembly, remanufacturing, and high-end equipment production, but the evidence ranges from conceptual proposals and simulations to controlled physical prototypes and a smaller number of pilot or industrial implementations. This section therefore distinguishes laboratory feasibility from sustained deployability and separates direct sustainability outcomes from indirect operational improvements.
Inspection is another important application. Vision systems and AI models can detect candidate defects, measure dimensions, and compare product states with digital references. Human inspectors remain important for ambiguous cases, root-cause interpretation, and final quality decisions. HMC can improve inspection by combining machine speed and consistency with human expertise. The interaction design matters: the system should highlight evidence, provide confidence scores, allow human correction, and record final decisions for model improvement. In sustainable manufacturing, better inspection reduces scrap, rework, warranty cost, and unnecessary material use [61,71].
Logistics and material handling benefit from collaboration between human workers, mobile robots, and warehouse management systems. Mobile robots can transport materials, while humans perform flexible picking, verification, or exception handling. In production lines, collaborative robots may deliver tools, position heavy parts, or support kit preparation. The main challenges include navigation in shared spaces, worker acceptance, traffic coordination, and integration with production schedules. HMC can improve ergonomic performance by reducing carrying and lifting, but it may also create new risks if robot motion is unpredictable or if workers must adapt to poorly designed traffic flows [34,62].
Maintenance and repair are suited to HMC because tasks are often complex, variable, and knowledge-intensive. Digital twins can provide equipment history, fault diagnosis, and disassembly instructions [72]. AR interfaces can guide technicians through procedures and convert expert knowledge into reusable digital guidance [73]. Robots and sensing systems can support positioning, measurement, remote manipulation, or hazardous access [35]. Human experts interpret abnormal situations and make decisions when documentation is incomplete. Collaborative maintenance can reduce downtime and extend equipment life.
Disassembly and remanufacturing are increasingly important for circular economy [74]. Unlike new product assembly, disassembly deals with uncertain product conditions, wear, deformation, contamination, and missing information [75]. Full automation is difficult because products may differ from design data and may require judgment during separation [76]. HMC is suitable because robots can perform repetitive or hazardous operations while humans make flexible decisions [36]. Digital twins and perception systems can identify components, estimate disassembly sequence, and support material recovery [37]. The sustainability value is direct: recovered parts and materials reduce resource consumption and waste.
High-end equipment manufacturing, such as aerospace and large-scale machinery, presents special challenges [77]. Workpieces are large, tolerances are strict, and operations often require complex alignment, drilling, fastening, inspection, and documentation [78]. HMC can support workers with robotic positioning, visual guidance, force assistance, and digital traceability [38,63]. However, deployment is difficult because production volumes are low, tasks are customized, and certification requirements are strict [64]. Digital-twin-supported planning and offline validation are especially useful in these scenarios because they reduce physical trial cost and improve confidence before execution [52,79].
Evaluation is a weak point in much of the HMC literature [80]. Many studies report algorithmic indicators such as recognition accuracy, planning time, or success rate, but do not evaluate production value or human impact [22]. A complete evaluation framework should include technical performance, safety performance, human factors, economic value, and sustainability impact [39,81]. Technical performance includes perception accuracy, latency, task completion time, planning robustness, and control stability. Safety performance includes minimum separation distance, collision rate, emergency stops, near-miss events, and compliance with standards.
Economic indicators include labor productivity, robot utilization, human idle time, rework cost, commissioning time, maintenance cost, and return on investment [65]. Sustainability indicators include energy consumption, material waste, defect reduction, equipment life extension, ergonomic risk reduction, and circularity contribution [82]. These indicators should be reported together because optimization of one metric can harm another [83]. For example, increasing robot speed may reduce cycle time but increase safety stops or operator stress [84]. Adding more sensors may improve perception but increase energy and maintenance requirements [85]. A balanced evaluation prevents misleading conclusions.
Benchmarking is difficult because collaborative tasks vary across industries [86]. A standard benchmark for object recognition may not represent assembly complexity, and a laboratory demonstration may not capture production disturbances [87]. Nevertheless, HMC research can improve comparability by reporting task details, hardware configuration, sensor placement, number of participants, product variation, cycle time baseline, safety constraints, and failure cases [88]. Negative results and recovery behavior should also be reported [89]. A system that succeeds in 90% of trials but fails unpredictably in 10% may be less valuable than a system with slightly lower nominal performance but predictable fallback behavior [66].
Longitudinal evaluation is particularly important. A short experiment can show feasibility, but industrial value depends on sustained operation. Over time, sensors drift, tools wear, operators learn, product variants change, and AI models may face distribution shift [90]. Human acceptance may improve with familiarity or decline if the system creates hidden workload [53]. Sustainability benefits may appear only after enough production cycles [67]. Therefore, future HMC studies should include longer deployments, repeated measures, and life-cycle indicators. Such evidence is necessary for moving from laboratory prototypes to industrial adoption.
Sustainability evidence is classified into three levels. Direct evidence is reported when a study measures indicators such as energy consumption, material waste, defect or rework reduction, ergonomic exposure, component recovery, or equipment-life extension. Indirect operational evidence includes cycle time, accuracy, downtime, utilization, or workload measures that may contribute to sustainability but do not by themselves establish an environmental or social outcome. Aspirational evidence consists of sustainability claims without corresponding measurements. Much of the current literature falls into the second or third category; direct multi-dimensional and longitudinal evidence remains limited.
Critical synthesis. Assembly and inspection dominate the literature because they offer bounded tasks, visible quality criteria, and accessible laboratory testbeds. Maintenance, disassembly, and remanufacturing provide stronger potential links to lifecycle sustainability, but they involve greater product uncertainty, incomplete information, and recovery complexity. High-end equipment applications benefit from offline verification and traceability, yet low production volumes and strict certification make economic and statistical validation difficult. Across applications, a controlled prototype is not equivalent to a deployable industrial architecture: sustained operation additionally requires robust interfaces, maintenance procedures, abnormal recovery, cybersecurity, operator training, and evidence under product, operator, and environmental variation.
Barathwaj et al. [41] optimized assembly-line balancing using a genetic algorithm. In addition, Cai et al. [59] proposed a task-allocation method for collaborative assembly lines that accounts for assembly complexity. Ji et al. [60] developed a visually guided robotic assembly method for spaceborne equipment. Further, Xu et al. [91] proposed a transformer-based multimodal-fusion framework for generative intelligent process design of aviation riveting. Li et al. [92] developed a physics-informed embodied-intelligence framework for smart garment manufacturing. Representative application and evaluation studies are summarized in Table 5.

8. Challenges and Future Trends in Digital-Twin-Enabled HMC System Development

The challenges and future trends below are framed around digital-twin-enabled HMC system development, spanning data interoperability, model consistency, real-time synchronization, human-digital-twin privacy, explainability of AI decisions, safety certification, and cross-scenario generalization. Although HMC has shown clear potential in smart manufacturing, its industrial deployment still faces several practical barriers. Generalization remains the first challenge: models trained in one workstation may become unreliable when lighting, camera position, tool geometry, product variants, or operator behavior changes [93]. Vision-language and cross-modal models provide a possible route for more transferable perception and semantic understanding [94,95]. For industrial use, however, these models still need uncertainty reporting, task-specific validation, and conservative fallback strategies [96].
Real-time reliability is another core barrier [97]. Collaborative systems must perceive human motion, infer task progress, update plans, verify safety constraints, and execute robot commands within limited cycle time [54]. End-to-end latency should therefore be measured across sensors, networks, AI inference, controllers, and interfaces rather than reported only for individual modules [55]. Human-centric digital twins can strengthen state synchronization and process verification [56]. Mixed-reality integration can improve guidance and remote collaboration, but it must not introduce excessive cognitive load [57]. Liu et al. [68] demonstrated that AR-assisted assembly can connect visual language understanding, deep reinforcement learning, task planning, and interactive guidance, while Zhu et al. [70] summarized how large language models extend industrial embodied intelligence across perception, reasoning, and action.
Human-state modeling and digital-twin integration also remain difficult. Fatigue, workload, attention, and trust are latent and individual-dependent states that require calibration, uncertainty reporting, and privacy-aware data governance. Liu et al. [58] showed that a vision-language-model-driven embodied multi-agent system can transform digital-twin insight into collaborative assembly action. Li et al. [92] further indicated that physics-informed embodied intelligence can improve adaptation when material properties, deformation behavior, and robot manipulation constraints must be considered together. These studies suggest that human-state models and production twins should be connected to actionable feedback rather than used only for visualization.
Adaptive control and safety assurance will continue to be decisive for physical HMC. Collaborative assembly, welding, inspection, and disassembly require robots to adjust paths, force, speed, and role allocation under uncertain human behavior and changing workpieces. AR-assisted VLM-DRL assembly systems provide one route for linking semantic task planning with motion-level execution [68]. Xu et al. [91] extended multimodal transformer fusion to generative process design in aviation riveting, showing how process knowledge and multimodal information can support intelligent planning. Wang et al. [40] proposed IndVisSGG for industrial scene graph generation, which is useful for representing spatial relations among objects, tools, and workspaces in collaborative manufacturing.
Human-AI collaborative decision-making introduces a further challenge: the system must combine machine reasoning with human responsibility. Large models and multimodal models can support task interpretation, planning, and communication, but their outputs must remain aligned with industrial constraints and human authority [44,45]. Perception-decision-execution coordination and machine knowledge integration show that autonomy should be embedded in a closed loop rather than treated as an isolated language interface [46,47]. Human–machine collaborative decision-making therefore requires clear authority boundaries, traceable evidence, explainable recommendations, and safe fallback strategies [20,48]. Digital twins can serve as verification layers for AI-generated task plans by checking feasibility, collision risk, resource state, and process history before physical execution [49,50].
From the application perspective, future research should move from isolated demonstrations to reusable, measurable, and transferable manufacturing solutions. Collaborative-robot deployment studies show that role division and dynamic task allocation need to be evaluated under real assembly constraints rather than only in simplified demonstrations [15,16]. Digital-twin-driven assembly studies further indicate that planning, monitoring, and feedback should be connected with process data and worker state [42,43]. Liu et al. [69] proposed an LLM-enhanced embodied multi-agent manufacturing system, suggesting a self-organizing production paradigm based on embodied perception, embodied analysis, and embodied decision. Wang et al. [98] also emphasized that cognitive intelligence can support product design by linking human-like reasoning, design knowledge, and AI-enabled decision processes. However, benchmarks should report not only recognition accuracy or task success rate, but also cycle time, failure recovery, human workload, safety events, energy use, material saving, and long-term maintainability.
Future HMC research can be summarized in four directions. First, trusted multimodal models should be developed for industrial scenes by combining domain data, synthetic data, uncertainty estimation, and operator feedback. Second, digital twins should evolve from visualization tools into runtime task-context interfaces that synchronize perception, planning, safety zones, process history, and sustainability indicators. Third, adaptive control should be safety-certified through layered verification, including simulation, staged physical tests, conservative fallback modes, and post-deployment monitoring. Fourth, HMC evaluation should integrate productivity, quality, ergonomics, energy, material use, circularity, and resilience rather than optimizing a single indicator.
Standardization, privacy, and worker acceptance should be treated as design requirements rather than secondary issues. HMC systems collect human motion, voice, posture, and sometimes physiological data; therefore, data minimization, access control, transparent purpose definition, and voluntary feedback mechanisms are necessary. Workers are more likely to accept collaboration when the system solves real production problems, reduces unsafe or repetitive work, provides understandable assistance, and preserves meaningful human intervention rights.
Cao et al. [29] examined AI-driven design of soft robots for adaptive interaction. In addition, Bayoudh et al. [99] surveyed multimodal hybrid deep learning for computer vision. Salichs et al. [72] developed a social-robot platform supporting natural human–robot interaction. Further, Min et al. [73] reviewed recent advances in natural language processing driven by large pre-trained language models. Xue et al. [74] improved multi-turn response selection using dual-view graph convolutions over BERT. In a related effort, Luo et al. [75] proposed a text-guided multi-task network for multimodal sentiment analysis.
Shafizadegan et al. [76] reviewed deep-learning-based multimodal human action recognition. In addition, Wang et al. [78] surveyed pre-trained language models in the biomedical domain. Huang et al. [81] surveyed human–computer interaction in mixed reality.
A central future challenge is to convert the digital twin from an architectural aspiration into an evidenced integration mechanism. Research should report the exact physical entities represented, model assumptions, synchronization rates and latency, state-update variables, model-update procedures, uncertainty handling, feedback authority, and failure-recovery logic. Without these details, it is difficult to determine whether a study presents a digital model, a one-way digital shadow, or a bidirectionally coupled twin. Comparative benchmarks should also separate module-level performance from end-to-end closed-loop performance and should identify the maturity level of each demonstration.
Recent advances in embodied intelligence and motion planning further indicate where digital-twin-enabled HMC systems are heading. Li and Yang [100] reviewed the transition from digital twins to embodied artificial intelligence and outlined future perspectives for intelligent robotic systems. In addition, He et al. [101] proposed MSAFNet for facial expression recognition in embodied AI systems, enriching human-state perception. Sun et al. [102] developed a phase-search-enhanced Bi-RRT algorithm for efficient path planning of mobile robots. Further, Cui et al. [103] developed a smooth and efficient motion-planning method for large-scale cooperative multi-arm tunnel-drilling robots, pointing toward the scaling of coordinated multi-robot systems. Sun et al. [104] reviewed advances in humanoid-robot dynamics and learning-based locomotion control, indicating emerging embodied platforms for future human–machine collaboration. Representative challenges and future directions are summarized in Table 6.

9. Conclusions

This review shows that digital twins can provide a valuable integration, coordination, and verification layer for human–machine collaboration in sustainable smart manufacturing, but the evidence does not support treating a fully integrated twin as an established reality across the field. The reviewed literature contains a continuum ranging from conceptual architectures and stand-alone simulations to synchronized laboratory prototypes and a smaller number of pilot or industrial implementations. The term digital-twin-enabled should therefore be reserved for systems with an identifiable physical counterpart, maintained virtual representation, explicit physical-to-virtual data connection, and a described use of synchronized information for analysis, verification, adaptation, or feedback.
The critical comparison highlights two technical requirements that are often conflated. State synchronization updates the current virtual representation of position, force, workload, task progress, or resource status, whereas model updating revises parameters, uncertainty, or model structure from newly observed physical behavior. Existing evidence is substantially stronger for state synchronization, virtual validation, and offline planning than for traceable online model updating, sustained bidirectional control, or cross-platform industrial integration. Hybrid physics–data models are promising, but they require explicit validation, uncertainty handling, and lifecycle maintenance.
From an application perspective, the strongest evidence concerns virtual commissioning, collision and feasibility checking, operator guidance, task coordination, and controlled collaborative assembly or inspection. Claims of industrial deployability require additional evidence on sustained synchronization, disturbance recovery, cybersecurity, maintainability, operator training, and safety certification. Sustainability claims must likewise distinguish direct measurements from indirect operational indicators and unmeasured aspirations. At present, direct longitudinal evidence that jointly evaluates environmental, economic, human, and technical outcomes remains limited.
Future research should prioritize reproducible digital-twin definitions, transparent model-update mechanisms, interoperable data contracts, staged virtual-to-physical validation, independently enforced safety constraints, explainable human–AI authority management, and long-term evaluation across product variants, operators, and production disturbances. Progress on these requirements would allow digital twins to move from a compelling architectural vision toward a reliable and measurable foundation for industrial HMC.

Author Contributions

Conceptualization, H.Z. and J.C.; methodology, H.Z. and J.C.; investigation, H.Z., J.C., G.L. and F.Y.; writing—original draft preparation, H.Z. and J.C.; writing—review and editing, J.C., G.L., F.Y. and H.G.; supervision, J.C.; funding acquisition, H.Z. Both H.Z. and J.C. contributed equally to this work and share first authorship. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Open Foundation of Industrial Perception and Intelligent Manufacturing Equipment Engineering Research Center of Jiangsu Province (Grant No. ZK22-05-01), and Start-up Fund for New Talented Researchers of Nanjing University of Industry Technology (Grant No. YK23-01-01).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Acknowledgments

The authors would like to thank their colleagues for the helpful discussions and suggestions during the preparation of this review.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Leng, J.; Sha, W.; Wang, B.; Zheng, P.; Zhuang, C.; Liu, Q.; Wuest, T.; Mourtzis, D.; Wang, L. Industry 5.0: Prospect and retrospect. J. Manuf. Syst. 2022, 65, 279–295. [Google Scholar] [CrossRef] [Scilit]
  2. Wang, B.; Tao, F.; Fang, X.; Liu, C.; Liu, Y.; Freiheit, T. Smart Manufacturing and Intelligent Manufacturing: A Comparative Review. Engineering 2021, 7, 738–757. [Google Scholar] [CrossRef] [Scilit]
  3. Li, S.; Wang, R.; Zheng, P.; Wang, L. Towards proactive human–robot collaboration: A foreseeable cognitive manufacturing paradigm. J. Manuf. Syst. 2021, 60, 547–552. [Google Scholar] [CrossRef] [Scilit]
  4. Hietanen, A.; Pieters, R.; Lanz, M.; Latokartano, J.; Kämäräinen, J.-K. AR-based interaction for human-robot collaborative manufacturing. Robot. Comput.-Integr. Manuf. 2020, 63, 101891. [Google Scholar] [CrossRef] [Scilit]
  5. Liau, Y.Y.; Ryu, K. Task Allocation in Human-Robot Collaboration (HRC) Based on Task Characteristics and Agent Capability for Mold Assembly. Procedia Manuf. 2020, 51, 179–186. [Google Scholar] [CrossRef] [Scilit]
  6. Duan, J.; Fang, Y.; Zhang, Q.; Qin, J. HRC for dual-robot intelligent assembly system based on multimodal perception. Proc. Inst. Mech. Eng. Part B J. Eng. Manuf. 2024, 238, 562–576. [Google Scholar] [CrossRef] [Scilit]
  7. Blankemeyer, S.; Wendorff, D.; Raatz, A. A hand-interaction model for augmented reality enhanced human-robot collaboration. CIRP Ann. 2024, 73, 17–20. [Google Scholar] [CrossRef] [Scilit]
  8. Baratta, A.; Cimino, A.; Longo, F.; Nicoletti, L. Digital twin for human-robot collaboration enhancement in manufacturing systems: Literature review and direction for future developments. Comput. Ind. Eng. 2024, 187, 109764. [Google Scholar] [CrossRef] [Scilit]
  9. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Rethlefsen, M.L.; Kirtley, S.; Waffenschmidt, S.; Ayala, A.P.; Moher, D.; Page, M.J.; Koffel, J.B.; PRISMA-S Group. PRISMA-S: An extension to the PRISMA Statement for Reporting Literature Searches in Systematic Reviews. Syst. Rev. 2021, 10, 39. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Kritzinger, W.; Karner, M.; Traar, G.; Henjes, J.; Sihn, W. Digital Twin in manufacturing: A categorical literature review and classification. IFAC-Pap. 2018, 51, 1016–1022. [Google Scholar] [CrossRef] [Scilit]
  12. Jones, D.; Snider, C.; Nassehi, A.; Yon, J.; Hicks, B. Characterising the Digital Twin: A systematic literature review. CIRP J. Manuf. Sci. Technol. 2020, 29, 36–52. [Google Scholar] [CrossRef] [Scilit]
  13. Ding, P.; Zhang, J.; Zheng, P.; Zhang, P.; Fei, B.; Xu, Z. Dynamic scenario-enhanced diverse human motion prediction network for proactive human–robot collaboration in customized assembly tasks. J. Intell. Manuf. 2025, 36, 4593–4612. [Google Scholar] [CrossRef] [Scilit]
  14. Liu, C.; Tang, D.; Zhu, H.; Zhang, Z.; Wang, L.; Zhang, Y. Vision language model-enhanced embodied intelligence for digital twin-assisted human-robot collaborative assembly. J. Ind. Inf. Integr. 2025, 48, 100943. [Google Scholar] [CrossRef] [Scilit]
  15. Liu, L.; Guo, F.; Zou, Z.; Duffy, V.G. Application, Development and Future Opportunities of Collaborative Robots (Cobots) in Manufacturing: A Literature Review. Int. J. Hum.-Comput. Interact. 2024, 40, 915–932. [Google Scholar] [CrossRef] [Scilit]
  16. Petzoldt, C.; Niermann, D.; Maack, E.; Sontopski, M.; Vur, B.; Freitag, M. Implementation and Evaluation of Dynamic Task Allocation for Human–Robot Collaboration in Assembly. Appl. Sci. 2022, 12, 12645. [Google Scholar] [CrossRef] [Scilit]
  17. Zanchettin, A.M.; Casalino, A.; Piroddi, L.; Rocco, P. Prediction of human activity patterns for human–robot collaborative assembly tasks. IEEE Trans. Ind. Inform. 2018, 15, 3934–3942. [Google Scholar] [CrossRef] [Scilit]
  18. Bruno, G.; Antonelli, D. Dynamic task classification and assignment for the management of human-robot collaborative teams in workcells. Int. J. Adv. Manuf. Technol. 2018, 98, 2415–2427. [Google Scholar] [CrossRef] [Scilit]
  19. Malik, A.A.; Bilberg, A. Complexity-based task allocation in human-robot collaborative assembly. Ind. Robot. Int. J. Robot. Res. Appl. 2019, 46, 471–480. [Google Scholar] [CrossRef] [Scilit]
  20. Kong, F.; Gao, T.; Li, H.; Lu, Z. Research on Human-robot Joint Task Assignment Considering Task Complexity. J. Mech. Eng. 2021, 57, 204–214. [Google Scholar] [CrossRef] [Scilit]
  21. Sleeman, W.C.; Kapoor, R.; Ghosh, P. Multimodal Classification: Current Landscape, Taxonomy and Future Directions. ACM Comput. Surv. 2022, 55, 150. [Google Scholar] [CrossRef] [Scilit]
  22. Xue, J.; Hu, B.; Li, L.; Zhang, J. Human–Machine augmented intelligence: Research and applications. Front. Inf. Technol. Electron. Eng. 2022, 23, 1139–1141. [Google Scholar] [CrossRef] [Scilit]
  23. Joo, T.; Jun, H.; Shin, D. Task Allocation in Human–Machine Manufacturing Systems Using Deep Reinforcement Learning. Sustainability 2022, 14, 2245. [Google Scholar] [CrossRef] [Scilit]
  24. Gao, Z.; Yang, R.; Zhao, K.; Yu, W.; Liu, Z.; Liu, L. Hybrid Convolutional Neural Network Approaches for Recognizing Collaborative Actions in Human–Robot Assembly Tasks. Sustainability 2024, 16, 139. [Google Scholar] [CrossRef] [Scilit]
  25. Mavsar, M.; Deni, M.; Nemec, B.; Ude, A. Intention Recognition with Recurrent Neural Networks for Dynamic Human-Robot Collaboration. In Proceedings of the 2021 20th International Conference on Advanced Robotics (ICAR), Ljubljana, Slovenia, 6–10 December 2021; pp. 208–215. [Google Scholar]
  26. Wang, P.; Liu, H.; Wang, L.; Gao, R.X. Deep learning-based human motion recognition for predictive context-aware human-robot collaboration. CIRP Ann. 2018, 67, 17–20. [Google Scholar] [CrossRef] [Scilit]
  27. Bandi, C.; Thomas, U. Skeleton-based Action Recognition for Human-Robot Interaction using Self-Attention Mechanism. In Proceedings of the 2021 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021), Jodhpur, India, 15–18 December 2021; pp. 1–8. [Google Scholar]
  28. Gao, X.; Yan, L.; Wang, G.; Gerada, C. Hybrid Recurrent Neural Network Architecture-Based Intention Recognition for Human–Robot Collaboration. IEEE Trans. Cybern. 2023, 53, 1578–1586. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Cao, Y.; Xu, B.; Li, B.; Fu, H. Advanced Design of Soft Robots with Artificial Intelligence. Nano-Micro Lett. 2024, 16, 214. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Hussain, A.; Khan, S.U.; Rida, I.; Khan, N.; Baik, S.W. Human centric attention with deep multiscale feature fusion framework for activity recognition in Internet of Medical Things. Inf. Fusion 2024, 106, 102211. [Google Scholar] [CrossRef] [Scilit]
  31. Zhang, Y.; Ding, K.; Hui, J.; Liu, S.; Guo, W.; Wang, L. Skeleton-RGB integrated highly similar human action prediction in human–robot collaborative assembly. Robot. Comput.-Integr. Manuf. 2024, 86, 102659. [Google Scholar] [CrossRef] [Scilit]
  32. Nadeem, M.; Sohail, S.S.; Javed, L.; Anwer, F.; Saudagar, A.K.J.; Muhammad, K. Vision-Enabled Large Language and Deep Learning Models for Image-Based Emotion Recognition. Cogn. Comput. 2024, 16, 2566–2579. [Google Scholar] [CrossRef] [Scilit]
  33. Hazmoune, S.; Bougamouza, F. Using transformers for multimodal emotion recognition: Taxonomies and state of the art review. Eng. Appl. Artif. Intell. 2024, 133, 108339. [Google Scholar] [CrossRef] [Scilit]
  34. Liu, C.; Wang, Y.; Yang, J. A transformer-encoder-based multimodal multi-attention fusion network for sentiment analysis. Appl. Intell. 2024, 54, 8415–8441. [Google Scholar] [CrossRef] [Scilit]
  35. Sun, T.; Feng, B.; Huo, J.; Xiao, Y.; Wang, W.; Peng, J.; Li, Z.; Du, C.; Wang, W.; Zou, G.; et al. Artificial Intelligence Meets Flexible Sensors: Emerging Smart Flexible Sensing Systems Driven by Machine Learning and Artificial Synapses. Nano-Micro Lett. 2023, 16, 14. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Li, Z.; Guo, Q.; Pan, Y.; Ding, W.; Yu, J.; Zhang, Y.; Liu, W.; Chen, H.; Wang, H.; Xie, Y. Multi-level correlation mining framework with self-supervised label generation for multimodal sentiment analysis. Inf. Fusion 2023, 99, 101891. [Google Scholar] [CrossRef] [Scilit]
  37. Wang, T.; Liu, Z.; Wang, L.; Li, M.; Wang, X.V. Data-efficient multimodal human action recognition for proactive human–robot collaborative assembly: A cross-domain few-shot learning approach. Robot. Comput.-Integr. Manuf. 2024, 89, 102785. [Google Scholar] [CrossRef] [Scilit]
  38. Yang, C.; Liu, Y.; Yin, C. More than a framework: Sketching out technical enablers for natural language-based source code generation. Comput. Sci. Rev. 2024, 53, 100637. [Google Scholar] [CrossRef] [Scilit]
  39. Balamurugan, K.; Sudhakar, G.; Xavier, K.F.; Bharathiraja, N.; Kaur, G. Human-machine interaction in mechanical systems through sensor enabled wearable augmented reality interfaces. Meas. Sens. 2025, 39, 101880. [Google Scholar] [CrossRef] [Scilit]
  40. Wang, Z.; Yan, Z.; Li, S.; Liu, J. IndVisSGG: VLM-based scene graph generation for industrial spatial intelligence. Adv. Eng. Inform. 2025, 65, 103107. [Google Scholar] [CrossRef] [Scilit]
  41. Barathwaj, N.; Raja, P.; Gokulraj, S. Optimization of assembly line balancing using genetic algorithm. J. Cent. South Univ. 2015, 22, 3957–3969. [Google Scholar] [CrossRef] [Scilit]
  42. Sun, X.; Zhang, R.; Liu, S.; Lv, Q.; Bao, J.; Li, J. A digital twin-driven human–robot collaborative assembly-commissioning method for complex products. Int. J. Adv. Manuf. Technol. 2022, 118, 3389–3402. [Google Scholar] [CrossRef] [Scilit]
  43. Wang, J.; Yan, Y.; Hu, Y.; Yang, X.; Zhang, L. A transfer reinforcement learning and digital-twin based task allocation method for human-robot collaboration assembly. Eng. Appl. Artif. Intell. 2025, 144, 110064. [Google Scholar] [CrossRef] [Scilit]
  44. Dimitropoulos, N.; Kaipis, M.; Giartzas, S.; Michalos, G. Generative AI for automated task modelling and task allocation in human robot collaborative applications. CIRP Ann. 2025, 74, 7–11. [Google Scholar] [CrossRef] [Scilit]
  45. Lim, J.; Patel, S.; Evans, A.; Pimley, J.; Li, Y.; Kovalenko, I. Enhancing Human-Robot Collaborative Assembly in Manufacturing Systems Using Large Language Models. In Proceedings of the 2024 IEEE 20th International Conference on Automation Science and Engineering (CASE), Bari, Italy, 28 August–1 September 2024; pp. 2581–2587. [Google Scholar]
  46. Chen, J.; Huang, S.; Wang, X.; Wang, P.; Zhu, J.; Xu, Z.; Wang, G.; Yan, Y.; Wang, L. Perception-decision-execution coordination mechanism driven dynamic autonomous collaboration method for human-like collaborative robot based on multimodal large language model. Robot. Comput.-Integr. Manuf. 2026, 98, 103167. [Google Scholar] [CrossRef] [Scilit]
  47. Chen, J.; Huang, S.; Xu, Z.; Yan, Y.; Wang, G. Human-robot Autonomous Collaboration Method of Smart Manufacturing Systems Based on Large Language Model and Machine Vision. J. Mech. Eng. 2025, 61, 130–141. [Google Scholar] [CrossRef] [Scilit]
  48. Liu, Z.; Peng, Y. Human-Machine Interactive Collaborative Decision-Making Based on Large Model Technology: Application Scenarios and Future Developments. In Proceedings of the 2025 2nd International Conference on Artificial Intelligence and Digital Technology (ICAIDT), Guangzhou, China, 28–30 April 2025; pp. 106–110. [Google Scholar]
  49. Bilberg, A.; Malik, A.A. Digital twin driven human–robot collaborative assembly. CIRP Ann. 2019, 68, 499–502. [Google Scholar] [CrossRef] [Scilit]
  50. Wang, Y.; Feng, J.; Liu, J.; Liu, X.; Wang, J. Digital Twin-based Design and Operation of Human-Robot Collaborative Assembly. IFAC-Pap. 2022, 55, 295–300. [Google Scholar] [CrossRef] [Scilit]
  51. Piardi, L.; Leitão, P.; Queiroz, J.; Pontes, J. Role of digital technologies to enhance the human integration in industrial cyber–physical systems. Annu. Rev. Control 2024, 57, 100934. [Google Scholar] [CrossRef] [Scilit]
  52. Ling, S.; Yuan, Y.; Yan, D.; Leng, Y.; Rong, Y.; Huang, G.Q. RHYTHMS: Real-time Data-driven Human-machine Synchronization for Proactive Ergonomic Risk Mitigation in the Context of Industry 4.0 and Beyond. Robot. Comput.-Integr. Manuf. 2024, 87, 102709. [Google Scholar] [CrossRef] [Scilit]
  53. Havard, V.; Jeanne, B.; Lacomblez, M.; Baudry, D. Digital twin and virtual reality: A co-simulation environment for design and assessment of industrial workstations. Prod. Manuf. Res. 2019, 7, 472–489. [Google Scholar] [CrossRef] [Scilit]
  54. Krupas, M.; Kajati, E.; Liu, C.; Zolotova, I. Towards a Human-Centric Digital Twin for Human–Machine Collaboration: A Review on Enabling Technologies and Methods. Sensors 2024, 24, 2232. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Zafar, M.H.; Langås, E.F.; Sanfilippo, F. Exploring the synergies between collaborative robotics, digital twins, augmentation, and industry 5.0 for smart manufacturing: A state-of-the-art review. Robot. Comput.-Integr. Manuf. 2024, 89, 102769. [Google Scholar] [CrossRef] [Scilit]
  56. Piardi, L.; Queiroz, J.; Pontes, J.; Leitão, P. Digital Technologies to Empower Human Activities in Cyber-Physical Systems. IFAC-Pap. 2023, 56, 8203–8208. [Google Scholar] [CrossRef] [Scilit]
  57. Choi, S.H.; Park, K.-B.; Roh, D.H.; Lee, J.Y.; Mohammed, M.; Ghasemi, Y.; Jeong, H. An integrated mixed reality system for safety-aware human-robot collaboration using deep learning and digital twin generation. Robot. Comput.-Integr. Manuf. 2022, 73, 102258. [Google Scholar] [CrossRef] [Scilit]
  58. Liu, C.; Qian, Y.; Tang, D.; Zhu, H.; Pang, J.; Cai, Q. From insight to action: Embodied multi-agent system integrating vision language model for digital twin-assisted human-robot collaborative assembly. J. Manuf. Syst. 2026, 85, 531–556. [Google Scholar] [CrossRef] [Scilit]
  59. Cai, M.; Wang, G.; Luo, X.; Xu, X. Task allocation of human-robot collaborative assembly line considering assembly complexity and workload balance. Int. J. Prod. Res. 2025, 63, 4749–4775. [Google Scholar] [CrossRef] [Scilit]
  60. Ji, X.; Wang, J.; Zhao, J.; Zhang, X.; Sun, Z. Intelligent Robotic Assembly Method of Spaceborne Equipment Based on Visual Guidance. J. Mech. Eng. 2018, 54, 63–72. [Google Scholar] [CrossRef] [Scilit]
  61. Liu, B.; Rocco, P.; Zanchettin, A.M.; Zhao, F.; Jiang, G.; Mei, X. A real-time hierarchical control method for safe human–robot coexistence. Robot. Comput.-Integr. Manuf. 2024, 86, 102666. [Google Scholar] [CrossRef] [Scilit]
  62. Xia, L.Q.; Li, C.X.; Zhang, C.B.; Liu, S.M.; Zheng, P. Leveraging error-assisted fine-tuning large language models for manufacturing excellence. Robot. Comput.-Integr. Manuf. 2024, 88, 102728. [Google Scholar] [CrossRef] [Scilit]
  63. Laplaza, J.; Moreno, F.; Sanfeliu, A. Enhancing Robotic Collaborative Tasks Through Contextual Human Motion Prediction and Intention Inference. Int. J. Soc. Robot. 2025, 17, 2077–2096. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  64. Wang, X.; Zhang, T. Reinforcement learning with imitative behaviors for humanoid robots navigation: Synchronous planning and control. Auton. Robot. 2024, 48, 5. [Google Scholar] [CrossRef] [Scilit]
  65. Ma, W.; Sun, B.; Zhao, S.; Dai, K.; Zhao, H.; Wu, J. Research on the Human-machine Hybrid Decision-making Strategy Basing on the Hybrid-augmented Intelligence. J. Mech. Eng. 2025, 61, 288–304. [Google Scholar] [CrossRef] [Scilit]
  66. Masehian, E.; Ghandi, S. Assembly sequence and path planning for monotone and nonmonotone assemblies with rigid and flexible parts. Robot. Comput.-Integr. Manuf. 2021, 72, 102180. [Google Scholar] [CrossRef] [Scilit]
  67. Zhang, X.; Bai, X.; Zhang, S.; He, W.; Wang, S.; Yan, Y.; Wang, P.; Liu, L. A novel mixed reality remote collaboration system with adaptive generation of instructions. Comput. Ind. Eng. 2024, 194, 110353. [Google Scholar] [CrossRef] [Scilit]
  68. Liu, C.; Tang, D.; Zhu, H.; Zhang, Z.; Wang, L.; Nie, Q. AR-assisted human-robot collaborative assembly system: Integrating visual language model and deep reinforcement learning for task planning and seamless interactive guidance. J. Manuf. Syst. 2026, 84, 40–67. [Google Scholar] [CrossRef] [Scilit]
  69. Liu, C.; Tang, D.; Zhu, H.; Wang, L.; Cai, Q.; Nie, Q. LLM-enhanced embodied multi-agent manufacturing system: A novel self-organizing production paradigm for embodied perception, embodied analysis and embodied decision. J. Manuf. Syst. 2026, 84, 357–382. [Google Scholar] [CrossRef] [Scilit]
  70. Zhu, J.; Huang, S.; Wang, P.; Xu, Z.; Liu, J.; Wang, B.; Zhao, Z.; Zheng, S.; Tao, Y.; Wang, G.; et al. A review on large language models for industrial embodied intelligence. Adv. Eng. Inform. 2026, 73, 104602. [Google Scholar] [CrossRef] [Scilit]
  71. Heydari, M.; Alinezhad, A.; Vahdani, B. A deep learning framework for quality control process in the motor oil industry. Eng. Appl. Artif. Intell. 2024, 133, 108554. [Google Scholar] [CrossRef] [Scilit]
  72. Salichs, M.A.; Castro-González, Á.; Salichs, E.; Fernández-Rodicio, E.; Maroto-Gómez, M.; Gamboa-Montero, J.J.; Marques-Villarroya, S.; Castillo, J.C.; Alonso-Martín, F.; Malfaz, M. Mini: A New Social Robot for the Elderly. Int. J. Soc. Robot. 2020, 12, 1231–1249. [Google Scholar] [CrossRef] [Scilit]
  73. Min, B.; Ross, H.; Sulem, E.; Veyseh, A.P.B.; Nguyen, T.H.; Sainz, O.; Agirre, E.; Heintz, I.; Roth, D. Recent Advances in Natural Language Processing via Large Pre-trained Language Models: A Survey. ACM Comput. Surv. 2023, 56, 30. [Google Scholar] [CrossRef] [Scilit]
  74. Xue, Z.; He, G.; Liu, J.; Jiang, Z.; Zhao, S.; Lu, W. Re-examining lexical and semantic attention: Dual-view graph convolutions enhanced BERT for academic paper rating. Inf. Process. Manag. 2023, 60, 103216. [Google Scholar] [CrossRef] [Scilit]
  75. Luo, Y.; Wu, R.; Liu, J.; Tang, X. A text guided multi-task learning network for multimodal sentiment analysis. Neurocomputing 2023, 560, 126836. [Google Scholar] [CrossRef] [Scilit]
  76. Shafizadegan, F.; Naghsh-Nilchi, A.R.; Shabaninia, E. Multimodal vision-based human action recognition using deep learning: A review. Artif. Intell. Rev. 2024, 57, 178. [Google Scholar] [CrossRef] [Scilit]
  77. Wang, Z.; Zhao, H.; Yang, Y.; Hu, D.; Bao, C.; Liu, M.; Di, K.; Dustdar, S.; Wang, Z.; Deng, S.J.A. OmniFuser: Adaptive Multimodal Fusion for Service-Oriented Predictive Maintenance. arXiv 2025, arXiv:2511.01320. [Google Scholar]
  78. Wang, B.; Xie, Q.; Pei, J.; Chen, Z.; Tiwari, P.; Li, Z.; Fu, J. Pre-trained Language Models in Biomedical Domain: A Systematic Survey. ACM Comput. Surv. 2023, 56, 55. [Google Scholar] [CrossRef] [Scilit]
  79. Yan, H.; Chen, Z.; Du, J.; Yan, Y.; Zhao, S. VisPower: Curriculum-Guided Multimodal Alignment for Fine-Grained Anomaly Perception in Power Systems. Electronics 2025, 14, 4747. [Google Scholar] [CrossRef] [Scilit]
  80. Li, J.; Wen, S.; Karimi, H.R. Zoom-Anomaly: Multimodal vision-Language fusion industrial anomaly detection with synthetic data. Inf. Fusion 2026, 127, 103910. [Google Scholar] [CrossRef] [Scilit]
  81. Huang, J.; Han, D.; Chen, Y.; Tian, F.; Wang, H.; Dai, G. A Survey on Human-Computer Interaction in Mixed Reality. J. Comput.-Aided Des. Comput. Graph. 2016, 28, 869–880. [Google Scholar]
  82. Bao, Y.; Liu, J.; Jia, X.; Qie, J. An assisted assembly method based on augmented reality. Int. J. Adv. Manuf. Technol. 2024, 135, 1035–1050. [Google Scholar] [CrossRef] [Scilit]
  83. Calderón-Sesmero, R.; Lozano-Hernández, A.; Frontela-Encinas, F.; Cabezas-López, G.; De-Diego-Moro, M. Human–Robot Interaction and Tracking System Based on Mixed Reality Disassembly Tasks. Robotics 2025, 14, 106. [Google Scholar] [CrossRef] [Scilit]
  84. Dong, H.; Zhou, X.; Li, J.; Liu, S.; Sun, J.; Gu, C. An Aircraft Part Assembly Based on Virtual Reality Technology and Mixed Reality Technology. In Proceedings of the 2021 International Conference on Big Data Analytics for Cyber-Physical System in Smart City, Singapore, 10 December 2021; pp. 1251–1263. [Google Scholar]
  85. Yuan, L.; Hongli, S.; Qingmiao, W. Research on AR assisted aircraft maintenance technology. J. Phys. Conf. Ser. 2021, 1738, 012107. [Google Scholar] [CrossRef] [Scilit]
  86. Yan, X.; Bai, G.; Tang, C. An Augmented Reality Tracking Registration Method Based on Deep Learning. In Proceedings of the 2022 5th International Conference on Artificial Intelligence and Pattern Recognition, Xiamen, China, 23–25 September 2022; pp. 360–365. [Google Scholar]
  87. Yang, H.; Li, S.; Zhang, X.; Shen, Q. Research on Satellite Cable Laying and Assembly Guidance Technology Based on Augmented Reality. In Proceedings of the 2021 40th Chinese Control Conference (CCC), Shanghai, China, 26–28 July 2021; pp. 6550–6555. [Google Scholar]
  88. Hamad, J.; Bianchi, M.; Ferrari, V. Integrated Haptic Feedback with Augmented Reality to Improve Pinching and Fine Moving of Objects. Appl. Sci. 2025, 15, 7619. [Google Scholar] [CrossRef] [Scilit]
  89. Kalkan, Ö.K.; Karabulut, Ş.; Höke, G. Effect of Virtual Reality-Based Training on Complex Industrial Assembly Task Performance. Arab. J. Sci. Eng. 2021, 46, 12697–12708. [Google Scholar] [CrossRef] [Scilit]
  90. Yan, Y.; Jiang, K.; Shi, X.; Zhang, J.; Ming, S.; He, Z. Application of Embodied Interaction Technology in Virtual Assembly Training. Packag. Eng. 2025, 46, 84–95. [Google Scholar] [CrossRef]
  91. Xu, H.-W.; Xing, H.-W.; Liu, L.-L.; Qin, W.; Lv, Y.-L.; Zhang, J. A framework for generative intelligent process design of aviation riveting based on transformer-based multimodal information fusion. J. Eng. Des. 2026, 1–41. [Google Scholar] [CrossRef] [Scilit]
  92. Li, S.; Dong, A.; Kong, R.W.M.; Liu, C. A physics-informed embodied intelligence framework for smart garment manufacturing. J. Manuf. Syst. 2026, 85, 307–317. [Google Scholar] [CrossRef] [Scilit]
  93. Seetohul, J.; Shafiee, M.; Sirlantzis, K. Augmented Reality (AR) for Surgical Robotic and Autonomous Systems: State of the Art, Challenges, and Solutions. Sensors 2023, 23, 6202. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  94. Gu, J.; Meng, X.; Lu, G.; Hou, L.; Minzhe, N.; Liang, X.; Yao, L.; Huang, R.; Zhang, W.; Jiang, X.; et al. Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark. Adv. Neural Inf. Process. Syst. 2022, 35, 26418–26431. [Google Scholar] [CrossRef] [Scilit]
  95. Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; Zhou, J. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv 2023, arXiv:2308.12966. [Google Scholar]
  96. Chen, Z.; Liu, L.; Wan, Y.; Chen, Y.; Dong, C.; Li, W.; Lin, Y. Improving BERT with local context comprehension for multi-turn response selection in retrieval-based dialogue systems. Comput. Speech Lang. 2023, 82, 101525. [Google Scholar] [CrossRef] [Scilit]
  97. Tam, Z.R.; Pai, Y.-T.; Lee, Y.-W.J.A. VisTW: Benchmarking Vision-Language Models for Traditional Chinese in Taiwan. arXiv 2025, arXiv:2503.10427. [Google Scholar]
  98. Wang, Z.; Liang, X.; Li, M.; Li, S.; Liu, J.; Zheng, L. Towards cognitive intelligence-enabled product design: The evolution, state-of-the-art, and future of AI-enabled product design. J. Ind. Inf. Integr. 2025, 43, 100759. [Google Scholar] [CrossRef] [Scilit]
  99. Bayoudh, K. A survey of multimodal hybrid deep learning for computer vision: Architectures, applications, trends, and challenges. Inf. Fusion 2024, 105, 102217. [Google Scholar] [CrossRef] [Scilit]
  100. Li, J.; Yang, S.X. Digital twins to embodied artificial intelligence: Review and perspective. Intell. Robot. 2025, 5, 202–227. [Google Scholar] [CrossRef] [Scilit]
  101. He, H.; Liao, R.; Li, Y. MSAFNet: A novel approach to facial expression recognition in embodied AI systems. Intell. Robot. 2025, 5, 313–332. [Google Scholar] [CrossRef] [Scilit]
  102. Sun, Y.; Zhu, H.; Liang, Z.; Liu, A.; Ni, H.; Wang, Y. A phase search-enhanced Bi-RRT path planning algorithm for mobile robots. Intell. Robot. 2025, 5, 404–418. [Google Scholar] [CrossRef] [Scilit]
  103. Cui, Y.; Pu, J.; Hu, N.; Cui, M. Smooth and efficient motion planning of large-scale and cooperative multi-arm tunnel drilling robot. Intell. Robot. 2025, 5, 450–473. [Google Scholar] [CrossRef] [Scilit]
  104. Sun, S.; Huang, H.; Li, C. Advancements in humanoid robot dynamics and learning-based locomotion control methods. Intell. Robot. 2025, 5, 631–660. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Analytical framework for examining digital-twin-enabled HMC, including enabling technologies, physical–virtual coupling, evidence maturity, deployment constraints, and sustainability outcomes.
Figure 1. Analytical framework for examining digital-twin-enabled HMC, including enabling technologies, physical–virtual coupling, evidence maturity, deployment constraints, and sustainability outcomes.
Electronics 15 03781 g001
Figure 2. Taxonomy of collaboration objects, levels, task coupling, and autonomy distribution defining the system boundary of digital-twin-enabled human–machine collaboration. Dashed arrows connect each review-focus box to the corresponding taxonomy dimension, while the solid bidirectional arrow in the legend denotes interaction among dimensions.
Figure 2. Taxonomy of collaboration objects, levels, task coupling, and autonomy distribution defining the system boundary of digital-twin-enabled human–machine collaboration. Dashed arrows connect each review-focus box to the corresponding taxonomy dimension, while the solid bidirectional arrow in the legend denotes interaction among dimensions.
Electronics 15 03781 g002
Figure 3. Human-centered design loop linking worker capability, task complexity, ergonomics, and adaptation in digital-twin-enabled HMC systems.
Figure 3. Human-centered design loop linking worker capability, task complexity, ergonomics, and adaptation in digital-twin-enabled HMC systems.
Electronics 15 03781 g003
Figure 4. Sustainability linkage map between digital-twin-enabled human–machine collaboration functions and manufacturing performance outcomes.
Figure 4. Sustainability linkage map between digital-twin-enabled human–machine collaboration functions and manufacturing performance outcomes.
Electronics 15 03781 g004
Figure 5. Sensor-fusion pipeline that can provide synchronized workspace and human-state information to a digital twin.
Figure 5. Sensor-fusion pipeline that can provide synchronized workspace and human-state information to a digital twin.
Electronics 15 03781 g005
Figure 6. Multimodal perception and human-state modeling as potential inputs to a purpose-specific human digital twin.
Figure 6. Multimodal perception and human-state modeling as potential inputs to a purpose-specific human digital twin.
Electronics 15 03781 g006
Figure 7. Intention understanding and bidirectional interaction loop linking operator signals, contextual interpretation, and machine response within the virtual–real loop. Solid arrows indicate the forward perception–inference–response flow, whereas dashed arrows indicate feedback and context-update pathways that close the bidirectional loop.
Figure 7. Intention understanding and bidirectional interaction loop linking operator signals, contextual interpretation, and machine response within the virtual–real loop. Solid arrows indicate the forward perception–inference–response flow, whereas dashed arrows indicate feedback and context-update pathways that close the bidirectional loop.
Electronics 15 03781 g007
Figure 8. Hierarchical task allocation and shared planning with digital-twin-based feasibility verification. Solid arrows show the forward planning and execution flow, whereas dashed arrows indicate execution feedback, learning/adaptation, and replanning loops.
Figure 8. Hierarchical task allocation and shared planning with digital-twin-based feasibility verification. Solid arrows show the forward planning and execution flow, whereas dashed arrows indicate execution feedback, learning/adaptation, and replanning loops.
Electronics 15 03781 g008
Figure 9. Reference architecture and critical integration points for digital-twin-enabled HMC. The complete closed loop represents a target configuration rather than a capability demonstrated by every reviewed study. Dashed teal arrows denote feedback/synchronization, as indicated in the legend.
Figure 9. Reference architecture and critical integration points for digital-twin-enabled HMC. The complete closed loop represents a target configuration rather than a capability demonstrated by every reviewed study. Dashed teal arrows denote feedback/synchronization, as indicated in the legend.
Electronics 15 03781 g009
Figure 10. Human digital twin modeling layers, data sources, and closed-loop adaptive feedback channels. Solid arrows indicate data/model/application flow, whereas dashed arrows indicate closed-loop feedback and continuous-improvement pathways.
Figure 10. Human digital twin modeling layers, data sources, and closed-loop adaptive feedback channels. Solid arrows indicate data/model/application flow, whereas dashed arrows indicate closed-loop feedback and continuous-improvement pathways.
Electronics 15 03781 g010
Table 1. Overview of conceptual architecture and taxonomy studies of digital-twin-enabled HMC systems.
Table 1. Overview of conceptual architecture and taxonomy studies of digital-twin-enabled HMC systems.
No.Authors (Year)Method
1Leng et al. (2022) [1]Industry 5.0 framing
2Wang et al. (2021) [2]Smart vs. intelligent manufacturing
3Li et al. (2021) [3]Proactive cognitive HRC
4Hietanen et al. (2020) [4]AR-assisted interaction
5Liau et al. (2020) [5]Capability-based task allocation
6Duan et al. (2024) [6]Multimodal dual-robot assembly
7Blankemeyer et al. (2024) [7]AR hand interaction
8Liu et al. (2024) [15]Cobot deployment review
9Petzoldt et al. (2022) [16]Dynamic task allocation
10Zanchettin et al. (2018) [17]Human-activity prediction
11Bruno and Antonelli (2018) [18]Dynamic task classification
12Malik and Bilberg (2019) [19]Complexity-based allocation
13Kong et al. (2021) [20]Joint task assignment
14Sleeman et al. (2022) [21]Multimodal taxonomy
15Xue et al. (2022) [22]Human–machine augmented intelligence
Table 2. Overview of enabling technologies for digital-twin-enabled HMC systems.
Table 2. Overview of enabling technologies for digital-twin-enabled HMC systems.
No.Authors (Year)Method
1Ding et al. (2025) [13]Human-motion prediction
2Gao et al. (2024) [24]CNN action recognition
3Mavsar et al. (2021) [25]RNN intention recognition
4Wang et al. (2018) [26]Human-motion recognition
5Bandi and Thomas (2021) [27]Skeleton action recognition
6Gao et al. (2023) [28]Hybrid recurrent intention model
7Cao et al. (2024) [29]AI-enabled soft robotics
8Hussain et al. (2024) [30]Human-activity recognition
9Zhang et al. (2024) [31]Skeleton–RGB action prediction
10Nadeem et al. (2024) [32]Vision/LLM emotion recognition
11Hazmoune and Bougamouza (2024) [33]Multimodal emotion review
12Liu et al. (2024) [34]Multimodal sentiment analysis
13Sun et al. (2023) [35]Flexible sensing
14Li et al. (2023) [36]Self-supervised multimodal learning
15Wang et al. (2024) [37]Human-action recognition
16Yang et al. (2024) [38]Natural-language HRC interface
17Balamurugan et al. (2025) [39]Wearable AR interface
18Wang et al. (2025) [40]VLM scene-graph generation
Table 3. Overview of digital-twin-enabled integration methods for human–machine collaboration systems.
Table 3. Overview of digital-twin-enabled integration methods for human–machine collaboration systems.
No.Authors (Year)Method
1Baratta et al. (2024) [8]Review of DT-enhanced HRC architectures
2Liu et al. (2025) [14]VLM perception in DT-assisted HRC
3Sun et al. (2022) [42]DT-driven collaborative assembly
4Wang et al. (2025) [43]Data-driven transfer between virtual and physical systems
5Dimitropoulos et al. (2025) [44]LLM-supported task modeling
6Bilberg and Malik (2019) [49]Workcell DT integration
7Wang et al. (2022) [50]DT-based HRC assembly design and operation
8Piardi et al. (2024) [51]Human integration in industrial CPS
9Ling et al. (2024) [52]Real-time human–machine synchronization
10Havard et al. (2019) [53]DT/VR co-simulation
11Krupas et al. (2024) [54]Human-centric DT review
12Zafar et al. (2024) [55]Industry 5.0/DT/cobot synthesis
13Piardi et al. (2023) [56]Digital technologies in CPS
14Choi et al. (2022) [57]Mixed-reality safety-aware HRC using DT
15Liu et al. (2026) [58]VLM-enabled embodied multi-agent DT assembly
Table 4. Overview of adaptive control, safety assurance, and human–AI decision-making methods in digital-twin-enabled HMC systems.
Table 4. Overview of adaptive control, safety assurance, and human–AI decision-making methods in digital-twin-enabled HMC systems.
No.Authors (Year)Method
1Joo et al. (2022) [23]DRL task allocation
2Lim et al. (2024) [45]LLM collaborative assembly
3Chen et al. (2026) [46]Perception–decision–execution coordination
4Chen et al. (2025) [47]LLM autonomous collaboration
5Liu et al. (2025) [48]Large-model human–AI decision-making
6Liu et al. (2024) [61]Real-time hierarchical safety control
7Xia et al. (2024) [62]LLM fine-tuning for manufacturing
8Laplaza et al. (2025) [63]Contextual motion/intention prediction
9Wang and Zhang (2024) [64]RL navigation and control
10Ma et al. (2025) [65]Human–machine hybrid decision
11Masehian and Ghandi (2021) [66]Assembly sequence/path planning
12Zhang et al. (2024) [67]Adaptive mixed-reality instructions
13Liu et al. (2026) [68]AR + VLM + DRL collaborative assembly
14Liu et al. (2026) [69]Embodied multi-agent manufacturing
15Zhu et al. (2026) [70]Industrial embodied-intelligence review
Table 5. Overview of industrial applications and evaluation studies of digital-twin-enabled HMC systems.
Table 5. Overview of industrial applications and evaluation studies of digital-twin-enabled HMC systems.
No.Authors (Year)Method
1Barathwaj et al. (2015) [41]Assembly-line balancing
2Cai et al. (2025) [59]Collaborative assembly allocation
3Ji et al. (2018) [60]Aerospace robotic assembly
4Heydari et al. (2024) [71]Quality inspection
5Wang et al. (2025) [77]Multimodal perception
6Yan et al. (2025) [79]Industrial anomaly detection
7Li et al. (2026) [80]VLM industrial perception
8Bao et al. (2024) [82]AR-assisted assembly
9Calderón-Sesmero et al. (2025) [83]Mixed-reality disassembly
10Dong et al. (2021) [84]Aerospace assembly with VR/MR
11Yuan et al. (2021) [85]AR aircraft maintenance
12Yan et al. (2022) [86]AR tracking registration
13Yang et al. (2021) [87]Satellite cable assembly guidance
14Hamad et al. (2025) [88]Haptic AR manipulation
15Kalkan et al. (2021) [89]VR industrial assembly training
16Yan et al. (2025) [90]Embodied virtual assembly training
17Xu et al. (2026) [91]Aviation riveting process design
18Li et al. (2026) [92]Smart garment manufacturing
Table 6. Overview of challenges and future directions for digital-twin-enabled HMC system development.
Table 6. Overview of challenges and future directions for digital-twin-enabled HMC system development.
No.Authors (Year)Method
1Cao et al. (2024) [29]AI-enabled soft robotics
2Bayoudh et al. (2024) [99]Multimodal fusion review
3Salichs et al. (2020) [72]Social-robot interaction
4Min et al. (2023) [73]Pre-trained language models
5Xue et al. (2023) [74]BERT-based assessment
6Luo et al. (2023) [75]Multimodal sentiment analysis
7Shafizadegan et al. (2024) [76]Human-action recognition review
8Wang et al. (2023) [78]Language-model survey
9Huang et al. (2016) [81]Mixed-reality HCI review
10Seetohul et al. (2023) [93]AR in robotic/autonomous systems
11Gu et al. (2022) [94]Cross-modal pre-training
12Bai et al. (2023) [95]General-purpose VLM
13Chen et al. (2023) [96]Context-aware dialogue modeling
14Tam et al. (2025) [97]VLM benchmarking
15Wang et al. (2025) [98]AI-enabled product-design review
16Li and Yang (2025) [100]DT-to-embodied-AI review
17He et al. (2025) [101]Facial-expression recognition
18Sun et al. (2025) [102]Bi-RRT path planning
19Cui et al. (2025) [103]Cooperative multi-arm planning
20Sun et al. (2025) [104]Humanoid locomotion review
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhang, H.; Chen, J.; Liu, G.; Yang, F.; Guo, H. Digital-Twin-Enabled Human–Machine Collaboration Systems in Sustainable Smart Manufacturing: System Architecture, Development Methods, Applications, and Future Trends. Electronics 2026, 15, 3781. https://doi.org/10.3390/electronics15173781

AMA Style

Zhang H, Chen J, Liu G, Yang F, Guo H. Digital-Twin-Enabled Human–Machine Collaboration Systems in Sustainable Smart Manufacturing: System Architecture, Development Methods, Applications, and Future Trends. Electronics. 2026; 15(17):3781. https://doi.org/10.3390/electronics15173781

Chicago/Turabian Style

Zhang, Haitao, Jingtao Chen, Gaoyu Liu, Fanyu Yang, and Hao Guo. 2026. "Digital-Twin-Enabled Human–Machine Collaboration Systems in Sustainable Smart Manufacturing: System Architecture, Development Methods, Applications, and Future Trends" Electronics 15, no. 17: 3781. https://doi.org/10.3390/electronics15173781

APA Style

Zhang, H., Chen, J., Liu, G., Yang, F., & Guo, H. (2026). Digital-Twin-Enabled Human–Machine Collaboration Systems in Sustainable Smart Manufacturing: System Architecture, Development Methods, Applications, and Future Trends. Electronics, 15(17), 3781. https://doi.org/10.3390/electronics15173781

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop