Next Article in Journal
The Energy-Environmental Kuznets Curve: Evidence from a Time-Varying Parametric Framework
Next Article in Special Issue
AI for Sustainable Cultural Industries: A Screenplay-Aware Knowledge-Enhanced State Space Model with LLM-Derived Narrative Features for Forecasting Film Industry Sustainability Across National Economies
Previous Article in Journal
Sustainable Entrepreneurship Orientation: Application of a Formative Measurement Model
Previous Article in Special Issue
Spatial Effects of Artificial Intelligence Innovation on Regional Carbon Intensity
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

Human–AI Collaboration Across Decision Support, Autonomous Systems, and LLM Agents: A Systematic Review and Collaboration Convergence Framework

1
Department of Management, Marketing and Operations, David B. O’Maley College of Business, Embry-Riddle Aeronautical University, Daytona Beach, FL 32114, USA
2
School of Graduate Studies, College of Aviation, Embry-Riddle Aeronautical University, Daytona Beach, FL 32114, USA
3
Department of Accounting, Economics, Finance & Information Sciences, David B. O’Maley College of Business, Embry-Riddle Aeronautical University, Daytona Beach, FL 32114, USA
4
Rockwell Automation, Milwaukee, WI 53204, USA
*
Author to whom correspondence should be addressed.
Sustainability 2026, 18(11), 5313; https://doi.org/10.3390/su18115313
Submission received: 25 March 2026 / Revised: 29 April 2026 / Accepted: 20 May 2026 / Published: 25 May 2026

Abstract

Across four decades of AI deployment, the same six human challenges (trust calibration, reliance behavior, cognitive engagement, skill retention, accountability, and transparency) recur, yet fragmentation across research communities obscures this continuity and limits knowledge transfer. Functionally similar phenomena are repeatedly relabeled (a jangle fallacy): what aviation researchers call “automation complacency,” decision scientists call “algorithm appreciation,” and LLM researchers describe as “over-reliance.” This systematic review synthesizes 152 papers spanning aviation, healthcare, manufacturing/supply chain, and cross-domain contexts across three AI technology generations: decision support systems, autonomous systems, and large language model (LLM) agents. We introduce the Collaboration Convergence Framework (CCF), a 6 × 3 matrix with solution-maturity indicators that maps each challenge across generations. The framework shows that Gen 3 designers can transfer decades of evidence from automation and decision support research (particularly reliance calibration, cognitive forcing, and skill maintenance) rather than rediscovering them. Cross-generational synthesis also isolates three Gen 3 phenomena without direct precedent in earlier generations: epistemia (attributing genuine knowledge to LLMs based on surface fluency), attribution ambiguity in co-creation, and motivational withdrawal. We distill twelve transferable design principles and propose ten research directions, prioritizing skill-retention interventions and accountability frameworks. These findings carry direct sustainability implications aligned with Industry 5.0: protecting workforce capability under increasing automation (SDG 8), reducing duplicated research effort through cross-generational knowledge reuse (SDG 9), and supporting responsible deployment by treating collaboration risks as predictable rather than novel (SDG 12). The CCF provides conceptual infrastructure for cumulative learning across AI generations and industries.

1. Introduction

In 2009, Air France Flight 447 crashed into the Atlantic Ocean, killing 228 people [1]. The accident involved a cascade of human–automation coordination failures: When icing disabled the pitot tubes and triggered an autopilot disconnection, the flight crew had to reassume manual control of an aircraft they had been monitoring passively. Loss of situation awareness after prolonged automated flight, combined with control-input coordination failures between the copilots, produced inappropriate responses that went uncorrected. The disaster reflected the interaction of automation design choices, degraded operator situation awareness, and crew-coordination breakdown, a cascade pattern that characterizes human–automation failures across domains.
Across domains, meta-analytic evidence shows that human–AI teams frequently fail to outperform either humans or AI alone: in 58% of effect sizes, the human–AI combination underperformed the better of human or AI alone [2], and formal analyses confirm that complementarity is harder to achieve than widely assumed [3,4]. In 2023, a field experiment with 758 consultants at a major firm showed that when tasks fell outside AI’s “jagged technological frontier,” consultants using AI were 19 percentage points less likely to produce correct solutions than those without AI, a risk compounded by the tendency of LLMs to generate incorrect but plausible outputs [5]. Three technologies, three decades, three domains, yet the collaboration breakdown is familiar: excessive trust in AI output paired with insufficient independent checking.
These vignettes reflect a persistent reality of AI deployment: many of the most consequential failures in human–AI systems are less about algorithmic capability than about human behavior under uncertainty, time pressure, and automation bias. Research on these challenges has expanded rapidly, but remains fragmented across three largely disconnected communities (those studying decision support systems, autonomous systems, and the newest generation of large language model (LLM) agents), each with distinct terminology, frameworks, and design traditions. Evidence accumulates in parallel rather than cumulatively, with functionally similar phenomena repeatedly renamed across generations (a jangle fallacy), limiting the transfer of validated design knowledge. Because sustainable technological transitions depend on institutions’ ability to accumulate and transfer knowledge across sectors and technological generations, this fragmentation imposes a concrete sustainability cost: lessons learned in one domain fail to inform others, slowing the institutional learning required for resilient and responsible AI deployment.
This paper presents a systematic review spanning three AI technology generations to identify recurring patterns in human collaboration challenges and to make cross-generational evidence usable for design. Synthesizing 152 papers across four domains (aviation, healthcare, manufacturing/supply chain, and cross-domain contexts), we show that the same six core challenges recur across generations: trust calibration, reliance behavior, cognitive engagement, skill retention, accountability, and transparency. We formalize this convergence via the Collaboration Convergence Framework (CCF), a 6 × 3 matrix mapping generation-specific terminology to unified constructs and indicating solution maturity. The framework yields transferable design principles and highlights where Gen 3 introduces genuinely new risks that require novel interventions.
Specifically, this study addresses four research questions:
RQ1: What are the dominant human–AI collaboration challenges identified in empirical research across three AI technology generations?
RQ2: What persists and what changes in human–AI collaboration challenges as AI systems evolve from advisory to autonomous to conversational interaction?
RQ3: What cross-domain patterns emerge when comparing collaboration findings across aviation, healthcare, manufacturing, and supply chain?
RQ4: What design principles can be transferred across domains and AI generations to improve human–AI collaboration outcomes and support sustainability objectives, particularly those associated with SDG 8 (Decent Work and Economic Growth), SDG 9 (Industry, Innovation and Infrastructure), and SDG 12 (Responsible Consumption and Production)?
The sustainability implications of this work are threefold and are aligned with the United Nations Sustainable Development Goals (SDGs). First, skill retention and cognitive engagement challenges threaten workforce capability, employee well-being, and meaningful human participation in increasingly automated systems, directly relating to SDG 8. Second, repeated reinvention and weak knowledge transfer across AI generations represent inefficient use of research and innovation resources, aligning with SDG 9. Third, reliable and responsible human–AI collaboration requires calibrated trust, transparency, and accountability, supporting SDG 12 by promoting sustainable technological deployment and reducing preventable operational failures.
This review makes four principal contributions. First, it provides a cross-generational synthesis of 152 studies spanning decision support, autonomous systems, and LLM agents, a scope that no prior review has attempted. Second, it introduces the CCF, a structured 6 × 3 matrix with solution maturity indicators that formalizes which human–AI challenges recur across generations and where knowledge transfer is most defensible. Third, it identifies three genuinely novel Gen 3 challenges, namely epistemia, attribution ambiguity, and motivational withdrawal, that lack precedent in earlier automation research. Fourth, it synthesizes twelve transferable design principles and ten prioritized research directions, grounded in convergent evidence and linked to sustainability outcomes.
The remainder of this paper is organized as follows. Section 2 reviews existing literature and positions this work relative to prior reviews. Section 3 describes the systematic review methodology. Section 4 presents findings organized by the six convergent challenges. Section 5 introduces the CCFand synthesizes cross-generational patterns. Section 6 discusses design principles, research directions, and limitations. Section 7 concludes with implications for research and practice.

2. Related Reviews

2.1. Human–AI Collaboration as a Research Field

Human–AI collaboration research has expanded rapidly since 2018, driven by advances in machine learning, autonomous systems, and generative AI. The field spans multiple parent disciplines, including human factors engineering, cognitive psychology, organizational behavior, and computer science, each contributing distinct theoretical commitments and methodological traditions. This interdisciplinarity is an advantage for understanding complex collaboration phenomena, but it also produces terminological fragmentation and siloed communities that rarely build on one another’s evidence base.

2.2. Existing Reviews and Their Limitations

Seventeen existing reviews and synthesis papers are directly relevant to human–AI collaboration. Taken together, they share three recurring structural limitations that collectively define the gap this paper addresses.
First, single-generation scope. Most prior reviews focus on one AI generation at a time. Lai et al. [6] and Schemmer et al. [7] emphasize Gen 1 advisory decision-making, synthesizing empirical work on reliance and explanations. Gomez et al. [8] extend this line with a taxonomy of interaction patterns in AI-assisted decisions. Romeo & Conti [9] review 35 studies of automation bias under XAI. O’Neill et al. [10] focus on Gen 2 human-autonomy teaming from a team-cognition perspective. Walker et al. [11] address trust in automated vehicles; Xia et al. [12] review air-traffic-management automation. Schmutz et al. [13] synthesize human–AI teaming from the perspective of coordination and shared cognition. Emerging surveys address Gen 3 LLM collaboration but without systematic connection to earlier generations. Design knowledge developed in one generation therefore rarely informs the next.
Second, single-domain focus. Domain-specific reviews (Kirwan [14] for aviation human-factors requirements, Dai & Abràmoff [15] for healthcare workflows, Wan et al. [16] for human–robot collaboration in robotic surgery, Samuels [17] and Kumar et al. [18] for supply chain) provide within-domain depth but rarely connect findings to functionally equivalent mechanisms in other domains, leaving cross-domain transfer opportunities unexploited.
Third, single-structural or single-construct focus. Several reviews isolate one construct or theoretical lens: trust measurement (Kohn et al. [19] synthesize measurement instruments; Wong et al. [20] trace a 30-year trust evolution), affordance theory (Bao et al. [21] apply an affordance lens), HAIT terminology scoping (Berretta et al. [22] define human–AI teaming from a human-centered perspective), or trust modeling (Do Khac & Leyer [23] propose a HAI trust model), rather than integrating multiple challenges under a unified framework. No existing review systematically maps challenges across all three AI generations to identify convergent patterns and transferable design principles.
The present review targets this three-fold gap by synthesizing evidence across Gen 1, Gen 2, and Gen 3, across four application domains, and across six human-factor constructs within a single framework (the CCF, Section 5). A detailed comparison of the 17 reviews against AI-generation coverage, domains, human-factor constructs, and key gaps is provided in Appendix A (Table A1).
Recent Industry 5.0 frameworks [24,25] emphasize human-centric collaboration design, but without cross-generational synthesis they cannot fully leverage decades of evidence on how people use, misuse, disuse, and abuse intelligent systems [26]. The present review targets this gap by aligning constructs across generations and domains, and by identifying mitigations that have robust support and can be transferred to newer technologies.

2.3. AI Technology Generations: A Working Taxonomy

To enable systematic cross-generational analysis, we define three AI technology generations based on interaction modality, human role, and system autonomy, not on the underlying computational technique [27,28]. This distinction is important. A widely used taxonomy in the computer science literature classifies AI systems by their technical architecture: rule-based/symbolic systems, machine-learning systems, and autonomous/general intelligence. Our taxonomy instead classifies systems by how humans collaborate with them, grounded in the human-factors tradition of levels-of-automation analysis [26,29]. Under this interaction-paradigm lens, a rule-based autopilot and a hypothetical neural-network autopilot both belong to the same generation because the human’s role (monitoring and intervening in autonomous operation) is identical. Similarly, large language models are classified separately from earlier machine-learning systems not because of algorithmic novelty, but because conversational co-creation introduces a qualitatively different collaboration dynamic. The taxonomy reflects a progression from advisory to supervisory to conversational collaboration, with each shift introducing qualitatively different risks of use, misuse, disuse, and abuse [26]. Importantly, these three generations are not sequential replacements; they coexist and overlap temporally (Figure 1). Decision support systems remain ubiquitous, autopilot continues in daily operation, and LLM agents are layered alongside both. The generational labels reflect the approximate chronological order in which each interaction paradigm became prominent in the research literature, not a hierarchy of technological sophistication.
Generation 1: AI as Decision Support. AI provides recommendations, predictions, or risk assessments; the human makes final decisions. Interaction is primarily one-directional (AI to human), and the human retains authority. Examples include clinical decision support systems, credit scoring, demand forecasting tools, and predictive maintenance alerts. This generation aligns with decision support traditions and corresponds to lower levels of automation (information acquisition/analysis) in Parasuraman et al.’s framework [29].
Generation 2: AI as Autonomous Partner. AI can act independently within defined scope; the human monitors, supervises, or intervenes. Control is shared via handoff protocols and variable autonomy. Examples include autopilot systems, surgical robots, autonomous vehicles, and warehouse AMRs. This generation corresponds to higher automation levels (decision selection and action implementation) [29].
Generation 3: AI as Cognitive Collaborator. AI engages in reasoning, dialogue, and content generation through conversational interaction. The human and AI jointly produce outputs through iterative co-creation. Examples include LLM copilots for analysis and writing, LLM agents for planning/execution, and multi-agent systems. This generation emerged with widely available large language models from 2022 onward [30]. Although anthropomorphism and conversational fluency create trust dynamics that differ from earlier automation, Gen 3 nonetheless reproduces many core reliance and engagement challenges in new interface forms [31].
Figure 1 visualizes this taxonomy and situates exemplar systems and review coverage across generations. We use these categories throughout the paper to (i) organize the evidence base, (ii) align terminology across communities, and (iii) assess which challenges and mitigations persist, transform, or newly emerge as interaction modalities change.
This taxonomy is intentionally stylized: real systems may blend characteristics (e.g., decision support embedded in semi-autonomous tools; LLM agents embedded in operational autonomy). Its purpose is analytic rather than ontological, to support cross-generational comparison and clarify which collaboration challenges persist versus newly emerge.

3. Methodology

3.1. Review Protocol and PRISMA Compliance

This systematic review is reported in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines [32]. A completed PRISMA 2020 checklist is provided as Supplementary Material. The review was not pre-registered; pre-registration was not pursued because (i) the cross-generational and cross-disciplinary scope made a stable a priori protocol difficult to specify without iterative refinement as the reviewed literatures were synthesized, and (ii) the rapidly evolving Gen 3 LLM-collaboration literature required protocol adjustments during the review window (2018–2026). The review protocol, search strategy, inclusion/exclusion criteria, screening procedure, data-extraction approach, and synthesis methods are documented below in full (Section 3.2, Section 3.3, Section 3.4, Section 3.5, Section 3.6, Section 3.7 and Section 3.8) to support transparency and replicability.
Because the review spans three AI technology generations, four application domains, and several human-factor constructs, we supplemented a master PRISMA-style search with targeted within-generation and within-domain searches to reduce undercoverage in fast-evolving areas (particularly Gen 3 LLM collaboration and high-impact healthcare/operations-management venues). Both search streams are described in Section 3.2; all sources are captured in the PRISMA 2020 flow diagram (Figure 2). The supplementary-search rationale and its associated confirmation-bias risk are discussed in the limitations (Section 6.5).

3.2. Search Strategy

Searches were conducted across four bibliographic databases: Web of Science (Core Collection), Scopus, PubMed (healthcare emphasis), and IEEE Xplore (engineering and autonomy emphasis). The final search across all four databases was conducted on 20 March 2026. The master search string combined three concept blocks using Boolean AND operators, with synonyms within each block combined using OR: (“human-AI” OR “human-artificial intelligence” OR “human–machine” OR “human-automation” OR “human–robot” OR “human-autonomy” OR “human-LLM” OR “human-agent”) AND (“collaboration” OR “cooperation” OR “teaming” OR “interaction” OR “decision making” OR “augmentation” OR “complementarity”) AND (“trust” OR “reliance” OR “performance” OR “cognitive load” OR “situation awareness” OR “skill” OR “accountability” OR “transparency” OR “explainability”).
Search fields were standardized across databases to title, abstract, and author/indexed keywords; full-text searches were not used, to limit recall to conceptually central studies. Database-specific syntax adaptations (Web of Science TS = field; Scopus TITLE-ABS-KEY; PubMed [tiab]; IEEE Xplore All Metadata) preserved the Boolean logic of the master string. No language filters were applied at the search stage; non-English records were excluded at screening (see Section 3.3).
The systematic timeframe was 2018–2026, chosen to capture the post-deep-learning inflection in human–AI collaboration research and the subsequent large-language-model acceleration. Pre-2018 works were included only if they satisfied both of two criteria: (i) widely cited as theoretical foundations in the reviewed domains (e.g., Bainbridge [33], Parasuraman et al. [26,29], Sarter & Woods [34], Lee & See [28], Hancock et al. [35], Beer et al. [27]) and (ii) directly cited as conceptual scaffolding by multiple included post-2018 studies. Seven foundational pre-2018 works met both criteria and were retained. Only English-language, peer-reviewed journal articles, peer-reviewed conference proceedings, and recognized accident/investigation reports were included; editorials, book reviews, and non-peer-reviewed preprints were excluded.
Master and supplementary searches. The master search returned the majority of included studies. Preliminary title/abstract screening revealed three areas of undercoverage that motivated targeted supplementary searches: (i) recent Gen 3 LLM-collaboration studies (2024–2026) indexed under diverse terminology (“LLM agents,” “foundation-model assistants,” “copilots”) not uniformly captured by the master string; (ii) domain-specific journals in aviation human factors and clinical informatics not fully reached by the cross-disciplinary master query; and (iii) recent healthcare-AI and operations-management studies published in top-tier venues (Nature Medicine, JAMA, Management Science, Organization Science, Information Systems Research, Production and Operations Management, MSOM) whose titles use field-specific terminology rather than explicit human-factor vocabulary. For each gap area, a targeted search string was constructed using domain- or generation-specific terms combined with the C3 human-factor block from the master query (e.g., for Gen 3 healthcare: (“LLM” OR “ChatGPT” OR “foundation model”) AND (“clinician” OR “physician” OR “radiologist”) AND (“reliance” OR “trust” OR “accuracy”)). Targeted-search records were integrated with the master-search pool prior to deduplication; the PRISMA flow diagram (Figure 2) reports the combined identified/screened/included counts. Risks introduced by the supplementary-search strategy, particularly confirmation bias toward findings consistent with the convergence hypothesis, are addressed through the inclusion of null and negative findings (Section 3.6) and discussed in Section 6.5.

3.3. Inclusion and Exclusion Criteria

Inclusion criteria (all of C1–C5 required):
C1.
Contribution type. Empirical study (experiment, field study, case study, survey, simulation), conceptual framework, systematic or scoping review, or recognized accident/investigation report.
C2.
Human–AI collaboration focus. Addresses human–AI collaboration, human–AI teaming, human–automation interaction, human–robot collaboration, or human–LLM/agent interaction, with a human present in the study design as user, subject, collaborator, or decision-maker.
C3.
Human-factor construct. Analyzes at least one of the following: trust, reliance, performance, cognitive load, situation awareness, skill retention/deskilling, accountability, transparency/explainability, cognitive engagement, or a closely related human-side collaboration construct.
C4.
Domain relevance. Aviation, healthcare, manufacturing/supply chain, or cross-domain human–AI collaboration context. Adjacent domains (e.g., education, finance, creative work) qualified only if findings were framed as cross-domain or explicitly transferable.
C5.
Peer-review status. Peer-reviewed journal article, peer-reviewed conference proceedings paper, recognized accident/investigation report, or foundational reference satisfying both pre-2018 criteria (Section 3.2).
Exclusion criteria (any of E1–E5 sufficient):
E1.
Purely technical AI paper. Model, algorithm, or architecture paper with no human-collaboration component in the study design.
E2.
Adoption/intention-only study. Behavioral-intention study (e.g., TAM, UTAUT, attitude or perception survey) without measurement of actual collaborative behavior, reliance, or trust-in-use.
E3.
No human-factor construct. Does not analyze any construct from C3.
E4.
Out-of-scope domain. Fully outside C4 with no cross-domain framing.
E5.
Not peer-reviewed. Preprint without subsequent publication, opinion piece, editorial, or book review.

3.4. Screening and Selection

The search yielded 1247 records across databases (combined master and supplementary searches). After deduplication (705 unique records, with 542 duplicates removed), title/abstract screening excluded 485 records and produced 220 candidate papers for full-text retrieval; 12 could not be retrieved (paywalls, inaccessible proceedings, or broken identifiers), so 208 were assessed at full text. Full-text screening against the criteria in Section 3.3 produced a final sample of 152 papers: 7 foundational taxonomy papers, 17 existing review papers (Table A1), 40 Gen 1 decision support papers, 27 Gen 2 autonomous-systems papers, 26 Gen 3 LLM/agent papers, 12 sustainability/Industry 5.0 papers, 11 additional top-journal papers (2024–2026) identified through targeted gap analysis, 3 papers added during manuscript drafting to cover identified gaps (1 aviation accident investigation report, 1 methodological guideline, 1 IoT/intralogistics study), and 9 healthcare and operations management papers added during late-stage drafting to strengthen clinical and OM coverage (including publications in Nature Medicine, JAMA, Organization Science, ISR, POM, and MSOM).
Screening reliability: verification re-screening. Primary title/abstract and full-text screening was conducted by the lead author against the criteria defined in Section 3.3. At the title/abstract stage, a deliberately inclusive threshold was applied: papers were retained whenever relevance to human–AI collaboration was plausible, to minimize premature exclusion. Borderline cases at the full-text stage ( n 25 ) were discussed with co-authors, with final inclusion decisions ratified by the author group during data extraction.
To assess screening reliability and support PRISMA 2020 Item 8 reporting, a verification re-screening of 40 randomly selected papers from the final included set was subsequently conducted by a second author (P.L.), blind to the original inclusion rationale. Papers were classified under a three-category scheme: Clear Include (all inclusion criteria unambiguously satisfied), Borderline Include (criteria met but with principled hesitation; e.g., peripheral human-factor construct, indirect domain fit, or edge-case contribution type), and Would Exclude (at least one inclusion criterion fails or at least one exclusion criterion applies). The second author confirmed inclusion for 38 of 40 papers (95.0%), with 10 of 40 (25.0%) flagged as borderline (primarily on C3 construct depth or C4 domain fit), and 2 of 40 (5.0%) flagged as Would Exclude under strict readings of E3 and E2, respectively. The two Would Exclude cases were re-evaluated by the full author group against fuller readings of the original texts (specifically: a firm-level human–AI coupling construct with a managerial-cognition moderator in one paper; a four-study design in the other whose qualitative sub-study analyzed trust, emotional bonds, and reliance in human–AI co-creation) and retained as borderline inclusions. Final post-resolution agreement was 40 of 40 papers (100%).
Formal inter-rater agreement (Cohen’s κ ) was not computed because the verification sample was drawn only from the included-paper set, producing a degenerate reference marginal under which chance-corrected agreement statistics are undefined. The verification screening should therefore be interpreted as a post hoc consistency check of inclusion decisions against the criteria in Section 3.3 rather than as classical dual-reviewer inter-rater reliability. This design choice and its implications are acknowledged as a limitation in Section 6.5.

3.5. Data Extraction and Coding

For each included paper, the following variables were extracted: bibliographic information (authors, year, venue); AI generation (Gen 1/2/3, applying the classification rules below); domain (aviation/healthcare/manufacturing/supply-chain/cross-domain); study type (experimental/observational/conceptual/review/accident-report); participant/sample characteristics; collaboration mode (advisory/supervisory/conversational-co-creative); human-factor constructs studied; key findings; and effect direction (positive/negative/mixed/null) where applicable. Data extraction was conducted by the lead author. Because the review synthesizes narratively rather than meta-analytically (Section 3.7), independent dual extraction with inter-rater reconciliation was not performed; extracted variables were cross-checked against each paper’s full text during iterative manuscript drafting and during the verification re-screening described in Section 3.4. The complete list of all 152 included studies, organized by AI generation and domain, is provided in Table 1 (Section 4.1).
Operationalization of AI generation classification. Papers were classified by AI generation using the interaction-paradigm taxonomy defined in Section 2.3, not by underlying computational technique (e.g., rule-based vs. neural vs. transformer-based). The operational criteria applied at coding were as follows:
  • Gen 1 (Decision Support). AI produces recommendations, predictions, scores, or risk assessments; the human evaluates the output and makes the final decision; interaction is primarily AI-to-human; AI does not take operational action without human ratification.
  • Gen 2 (Autonomous Partner). AI acts independently within a defined operational scope; the human monitors, supervises, or intervenes; authority transfer occurs via handoff protocols or variable autonomy; physical or operational consequences of AI actions may occur before human intervention.
  • Gen 3 (Cognitive Collaborator). Interaction is dialogic and generative; human and AI jointly produce outputs through iterative co-creation; outputs are primarily language or multi-modal content; the AI component is explicitly a large language model, multi-modal LLM, or LLM-based agent.
For papers spanning generations (e.g., LLM-assisted clinical decision support combining Gen 1 advisory output with Gen 3 conversational interaction), the generation corresponding to the dominant empirical focus of the study was assigned; dual-generation papers are noted in Section 4 where relevant. Boundary cases (for example, highly automated decision support systems approximating Gen 2 supervisory dynamics) were resolved through author-group discussion against the dominant-interaction-modality criterion. Seventeen review/foundational papers (Table A1) were not assigned to a single generation because they span multiple generations by design.

3.6. Risk of Bias and Evidence Consistency

Study-level risk of bias. Formal study-level risk of bias assessment using a standardized instrument (e.g., Cochrane RoB 2, ROBINS-I, CASP) was not conducted. The methodological heterogeneity of the corpus (spanning randomized experiments, field experiments, qualitative case studies, conceptual frameworks, systematic reviews, simulations, and accident/investigation reports across four disciplines) precluded the application of any single risk-of-bias instrument to the full set of included studies. In place of formal RoB scoring, three quality-filtering mechanisms were applied. First, the inclusion criterion C5 (Section 3.3) restricted the corpus to peer-reviewed journal articles, peer-reviewed conference proceedings, recognized accident/investigation reports, and widely cited foundational works. Second, the 152-paper corpus is drawn substantially from high-impact venues including Nature Medicine, JAMA, Management Science, Organization Science, Information Systems Research, Production and Operations Management, MSOM, Human Factors, European Journal of Information Systems, ACM CHI/CSCW/TOCHI/IUI/FAccT, and Automatica, where editorial review provides an implicit quality threshold. Third, we used the solution-maturity indicators of the CCF (Table 2; Known, Emerging, Unsolved) as a cross-study evidence-consistency signal: a Known classification required convergent findings across at least two independent studies and at least two domains (operational criteria in Section 5.2). This substitution of evidence-consistency coding for formal study-level RoB is a methodological trade-off that preserves comprehensive cross-generational scope at the cost of quantitative bias scoring; the trade-off is acknowledged as a limitation in Section 6.5.
Reporting and publication bias. Publication bias may over-represent positive collaboration findings, particularly for Gen 3 where research communities are highly motivated to demonstrate LLM benefits. This risk was addressed by actively seeking and including studies reporting null, negative, or mixed findings (e.g., Vaccaro et al. [2] on human–AI teams underperforming better individual agents; Dell’Acqua et al. [5] on jagged-frontier failures; Jabbour et al. [36] on the insensitivity of biased-AI harms to explanation provision; Bansal et al. [37] and Schemmer et al. [7] on explanations failing to improve discrimination). Residual publication-bias risk is acknowledged in Section 6.5.

3.7. Analysis Approach

Given the heterogeneity in methods, systems, and outcomes across the 152 papers, formal meta-analysis was not appropriate. We therefore used narrative synthesis with thematic analysis along two axes: (a) within-generation synthesis (Section 4), grouping findings by recurring sub-challenges within each generation; and (b) cross-generational synthesis (Section 5), where the CCF maps recurring challenges to identify convergent patterns and transferable design principles. Solution maturity indicators were assigned based on the strength and consistency of evidence for interventions addressing each challenge.

3.8. Analytical Flow

Figure 3 maps the four research questions (Section 1) onto the corresponding methodological steps and results sections. The analytical logic moves from within-generation synthesis (RQ1, Section 4), to cross-generational convergence and the construction of the CCF (RQ2, Section 5.1), to cross-domain comparative analysis (RQ3, Section 5.2 and Section 5.4), to design-principle distillation (RQ4, Section 5.3). Solution-maturity coding (Section 3.6) operates across the synthesis steps, informing both the CCF cells (Table 2) and the prioritization of research directions (Section 6.4).

4. Findings: Dominant Human–AI Collaboration Challenges Across Three AI Generations (RQ1)

This section synthesizes empirical findings on human–AI collaboration organized by AI technology generation. For each generation, we summarize dominant collaboration patterns and recurring human-side challenges. The synthesis draws on 152 papers across four domains (aviation, healthcare, manufacturing/supply chain, and cross-domain contexts) published between 2018 and 2026, supplemented with seminal foundational works. Sustainability and Industry 5.0 perspectives are integrated where relevant because collaboration quality directly affects resilience, workforce well-being, and equitable deployment outcomes [38].

4.1. Overview of Included Studies

Table 1 lists the 152 included studies organized by AI generation (Foundational, Review, Gen 1, Gen 2, Gen 3, and Cross-Gen for studies spanning multiple generations) and, within each generation, by primary application domain. Each study is reported by its first author and year, with the corresponding reference number, primary AI generation classification, application domain, study type (empirical, conceptual, review, or accident/investigation report), and the human-factor construct(s) most directly addressed. Detailed within-generation findings are presented in Section 4.2, Section 4.3 and Section 4.4; cross-domain analyses are reported in Section 5.4. The 17 prior reviews listed here are characterized in greater detail in Appendix A (Table A1). Note: Section 3.4 enumerates the 152 papers by source category (foundational, prior reviews, generation-specific contributions, and supplementary additions for sustainability, gap-filling, and revision-stage coverage); Table 1 regroups those source categories into the six analytic generation groups used throughout Section 4 and Section 5. Both breakdowns sum to 152.

4.2. Generation 1: AI as Decision Support

Generation 1 systems provide recommendations, predictions, or risk assessments that humans evaluate before making final decisions. This advisory relationship remains the most studied form of human–AI collaboration. We identified 40 papers spanning clinical decision support, demand forecasting, credit scoring, judicial risk assessment, and organizational decision-making.

4.2.1. Trust and Reliance: The Aversion-Appreciation Paradox

Do users trust algorithms too much, or too little? The Gen 1 literature gives both answers. Dietvorst et al. [64] demonstrated that seeing an algorithm err reduced participants’ willingness to use it from 65% to 26%, even when human forecasters produced 15–29% more errors than the algorithm. Yet in a contrasting series of experiments, Logg et al. [79] found systematic “algorithm appreciation,” with participants preferring algorithmic advice to human advice under identical conditions. Jussupow et al. [72] reconciled these contradictory findings by identifying moderators, including task type, perceived agency, and framing, that determine when users tip from aversion to appreciation. Earlier work by the same group [47] traced the underlying cognitive process: physicians evaluating AI-generated diagnoses rely on metacognitive monitoring, both self-monitoring of their own reasoning and system-monitoring of the AI, and failures in either metacognitive channel lead to incorrect decisions regardless of AI accuracy.
The resolution appears to be developmental. Horowitz et al. [70] documented an inverted-U relationship across nine countries, where low AI exposure produces aversion, moderate exposure produces over-reliance, and high exposure reduces bias toward more calibrated use. Motivational factors compound these cognitive dynamics: Bockstedt and Buckman [58] showed that loss framing eliminates algorithm aversion, increasing AI delegation by 44.9% through enhanced situational awareness, a rare practical intervention in a literature dominated by problem documentation. Goergen et al. [66] demonstrated that mere awareness of AI assessment changes human behavior independent of actual AI performance, adding a motivational layer to what had previously been treated as a purely cognitive calibration problem.
Confidence dynamics further complicate the picture. Li et al. [78] showed that human self-confidence aligns with AI confidence during collaboration, and this alignment persists even after the AI is removed. Because most participants were overconfident yet less confident than the AI, alignment actually worsened their calibration, pushing self-confidence further from actual accuracy. Only real-time accuracy feedback disrupted this anchoring effect. Healthcare settings reveal a counterweight: Küper et al. [48] found that AI assistance improved dermatological diagnostic accuracy by approximately one percentage point on average, but experienced clinicians displayed protective scepticism, with medical experience predicting lower trust and higher self-reliance [42,51]. Hou et al. [46] confirmed this pattern in a field experiment on a healthcare platform, finding that smarter AI increased physician adoption rates by 32.6%, but that transparency about AI capabilities modulated the effect, with more experienced physicians showing more discriminating adoption behavior. Domain expertise, in other words, acts as a natural calibration mechanism, one that takes time to develop. Kahr et al. [73] and Dang and Li [62] confirmed that trust develops through repeated interaction rather than forming at first exposure, with organizational context and cultural dynamics shaping the trajectory [85]. Xu et al. [86] further showed that in human–machine teams, AI recommendation accuracy and task difficulty jointly determine whether AI influence converges or fragments group decision-making, reinforcing the view that trust calibration is context-dependent rather than a stable trait.
The “AI” label itself turns out to be a trust amplifier. Klingbeil et al. [75] gave participants identical ChatGPT (GPT-4)-generated advice, varying only the source label: cooperation rates reached 81% when the advice was labeled “AI” but only 67% when labeled “expert,” contradicting the long-held assumption that humans inherently prefer human judgment [81]. These labeling effects interact with recommendation diversity and social context [74], which means that calibration interventions targeting only accuracy feedback will miss important drivers of reliance. Holstein et al. [69] added another layer: reliance behavior differs depending on the type of uncertainty. Humans respond differently to reducible (epistemic) versus irreducible (aleatoric) uncertainty in AI recommendations, a distinction that most calibration research has treated as irrelevant [82].
Who is most vulnerable to automation bias? Romeo and Conti’s [9] systematic review of 35 studies points to an uncomfortable answer: not novices, but users with moderate AI knowledge. Professional experience is the strongest protective factor, while the Dunning–Kruger effect leaves moderately knowledgeable users paradoxically most susceptible to over-reliance. Organizational adoption data from healthcare [39] and financial services [61] mirror this pattern, reinforcing the case that calibration interventions should be developmentally staged to match users’ evolving exposure and expertise.

4.2.2. The Transparency Paradox: When More Explanation Does Not Help

If transparency helps users calibrate reliance, then more explanation should produce better decisions. It does not. Bansal et al. [37] tested this across three tasks and found that while human–AI teams achieved complementary performance in every condition, explanations did not improve complementarity beyond simpler baselines such as displaying confidence scores; explanations made participants more compliant, agreeing with AI regardless of correctness, rather than more discerning. Schemmer et al. [7] replicated the asymmetry: explanations increased following of correct AI advice but did not help participants reject incorrect advice. Guo et al. [67] reanalyzed these data and quantified the gap, showing that “discrimination loss,” the inability to distinguish when to rely on AI, far exceeded “reliance loss.” In other words, explanations fail precisely where they are needed most: helping users identify AI errors.
Healthcare provides a particularly striking illustration. Jabbour et al. [36] conducted a randomized clinical vignette study and found that systematically biased AI reduced physician diagnostic accuracy by 11.3 percentage points; critically, providing AI explanations did not mitigate this harm. Lebovitz et al. [49] documented an even more fundamental form of the paradox: radiologists dealing with opaque AI-based diagnostic tools sometimes chose not to engage with AI outputs at all, preferring to preserve their independent reasoning. This “unengaged augmentation” strategy represents a rational professional response to opacity but one that forfeits the potential benefits of AI assistance entirely.
A sharper paradox appears when transparency functions as reassurance. Harbarth et al. [68] found that high transparency reduced complacency potential (self-reported intention) while increasing complacency behavior (greater following of incorrect recommendations). Starke et al. [84] observed related effects in algorithmic fairness: revealing decision processes can reduce perceived fairness when the process itself appears unjust. In clinical contexts, systematic reviews of explainable AI [41] similarly conclude that explanations rarely achieve intended calibration effects, in part due to workflow friction, liability concerns, and domain complexity. Overall, transparency appears necessary but insufficient; explanation must be designed to support discrimination (identifying when the AI is wrong), not only comprehension.

4.2.3. Complementarity: When Do Human–AI Teams Outperform Individuals?

Complementarity, the condition where a human–AI team outperforms both the human alone and the AI alone, is the field’s central aspiration, but the evidence shows it is harder to achieve than widely assumed. Hemmer et al. [3] formalized two sources: information asymmetry and capability asymmetry. In both cases, complementarity required participants to accurately assess their own capability relative to the AI’s, a metacognitive demand that most users fail to meet [80]. Fügener et al. [4] proposed an elegant workaround: inversion, where the AI delegates uncertain cases to humans rather than the reverse. Under inversion, teams achieved complementarity; under standard delegation, they did not.
Supply chain forecasting provides comparatively strong positive evidence when integration is structured. Nair et al. [53] showed integrated human–machine forecasts reduced error by 35% relative to judgmental forecasts and by 14% relative to machine learning alone; related findings emphasize the need for protocols and structured integration [52,55]. Field evidence also cautions that fit matters: tailored AI improved outcomes substantially, whereas untailored AI reduced performance [76], reinforcing that human–AI alignment can matter as much as raw model quality [35,63]. Dynamic delegation models [71] are consistent with the inversion insight: systems should determine delegation under uncertainty rather than relying on user judgment. Dai and Singh [44] formalized this design question in healthcare, modeling whether AI should serve as a gatekeeper (screening patients first) or as a second opinion (reviewing after specialist consultation); their analysis showed that the gatekeeper approach is preferable in low-risk settings, while the second-opinion role suits high-risk patients, and notably that intermediate-risk patients may be better served without AI at all, challenging the premise that AI is most useful where uncertainty is highest. At scale, organizational constraints become central, including workflow design and accountability arrangements in high-stakes environments [45,57,60,77]. Across these contexts, complementarity failures often trace back to the same core barriers: weak reliance calibration and reduced cognitive engagement.

4.2.4. Invisible Bias Transfer

Biases can also transfer from AI to human judgment without either party noticing. Glickman et al. [65] demonstrated that biased AI shifted human judgments and that this shift persisted after AI removal, creating a self-reinforcing loop: biased algorithm → human absorbs bias → humans retrain AI → amplified bias. What makes this finding especially troubling is that bias transfer was independent of perceived source (human vs. AI label) and did not reduce accuracy; it was invisible to standard performance metrics. Influence is not uniformly negative: Shin et al. [83] found that superhuman AI can increase the novelty of solutions considered, improving decisions when the AI expands rather than narrows the decision space. In operations contexts, responsible AI governance [56] and employee-centric outcome attention [54] appear to moderate whether AI collaboration enhances or erodes judgment quality. In healthcare, Ratwani et al. [50] argued that addressing AI algorithmic bias requires a shared responsibility model among healthcare facilities, AI developers, and regulators, highlighting that bias transfer in clinical contexts creates compounding risks that no single stakeholder can mitigate alone. Bias transfer links directly to trust calibration and reliance behavior, and its invisibility suggests elevated risk in Gen 3 settings where fluent outputs can bypass scepticism.
Across Gen 1 decision support, the dominant pattern is a developmental trajectory from under-reliance to calibrated reliance, moderated by domain expertise and AI exposure. Transparency and explanation fail to improve discrimination; complementarity is achievable but requires metacognitive accuracy and often inverted delegation. Motivational interventions (e.g., loss framing) represent rare practical levers in a literature dominated by problem documentation. Gen 1 decision support failures concentrate on workforce-capability (SDG 8) and responsible-production (SDG 12) dimensions. Miscalibrated reliance in healthcare and financial advisory contexts [9,39,61] translates directly into diagnostic and decision-quality externalities; the developmental character of trust calibration [70,73] implies that workforce development (not one-off training) is the primary lever for safe deployment. The transparency paradox documented across Gen 1 studies also implies that “explainability” alone is insufficient for responsible deployment (SDG 12), shifting responsibility toward interaction-design interventions such as cognitive forcing and inverted delegation.

4.3. Generation 2: AI as Autonomous Partner

Generation 2 systems act independently within defined scope, while humans monitor, supervise, or intervene. This shared-control paradigm introduces physical safety risk, real-time authority transfer, and transition-of-control challenges largely absent from Gen 1. We identified 27 papers across aviation autopilot, air traffic management, surgical robotics, autonomous vehicles, collaborative manufacturing robots, and warehouse automation.

4.3.1. The Automation Paradox: Better Routine Performance, Worse Failure Performance

Onnasch et al.’s [107] meta-analysis quantified what practitioners had long suspected: higher automation improves routine performance but monotonically degrades failure performance. They called this the “lumberjack effect”: the taller the tree, the harder it falls. The systems designed to be most helpful during normal operations become most dangerous during failures, and the effect replicates across autonomous vehicles [105,109], where standardized takeover metrics remain elusive despite decades of research [102]. Bainbridge [33] anticipated this over four decades ago: by automating the easy parts, designers leave operators with the “boring but critical” monitoring role, where vigilance demands remain high while cognitive engagement drops. Emerging cognitive neuroscience evidence on human–AI interaction [104] reinforces the behavioral observations: monitoring high-reliability systems taxes attentional resources in ways that standard performance metrics fail to capture.

4.3.2. The Skill Retention Challenge

Automation does not simply deskill; it reshapes how people think. Kim et al. [129] proposed a theoretical framework linking sustained GenAI use to cognitive offloading, emotional dependence, and skill atrophy, arguing that these mechanisms may be particularly acute for younger users whose cognitive and social skills are still developing. Treiman et al. [110] traced the mechanism: humans trained with AI assistance develop AI-aligned decision strategies that transfer poorly to novel situations. The damage extends beyond cognition to motivation: when collaborative AI is removed, users may experience reduced willingness to re-engage with manual performance [129]. Manufacturing and warehouse settings, where Pasparakis et al. [96] studied collaborative order-picking, Segura et al. [98] examined quality inspection, and Ferdman [90] analyzed structural deskilling (see also [124]), face particularly acute versions of this challenge. Endsley’s [101] naturalistic study of Tesla drivers crystallized the explanation: operators relying on automation develop inaccurate and incomplete mental models of system state, leaving them without adequate representations when automation fails and full situation awareness must suddenly be reconstructed.

4.3.3. Authority Transfer and the Mode Error

Mode errors, situations in which an operator believes automation is in one mode but it is actually in another, remain a critical failure mechanism in supervisory control. Sarter and Woods [34] documented mode errors as a major contributor to aviation incidents, arguing that the problem stems from the opacity of multi-mode systems under workload, not simply poor interface labeling. Related issues appear in autonomous vehicle handoffs [111] and in human–robot collaboration in manufacturing [92,94], where ergonomic evidence shows effects on cognitive workload and physical postural control [88]. Attitudes toward collaborative robots vary across contexts [97], and effective task allocation must account for cognitive, physical, and skill-maintenance demands [93,100].
Mode errors persist despite decades of attention. Rodak et al. [108] showed that human–machine interface design for authority transfer remains challenging, with haptic feedback producing the fastest reaction times and lowest mental workload, yet no single modality eliminates the problem. The underlying challenge is structural: discrete automation modes map imperfectly onto continuous human mental models. Multimodal feedback designs show promise [91,106], and adaptive authority allocation approaches offer more systematic solutions, but mode errors have not been eliminated.
Across Gen 2 autonomous systems, automation improves routine performance at the cost of failure-response capability (the lumberjack effect) and reshapes mental models rather than merely atrophying specific skills. Mode errors and authority-transfer failures persist despite decades of attention: discrete automation modes map imperfectly onto continuous human mental models, and no single feedback modality has eliminated them. The structural, organizational dimension of deskilling is most salient in manufacturing. Gen 2 autonomous systems raise acute workforce-retention concerns (SDG 8) because skill atrophy under supervisory control compromises failure-response capability over career timescales [90,96,101,107]. These dynamics also bear on Industry 5.0 resilience (SDG 9): organizations that optimize for routine performance without planning for failure-response preservation (for example, via DP2 unassisted intervals or periodic manual-mode exercises) accumulate latent safety debt that is difficult to reverse. Authority-transfer research [34,108,111] implies that multimodal handoff design is a prerequisite for safe human–AI integration in safety-critical sectors.

4.4. Generation 3: AI as Cognitive Collaborator

Generation 3 systems engage in reasoning, dialogue, and content generation through conversational interaction. Humans and AI jointly produce outputs through iterative co-creation. We identified 26 papers spanning clinical decision support with LLMs, creative writing and design collaboration, code generation [135], analytical reporting, and multi-agent LLM systems.

4.4.1. Epistemia: The Illusion of Knowledge from Surface Plausibility

Epistemia denotes a user’s attribution of genuine understanding or accurate knowledge to an LLM based on surface features of its output (fluency, coherence, syntactic correctness, confident tone, and plausible domain vocabulary) in the absence of corresponding verification of factual or inferential accuracy. Epistemia is a human-side response to specific linguistic cues produced by Gen 3 systems; as such, it is distinct from three neighboring concepts.
First, epistemia is not hallucination. Hallucination refers to AI-side output behavior: the model generates content unmoored from factual basis. Epistemia is the complementary human-side phenomenon: the user treats hallucinated or partially correct output as if it were knowledge, because the output’s surface properties cue knowledge-like judgment. A system that hallucinates creates the conditions for epistemia; epistemia is the user-side failure to catch it.
Second, epistemia is not general overtrust or automation bias. Overtrust is a broad calibration failure that can arise for any AI system regardless of output modality. Automation bias presupposes a reliability history against which users miscalibrate. Epistemia is narrower and mechanism-specific: it arises when outputs carry linguistic-fluency cues that are decoupled from the underlying reliability of the claims conveyed, and it can operate on each individual output independently of any accumulated reliability signal.
Third, epistemia is not a belief state about AI capabilities but a per-output inferential slip. A user may hold accurate meta-beliefs about LLM limitations in general and nonetheless succumb to epistemia on a specific fluent output, because the fluency cues operate locally on the judgment being made. This local, fluency-cued character is what distinguishes epistemia as a distinct human–AI collaboration challenge warranting the label.
Gen 1 errors are often detectable as statistically implausible. Gen 2 errors have physical consequences that make them immediately apparent. LLM errors, by contrast, can be internally coherent, contextually appropriate, and factually wrong all at the same time. Omar et al. [118] found that LLMs are “highly vulnerable to adversarial hallucination attacks” in clinical decision support, producing plausible-sounding diagnoses that are medically incorrect. The vulnerability is not confined to text: multi-modal LLM agents exhibit analogous failure modes in clinical calculation tasks [114], and behavioral experiments reveal that LLM agents can diverge systematically from human judgment patterns despite producing fluent outputs [141].
Epistemia extends beyond hallucination to systematic bias and mutual reinforcement. Mahajan et al. [117] found recognizable diagnostic biases in LLM outputs (e.g., anchoring and availability heuristic). Chen et al. [112] reported that errors can become correlated in human–LLM collaboration: both professionals and models exhibit suggestibility and can reinforce each other’s mistakes. Clinical studies confirm this pattern across different physician workflows [113,120]. Work on social reasoning suggests that conversational fluency should not be equated with genuine understanding: LLMs may simulate interpersonal trust dynamics without robust social prediction capability [137,140]. Emerging evidence suggests that repeated exposure to plausible-sounding errors may degrade calibration over time, slowly eroding the expertise needed to detect errors.

4.4.2. Cognitive Offloading and Skill Atrophy

Lee et al. [130] documented a phenomenon that many educators and managers have observed anecdotally: knowledge workers reported that GenAI use reduced the cognitive effort they invested in analytical tasks, with higher confidence in GenAI predicting less critical thinking. The final outputs were often acceptable; this is not a performance problem but a thinking problem. Brynjolfsson et al. [125] and Noy and Zhang [134] documented similar patterns in customer support and professional writing tasks, respectively, where LLM availability compressed the quality distribution (weaker workers improved, stronger workers plateaued) rather than raising all performance. The motivational consequences compound the cognitive ones. Kim et al. [129] argued that sustained GenAI use fosters emotional dependence alongside cognitive offloading, such that removal of AI support may reduce willingness to re-engage with manual analytical tasks. This motivational dimension resonates with broader workforce disruption patterns reported across employee well-being studies [153], innovation management [150], supply chain operations [123,146], and Industry 5.0 workforce research [144,145], breadth that suggests a fundamental human–AI interaction dynamic rather than a domain-specific artifact.

4.4.3. Co-Creation and Attribution Ambiguity

Who wrote this paragraph, the human or the LLM? In Gen 3, the question often has no clear answer. Rafner et al. [136] found that users reported fluctuating creative agency when working with LLMs, experiencing reduced ownership and control despite actively engaging in co-creation. The paradox stems from authorship ambiguity: users could not tell which creative choices were theirs and which were the model’s. Wang et al. [139] and McGuire et al. [133] observed the same dynamic in collaborative design, while Dhillon et al. [126] and Liu et al. [131] documented it in co-writing contexts. The pattern appears intrinsic to LLM co-creation. This ambiguity has direct accountability implications. In Gen 1, responsibility is typically assigned to the human because the AI advises. In Gen 2, responsibility is negotiated through predefined authority arrangements. In Gen 3, responsibility becomes shared and difficult to allocate because both human and LLM contribute essential generative functions, yet accountability for poor outcomes remains unresolved [116,128,138].
Across Gen 3 LLM collaboration, conversational fluency creates new failure modes (epistemia, correlated human–LLM errors) that can accumulate over time as error-detection expertise degrades. Cognitive offloading compresses quality distributions rather than raising all performance, with motivational dependence emerging as a distinct risk. Co-creation produces diffuse accountability that neither Gen 1 advisory nor Gen 2 supervisory frameworks resolve. Gen 3 LLM collaboration raises novel workforce-sustainability concerns (SDG 8) that differ in kind from earlier generations. Cognitive offloading and motivational withdrawal [129,130,153] suggest that workers may remain productively engaged with AI while cognitively disengaging, a failure mode invisible to conventional output-based productivity metrics. Quality-distribution compression [125,134] carries distributional equity implications: if Gen 3 raises the floor but lowers the ceiling, its net sustainability effect on workforce capability depends critically on organizational context and who benefits from compression. Responsible production (SDG 12) is also implicated through attribution ambiguity in co-creation [136,139]: without resolvable accountability, liability for LLM-mediated harms defaults to whichever party is institutionally nearest to the output, which may not correspond to who actually contributed the flawed reasoning.

5. Cross-Generational Synthesis: Persistence and Change in the CCF (RQ2)

5.1. Identifying Convergent Patterns

Section 4 documented trust failures in clinical decision support, vigilance decrements in aviation, and cognitive offloading in LLM-assisted work, findings that are often treated as domain-specific and generation-specific. Yet the underlying behavioral mechanism is consistent: as automation becomes more capable, users tend to reduce independent cognitive effort, with predictable consequences for error detection, skill maintenance, and accountability. Aviation researchers describe this as “automation complacency”; decision scientists discuss “algorithm appreciation”; and LLM studies report “over-reliance.” Berretta et al. [22] identify this terminological fragmentation as a jangle fallacy (different labels for functionally equivalent phenomena), and our cross-generational synthesis indicates that the same jangle fallacy extends across all six challenges reviewed in Section 4. Janhunen et al. [148], in a systematic review of trust in digital human–AI team collaboration, independently document terminological fragmentation of trust constructs, corroborating the jangle fallacy as a structural feature of human–AI collaboration research rather than an artifact of any single construct or domain.
To make this convergence explicit, we propose the Collaboration Convergence Framework (Table 2). The CCF maps six recurring human-side challenges across three AI technology generations (Gen 1 decision support, Gen 2 autonomous systems, Gen 3 LLM agents). For each challenge, the framework provides (i) a generation-agnostic label capturing the underlying functional risk, (ii) the dominant terminology and manifestations emphasized within each generation, and (iii) a solution maturity indicator (known, emerging, unsolved) summarizing the strength and consistency of evidence for effective design interventions.
Relationship to prior frameworks. The CCF complements rather than substitutes for existing theoretical frameworks in human factors and information systems. Parasuraman et al.’s [29] levels-of-automation taxonomy classifies systems by the stage and degree of function allocation between human and machine; the CCF classifies human challenges by their invariance across technology generations. Bainbridge’s [33] ironies of automation articulate the mechanism by which automation generates challenges for human operators; the CCF operates at a different level of abstraction, mapping which ironies recur across qualitatively different AI paradigms and which are genuinely novel to each. Lee and See’s [28] trust-in-automation framework specifies the psychological basis for a single construct (trust); the CCF organizes six such constructs by their cross-generational trajectory. Hancock et al.’s [35] meta-analytic synthesis of trust in human–robot interaction characterizes moderators within one AI paradigm; the CCF asks when similar moderators recur across paradigms under different names. The CCF is therefore best understood as a cross-generational organizing structure that runs alongside these mechanism-level theories: it accepts their substantive content and organizes it for transfer. Where prior frameworks answer “what is automation/trust/irony?”, the CCF answers “which of these phenomena recur across three AI generations, and with what solution maturity?”
The maturity indicators are assigned according to the following operational criteria:
  • Known: At least one design intervention has been empirically validated in two or more independent studies across at least two domains, with consistent direction of effect (e.g., cognitive forcing functions for over-reliance calibration).
  • Emerging: At least two studies document the phenomenon and propose or pilot interventions, but replication across domains or generations remains incomplete (e.g., structured prompting protocols for LLM-assisted reasoning).
  • Unsolved: The challenge is identified and documented, but no intervention has been empirically tested in a controlled study (e.g., attribution frameworks for human–LLM co-created outputs).
These thresholds reflect the cross-generational scope of this review; we acknowledge they represent a pragmatic classification rather than a formal meta-analytic standard, and we flag borderline cases in the discussion of individual challenges.
Figure 4 visualizes the CCF and highlights its central implication: novelty in interaction modality and autonomy does not eliminate the same six human challenges, but it does rename them in ways that impede knowledge transfer. Figure 5 complements this synthesis by summarizing the relative evidence volume by generation and challenge. Together, these views clarify two priorities: critical gaps where evidence remains thin across all three generations (notably Skill Retention and Accountability), and transfer opportunities where robust Gen 1/Gen 2 evidence can accelerate Gen 3 design.

5.2. Cross-Domain Patterns and Domain-Specific Variations (RQ3)

While the six challenges recur across domains, their salience, manifestation, and tractable interventions are shaped by domain context. Addressing RQ3 (cross-domain patterns) and the “what changes” component of RQ2, we highlight three contextual dimensions that systematically moderate collaboration risks: (i) whether AI errors are physically visible or remain latent, (ii) whether accountability is legally or institutionally codified, and (iii) whether the human operator is a trained professional or a general user.

5.3. Transferable Design Principles Across Generations and Domains (RQ4)

Building on the CCF and the cross-generational evidence base, we distill twelve transferable design principles (DP1–DP12) that have demonstrated effectiveness across multiple generations and domains. These principles are not new to the broader human factors and decision support literature, as many were articulated in earlier work on aviation automation and decision aids, but their relevance to Gen 3 LLM collaboration has not been systematically consolidated. Table 3 summarizes each principle, the challenge(s) it targets, the supporting evidence base, and a concrete LLM-oriented implementation.
Conflicts and trade-offs among design principles. The twelve principles are not fully independent, and in some deployment contexts they actively compete. Four conflicts merit explicit discussion. (i) DP2 (unassisted intervals) and DP9 (adaptive task allocation) can conflict: the former preserves skill by scheduling AI-free work periods, while the latter dynamically routes work to whoever is best placed at a given moment. Resolution is typically temporal: adaptive allocation within shifts, unassisted intervals across shifts or work-weeks [96,101]. (ii) DP3 (partial explanations to preserve engagement) and DP1 (calibration feedback on AI reliability) pull in opposite directions: one deliberately withholds information to preserve critical engagement, the other provides reliability signals. In practice these operate at different information layers: meta-level reliability (e.g., calibration curves, historical accuracy) under DP1 versus per-decision explanation content under DP3 [7,37,68]. (iii) DP4 (initial independent judgment) and DP7 (structured prompting protocols for LLM-assisted reasoning) can conflict because DP7 risks framing the problem in ways that anchor subsequent judgment, partially defeating DP4; resolution is typically sequential, with independent judgment recorded before structured AI engagement begins. (iv) DP11 (sequential teaming: human judges first, AI reviews, human decides) and DP4 (initial independent judgment) are complementary at a high level but conflict in their ordering assumptions in some healthcare and operations contexts, where inversion (AI screens first; human judges AI-flagged uncertain cases [3,4]) produces better complementarity than human-first sequencing, at the cost of DP4’s metacognitive-integrity benefit. These conflicts do not invalidate the principles but clarify that responsible deployment requires context-specific sequencing and trade-off decisions. This observation also motivates Research Direction 9 (adaptive intervention systems) in Section 6.4.3, where context-aware switching between competing principles is identified as a key open problem.

5.4. Sustainability Dimensions of the Framework

Each of the six convergent challenges has a direct sustainability analogue. Trust miscalibration undermines operational resilience: over- or under-trust can trigger cascading safety and economic losses that compound over system lifecycles. Reliance failures raise responsible production concerns, as miscalibrated reliance contributes to quality defects, waste, and rework in manufacturing [98] and to diagnostic errors in healthcare [48]. Cognitive disengagement threatens workforce well-being: across Gen 2 and Gen 3, the “boring but responsible” monitoring role can erode engagement and satisfaction [144]. Skill degradation is the most acute workforce sustainability risk, reducing organizational adaptability and individual career resilience [90,124]. Accountability gaps create governance sustainability risks through unclear responsibility allocation, while transparency failures undermine social trust, generating public resistance that can stall beneficial deployment [38]. These mappings position the CCF within Industry 5.0’s emphasis on human-centric, sustainable, and resilient technology deployment [24,25,143]; prescriptive implications are developed in Section 6.

5.4.1. Aviation

Aviation provides the deepest evidence base for cross-generational convergence, having grappled with human–automation collaboration for over four decades. Bainbridge’s [33] ironies of automation were formulated around flight deck automation; Sarter and Woods’ [34] mode error analysis drew on cockpit incident data; and Parasuraman et al.’s [29] levels-of-automation taxonomy was grounded in military and civil operations. This heritage makes aviation research the most theoretically mature of the reviewed domains, and therefore a rich source of transferable design knowledge.
Three challenges are particularly acute in aviation. Skill retention is dominant: pilots who rely on autopilot for routine operations develop inaccurate and incomplete mental models of aircraft state [101], while what Onnasch et al. termed the “lumberjack” analogy [107] delays the visibility of degradation until rare emergencies, precisely when stakes are highest. Xia et al.’s [12] systematic review of air traffic management confirmed that situation awareness and trust calibration remain among the most studied human factors in aviation automation research, but also noted that longitudinal studies of skill retention under realistic operational tempo are largely absent.
Aviation research is also advanced in formalizing authority allocation for human–AI teaming. Kirwan [14] mapped human factors requirements for aviation teaming across all six challenges and argued that the transition from Gen 2 (autopilot, flight management systems) to Gen 3 (AI-assisted flight planning, natural-language ATC decision support) will require updated authority allocation frameworks. Extending classic function allocation (HABA–MABA) to accommodate adaptive AI authority levels represents one promising direction for future authority allocation frameworks.
Finally, aviation offers the strongest existing accountability model via the Just Culture framework, which distinguishes honest errors, at-risk behavior, and reckless behavior in automation failures. No other reviewed domain has an equivalent mechanism for allocating responsibility when humans and AI share control. However, Just Culture was developed primarily for Gen 2 scenarios (pilot vs. autopilot), and extending it to Gen 3 settings, where an LLM-based planning tool may have contributed reasoning that shaped an operational decision, remains an open challenge.
What aviation can transfer to other domains is substantial. The lumberjack effect [107], mode error taxonomies [34], automation complacency findings [87], and the distinction between automation complacency and automation bias [26] translate directly to healthcare decision support, manufacturing cobots, and LLM-assisted analytical work. The CCF formalizes where such transfer is most plausible and where genuinely new evidence is required.

5.4.2. Healthcare

Healthcare is the domain where all six challenges are simultaneously salient and where collaboration failures are most directly measured in patient outcomes. AI in healthcare has advanced rapidly, with over 1200 FDA-cleared AI systems [142] and growing evidence that AI can improve clinical outcomes when well integrated [40]. Küper et al. [48] showed that AI assistance improved dermatological diagnostic accuracy by approximately one percentage point on average, but the gains were uneven: experienced clinicians exhibited protective scepticism, while less-experienced clinicians showed greater over-reliance [42,51]. Hou et al. [46] confirmed this expertise-mediated pattern in a field experiment: smarter AI increased physician adoption rates, but more experienced physicians showed more discriminating adoption behavior. Jussupow et al. [47] traced the underlying mechanism, showing that physicians rely on metacognitive monitoring—both self-monitoring of their own reasoning and system-monitoring of the AI—to evaluate AI advice, and that breakdowns in either metacognitive channel lead to incorrect decisions. This expertise-mediated trust trajectory parallels the developmental pattern documented in aviation, reinforcing the convergence claim.
Healthcare’s distinctive pressure point is epistemia. In aviation, automation failures are often physically apparent (the aircraft is not where it should be); in clinical contexts, an LLM-generated differential diagnosis can be internally coherent, confidently stated, and clinically wrong. Omar et al. [118] demonstrated that clinical LLMs are highly vulnerable to adversarial hallucination attacks, and Chen et al. [112] showed that LLM–clinician collaboration can produce correlated rather than independent errors. Jabbour et al. [36] provided direct clinical evidence that AI explanations fail to mitigate the harms of biased diagnostic AI (Section 4), reinforcing the transparency paradox in clinical contexts. Lebovitz et al. [49] documented a complementary phenomenon among radiologists, who sometimes chose not to engage with opaque AI diagnostic outputs at all in order to preserve reasoning independence. Romeo and Conti [9] further confirmed, across healthcare, national security, and other domains, that automation bias is moderated by professional experience, providing extensive evidence that Dunning–Kruger dynamics can amplify over-reliance in decision support contexts including clinical settings.
An operations management perspective highlights that successful healthcare AI deployment depends on workflow integration, not merely technical accuracy. Dai and Tayur [43] introduced a “4Ps” framework—physician buy-in, patient acceptance, provider investment, and payer support—arguing that AI-augmented healthcare delivery systems must be designed around stakeholder adoption dynamics rather than algorithmic performance alone. Dai and Singh [44] formalized a key design decision: whether AI should serve as a gatekeeper (screening patients before specialist consultation) or as a second opinion (reviewing after specialist evaluation). Their analysis showed that the gatekeeper approach is preferable in low-risk settings, while the second-opinion role suits high-risk patients, directly mapping onto the authority allocation and adaptive task allocation principles (DP5, DP9). Adams et al. [40] demonstrated what well-integrated Gen 1 healthcare AI can achieve: the TREWS machine learning-based early warning system for sepsis, deployed prospectively across multiple hospital sites, reduced mortality when alerts were confirmed promptly, providing a concrete success benchmark for the field. Dai and Abràmoff [15] synthesized the broader operations research perspective on incorporating AI into healthcare workflows, identifying models and insights that bridge clinical and operational considerations.
The most consequential gap in healthcare is accountability. Unlike aviation’s Just Culture model, healthcare lacks an established framework for allocating responsibility when a clinician acts on an LLM-generated recommendation that later proves incorrect [116]. Ratwani et al. [50] argued that addressing AI algorithmic bias in healthcare requires a shared responsibility model among healthcare facilities, AI developers, and regulators, highlighting that neither technical debiasing nor individual clinician vigilance alone is sufficient. Moreover, the combination of high hallucination vulnerability [118] and correlated errors [112] suggests that sustained LLM use may gradually erode the clinical expertise that currently serves as a primary safeguard against AI errors.

5.4.3. Manufacturing and Supply Chain

Manufacturing and supply chain settings add two dimensions less central in aviation and healthcare: physical human–robot proximity and diffuse organizational accountability. Collaborative robots (cobots) introduce collaboration risks that span both cognitive and physical ergonomics. Bibbo et al. [88] showed that cobot collaboration affects postural control; Menanno et al. [95] developed an ergonomic risk assessment system combining 3D human pose estimation with cobot motion planning, representing one concrete instantiation of the multi-modal feedback design principle (DP6) transferred from aviation into manufacturing contexts. Liu et al. [94] analyzed human error patterns in human–robot collaboration accidents, finding authority-transfer and mode-confusion failures functionally equivalent to the Gen 2 mode errors documented in aviation by Sarter and Woods [34], a direct cross-generational parallel. Attitudes toward cobots also vary across manufacturing, warehouse logistics, and agriculture [97], complicating the direct transplantation of standardized interventions. The growing deployment of Internet of Things (IoT) technology in intralogistics further reshapes human–robot task allocation: Shi et al. [99] developed online optimization algorithms for IoT-enabled human–robot hybrid sortation in parcel centers, while de Koster et al. [89] survey emerging IoT applications across warehousing, manufacturing, and port operations, highlighting how real-time data streams create new demands on human supervisory roles.
In this domain, cognitive engagement and skill retention are especially salient. Pasparakis et al. [96] documented impacts of collaborative order-picking systems on warehouse workers, finding that automation-assisted workers developed narrower task strategies. Ferdman [90] argued that AI deskilling is structural rather than individual, implying that remedies must be organizational as well as interface-level. At the same time, supply chain forecasting provides strong evidence for complementarity: Nair et al. [53] showed that integrated human–machine forecasts can outperform either alone, and Brau et al. [52] demonstrated that structured protocols mediate integration quality. Responsible AI governance [56] and employee-centric outcomes [54] appear to moderate whether AI collaboration enhances or erodes judgment quality in operations.
Looking forward, Boone et al. [122] identified generative AI as a transformative force for supply chain resilience, arguing that GenAI capabilities in demand planning, procurement, and logistics represent a qualitative shift from traditional AI—one that introduces the same hallucination risks and cognitive offloading concerns documented in Gen 3 research to operational supply chain contexts.
The primary evidence gap is longitudinal. No published study has tracked manufacturing worker skill and motivation under sustained cobot collaboration for periods exceeding six months [93,98,100], leaving open foundational questions about long-run learning trajectories and the dose–response of mitigation strategies.

5.4.4. Cross-Domain Patterns

Cross-domain comparison clarifies which collaboration dynamics are invariant and which are contingent on institutional and task context, directly addressing RQ3.
Domain-invariant patterns. Trust calibration exhibits a common developmental trajectory across domains: early aversion or appreciation, moderated by experience, gradually resolving toward calibrated reliance with repeated interaction [70,73]. Cognitive offloading appears across decision support, autopilot, and LLM writing assistance [87,101,130]. The transparency paradox, in which explanations increase compliance without improving discrimination, also replicates across clinical [41], autonomous driving [108], and general decision-making [7,37] contexts.
Domain-contingent patterns. Error visibility varies systematically: aviation errors are often physically manifest, healthcare errors may surface only after delay, and LLM analytical errors may never be detected. This visibility gradient shapes how quickly recalibration occurs. Accountability maturity also differs: aviation has Just Culture, healthcare has emerging malpractice frameworks, and manufacturing and general-user LLM contexts lack established models. Finally, professional expertise serves as a protective factor in aviation [14] and healthcare [48] but is less available in general-user LLM contexts. These contingent features help explain why the same underlying challenge (e.g., skill retention) often requires different implementations of the same design principle: the CCF identifies the invariant challenge, while domain context determines the most feasible and credible intervention.

6. Discussion

This section interprets the cross-generational convergence findings, develops theoretical and practical implications, articulates sustainability relevance, proposes a research agenda to address critical gaps, and acknowledges limitations.

6.1. Theoretical Implications

Across Gen 1 decision support, Gen 2 autonomous systems, and Gen 3 LLM agents, the same six human challenges recur, namely trust calibration, reliance behavior, cognitive engagement, skill retention, accountability, and transparency, often under different labels that obscure their functional equivalence. This convergence supports Bainbridge’s [33] foundational argument that the ironies of automation reflect stable properties of human cognition rather than idiosyncrasies of specific technologies. The evidence reviewed here extends Bainbridge’s insight to two generations she could not have anticipated: (i) autonomous systems that act without human initiation and (ii) LLM agents that produce natural language outputs that can appear indistinguishable from human reasoning. That these challenges persist across such different interaction modalities suggests they are rooted in cognitive architecture, specifically how humans allocate attention, calibrate confidence, maintain expertise, and assign responsibility, rather than in the surface features of any particular AI system.
At the same time, the Collaboration Convergence Framework makes clear what is genuinely novel in each generation. Gen 2 elevated physical safety risk and dynamic authority transfer beyond what is typical for advisory systems. Gen 3 introduces epistemia (the illusion of knowledge arising from surface plausibility), co-creation dynamics, motivational withdrawal, and scalable error propagation—features that do not map cleanly onto earlier decision support or automation paradigms [132,151]. Recognizing both convergence and novelty is essential: a reductive “it’s all the same” stance overlooks epistemia and co-creation, while an ahistorical “this is unprecedented” stance discards decades of transferable evidence from prior generations.
The jangle fallacy documented throughout this review has a direct methodological consequence. When functionally equivalent phenomena are labeled differently across communities, meta-analysis becomes infeasible, replication is obscured, and design knowledge fails to transfer. Skill retention provides a clear example: aviation emphasizes “deskilling,” decision support research emphasizes “expertise decay,” and LLM research emphasizes “cognitive atrophy.” All three describe the same underlying process, accelerated skill atrophy and metacognitive decline after sustained AI collaboration [129], yet they draw on different theoretical traditions and propose interventions in isolation. The CCF addresses this fragmentation by organizing theory around invariant human challenges rather than technology categories. This reframing shifts the theoretical center of gravity from “what can AI do?” to “what do humans consistently struggle with when working alongside intelligent systems?”, a question whose answer has proven stable across four decades of technological change, and one that positions the field to anticipate (rather than merely react to) the collaboration challenges of future AI generations.

6.2. Practical Implications for Organizations

The convergence evidence yields four concrete actions organizations should take when deploying AI across any of the three generations:
(a)
Conduct a collaboration audit before deployment. Map the intended human–AI workflow onto the six CCF challenge dimensions to identify the most salient risks and the corresponding design principles (Table 3) that should be prioritized.
(b)
Design cross-generational training. Training programs should draw on evidence from all applicable generations. For example, a healthcare organization deploying an LLM-based clinical decision support tool should combine Gen 3 evidence on hallucination detection with Gen 1 evidence on cognitive forcing interventions [59] and Gen 2 evidence on authority-transfer protocols [34].
(c)
Monitor reliance behavior directly, not only self-reported trust. Monitoring frameworks should include behavioral measures of reliance (e.g., override rates, confidence calibration, time-to-verification) alongside self-reported trust measures, given the trust–reliance dissociation documented across all three generations.
(d)
Implement unassisted intervals as a baseline condition. Plan for long-run capability by embedding “unassisted intervals” (DP2) into AI-augmented workflows to mitigate the skill atrophy observed across all three generations, even when this introduces short-term efficiency costs.
Manufacturing and supply chain organizations face particularly acute challenges as they transition from Industry 4.0 toward Industry 5.0. Rather than treating each new AI deployment as a novel and disconnected training requirement, the CCF provides a structured approach: identify which of the six recurring challenges are most salient for the workforce and implement evidence-backed interventions from the transferable design principles.

6.3. Sustainability Implications

Section 5.3 mapped each convergent collaboration challenge to a sustainability dimension. Here, we translate that mapping into prescriptive implications, specifically what organizations and policymakers should do differently given the convergence evidence. The central point is that human–AI collaboration risks are not isolated technical failures; they shape workforce capability, institutional learning, and long-run system resilience, and therefore directly affect sustainability outcomes.
Social sustainability (SDG 8: Decent Work). The convergence of skill atrophy across all three generations [59,101,129], together with the risk of cognitive offloading and reduced engagement when AI support becomes routine [129], motivates proactive workforce safeguards. Three design principles offer immediately actionable countermeasures. DP2 (unassisted intervals) preserves capability by scheduling AI-free work periods; DP4 (initial independent judgment) sustains cognitive engagement by requiring human reasoning before AI exposure; and DP7 (structured prompting) reduces passive offloading by embedding reasoning steps into interaction protocols. These should be treated as baseline deployment conditions rather than optional “best practices” [147,149,154]. Otherwise, systems optimized for short-run throughput risk producing work that is simultaneously less engaging and more brittle, with elevated failure consequences in rare, high-stakes events [144].
Economic sustainability (SDG 9: Industry, Innovation and Infrastructure). The terminological fragmentation identified in this review represents a recurring and costly inefficiency in the human–AI collaboration literature. LLM research is actively rediscovering phenomena that are well established in earlier automation and decision support traditions, including cognitive offloading, verification failure, and skill atrophy, without systematically leveraging the corresponding design knowledge. Funding agencies and reviewers can reduce this reinvention cost by incentivizing cross-generational synthesis as a prerequisite for new human–AI collaboration programs, and by encouraging authors to position contributions against the six invariant challenges rather than generation-specific terminology. The twelve transferable design principles (Table 3) provide a concrete mechanism for knowledge reuse, while the CCF solution-maturity indicators clarify where transfer is most defensible (“Known”) and where new investment is required (“Unsolved”).
Operational sustainability (SDG 12: Responsible Consumption and Production). Organizations that deploy new AI technologies without drawing on evidence from prior generations risk repeating preventable failures. The practical prescription is therefore simple: before deployment, map the intended human–AI workflow to the six CCF challenge dimensions, identify the most salient risks for the specific operational context, and implement the corresponding design principles. The CCF solution-maturity indicators provide triage guidance: adopt validated interventions where evidence is strong (“Known”), pilot and evaluate promising approaches where evidence is incomplete (“Emerging”), and invest in governance and research partnerships where the field lacks robust answers (“Unsolved”) [143]. In this sense, responsible deployment is not only a technical problem but also a sustainability practice: minimizing rework, defects, and downstream harm by treating collaboration risks as predictable and manageable rather than novel surprises.

6.4. Research Agenda

This review has identified substantial gaps between current knowledge and practitioner needs [21,23]. We propose ten specific research directions organized by type.

6.4.1. Empirical Directions

RD 1: Skill atrophy dose–response. Randomized controlled trial (N ≥ 120) testing whether “unassisted intervals” (DP2) preserve skills during sustained (8+ week) LLM collaboration, varying interval schedules. No study has tested this directly despite evidence that sustained GenAI use can produce measurable cognitive offloading [129].
RD 2: Negative feedback for hallucination detection. Controlled experiment (N ≥ 200 healthcare professionals) testing whether exposure to LLM errors before correct outputs improves hallucination detection, building on evidence that negative examples can reduce automation bias [9,130] and that cognitive forcing interventions can recalibrate reliance [59].
RD 3: Cross-generational transfer. Randomized trial (N ≥ 160) applying aviation skill-maintenance protocols [101] to LLM-assisted analytical work, comparing standard training, aviation-style manual practice intervals, and aviation-style practice plus cognitive forcing functions over 12 weeks.
RD 4: Intervention habituation. Longitudinal study (N ≥ 200; 20 weeks) testing whether trust-calibration feedback mechanisms [70] lose effectiveness as users habituate to repeated accuracy information.
RD 5: Accountability models for co-creation. Legal and organizational analysis combined with experimental vignettes (N ≥ 100 professionals) to develop accountability allocation rules for outputs that are jointly produced by humans and LLM systems.

6.4.2. Methodological Directions

RD 6: Standardized reliance calibration measures. Development and validation of a “reliance calibration index” measuring discrimination and reliance accuracy independently, with comparability across all three generations (N ≥ 500).
RD 7: Skill atrophy meta-analysis. Systematic meta-analysis of skill-retention trajectories, effect sizes, and moderators across automation contexts, extending the work of Onnasch et al. [107].
RD 8: Organizational accountability measurement. Mixed-methods study combining qualitative case studies (N ≥ 15 organizations) with survey validation (N ≥ 200 organizations) to develop accountability measurement frameworks for LLM-assisted decision contexts.

6.4.3. Design Directions

RD 9: Adaptive intervention systems. Reinforcement-learning-based systems that modulate explanation depth, uncertainty display, and prompt structure based on observed reliance patterns, evaluated against static baselines.
RD 10: Minimum viable unassisted interval. Dose–response experiment varying unassisted-interval frequency and duration during LLM-assisted work, establishing evidence-based guidelines for the most underutilized transfer principle (DP2).

6.5. Limitations

This review has several limitations. First, the English-language restriction may exclude relevant studies from non-English-speaking research communities, particularly in manufacturing automation (e.g., German and Japanese traditions) and healthcare AI (e.g., Chinese clinical studies). Second, the 2018–2026 timeframe, while capturing the post-deep-learning and post-LLM acceleration, may miss relevant pre-2018 studies beyond the seminal works included. Third, the three-generation taxonomy is necessarily simplified; real-world systems increasingly span generations, and the boundaries between Gen 2 and Gen 3 are blurring as LLMs are embedded in autonomous systems [115,119]. Fourth, the narrative synthesis approach, while appropriate for the heterogeneity of study designs and outcome measures, does not provide the quantitative precision of a formal meta-analysis; future meta-analyses on skill atrophy, trust calibration, and reliance behavior are warranted [152]. Finally, publication bias may over-represent positive collaboration findings, particularly for Gen 3 where the research community is heavily invested in demonstrating LLM benefits (e.g., the complementarity evidence reported by Zöller et al. [121] is notable precisely because it is one of few well-powered demonstrations).
The review’s screening procedure also has a methodological limitation that warrants explicit acknowledgment. Primary title/abstract and full-text screening was conducted by the lead author rather than by dual independent reviewers; the verification screening reported in Section 3.4 (second author re-screened 40 papers from the included set; 38/40 confirmed; 40/40 post-resolution) serves as a post hoc consistency check of inclusion decisions against the criteria in Section 3.3 rather than as classical dual-reviewer inter-rater reliability. Because the verification sample was drawn only from the included-paper set, Cohen’s κ is undefined (degenerate reference marginal), and the 95%/100% agreement statistics reported in Section 3.4 should therefore be interpreted as a one-sided consistency indicator, not as a chance-corrected agreement estimate. A fully dual-reviewed design with independent title/abstract screening of the full search pool would have provided stronger reliability evidence at the cost of substantially slower synthesis across the cross-disciplinary corpus; future cross-generational reviews should plan dual-reviewer screening from the outset. Separately, the supplementary-search strategy described in Section 3.2, while documented transparently, carries a residual confirmation-bias risk that the Section 3.6 mitigations (deliberate inclusion of null and negative findings; high-impact-venue sampling) reduce but cannot eliminate.
Despite these limitations, the evidence base is sufficiently consistent to support a clear cross-generational conclusion: a small set of human-side collaboration challenges recurs reliably as AI capabilities evolve, while terminology and research silos impede cumulative learning. The following section summarizes the resulting contributions and the practical significance of addressing this fragmentation.

7. Conclusions

This paper presented the first systematic review spanning all three AI technology generations, namely decision support systems, autonomous systems, and LLM-based agents, to identify recurring patterns in human–AI collaboration challenges. Synthesizing 152 papers across four domains (aviation, healthcare, manufacturing/supply chain, and cross-domain contexts), we find a stable cross-generational result: the same six human-side challenges recur as AI capabilities evolve, namely trust calibration, reliance behavior, cognitive engagement, skill retention, accountability, and transparency, even as their terminology and surface manifestations change.
The review yields three primary contributions. First, the Collaboration Convergence Framework formalizes this result in a 6 × 3 matrix, making visible how functionally equivalent challenges are repeatedly renamed across generations and thereby obscured. Second, we distill twelve transferable design principles from cross-generational evidence, highlighting immediate transfer opportunities for Gen 3 LLM collaboration design, including unassisted intervals (DP2, from Gen 2 aviation), initial independent judgment (DP4, from Gen 1 healthcare), partial rather than full explanations (DP3), and structured prompting protocols (DP7). Third, we propose ten research directions targeting the most consequential gaps, with particular urgency around longitudinal skill retention interventions, intervention habituation, and accountability framework development.
As AI systems assume greater autonomy and diffuse across more domains, the cost of relearning collaboration lessons in isolation increases. Aviation spent decades establishing that automation complacency requires active countermeasures; healthcare continues to show that explanations do not automatically produce appropriate reliance; and LLM deployments are revealing how conversational fluency can invite dangerous cognitive offloading. The CCF is intended to bridge these silos by providing a shared language and cumulative evidence base, enabling designers, organizations, and researchers to build on, rather than rediscover, decades of knowledge about how humans and intelligent systems can work together effectively, safely, and sustainably. In this sense, fragmentation in human–AI collaboration research represents not only a scholarly challenge but also a sustainability risk, as industries may repeatedly encounter the same behavioral failures without benefiting from knowledge generated elsewhere.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/su18115313/s1. PRISMA 2020 Checklist.

Author Contributions

Conceptualization, A.D. and P.L.; methodology, A.D. and P.L.; investigation, A.D.; data curation, A.D. and Y.C.; writing—original draft preparation, A.D.; writing—review and editing, P.L., L.Z., S.G. and M.H.; visualization, A.D.; formal analysis, A.D. and M.H.; supervision, P.L.; project administration, P.L. L.Z. contributed domain expertise on sustainability dimensions. S.G. provided insights on aviation human factors. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article as it is a systematic literature review.

Conflicts of Interest

Author Meiling He was employed by the company Rockwell Automation. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
AMRAutonomous Mobile Robot
ATCAir Traffic Control
CCFCollaboration Convergence Framework
CDSSClinical Decision Support System
DPDesign Principle
LLMLarge Language Model
PRISMAPreferred Reporting Items for Systematic Reviews and Meta-Analyses
RDResearch Direction
RQResearch Question
SASituation Awareness
SDGSustainable Development Goal
XAIExplainable Artificial Intelligence

Appendix A. Prior Reviews: Full Comparison Table

This appendix provides the detailed comparison of 17 prior reviews and synthesis papers directly relevant to human–AI collaboration, moved here from Section 2.2 to support readability of the main-text argument. Section 2.2 summarizes the three structural shortcomings these reviews share; the table below supplies the per-review detail.
Table A1. Comparison of Existing Reviews of Human–AI Collaboration (moved from Section 2.2 per reviewer feedback).
Table A1. Comparison of Existing Reviews of Human–AI Collaboration (moved from Section 2.2 per reviewer feedback).
Author(s)/YearScope/FocusAI GenDomainsHuman FactorsKey Gap
Kohn et al. [19]Trust measurement reviewGen 1–2CrossTrust measureNo behavioral/Gen 3
Bao et al. [21]Affordance theory reviewGen 1CrossTrust, relianceNo Gen 2–3
Berretta et al. [22]HAIT definitions scopingGen 1–2CrossTrust, explainabilityTerminological
Walker et al. [11]Trust in automated vehiclesGen 2AutoTrust calibrationSingle domain
Samuels [17]AI in supply chainGen 1–2SupplyWorkforce skillsWeak methods
Gomez et al. [8]Interaction patternsGen 1CrossRelianceNo Gen 2–3
Schmutz et al. [13]AI-teaming reviewGen 2CrossTrust, coordinationNo Gen 1 or 3
Romeo & Conti [9]Automation bias & XAIGen 1CrossRelianceNo Gen 2–3
Do Khac & Leyer [23]HAI trust modelGen 1–2CrossTrustPreprint; no Gen 3
Lai et al. [6]AI-assisted decisionsGen 1CrossReliance, explanationsNo Gen 2–3
Schemmer et al. [7]XAI and relianceGen 1CrossRelianceNo Gen 2–3
O’Neill et al. [10]Human–autonomy teamingGen 2CrossTrust, SANo Gen 1 or 3
Xia et al. [12]ATM automationGen 2AviationSA, trustSingle domain
Kirwan [14]AI in aviationGen 2–3AviationAll sixSingle domain
Kumar et al. [18]Human–AI supply chainGen 1–2SupplyTrustNo Gen 3
Wong et al. [20]30-year trust reviewGen 1–2CrossTrust evolutionLimited Gen 3
Dai & Abràmoff [15]AI in healthcare workflowsGen 1–2HealthWorkflow integrationNo Gen 3
This studyCross-generational synthesisGen 1–3CrossAll six CCF challenges
Notes: Cross = Cross-domain; Supply = Supply Chain; Auto = Automotive/Transport; Mfg = Manufacturing; SA = Situation Awareness; XAI = Explainable AI. Gen 1 = Decision Support; Gen 2 = Autonomous Systems; Gen 3 = LLM Agents.

References

  1. Bureau d’Enquêtes et d’Analyses pour la Sécurité de L’aviation Civile (BEA). Final Report on the Accident on 1st June 2009 to the Airbus A330-203 Registered F-GZCP Operated by Air France Flight AF 447 Rio de Janeiro–Paris; BEA: Le Bourget, France, 2012.
  2. Vaccaro, M.; Almaatouq, A.; Malone, T. When combinations of humans and AI are useful: A systematic review and meta-analysis. Nat. Hum. Behav. 2024, 8, 2293–2303. [Google Scholar] [CrossRef]
  3. Hemmer, P.; Schemmer, M.; Kühl, N.; Vössing, M.; Satzger, G. Complementarity in Human-AI Collaboration: Concept, Sources, and Evidence. Eur. J. Inf. Syst. 2025, 34, 979–1002. [Google Scholar] [CrossRef]
  4. Fügener, A.; Grahl, J.; Gupta, A.; Ketter, W. Cognitive Challenges in Human–Artificial Intelligence Collaboration: Investigating the Path Toward Productive Delegation. Inf. Syst. Res. 2022, 33, 678–696. [Google Scholar] [CrossRef]
  5. Dell’Acqua, F.; McFowland, E., III; Mollick, E.; Lifshitz, H.; Kellogg, K.C.; Rajendran, S.; Krayer, L.; Candelon, F.; Lakhani, K.R. Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Organ. Sci. 2026, 37, 403–423. [Google Scholar] [CrossRef]
  6. Lai, V.; Chen, C.; Liao, Q.V.; Smith-Renner, A.; Tan, C. Towards a Science of Human-AI Decision Making: An Overview of Design Space in Empirical Human-Subject Studies. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, Chicago, IL, USA, 12–15 June 2023; Association for Computing Machinery: New York, NY, USA, 2023. [Google Scholar]
  7. Schemmer, M.; Kühl, N.; Benz, C.; Bartos, A.; Satzger, G. Appropriate Reliance on AI Advice: Conceptualization and the Effect of Explanations. In Proceedings of the 28th International Conference on Intelligent User Interfaces, Sydney, Australia, 27–31 March 2023; Association for Computing Machinery: New York, NY, USA, 2023; pp. 410–421. [Google Scholar]
  8. Gomez, C.; Cho, S.M.; Ke, S.; Huang, C.-M.; Unberath, M. Human-AI collaboration is not very collaborative yet: A taxonomy of interaction patterns in AI-assisted decision making from a systematic review. Front. Comput. Sci. 2025, 6, 1521066. [Google Scholar] [CrossRef]
  9. Romeo, G.; Conti, D. Exploring automation bias in human–AI collaboration: A review and implications for explainable AI. AI Soc. 2026, 41, 259–278. [Google Scholar] [CrossRef]
  10. O’Neill, T.; McNeese, N.; Barron, A.; Schelble, B. Human–autonomy teaming: A review and analysis of the empirical literature. Hum. Factors 2022, 64, 904–938. [Google Scholar] [CrossRef] [PubMed]
  11. Walker, F.; Forster, Y.; Hergeth, S.; Kraus, J.; Payre, W.; Wintersberger, P.; Martens, M. Trust in automated vehicles: Constructs, psychological processes, and assessment. Front. Psychol. 2023, 14, 1279271. [Google Scholar] [CrossRef]
  12. Xia, Z.; Chen, C.-H.; Hsieh, M.-H.; Ma, G.; Liu, B.; Yu, X.; Xin, S.; Dong, L. A systematic review on human-AI hybrid systems and human factors in air traffic management. J. Eng. Des. 2025, 36, 871–919. [Google Scholar] [CrossRef]
  13. Schmutz, J.B.; Outland, N.; Kerstan, S.; Georganta, E.; Ulfert, A.-S. AI-teaming: Redefining collaboration in the digital era. Curr. Opin. Psychol. 2024, 58, 101837. [Google Scholar] [CrossRef]
  14. Kirwan, B. Human Factors Requirements for Human-AI Teaming in Aviation. Future Transp. 2025, 5, 42. [Google Scholar] [CrossRef]
  15. Dai, T.; Abràmoff, M.D. Incorporating Artificial Intelligence into Healthcare Workflows: Models and Insights. In Tutorials in Operations Research: Advancing the Frontiers of OR/MS; INFORMS: Catonsville, MD, USA, 2023; pp. 133–155. [Google Scholar]
  16. Wan, Q.; Shi, Y.; Xiao, X.; Li, X.; Mo, H. Human–robot collaboration in robotic surgery: A review. Adv. Intell. Syst. 2025, 7, 2400319. [Google Scholar] [CrossRef]
  17. Samuels, A. Examining the integration of artificial intelligence in supply chain management from Industry 4.0 to 6.0: A systematic literature review. Front. Artif. Intell. 2025, 7, 1477044. [Google Scholar] [CrossRef]
  18. Kumar, N.; Kumar, R.R. Human–AI collaboration in operations and supply chain management: A systematic literature review. Manag. Rev. Q. 2025. online first. [Google Scholar] [CrossRef]
  19. Kohn, S.C.; de Visser, E.J.; Wiese, E.; Lee, Y.-C.; Shaw, T.H. Measurement of Trust in Automation: A Narrative Review and Reference Guide. Front. Psychol. 2021, 12, 604977. [Google Scholar] [CrossRef] [PubMed]
  20. Wong, K.K.L.; Han, Y.; Cai, Y.; Ouyang, W.; Du, H.; Liu, C. From Trust in Automation to Trust in AI in Healthcare: A 30-Year Longitudinal Review and an Interdisciplinary Framework. Bioengineering 2025, 12, 1070. [Google Scholar] [CrossRef]
  21. Bao, Y.; Gong, W.; Yang, K. A Literature Review of Human–AI Synergy in Decision Making: From the Perspective of Affordance Actualization Theory. Systems 2023, 11, 442. [Google Scholar] [CrossRef]
  22. Berretta, S.; Tausch, A.; Ontrup, G.; Gilles, B.; Peifer, C.; Kluge, A. Defining human-AI teaming the human-centered way: A scoping review and network analysis. Front. Artif. Intell. 2023, 6, 1250725. [Google Scholar] [CrossRef]
  23. Do Khac, L.T.; Leyer, M. Towards an integrative model of organizational human-AI collaboration: A semi-systematic review of the current state of the art. Technol. Soc. 2026, 84, 103064. [Google Scholar] [CrossRef]
  24. van Erp, T.; Carvalho, N.G.P.; Gerolamo, M.C.; Gonçalves, R.; Rytter, N.G.M.; Gladysz, B. Industry 5.0: A new strategy framework for sustainability management and beyond. J. Clean. Prod. 2024, 461, 142271. [Google Scholar] [CrossRef]
  25. Tóth, A.; Nagy, L.; Kennedy, R.; Bohuš, B.; Abonyi, J.; Ruppert, T. The human-centric Industry 5.0 collaboration architecture. MethodsX 2023, 11, 102260. [Google Scholar] [CrossRef]
  26. Parasuraman, R.; Riley, V. Humans and automation: Use, misuse, disuse, abuse. Hum. Factors 1997, 39, 230–253. [Google Scholar] [CrossRef]
  27. Beer, J.M.; Fisk, A.D.; Rogers, W.A. Toward a framework for levels of robot autonomy in human–robot interaction. J. Hum.-Robot Interact. 2014, 3, 74–99. [Google Scholar] [CrossRef]
  28. Lee, J.D.; See, K.A. Trust in automation: Designing for appropriate reliance. Hum. Factors 2004, 46, 50–80. [Google Scholar] [CrossRef]
  29. Parasuraman, R.; Sheridan, T.B.; Wickens, C.D. A Model for Types and Levels of Human Interaction with Automation. IEEE Trans. Syst. Man Cybern.-Part A Syst. Hum. 2000, 30, 286–297. [Google Scholar] [CrossRef]
  30. Luo, J.; Zhang, W.; Yuan, Y.; Zhao, Y.; Yang, J.; Gu, Y.; Wu, B.; Chen, B.; Qiao, Z.; Long, Q.; et al. Large Language Model Agent: A Survey on Methodology, Applications and Challenges. arXiv 2025, arXiv:2503.21460. [Google Scholar] [CrossRef]
  31. Ansari, B.; Sameer, A. A TCCM-Based Systematic Literature Review of Anthropomorphism in Human-AI Interaction and Future Research Agenda. J. Technol. Behav. Sci. 2026. online first. [Google Scholar] [CrossRef]
  32. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [PubMed]
  33. Bainbridge, L. Ironies of Automation. Automatica 1983, 19, 775–779. [Google Scholar] [CrossRef]
  34. Sarter, N.B.; Woods, D.D. How in the world did we ever get into that mode? Mode error and awareness in supervisory control. Hum. Factors 1995, 37, 5–19. [Google Scholar] [CrossRef]
  35. Hancock, P.A.; Billings, D.R.; Schaefer, K.E.; Chen, J.Y.C.; de Visser, E.J.; Parasuraman, R. A meta-analysis of factors affecting trust in human–robot interaction. Hum. Factors 2011, 53, 517–527. [Google Scholar] [CrossRef]
  36. Jabbour, S.; Fouhey, D.; Shepard, S.; Valley, T.S.; Kazerooni, E.A.; Banovic, N.; Wiens, J.; Sjoding, M.W. Measuring the Impact of AI in the Diagnosis of Hospitalized Patients: A Randomized Clinical Vignette Survey Study. JAMA 2023, 330, 2275–2284. [Google Scholar] [CrossRef]
  37. Bansal, G.; Wu, T.; Zhou, J.; Fok, R.; Nushi, B.; Kamar, E.; Ribeiro, M.T.; Weld, D.S. Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, Yokohama, Japan, 8–13 May 2021; Association for Computing Machinery: New York, NY, USA, 2021. [Google Scholar]
  38. Almusharraf, A.I. Automation and Its Influence on Sustainable Development: Economic, Social, and Environmental Dimensions. Sustainability 2025, 17, 1754. [Google Scholar] [CrossRef]
  39. Abbas, Q.; Jeong, W.; Lee, S.W. Explainable AI in Clinical Decision Support Systems: A Meta-Analysis of Methods, Applications, and Usability Challenges. Healthcare 2025, 13, 2154. [Google Scholar] [CrossRef] [PubMed]
  40. Adams, R.; Henry, K.E.; Sridharan, A.; Soleimani, H.; Zhan, A.; Rawat, N.; Johnson, L.; Hager, D.N.; Cosgrove, S.E.; Markowski, A.; et al. Prospective, Multi-Site Study of Patient Outcomes After Implementation of the TREWS Machine Learning-Based Early Warning System for Sepsis. Nat. Med. 2022, 28, 1455–1460. [Google Scholar] [CrossRef]
  41. Antoniadi, A.M.; Du, Y.; Guendouz, Y.; Wei, L.; Mazo, C.; Becker, B.A.; Mooney, C. Current Challenges and Future Opportunities for XAI in Machine Learning-Based Clinical Decision Support Systems: A Systematic Review. Appl. Sci. 2021, 11, 5088. [Google Scholar] [CrossRef]
  42. Burgess, E.R.; Jankovic, I.; Austin, M.; Cai, N.; Kapuścińska, A.; Currie, S.T.; Overhage, J.M.; Poole, E.S.; Kaye, J. Healthcare AI Treatment Decision Support: Design Principles to Enhance Clinician Adoption and Trust. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, Hamburg, Germany, 23–28 April 2023; Association for Computing Machinery: New York, NY, USA, 2023. [Google Scholar]
  43. Dai, T.; Tayur, S. Designing AI-Augmented Healthcare Delivery Systems for Physician Buy-In and Patient Acceptance. Prod. Oper. Manag. 2022, 31, 4443–4451. [Google Scholar] [CrossRef]
  44. Dai, T.; Singh, S. Using Artificial Intelligence as Gatekeeper or Second Opinion: Designing Patient Pathways for Artificial Intelligence Augmented Healthcare. Prod. Oper. Manag. 2025. articles in advance. [Google Scholar] [CrossRef]
  45. Dean, A.; Zhalechian, M.; Attari, I.; Van Oyen, M.; Garifullin, M.; Pardo, J. From Implementation to Impact: Human-Algorithm Collaboration in Hospital Bed Placement. Manag. Sci. 2025. articles in advance. [Google Scholar]
  46. Hou, T.; Li, M.; Tan, Y.; Zhao, H. Physician Adoption of AI Assistant. Manuf. Serv. Oper. Manag. 2024, 26, 1639–1655. [Google Scholar] [CrossRef]
  47. Jussupow, E.; Spohrer, K.; Heinzl, A.; Gawlitza, J. Augmenting Medical Diagnosis Decisions? An Investigation into Physicians’ Decision-Making Process with Artificial Intelligence. Inf. Syst. Res. 2021, 32, 713–735. [Google Scholar] [CrossRef]
  48. Küper, A.; Lodde, G.C.; Livingstone, E.; Schadendorf, D.; Krämer, N.C. Psychological Factors Influencing Appropriate Reliance on AI-Enabled Clinical Decision Support Systems: Experimental Web-Based Study Among Dermatologists. J. Med. Internet Res. 2025, 27, e58660. [Google Scholar] [CrossRef]
  49. Lebovitz, S.; Lifshitz-Assaf, H.; Levina, N. To Engage or Not to Engage with AI for Critical Judgments: How Professionals Deal with Opacity When Using AI for Medical Diagnosis. Organ. Sci. 2022, 33, 126–148. [Google Scholar] [CrossRef]
  50. Ratwani, R.M.; Sutton, K.; Galarraga, J.E. Addressing AI Algorithmic Bias in Health Care. JAMA 2024, 332, 1051–1052. [Google Scholar] [CrossRef] [PubMed]
  51. Tun, H.M.; Abdul Rahman, H.; Naing, L.; Malik, O.A. Trust in Artificial Intelligence–Based Clinical Decision Support Systems Among Health Care Workers: Systematic Review. J. Med. Internet Res. 2025, 27, e69678. [Google Scholar] [CrossRef]
  52. Brau, R.; Aloysius, J.A.; Siemsen, E. Demand Planning for the Digital Supply Chain: How to Integrate Human Judgment and Predictive Analytics. J. Oper. Manag. 2023, 69, 965–982. [Google Scholar] [CrossRef]
  53. Nair, D.; Huchzermeier, A. Predictably Unpredictable? How Judgmental and Machine Learning Forecasts Complement Each Other. Prod. Oper. Manag. 2024, 33, 1214–1234. [Google Scholar] [CrossRef]
  54. Omoush, M.M. Human–AI Collaboration in HRM and Employee-Centric Outcomes: Evidence from E-Supply Chain Management. Hum. Syst. Manag. 2026, 45, 343–358. [Google Scholar] [CrossRef]
  55. Punia, S. Medium-to Long-Term Demand Forecasting in Retail and Manufacturing Organizations: Integration of Machine Learning, Human Judgment, and Interval Variable. J. Forecast. 2026, 45, 122–134. [Google Scholar] [CrossRef]
  56. Yaroson, E.V.; Abadie, A.; Roux, M. Human-Artificial Intelligence Collaboration in Supply Chain Outcomes: The Mediating Role of Responsible Artificial Intelligence. Ann. Oper. Res. 2025, 354, 35–69. [Google Scholar] [CrossRef]
  57. Ben-Michael, E.; Greiner, D.J.; Huang, M.; Imai, K.; Jiang, Z.; Shin, S. Does AI Help Humans Make Better Decisions? A Statistical Evaluation Framework for Experimental and Observational Studies. arXiv 2024, arXiv:2403.12108. [Google Scholar] [CrossRef] [PubMed]
  58. Bockstedt, J.C.; Buckman, J.R. Humans’ Use of AI Assistance: The Effect of Loss Aversion on Willingness to Delegate Decisions. Manag. Sci. 2026, 72, 323–342. [Google Scholar] [CrossRef]
  59. Buçinca, Z.; Malaya, M.B.; Gajos, K.Z. To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proc. ACM Hum.-Comput. Interact. 2021, 5, 188. [Google Scholar] [CrossRef]
  60. Caro, F.; Colliard, J.-E.; Katok, E.; Ockenfels, A.; Stier-Moses, N.; Tucker, C.; Wu, D.J. Introduction to the Special Issue on the Human-Algorithm Connection. Manag. Sci. 2025, 72, 1–13. [Google Scholar] [CrossRef]
  61. Choudhury, A.; Shamszare, H. The Impact of Performance Expectancy, Workload, Risk, and Satisfaction on Trust in ChatGPT: Cross-Sectional Survey Analysis. JMIR Hum. Factors 2024, 11, e55399. [Google Scholar] [CrossRef] [PubMed]
  62. Dang, Q.; Li, G. Unveiling trust in AI: The interplay of antecedents, consequences, and cultural dynamics. AI Soc. 2026, 41, 669–692. [Google Scholar] [CrossRef]
  63. de Jong, S.; Paananen, V.; Tag, B.; van Berkel, N. Cognitive Forcing for Better Decision-Making: Reducing Overreliance on AI Systems Through Partial Explanations. Proc. ACM Hum.-Comput. Interact. CSCW 2025, 9, CSCW048. [Google Scholar] [CrossRef]
  64. Dietvorst, B.J.; Simmons, J.P.; Massey, C. Algorithm Aversion: People Erroneously Avoid Algorithms After Seeing Them Err. J. Exp. Psychol. Gen. 2015, 144, 114–126. [Google Scholar] [CrossRef]
  65. Glickman, M.; Sharot, T. How Human–AI Feedback Loops Alter Human Perceptual, Emotional and Social Judgements. Nat. Hum. Behav. 2025, 9, 345–359. [Google Scholar] [CrossRef]
  66. Goergen, J.; de Bellis, E.; Klesse, A.-K. AI assessment changes human behavior. Proc. Natl. Acad. Sci. USA 2025, 122, e2425439122. [Google Scholar] [CrossRef] [PubMed]
  67. Guo, Z.; Wu, Y.; Hartline, J.; Hullman, J. A Decision Theoretic Framework for Measuring AI Reliance. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, Rio de Janeiro, Brazil, 3–6 June 2024; Association for Computing Machinery: New York, NY, USA, 2024. [Google Scholar]
  68. Harbarth, L.; Gößwein, E.; Bodemer, D.; Schnaubert, L. (Over)Trusting AI Recommendations: How System and Person Variables Affect Dimensions of Complacency. Int. J. Hum.-Interact. 2025, 41, 391–410. [Google Scholar] [CrossRef]
  69. Holstein, J.; Böcking, L.; Spitzer, P.; Kühl, N.; Vössing, M.; Satzger, G. Balancing the Unknown: Exploring Human Reliance on AI Advice under Aleatoric and Epistemic Uncertainty. ACM Trans. Comput.-Hum. Interact. 2025, 32, 64. [Google Scholar] [CrossRef]
  70. Horowitz, M.C.; Kahn, L. Bending the Automation Bias Curve: A Study of Human and AI-Based Decision Making in National Security Contexts. Int. Stud. Q. 2024, 68, sqae020. [Google Scholar] [CrossRef]
  71. Hu, M.; Liu, J.; Yue, W.T. Dynamic Delegation in Human-AI Collaboration: Heuristic-Based and Instance-Based Decision-Making. Manag. Sci. 2025. articles in advance. [Google Scholar]
  72. Jussupow, E.; Benbasat, I.; Heinzl, A. An Integrative Perspective on Algorithm Aversion and Appreciation in Decision-Making. MIS Q. 2024, 48, 1575–1590. [Google Scholar] [CrossRef]
  73. Kahr, P.K.; Rooks, G.; Willemsen, M.C.; Snijders, C.C.P. Understanding Trust and Reliance Development in AI Advice: Assessing Model Accuracy, Model Explanations, and Experiences from Previous Interactions. ACM Trans. Interact. Intell. Syst. 2024, 14, 29. [Google Scholar] [CrossRef]
  74. Kahr, P.; Rooks, G.; Snijders, C.; Willemsen, M.C. Good Performance Is not Enough to Trust AI: Lessons from Logistics Experts on Their Long-Term Collaboration with an AI Planning System. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Yokohama, Japan, 26 April–1 May 2025; Association for Computing Machinery: New York, NY, USA, 2025. [Google Scholar]
  75. Klingbeil, A.; Grützner, C.; Schreck, P. Trust and Reliance on AI—An Experimental Study on the Extent and Costs of Overreliance on AI. Comput. Hum. Behav. 2024, 160, 108352. [Google Scholar] [CrossRef]
  76. Krakowski, S.; Haftor, D.; Luger, J.; Pashkevich, N.; Raisch, S. Human-Centered Artificial Intelligence: A Field Experiment. Manag. Sci. 2026, 72, 57–72. [Google Scholar] [CrossRef]
  77. Legros, B.; de Véricourt, F. Human–AI Interaction in Congested Service Systems: When It Improves Performance, and When It Doesn’t. Manag. Sci. 2026. articles in advance. [Google Scholar] [CrossRef]
  78. Li, J.; Yang, Y.; Liao, Q.V.; Zhang, J.; Lee, Y.-C. As Confidence Aligns: Understanding the Effect of AI Confidence on Human Self-confidence in Human-AI Decision Making. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Yokohama, Japan, 26 April–1 May 2025; Association for Computing Machinery: New York, NY, USA, 2025. [Google Scholar]
  79. Logg, J.M.; Minson, J.A.; Moore, D.A. Algorithm Appreciation: People Prefer Algorithmic to Human Judgment. Organ. Behav. Hum. Decis. Process. 2019, 151, 90–103. [Google Scholar] [CrossRef]
  80. Ma, S.; Chen, Q.; Wang, X.; Zheng, C.; Peng, Z.; Yin, M.; Ma, X. Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Yokohama, Japan, 26 April–1 May 2025; Association for Computing Machinery: New York, NY, USA, 2025. [Google Scholar]
  81. Rieger, T.; Kugler, L.; Manzey, D.; Roesler, E. The (Im)perfect Automation Schema: Who Is Trusted More, Automated or Human Decision Support? Hum. Factors 2024, 66, 1995–2007. [Google Scholar] [CrossRef] [PubMed]
  82. Scharowski, N.; Perrig, S.A.C.; Svab, M.; Opwis, K.; Brühlmann, F. Exploring the Effects of Human-Centered AI Explanations on Trust and Reliance. Front. Comput. Sci. 2023, 5, 1151150. [Google Scholar] [CrossRef]
  83. Shin, M.; Kim, J.; van Opheusden, B.; Griffiths, T.L. Superhuman artificial intelligence can improve human decision-making by increasing novelty. Proc. Natl. Acad. Sci. USA 2023, 120, e2214840120. [Google Scholar] [CrossRef]
  84. Starke, C.; Baleis, J.; Keller, B.; Marcinkowski, F. Fairness Perceptions of Algorithmic Decision-Making: A Systematic Review of the Empirical Literature. Big Data Soc. 2022, 9, 20539517221115189. [Google Scholar] [CrossRef]
  85. Vining, R.; Karimian, S.; McDonald, N.; Ebiele, M.; Doyle, B.; McKenna, L.; Ward, M.E.; Bendechache, M.; Brennan, R. Trustworthy Artificial Intelligence and Organisational Trust: A Scoping Review Using Socio-Technical Systems Analysis. Ergonomics 2025, 68, 1–20. [Google Scholar] [CrossRef]
  86. Xu, W.; Xu, X.; Qin, Y.; Zhao, C. Exploring AI’s Role in Influencing Individual Decision-Making Within Human–Machine Teams. Comput. Ind. Eng. 2025, 211, 111658. [Google Scholar] [CrossRef]
  87. Singh, I.L.; Molloy, R.; Parasuraman, R. Automation-induced “complacency”: Development of the Complacency-Potential Rating Scale. Int. J. Aviat. Psychol. 1993, 3, 111–122. [Google Scholar] [CrossRef]
  88. Bibbo, D.; Corvini, G.; Schmid, M.; Ranaldi, S.; Conforto, S. The Impact of Human-Robot Collaboration Levels on Postural Stability During Working Tasks Performed While Standing: Experimental Study. JMIR Hum. Factors 2025, 12, e64892. [Google Scholar] [CrossRef]
  89. de Koster, R.; Roy, D.; Lim, Y.F.; Kumar, S. Internet of Things in Intralogistics: Applications and Emerging Research. Prod. Oper. Manag. 2025, 34, 3343–3363. [Google Scholar] [CrossRef]
  90. Ferdman, A. AI deskilling is a structural problem. AI Soc. 2026, 41, 3001–3013. [Google Scholar] [CrossRef]
  91. Haghighi, A.; Cheraghi, M.; Pocachard, J.; Botta-Genoulaz, V.; Jocelyn, S.; Pourzarei, H. A Comprehensive Review and Bibliometric Analysis on Collaborative Robotics for Industry: Safety Emerging as a Core Focus. Front. Robot. AI 2025, 12, 1605682. [Google Scholar] [CrossRef]
  92. Konstant, A.; Orr, N.; Hagenow, M.; Gundrum, I.; Hu, Y.H.; Mutlu, B.; Zinn, M.; Gleicher, M.; Radwin, R.G. Human–Robot Collaboration with a Corrective Shared Controlled Robot in a Sanding Task. Hum. Factors 2025, 67, 246–263. [Google Scholar] [CrossRef] [PubMed]
  93. Liu, L.; Schoen, A.J.; Henrichs, C.; Li, J.; Mutlu, B.; Zhang, Y.; Radwin, R.G. Human Robot Collaboration for Enhancing Work Activities. Hum. Factors 2024, 66, 158–179. [Google Scholar] [CrossRef]
  94. Liu, L.; Sheng, S.; Li, J.; Man, S.S. Human Error Identification and Risk Prioritization in Human–Robot Collaboration in Manufacturing. Hum. Factors Ergon. Manuf. Serv. Ind. 2025, 35, e70012. [Google Scholar] [CrossRef]
  95. Menanno, M.; Riccio, C.; Benedetto, V.; Gissi, F.; Savino, M.M.; Troiano, L. An Ergonomic Risk Assessment System Based on 3D Human Pose Estimation and Collaborative Robot. Appl. Sci. 2024, 14, 4823. [Google Scholar] [CrossRef]
  96. Pasparakis, A.; de Vries, J.; de Koster, R. Assessing the impact of human–robot collaborative order picking systems on warehouse workers. Int. J. Prod. Res. 2023, 61, 7776–7790. [Google Scholar] [CrossRef]
  97. Pietrantoni, L.; Favilla, M.; Fraboni, F.; Mazzoni, E.; Morandini, S.; Benvenuti, M.; De Angelis, M. Integrating Collaborative Robots in Manufacturing, Logistics, and Agriculture: Expert Perspectives on Technical, Safety, and Human Factors. Front. Robot. AI 2024, 11, 1342130. [Google Scholar] [CrossRef]
  98. Segura, P.; Lobato-Calleros, O.; Soria-Arguello, I.; Hernández-Martínez, E.G. Work Roles in Human–Robot Collaborative Systems: Effects on Cognitive Ergonomics for the Manufacturing Industry. Appl. Sci. 2025, 15, 744. [Google Scholar] [CrossRef]
  99. Shi, Y.; Yu, H.; Yu, Y.; Yue, X. Analytics for IoT-Enabled Human–Robot Hybrid Sortation: An Online Optimization Approach. Prod. Oper. Manag. 2025, 34, 3364–3380. [Google Scholar] [CrossRef]
  100. Wang, T.; Zheng, P.; Li, S.; Wang, L. Multimodal human–robot interaction for human-centric smart manufacturing: A survey. Adv. Intell. Syst. 2024, 6, 2300359. [Google Scholar] [CrossRef]
  101. Endsley, M.R. Autonomous driving systems: A preliminary naturalistic study of the Tesla Model S. J. Cogn. Eng. Decis. Mak. 2017, 11, 225–238. [Google Scholar] [CrossRef]
  102. Han, D.W.; Chung, H.; Cao, Y.; Zhou, F.; Molnar, L.; Robert, L.P.; Tilbury, D.M.; Yang, X.J. A Systematic Review of Metrics Measuring Takeover Performance in Conditionally Automated Driving. Int. J. Hum.-Comput. Interact. 2025, 42, 6331–6360. [Google Scholar] [CrossRef]
  103. Holland, C.; Neyedli, H.F. The influence of performance feedback on trust and self-confidence in dynamically reliable automation. Proc. Hum. Factors Ergon. Soc. Annu. Meet. 2025, 69, 403–407. [Google Scholar] [CrossRef] [PubMed]
  104. Jiang, T.; Wu, J.; Leung, S.C.H. The Cognitive Impacts of Large Language Model Interactions on Problem Solving and Decision Making Using EEG Analysis. Front. Comput. Neurosci. 2025, 19, 1556483. [Google Scholar] [CrossRef]
  105. Merlhiot, G.; Bueno, M. How drowsiness and distraction can interfere with take-over performance: A systematic and meta-analytic review. Accid. Anal. Prev. 2022, 170, 106536. [Google Scholar] [CrossRef]
  106. Millard, A.S.; Greenlee, E.T. Effects of automation reliability on driver vigilance. In Proceedings of the Human Factors and Ergonomics Society Annual Meeting; SAGE Publications: Los Angeles, CA USA, 2024; Volume 68, pp. 1435–1436. [Google Scholar]
  107. Onnasch, L.; Wickens, C.D.; Li, H.; Manzey, D. Human performance consequences of stages and levels of automation: An integrated meta-analysis. Hum. Factors 2014, 56, 476–488. [Google Scholar] [CrossRef]
  108. Rodak, A.; Kruszewski, M.; Niedzicka, A. HMI Efficiency, Usability and Workload During Take-Over in AVs. Sci. Rep. 2025, 15, 15176. [Google Scholar] [CrossRef]
  109. Soares, S.; Lobo, A.; Ferreira, S.; Cunha, L.; Couto, A. Takeover Performance Evaluation Using Driving Simulation: A Systematic Review and Meta-Analysis. Eur. Transp. Res. Rev. 2021, 13, 47. [Google Scholar] [CrossRef]
  110. Treiman, L.S.; Ho, C.-J.; Kool, W. The consequences of AI training on human decision-making. Proc. Natl. Acad. Sci. USA 2024, 121, e2408731121. [Google Scholar] [CrossRef]
  111. Wang, A.; Wang, J.; Huang, C.; He, D.; Yang, H. Exploring How Physio-Psychological States Affect Drivers’ Takeover Performance in Conditional Automated Vehicles. Accid. Anal. Prev. 2025, 216, 108022. [Google Scholar] [CrossRef]
  112. Chen, M.; Wu, Y.; Ma, J.; Jia, X.; Gao, C.; Zhao, F.; Qiao, Y. Independent and collaborative performance of large language models and healthcare professionals in diagnosis and triage. npj Digit. Med. 2026, 9, 222. [Google Scholar] [CrossRef]
  113. Goh, E.; Bunning, B.; Khoong, E.C.; Gallo, R.J.; Milstein, A.; Centola, D.; Chen, J.H. Physician clinical decision modification and bias assessment in a randomized controlled trial of AI assistance. Commun. Med. 2025, 5, 59. [Google Scholar] [CrossRef]
  114. Goodell, A.J.; Chu, S.N.; Rouholiman, D.; Chu, L.F. Large language model agents can use tools to perform clinical calculations. npj Digit. Med. 2025, 8, 163. [Google Scholar] [CrossRef]
  115. Liu, Y.; Zhang, Y. ChatGPT as a clinical support tool: A comprehensive review of applications, assessment, and implementation challenges. Physiother. Pract. Res. 2025, 47, 160–172. [Google Scholar] [CrossRef]
  116. Liu, P.; Zhang, J.; Chen, S.; Chen, S. Human-AI teaming in healthcare: 1 + 1 > 2? npj Artif. Intell. 2025, 1, 47. [Google Scholar] [CrossRef]
  117. Mahajan, A.; Obermeyer, Z.; Daneshjou, R.; Lester, J.; Powell, D. Cognitive bias in clinical large language models. npj Digit. Med. 2025, 8, 428. [Google Scholar] [CrossRef] [PubMed]
  118. Omar, M.; Sorin, V.; Collins, J.D.; Reich, D.; Freeman, R.; Gavin, N.; Charney, A.; Stump, L.; Bragazzi, N.L.; Nadkarni, G.N.; et al. Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Commun. Med. 2025, 5, 330. [Google Scholar] [CrossRef] [PubMed]
  119. Oniani, D.; Wu, X.; Visweswaran, S.; Kapoor, S.; Kooragayalu, S.; Polanska, K.; Wang, Y. Enhancing Large Language Models for Clinical Decision Support by Incorporating Clinical Practice Guidelines. In Proceedings of the 2024 IEEE 12th International Conference on Healthcare Informatics (ICHI), Orlando, FL, USA, 3–6 June 2024; IEEE: New York, NY, USA, 2024; pp. 694–702. [Google Scholar]
  120. Siden, R.; Kerman, H.; Gallo, R.J.; Cool, J.A.; Hom, J.; Goh, E.; Ahuja, N.; Heidenreich, P.; Shieh, L.; Yang, D.; et al. A typology of physician input approaches to using AI chatbots for clinical decision-making. npj Digit. Med. 2026, 9, 14. [Google Scholar] [CrossRef]
  121. Zöller, N.; Berger, J.; Lin, I.; Fu, N.; Komarneni, J.; Barabucci, G.; Laskowski, K.; Shia, V.; Harack, B.; Chu, E.A.; et al. Human-AI collectives most accurately diagnose clinical vignettes. Proc. Natl. Acad. Sci. USA 2025, 122, e2426153122. [Google Scholar] [CrossRef]
  122. Boone, T.; Fahimnia, B.; Ganeshan, R.; Herold, D.M.; Sanders, N.R. Generative AI: Opportunities, Challenges, and Research Directions for Supply Chain Resilience. Transp. Res. Part E Logist. Transp. Rev. 2025, 199, 104135. [Google Scholar] [CrossRef]
  123. Jackson, I.; Ivanov, D.; Dolgui, A.; Namdar, J. Generative artificial intelligence in supply chain and operations management: A capability-based framework for analysis and implementation. Int. J. Prod. Res. 2024, 62, 6120–6145. [Google Scholar] [CrossRef]
  124. Adiasto, K. SustAInable employability: Sustainable employability in the age of generative artificial intelligence. Group Organ. Manag. 2024, 49, 1338–1348. [Google Scholar] [CrossRef]
  125. Brynjolfsson, E.; Li, D.; Raymond, L.R. Generative AI at Work. Q. J. Econ. 2025, 140, 889–942. [Google Scholar] [CrossRef]
  126. Dhillon, P.S.; Molaei, S.; Li, J.; Golub, M.; Zheng, S.; Robert, L.P. Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA, 11–16 May 2024; Association for Computing Machinery: New York, NY, USA, 2024. [Google Scholar]
  127. Gerlich, M. From Offloading to Engagement: An Experimental Study on Structured Prompting and Critical Reasoning with Generative AI. Data 2025, 10, 172. [Google Scholar] [CrossRef]
  128. He, G.; Demartini, G.; Gadiraju, U. Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents as A Daily Assistant. In Proceedings of the CHI Conference on Human Factors in Computing Systems, Yokohama, Japan, 26 April–1 May 2025; Association for Computing Machinery: New York, NY, USA, 2025. [Google Scholar]
  129. Kim, T.W.; Usman, U.; Garvey, A.M.; Duhachek, A. From algorithm aversion to AI dependence: Deskilling, upskilling, and emerging addictions in the GenAI age. Consum. Psychol. Rev. 2026, 9, 142–164. [Google Scholar] [CrossRef]
  130. Lee, H.P.; Sarkar, A.; Tankelevitch, L.; Drosos, I.Z.; Rintel, S.; Banks, R.; Wilson, N. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects from a Survey of Knowledge Workers. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Yokohama, Japan, 26 April–1 May 2025; Association for Computing Machinery: New York, NY, USA, 2025. [Google Scholar]
  131. Liu, Y.; Yang, Y.; Xu, H. From humans to AI: Understanding why AI is perceived as the preferred co-creation partner. Front. Psychol. 2025, 16, 1695532. [Google Scholar] [CrossRef] [PubMed]
  132. Loru, E.; Nudo, J.; Di Marco, N.; Santirocchi, A.; Atzeni, R.; Cinelli, M.; Cestari, V.; Rossi-Arnaud, C.; Quattrociocchi, W. The simulation of judgment in LLMs. Proc. Natl. Acad. Sci. USA 2025, 122, e2518443122. [Google Scholar] [CrossRef] [PubMed]
  133. McGuire, J.; De Cremer, D.; Van de Cruys, T. Establishing the importance of co-creation and self-efficacy in creative collaboration with artificial intelligence. Sci. Rep. 2024, 14, 18525. [Google Scholar] [CrossRef]
  134. Noy, S.; Zhang, W. Experimental evidence on the productivity effects of generative artificial intelligence. Science 2023, 381, 187–192. [Google Scholar] [CrossRef] [PubMed]
  135. Peng, S.; Kalliamvakou, E.; Cihon, P.; Demirer, M. The impact of AI on developer productivity: Evidence from GitHub Copilot. arXiv 2023, arXiv:2302.06590. [Google Scholar] [CrossRef]
  136. Rafner, J.; Zana, B.; Hansen, I.B.; Ceh, S.; Sherson, J.; Benedek, M.; Lebuda, I. Agency in Human-AI Collaboration for Image Generation and Creative Writing: Preliminary Insights from Think-Aloud Protocols. Creat. Res. J. 2025. online first. [Google Scholar] [CrossRef]
  137. Sakamoto, Y.; Uchida, T.; Ishiguro, H. Value-based large language model agent simulation for mutual evaluation of trust and interpersonal closeness. Sci. Rep. 2025, 15, 41653. [Google Scholar] [CrossRef]
  138. Sidra, S.; Mason, C. Generative AI in Human-AI Collaboration: Validation of the Collaborative AI Literacy and Collaborative AI Metacognition Scales for Effective Use. Int. J. Hum.-Comput. Interact. 2025, 42, 5084–5108. [Google Scholar] [CrossRef]
  139. Wang, N.; Kim, H.; Peng, J.; Wang, J. Exploring creativity in human–AI co-creation: A comparative study across design experience. Front. Comput. Sci. 2025, 7, 1672735. [Google Scholar] [CrossRef]
  140. Xiao, F.; Wang, X.T. Evaluating the ability of large language models to predict human social decisions. Sci. Rep. 2025, 15, 32290. [Google Scholar] [CrossRef]
  141. Xie, C.; Chen, C.; Jia, F.; Ye, Z.; Lai, S.; Shu, K.; Gu, J.; Bibi, A.; Hu, Z.; Jurgens, D.; et al. Can Large Language Model Agents Simulate Human Trust Behavior? In Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, BC, Canada, 10–15 December 2024; Curran Associates Inc.: Red Hook, NY, USA, 2024. [Google Scholar]
  142. Rajpurkar, P.; Chen, E.; Banerjee, O.; Topol, E.J. AI in Health and Medicine. Nat. Med. 2022, 28, 31–38. [Google Scholar] [CrossRef]
  143. Ivanov, D. The Industry 5.0 framework: Viability-based integration of the resilience, sustainability, and human-centricity perspectives. Int. J. Prod. Res. 2023, 61, 1683–1695. [Google Scholar] [CrossRef]
  144. Passalacqua, M.; Pellerin, R.; Magnani, F.; Doyon-Poulin, P.; Del-Aguila, L.; Boasen, J.; Léger, P.-M. Human-centred AI in industry 5.0: A systematic review. Int. J. Prod. Res. 2025, 63, 2638–2669. [Google Scholar] [CrossRef]
  145. Shabur, M.A.; Shahriar, A.; Ara, M.A. From automation to collaboration: Exploring the impact of industry 5.0 on sustainable manufacturing. Discov. Sustain. 2025, 6, 341. [Google Scholar] [CrossRef]
  146. Sun, X.; Song, Y. Unlocking the synergy: Increasing productivity through Human-AI collaboration in the industry 5.0 Era. Comput. Ind. Eng. 2025, 200, 110657. [Google Scholar] [CrossRef]
  147. Huang, J.; Chen, Y. Human-artificial intelligence collaboration and green collaborative innovation: The roles of knowledge management and CEO green experience. J. Clean. Prod. 2025, 535, 147117. [Google Scholar] [CrossRef]
  148. Janhunen, E.; Toivikko, T.; Blomqvist, K.; Siemon, D. Trust in Digital Human-AI Team Collaboration: A Systematic Review. In Proceedings of the 30th Americas Conference on Information Systems (AMCIS), Salt Lake City, UT, USA, 15–17 August 2024. [Google Scholar]
  149. Lin, Z. Human-AI complementarity needs augmentation, not emulation. Nat. Rev. Psychol. 2026, 5, 228–229. [Google Scholar] [CrossRef]
  150. Mancuso, I.; Messeni Petruzzelli, A.; Panniello, U.; Vaia, G. The bright and dark sides of AI innovation for sustainable development: Understanding the paradoxical tension between value creation and value destruction. Technovation 2025, 143, 103232. [Google Scholar] [CrossRef]
  151. Rainey, P.B.; Hochberg, M.E. Could humans and AI become a new evolutionary individual? Proc. Natl. Acad. Sci. USA 2025, 122, e2509122122. [Google Scholar] [CrossRef] [PubMed]
  152. Shin, Y. Toward Human-Centered Artificial Intelligence for Users’ Digital Well-Being: Systematic Review, Synthesis, and Future Directions. JMIR Hum. Factors 2025, 12, e69533. [Google Scholar] [CrossRef] [PubMed]
  153. Valtonen, A.; Saunila, M.; Ukko, J.; Treves, L.; Ritala, P. AI and employee wellbeing in the workplace: An empirical study. J. Bus. Res. 2025, 199, 115584. [Google Scholar] [CrossRef]
  154. Xu, C.; Cho, S.-E. Factors Affecting Human–AI Collaboration Performances in Financial Sector: Sustainable Service Development Perspective. Sustainability 2025, 17, 4335. [Google Scholar] [CrossRef]
Figure 1. Three AI technology generations defined by interaction modality, human role, and system autonomy. Paper counts indicate the number of generation-specific studies reviewed (Ntotal = 93, plus 59 foundational, review, sustainability, and gap-filling papers). Temporal spans overlap: real-world systems increasingly combine elements across generations.
Figure 1. Three AI technology generations defined by interaction modality, human role, and system autonomy. Paper counts indicate the number of generation-specific studies reviewed (Ntotal = 93, plus 59 foundational, review, sustainability, and gap-filling papers). Temporal spans overlap: real-world systems increasingly combine elements across generations.
Sustainability 18 05313 g001
Figure 2. PRISMA 2020 flow diagram of the systematic review process. Searches were conducted across Web of Science, Scopus, PubMed, and IEEE Xplore (final search 20 March 2026; timeframe 2018–2026), supplemented with seminal pre-2018 works (Section 3.2) and targeted gap-filling searches described in Section 3.2. Record counts shown in the diagram are exact integers and combine master and supplementary searches; the supplementary-search rationale and integration are described in Section 3.2.
Figure 2. PRISMA 2020 flow diagram of the systematic review process. Searches were conducted across Web of Science, Scopus, PubMed, and IEEE Xplore (final search 20 March 2026; timeframe 2018–2026), supplemented with seminal pre-2018 works (Section 3.2) and targeted gap-filling searches described in Section 3.2. Record counts shown in the diagram are exact integers and combine master and supplementary searches; the supplementary-search rationale and integration are described in Section 3.2.
Sustainability 18 05313 g002
Figure 3. Analytical flow of the review. Each research question (RQ1–RQ4) maps to a specific methodological step and a results section. Solution-maturity coding (described in Section 3.6 and operationalized in Section 5.2) is applied across RQ2–RQ4 synthesis, informing both the CCF (Table 2) and the prioritization of research directions (Section 6.4).
Figure 3. Analytical flow of the review. Each research question (RQ1–RQ4) maps to a specific methodological step and a results section. Solution-maturity coding (described in Section 3.6 and operationalized in Section 5.2) is applied across RQ2–RQ4 synthesis, informing both the CCF (Table 2) and the prioritization of research directions (Section 6.4).
Sustainability 18 05313 g003
Figure 4. The Collaboration Convergence Framework: a 6 × 3 matrix mapping six invariant human challenges across three AI technology generations. Each challenge recurs under generation-specific terminology (the “jangle fallacy”). Solution maturity reflects the strength of evidence for effective design interventions: green = known solutions with validated interventions, amber = emerging evidence with promising approaches, red = unsolved challenges requiring urgent research. Arrows indicate terminological evolution; functionally equivalent phenomena receive different labels as technology changes.
Figure 4. The Collaboration Convergence Framework: a 6 × 3 matrix mapping six invariant human challenges across three AI technology generations. Each challenge recurs under generation-specific terminology (the “jangle fallacy”). Solution maturity reflects the strength of evidence for effective design interventions: green = known solutions with validated interventions, amber = emerging evidence with promising approaches, red = unsolved challenges requiring urgent research. Arrows indicate terminological evolution; functionally equivalent phenomena receive different labels as technology changes.
Sustainability 18 05313 g004
Figure 5. Evidence landscape and research priorities across the Collaboration Convergence Framework. Bar lengths represent the relative volume of empirical evidence for each challenge within each AI generation. Thin evidence across all generations (Skill Retention, Accountability) marks critical gaps targeted by research directions RD1–10. Strong Gen 1/Gen 2 evidence with weak Gen 3 coverage indicates transfer opportunities where established design principles (DP1–12) can accelerate Gen 3 system design.
Figure 5. Evidence landscape and research priorities across the Collaboration Convergence Framework. Bar lengths represent the relative volume of empirical evidence for each challenge within each AI generation. Thin evidence across all generations (Skill Retention, Accountability) marks critical gaps targeted by research directions RD1–10. Strong Gen 1/Gen 2 evidence with weak Gen 3 coverage indicates transfer opportunities where established design principles (DP1–12) can accelerate Gen 3 system design.
Sustainability 18 05313 g005
Table 1. Characteristics of all included studies (N = 152), organized by AI generation and domain. Reference numbers correspond to entries in the bibliography.
Table 1. Characteristics of all included studies (N = 152), organized by AI generation and domain. Reference numbers correspond to entries in the bibliography.
#Study [Ref]GenerationDomainStudy TypeKey Construct
Foundational (n = 7)
1Bainbridge (1983) [33]FoundationalAviationConceptualSkill retention
2Bureau (2012) [1]FoundationalAviationReportSafety/failure
3Parasuraman (1997) [26]FoundationalAviationConceptualTrust/reliance
4Parasuraman (2000) [29]FoundationalAviationConceptualAutonomy/teaming
5Sarter (1995) [34]FoundationalAviationConceptualAuthority transfer
6Hancock (2011) [35]FoundationalCross-domainReviewTrust/reliance
7Lee (2004) [28]FoundationalCross-domainConceptualTrust/reliance
Review (n = 17)
8Kirwan (2025) [14]ReviewAviationReviewMultiple
9Xia (2025) [12]ReviewAviationReviewMultiple
10Dai (2023) [15]ReviewHealthcareReviewCollaboration
11Kumar (2025) [18]ReviewManuf./SCReviewTrust/reliance
12Samuels (2025) [17]ReviewManuf./SCReviewSustainability/I5.0
13Bao (2023) [21]ReviewCross-domainReviewTrust/reliance
14Berretta (2023) [22]ReviewCross-domainReviewAutonomy/teaming
15Do (2025) [23]ReviewCross-domainReviewTrust/reliance
16Gomez (2025) [8]ReviewCross-domainReviewCollaboration
17Kohn (2021) [19]ReviewCross-domainReviewTrust/reliance
18Lai (2023) [6]ReviewCross-domainReviewTrust/reliance
19O’Neill (2022) [10]ReviewCross-domainReviewAutonomy/teaming
20Romeo (2026) [9]ReviewCross-domainReviewTrust/reliance
21Schemmer (2023) [7]ReviewCross-domainReviewTransparency
22Schmutz (2024) [13]ReviewCross-domainReviewAutonomy/teaming
23Walker (2023) [11]ReviewCross-domainReviewTrust/reliance
24Wong (2025) [20]ReviewCross-domainReviewTrust/reliance
Gen 1 (n = 52)
25Abbas (2025) [39]Gen 1HealthcareReviewTransparency
26Adams (2022) [40]Gen 1HealthcareEmpiricalSafety/failure
27Antoniadi (2021) [41]Gen 1HealthcareReviewTransparency
28Burgess (2023) [42]Gen 1HealthcareEmpiricalTrust/reliance
29Dai (2022) [43]Gen 1HealthcareConceptualTrust/reliance
30Dai (2025) [44]Gen 1HealthcareEmpiricalAuthority transfer
31Dean (2025) [45]Gen 1HealthcareEmpiricalCollaboration
32Hou (2024) [46]Gen 1HealthcareEmpiricalTrust/reliance
33Jabbour (2023) [36]Gen 1HealthcareEmpiricalTransparency
34Jussupow (2021) [47]Gen 1HealthcareEmpiricalTrust/reliance
35Küper (2025) [48]Gen 1HealthcareEmpiricalTrust/reliance
36Lebovitz (2022) [49]Gen 1HealthcareEmpiricalTransparency
37Ratwani (2024) [50]Gen 1HealthcareConceptualAccountability
38Tun (2025) [51]Gen 1HealthcareReviewTrust/reliance
39Brau (2023) [52]Gen 1Manuf./SCEmpiricalCollaboration
40Nair (2024) [53]Gen 1Manuf./SCEmpiricalCollaboration
41Omoush (2025) [54]Gen 1Manuf./SCEmpiricalCollaboration
42Punia (2026) [55]Gen 1Manuf./SCEmpiricalCollaboration
43Yaroson (2025) [56]Gen 1Manuf./SCEmpiricalAccountability
44Bansal (2021) [37]Gen 1Cross-domainEmpiricalTransparency
45Ben-Michael (2024) [57]Gen 1Cross-domainConceptualCollaboration
46Bockstedt (2026) [58]Gen 1Cross-domainEmpiricalTrust/reliance
47Buçinca (2021) [59]Gen 1Cross-domainEmpiricalCognitive engagement
48Caro (2026) [60]Gen 1Cross-domainConceptualCollaboration
49Choudhury (2024) [61]Gen 1Cross-domainEmpiricalTrust/reliance
50Dang (2026) [62]Gen 1Cross-domainEmpiricalTrust/reliance
51de Jong (2025) [63]Gen 1Cross-domainEmpiricalCognitive engagement
52Dietvorst (2015) [64]Gen 1Cross-domainEmpiricalTrust/reliance
53Fügener (2022) [4]Gen 1Cross-domainEmpiricalAuthority transfer
54Glickman (2025) [65]Gen 1Cross-domainEmpiricalAccountability
55Goergen (2025) [66]Gen 1Cross-domainEmpiricalTrust/reliance
56Guo (2024) [67]Gen 1Cross-domainConceptualTrust/reliance
57Harbarth (2025) [68]Gen 1Cross-domainEmpiricalTrust/reliance
58Hemmer (2025) [3]Gen 1Cross-domainConceptualCollaboration
59Holstein (2025) [69]Gen 1Cross-domainEmpiricalTrust/reliance
60Horowitz (2024) [70]Gen 1Cross-domainEmpiricalTrust/reliance
61Hu (2025) [71]Gen 1Cross-domainEmpiricalAuthority transfer
62Jussupow (2024) [72]Gen 1Cross-domainConceptualTrust/reliance
63Kahr (2024) [73]Gen 1Cross-domainEmpiricalTrust/reliance
64Kahr (2025) [74]Gen 1Cross-domainEmpiricalTrust/reliance
65Klingbeil (2024) [75]Gen 1Cross-domainEmpiricalTrust/reliance
66Krakowski (2026) [76]Gen 1Cross-domainEmpiricalCollaboration
67Legros (2025) [77]Gen 1Cross-domainEmpiricalAuthority transfer
68Li (2025) [78]Gen 1Cross-domainEmpiricalTrust/reliance
69Logg (2019) [79]Gen 1Cross-domainEmpiricalTrust/reliance
70Ma (2025) [80]Gen 1Cross-domainEmpiricalCognitive engagement
71Rieger (2024) [81]Gen 1Cross-domainEmpiricalTrust/reliance
72Scharowski (2023) [82]Gen 1Cross-domainEmpiricalTransparency
73Shin (2023) [83]Gen 1Cross-domainEmpiricalCognitive engagement
74Starke (2022) [84]Gen 1Cross-domainReviewAccountability
75Vining (2025) [85]Gen 1Cross-domainReviewTrust/reliance
76Xu (2025) [86]Gen 1Cross-domainEmpiricalTrust/reliance
Gen 2 (n = 26)
77Singh (1993) [87]Gen 2AviationEmpiricalTrust/reliance
78Bibbo (2025) [88]Gen 2Manuf./SCEmpiricalSafety/failure
79de Koster (2025) [89]Gen 2Manuf./SCEmpiricalCollaboration
80Ferdman (2025) [90]Gen 2Manuf./SCConceptualSkill retention
81Haghighi (2025) [91]Gen 2Manuf./SCReviewSafety/failure
82Konstant (2025) [92]Gen 2Manuf./SCEmpiricalSafety/failure
83Liu (2024) [93]Gen 2Manuf./SCEmpiricalCollaboration
84Liu (2025) [94]Gen 2Manuf./SCEmpiricalSafety/failure
85Menanno (2024) [95]Gen 2Manuf./SCEmpiricalSafety/failure
86Pasparakis (2023) [96]Gen 2Manuf./SCEmpiricalSkill retention
87Pietrantoni (2024) [97]Gen 2Manuf./SCEmpiricalSafety/failure
88Segura (2025) [98]Gen 2Manuf./SCEmpiricalCognitive engagement
89Shi (2025) [99]Gen 2Manuf./SCEmpiricalCollaboration
90Wang (2024) [100]Gen 2Manuf./SCReviewCollaboration
91Beer (2014) [27]Gen 2Cross-domainConceptualAutonomy/teaming
92Endsley (2017) [101]Gen 2Cross-domainEmpiricalSkill retention
93Han (2025) [102]Gen 2Cross-domainReviewAuthority transfer
94Holland (2025) [103]Gen 2Cross-domainEmpiricalTrust/reliance
95Jiang (2025) [104]Gen 2Cross-domainEmpiricalCognitive engagement
96Merlhiot (2022) [105]Gen 2Cross-domainReviewSafety/failure
97Millard (2024) [106]Gen 2Cross-domainEmpiricalSkill retention
98Onnasch (2014) [107]Gen 2Cross-domainReviewAuthority transfer
99Rodak (2025) [108]Gen 2Cross-domainEmpiricalAuthority transfer
100Soares (2021) [109]Gen 2Cross-domainReviewAuthority transfer
101Treiman (2024) [110]Gen 2Cross-domainEmpiricalSkill retention
102Wang (2025) [111]Gen 2Cross-domainEmpiricalAuthority transfer
Gen 3 (n = 32)
103Chen (2026) [112]Gen 3HealthcareEmpiricalHallucination/epistemia
104Goh (2025) [113]Gen 3HealthcareEmpiricalHallucination/epistemia
105Goodell (2025) [114]Gen 3HealthcareEmpiricalHallucination/epistemia
106Liu (2025) [115]Gen 3HealthcareReviewMultiple
107Liu (2025) [116]Gen 3HealthcareEmpiricalCollaboration
108Mahajan (2025) [117]Gen 3HealthcareEmpiricalHallucination/epistemia
109Omar (2025) [118]Gen 3HealthcareEmpiricalHallucination/epistemia
110Oniani (2024) [119]Gen 3HealthcareEmpiricalHallucination/epistemia
111Siden (2026) [120]Gen 3HealthcareEmpiricalTrust/reliance
112Zöller (2025) [121]Gen 3HealthcareEmpiricalCollaboration
113Boone (2025) [122]Gen 3Manuf./SCConceptualSustainability/I5.0
114Jackson (2024) [123]Gen 3Manuf./SCConceptualCollaboration
115Adiasto (2024) [124]Gen 3Cross-domainEmpiricalSustainability/I5.0
116Brynjolfsson (2025) [125]Gen 3Cross-domainEmpiricalSkill retention
117Dell’Acqua (2023) [5]Gen 3Cross-domainEmpiricalCognitive engagement
118Dhillon (2024) [126]Gen 3Cross-domainEmpiricalCollaboration
119Gerlich (2025) [127]Gen 3Cross-domainEmpiricalCognitive engagement
120He (2025) [128]Gen 3Cross-domainEmpiricalTrust/reliance
121Kim (2026) [129]Gen 3Cross-domainConceptualSkill retention
122Lee (2025) [130]Gen 3Cross-domainEmpiricalCognitive engagement
123Liu (2025) [131]Gen 3Cross-domainEmpiricalCollaboration
124Loru (2025) [132]Gen 3Cross-domainEmpiricalHallucination/epistemia
125Luo (2025) [30]Gen 3Cross-domainReviewAutonomy/teaming
126McGuire (2024) [133]Gen 3Cross-domainEmpiricalCollaboration
127Noy (2023) [134]Gen 3Cross-domainEmpiricalCognitive engagement
128Peng (2023) [135]Gen 3Cross-domainEmpiricalCognitive engagement
129Rafner (2025) [136]Gen 3Cross-domainEmpiricalCollaboration
130Sakamoto (2025) [137]Gen 3Cross-domainEmpiricalTrust/reliance
131Sidra (2025) [138]Gen 3Cross-domainEmpiricalCollaboration
132Wang (2025) [139]Gen 3Cross-domainEmpiricalCollaboration
133Xiao (2025) [140]Gen 3Cross-domainEmpiricalHallucination/epistemia
134Xie (2024) [141]Gen 3Cross-domainEmpiricalTrust/reliance
Cross-Gen (n = 18)
135Rajpurkar (2022) [142]Cross-GenHealthcareConceptualMultiple
136Ivanov (2023) [143]Cross-GenManuf./SCConceptualSustainability/I5.0
137Passalacqua (2025) [144]Cross-GenManuf./SCReviewSustainability/I5.0
138Shabur (2025) [145]Cross-GenManuf./SCConceptualSustainability/I5.0
139Sun (2025) [146]Cross-GenManuf./SCEmpiricalSustainability/I5.0
140Tóth (2023) [25]Cross-GenManuf./SCConceptualSustainability/I5.0
141van Erp (2024) [24]Cross-GenManuf./SCConceptualSustainability/I5.0
142Almusharraf (2025) [38]Cross-GenCross-domainConceptualSustainability/I5.0
143Ansari (2026) [31]Cross-GenCross-domainReviewMultiple
144Huang (2025) [147]Cross-GenCross-domainEmpiricalSustainability/I5.0
145Janhunen (2024) [148]Cross-GenCross-domainReviewTrust/reliance
146Lin (2026) [149]Cross-GenCross-domainConceptualCollaboration
147Mancuso (2025) [150]Cross-GenCross-domainConceptualSustainability/I5.0
148Rainey (2025) [151]Cross-GenCross-domainConceptualCollaboration
149Shin (2025) [152]Cross-GenCross-domainReviewMultiple
150Vaccaro (2024) [2]Cross-GenCross-domainReviewCollaboration
151Valtonen (2025) [153]Cross-GenCross-domainEmpiricalSustainability/I5.0
152Xu (2025) [154]Cross-GenCross-domainEmpiricalSustainability/I5.0
Table 2. The Collaboration Convergence Framework: Six Recurring Challenges Across Three AI Generations, with Illustrative Failure Cases.
Table 2. The Collaboration Convergence Framework: Six Recurring Challenges Across Three AI Generations, with Illustrative Failure Cases.
ChallengeGen 1: Decision SupportGen 2: Autonomous SystemsGen 3: LLM AgentsSolution MaturityIllustrative Failure
1. Trust CalibrationAlgorithm aversion vs. appreciation (inverted-U)Automation complacency vs. scepticism (Bainbridge paradox)Over-reliance on epistemic fluency; distrust of LLM transparencyKnown: feedback mechanisms, gradual exposure, professional expertiseAir France 447: autopilot disconnect, SA loss [1]
2. Reliance BehaviorDiscrimination loss: unable to identify when AI errsMode errors; authority transfer failuresHallucination acceptance; inability to verify claim truthEmerging: cognitive forcing, partial explanationsDell’Acqua 2023: consultants using AI 19 pp worse off-frontier [5]
3. Cognitive EngagementPassive verification: monitoring AI without active reasoningVigilance decrement: bored monitoring of high-reliability automationCognitive offloading: delegating thinking to the LLMKnown: unassisted intervals, sequential teaming, structured promptingBuçinca 2021: users accept AI advice absent of cognitive forcing [59]
4. Skill RetentionExpertise loss from disuse of judgment (rare decision types)Deskilling from continuous automation; reduced mental modelsAtrophy of analytical thinking; narrowed strategiesUnsolved: dose–response of unassisted intervals; transfer of trainingEndsley 2017: Tesla drivers, inaccurate mental models [101]
5. AccountabilityImplicit human responsibility for AI advisory errorsJust Culture frameworks; automation blame diffusionDiffuse responsibility across creators, deployers, usersUnsolved: robust responsibility allocation frameworksRatwani 2024: no shared-responsibility model in healthcare [50]
6. Transparency & EpistemiaBlack-box opacity; inability to interpret outputsExplainability paradox; over-trust from explanationsSurface plausibility: LLM outputs “look right” (epistemia)Emerging: negative explanations, uncertainty displayJabbour 2023: biased AI cut accuracy 11.3 pp; XAI did not mitigate [36]
Table 3. Twelve Transferable Design Principles (DP1–DP12) for Human–AI Collaboration.
Table 3. Twelve Transferable Design Principles (DP1–DP12) for Human–AI Collaboration.
PrincipleChallengeEvidence BaseLLM Application
DP1: Feedback on AI reliabilityTrust calibrationHorowitz et al. [70]; Holland et al. [103]Real-time accuracy feedback on LLM outputs; confidence calibration displays
DP2: Unassisted intervalsCognitive engagement; skill retentionEndsley [101]; Singh et al. [87]; Kirwan [14]Periodic “AI-free” work intervals; enforced manual analysis phases
DP3: Partial rather than full explanationsReliance calibration; epistemiaGuo et al. [67]; Schemmer et al. [7]Provide explanation uncertainty; highlight missing information
DP4: Initial independent judgmentCognitive engagement; discrimination lossLai et al. [6]; Buçinca et al. [59]Users generate independent analysis before LLM exposure
DP5: Clear authority boundariesReliance behavior; accountabilitySarter & Woods [34]; Hancock et al. [35]; Dai & Singh [44]Explicit rules: “LLM advises here; human decides here”
DP6: Multi-modal alerts for failuresSkill retention; mode errorsRodak et al. [108]; Endsley [101]Multimodal alerts for hallucinations; forcing functions
DP7: Structured prompting protocolsCognitive engagement; skill retentionGerlich [127]; Lee et al. [130]Template-based prompts that enforce human reasoning steps
DP8: Domain-expert involvementTrust calibration; accountabilityKüper et al. [48]; Hancock et al. [35]; Dai & Tayur [43]; Hou et al. [46]Involve domain experts in system design and validation
DP9: Adaptive task allocationComplementarity; cognitive loadFügener et al. [4]; Hemmer et al. [3]; Dai & Singh [44]AI handles high-volume tasks; humans handle high-judgment tasks
DP10: Negative feedback emphasisReliance behavior; epistemiaRomeo & Conti [9]; Bansal et al. [37]; Jabbour et al. [36]Prominently show LLM errors; counter-examples before AI advice
DP11: Sequential teamingCognitive engagement; accountabilityHemmer et al. [3]; Nair et al. [53]Humans and AI contribute sequentially; human sees AI before deciding
DP12: Skill preservation contractsSkill retention; motivationKim et al. [129]; Endsley [101]Explicit commitment to maintain manual skills; routine practice
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Dong, A.; Li, P.; Chen, Y.; Gibson, S.; Zhao, L.; He, M. Human–AI Collaboration Across Decision Support, Autonomous Systems, and LLM Agents: A Systematic Review and Collaboration Convergence Framework. Sustainability 2026, 18, 5313. https://doi.org/10.3390/su18115313

AMA Style

Dong A, Li P, Chen Y, Gibson S, Zhao L, He M. Human–AI Collaboration Across Decision Support, Autonomous Systems, and LLM Agents: A Systematic Review and Collaboration Convergence Framework. Sustainability. 2026; 18(11):5313. https://doi.org/10.3390/su18115313

Chicago/Turabian Style

Dong, Aqi, Peng Li, Yanbing Chen, Shanan Gibson, Lin Zhao, and Meiling He. 2026. "Human–AI Collaboration Across Decision Support, Autonomous Systems, and LLM Agents: A Systematic Review and Collaboration Convergence Framework" Sustainability 18, no. 11: 5313. https://doi.org/10.3390/su18115313

APA Style

Dong, A., Li, P., Chen, Y., Gibson, S., Zhao, L., & He, M. (2026). Human–AI Collaboration Across Decision Support, Autonomous Systems, and LLM Agents: A Systematic Review and Collaboration Convergence Framework. Sustainability, 18(11), 5313. https://doi.org/10.3390/su18115313

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop