Next Article in Journal
Does Policy Uncertainty Distort Green Innovation? Evidence from Heavy-Polluting Firms in China
Previous Article in Journal
Adaptive Evolution of Community Response Organizations in Public Health Emergencies: A Grounded System Dynamics Case Study from China
Previous Article in Special Issue
Evolution and Ecological Activation Mechanisms of Chinese Electric Vehicles’ International Image: A Complex Adaptive Systems Perspective
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

When AI Joins the Diagnosis: How Doctor–AI Collaboration Shapes Perceived Doctor Responsibility Under Perceived Diagnostic Errors

1
School of Business Administration, Huaqiao University, Quanzhou 362021, China
2
School of Business Administration, Zhongnan University of Economics and Law, Wuhan 430073, China
*
Author to whom correspondence should be addressed.
Systems 2026, 14(9), 1081; https://doi.org/10.3390/systems14091081
Submission received: 12 June 2026 / Revised: 14 August 2026 / Accepted: 27 August 2026 / Published: 2 September 2026

Highlights

Please indicate how your work links to systems science via your contributions to systems practice, theory, and/or methodology.
  • This research situates doctor–AI collaborative diagnosis within the healthcare sociotechnical system and shows how the structure of human–AI participation shapes perceived doctor responsibility.
  • It extends responsibility attribution theory by proposing two connected stages—agent identification and responsibility allocation—and identifies perceived shared agency as the psychological mechanism linking diagnostic mode to perceived doctor responsibility.
What are the main findings and/or implications of the main findings?
  • Across five studies using patient and observer perspectives, doctor–AI collaborative diagnosis reduced perceived doctor responsibility relative to doctor-only diagnosis when a diagnostic error was perceived.
  • This responsibility-reducing effect was weaker when doctors rejected correct AI advice than when they accepted incorrect AI advice, highlighting the importance of recognizing AI’s substantive involvement in the diagnostic process for responsibility governance.

Abstract

Artificial intelligence (AI), especially generative AI, is becoming deeply embedded in healthcare, making doctor–AI collaborative diagnosis a common mode of medical decision-making. This development raises a critical question about responsibility attribution: When people perceive that a diagnostic error has occurred, how does AI involvement shape patients’ and observers’ judgments of the doctor’s responsibility? Drawing on responsibility attribution theory, we examine this question across five studies—one event-related potential (ERP) experiment and four scenario experiments. We find that when a diagnostic error is perceived, doctor–AI collaborative diagnosis (vs. doctor-only diagnosis) reduces perceived doctor responsibility by increasing perceived shared agency. This responsibility-reducing effect is weaker when the doctor rejects correct AI advice than when the doctor accepts incorrect AI advice. Theoretically, our findings show that responsibility attribution in human–AI collaboration involves two stages: agent identification and responsibility allocation. This account extends responsibility attribution theory to human–AI collaboration and identifies perceived shared agency as a key psychological mechanism underlying responsibility judgments in these settings. Practically, the findings can inform technology deployment, responsibility communication, and governance mechanisms in hospitals, AI firms, and regulatory agencies.

1. Introduction

In the post-pandemic era, countries around the world increasingly view AI as an important means of addressing shortages in healthcare resources and personnel, improving service efficiency, and expanding access to high-quality care [1,2,3]. Contemporary AI systems, particularly generative AI, possess powerful natural-language interaction and content-generation capabilities [4,5], enabling them to participate in multiple stages of healthcare delivery. For example, Stanford Health Care and Mass General Brigham in the United States have used or evaluated AI tools to assist with clinical-note generation, while Beijing Tiantan Hospital in China has used the large language model Longying for medical-image diagnosis. Among the most important applications is support for doctors throughout the diagnostic process, including question answering, medical-record generation, image-interpretation support, and the generation of diagnostic and treatment recommendations [4,6]. Doctor–AI collaboration in diagnosis is therefore becoming an increasingly salient feature of clinical practice. People are also becoming aware that a doctor’s diagnostic opinion may originate either from the doctor alone or from a judgment jointly shaped by the doctor and AI.
The introduction of AI into healthcare does not, however, eliminate diagnostic errors. Prior research indicates that large language models used in medicine may still hallucinate, produce factual errors, misunderstand context, or fail to adapt adequately to clinical settings. Thus, despite their strong overall performance, they may generate inaccurate or unreliable recommendations [5,7]. More importantly, when adverse medical outcomes occur in real-world care, people, including patients and observers, typically lack complete medical information and professional expertise. They may therefore be unable to determine whether an adverse outcome resulted from a diagnostic error, the inherent risks of the disease, or other medical factors [8,9,10,11]. Under such conditions, people may interpret an adverse outcome, based on limited contextual cues, as evidence that a diagnostic error has occurred. This perceived diagnostic error can then trigger responsibility judgments and subsequent evaluative and behavioral responses [10,11,12,13]. What often shapes social evaluations and behavioral responses is therefore not objective responsibility established through professional or legal procedures, but people’s subjective perceptions of the diagnostic error and the doctor’s responsibility [10]. Two pretests conducted from the patient and observer perspectives further confirmed that participants readily perceived the adverse outcomes as diagnostic errors (see Appendix A).
This raises an important but insufficiently answered question: When people perceive a diagnostic error, does AI involvement increase or decrease perceived doctor responsibility? If people regard AI as an advanced diagnostic tool used by the doctor, its use may increase perceptions of the doctor’s diagnostic capability, control over the outcome, and ability to avoid the error. Responsibility attribution theory suggests that when actors possess greater control and more resources for preventing an error, people may believe that they should have avoided the adverse outcome and consequently assign them greater responsibility [13,14]. Alternatively, if people view AI as a co-actor that directly participates in diagnostic judgment and influences the resulting decision, the diagnosis is no longer seen solely as the product of the doctor’s actions. AI involvement may then diffuse responsibility that would otherwise be concentrated on the doctor [15,16,17].
This question is especially important in healthcare, a context characterized by high stakes, substantial uncertainty, and heavy reliance on trust. Patients’ and observers’ perceived doctor responsibility may affect their evaluations of the doctor, trust repair, complaint intentions, and willingness to continue seeking care. Such judgments may also shape public acceptance of medical AI and influence how hospitals and regulators design AI oversight, information-disclosure, and responsibility-governance mechanisms [10,18,19]. Understanding how AI involvement in diagnosis changes perceived doctor responsibility therefore has both theoretical significance and important implications for healthcare governance.
Existing research has not clarified how or why AI participation changes perceived doctor responsibility when people perceive a diagnostic error. First, research on medical AI has focused primarily on whether patients accept or resist AI diagnosis and on trust repair and forgiveness following AI failures [20,21,22]. It has paid less attention to how the incorporation of AI into doctors’ decision processes changes perceived doctor responsibility. Second, although research on AI responsibility has extensively debated whether AI can qualify as a legally or ethically responsible entity, it has focused mainly on objective responsibility allocation rather than the psychological rules governing perceived responsibility in specific human–AI collaborative settings [19,23,24]. Third, research on public attitudes toward medical AI shows that patients are highly concerned with AI safety, human oversight, and whether doctors retain final discretion [18]. Yet it remains unclear how these judgments about AI’s role translate into perceived doctor responsibility when an adverse medical outcome occurs.
Drawing on responsibility attribution theory, this research examines how AI changes perceived doctor responsibility among patients and observers when they perceive a diagnostic error. Importantly, the focal construct is perceived responsibility rather than legal responsibility allocation. This research makes three contributions. First, it advances the explanatory logic of responsibility attribution theory in human–AI collaboration. Prior work generally assumes that the potential responsible parties are already known and examines how causal contribution, control, avoidability, and normative obligation shape responsibility allocation [13,14,25,26]. We show that in human–AI collaboration, people first determine whether AI is merely a passive tool or a co-actor that participates in producing the outcome. We thus propose a two-stage account comprising agent identification and responsibility allocation, and identify perceived shared agency as the key psychological mechanism through which AI shapes perceived human responsibility [27]. Second, this research advances understanding of responsibility judgments concerning human professionals in AI-enabled settings. In contrast to work that asks whether AI itself should bear legal, ethical, or moral responsibility, we show that even when AI does not qualify as a legally responsible entity, people may psychologically include it within the set of acting agents and accordingly adjust their judgments of a human professional’s responsibility. Third, this research broadens the outcomes examined in medical AI research. Moving beyond AI adoption, algorithm aversion, trust, and failure recovery, we reveal how AI involvement in diagnosis changes perceived doctor responsibility and offer a new perspective on the responsibility implications and governance boundaries of medical AI.

2. Theoretical Background and Hypothesis Development

2.1. Perceived Responsibility and Responsibility Attribution Theory

Responsibility has both normative and psychological meanings. The normative meaning concerns who should bear obligations and consequences under legal, institutional, or ethical standards, whereas the psychological meaning concerns how people determine who should be held responsible for an outcome in everyday judgment [28,29]. This distinction is particularly important in healthcare. Objective responsibility involves professional assessment, legal attribution, and institutional accountability, whereas perceived doctor responsibility reflects people’s subjective judgments of how much responsibility a doctor should bear [30,31,32,33]. Because this research examines such subjective judgments when a diagnostic error is perceived, it does not address objective medical accountability or determine whether AI constitutes a legally responsible entity. To avoid conflating these concepts, we refer to the focal construct as perceived doctor responsibility.
Specifically, perceived doctor responsibility refers to the extent to which an individual believes that a doctor should be held responsible after perceiving that a negative treatment outcome has occurred [25,31,34]. The construct is fundamentally an outcome of responsibility attribution. Although it does not correspond directly to a legal ruling, it can meaningfully affect evaluations, trust, complaints, intentions to seek accountability, and subsequent behavior [12,13,14]. More broadly, attribution research conceptualizes perceived responsibility as a subjective judgment that integrates an actor’s causal contribution, control, and moral blameworthiness following a negative outcome [26], with important consequences for subsequent behavior [11,35,36,37]. In healthcare, perceived doctor responsibility therefore captures the psychological responses of patients and observers more directly than objective responsibility and better explains attribution-related outcomes such as relationship rupture, declining trust, and persistent demands for accountability.
Responsibility attribution theory provides a foundational framework for understanding these judgments. When the acting party is known, the theory explains how the actor’s causal contribution, control, foreseeability, and degree of fault shape attributed responsibility [12,13,14,26,28]. In other words, when a negative outcome occurs, people do not mechanically assign all responsibility to the actor associated with that outcome. Instead, they consider whether the actor contributed to the error, could have prevented it, and played a pivotal role in the decision chain. With highly capable AI, particularly generative AI, the relevant set of actors in a collaborative decision is not always apparent in advance. People must first determine whether the final decision was made independently by the human actor or jointly by the human and AI. We therefore propose that when people perceive a diagnostic error, their judgment of doctor responsibility follows a two-stage process of agent identification and responsibility allocation. People first determine whether the final diagnosis represents the doctor’s sole agency or the shared agency of the doctor and AI, and then use that perceived agency structure to decide how much responsibility to assign to the doctor.

2.2. How AI Differs from Traditional Computerized Diagnostic Programs

Traditional computerized diagnostic programs generally rely on predefined rules, fixed procedures, or narrow task models to execute specified instructions and produce standardized recommendations. Their function is therefore closer to that of a programmed tool [38]. Although such programs can influence doctors’ judgments, their operating logic and output range are typically constrained by rules established in advance, making them easier to understand as tools invoked by doctors to assist with a task [39]. By contrast, contemporary AI systems, especially generative AI systems, can rely on large-scale pretraining and context-sensitive generation. They can continuously receive natural-language input, integrate information from multiple sources, generate diagnostic explanations, and revise or supplement their recommendations in response to follow-up questions [5]. The critical difference therefore lies not merely in processing capacity or output accuracy, but in how the systems participate in a task. Traditional programs primarily execute human-initiated operations under predefined rules, whereas such AI systems can be delegated relatively open-ended tasks and can formulate context-specific judgments under uncertainty [40].
This change in the mode of participation can also alter how people understand the system’s role. People need not believe that AI literally possesses human consciousness to perceive it as agentic. When a system responds specifically to a particular problem, generates coherent explanations, revises its recommendations across multiple rounds of interaction, and displays a relatively independent form of judgment, these features serve as social cues from which people infer thought, intention, and agency. The system is consequently more likely to be viewed as an actor with some degree of agency rather than as a passive technical device that merely executes instructions [41,42]. Here, describing AI as an acting agent does not imply that it has attained legal standing or moral responsibility. Rather, it means that people psychologically perceive AI as an entity that participates in judgment and influences the diagnostic outcome.
Medical diagnosis makes this agent-identification process especially salient. Diagnosis is not simply information retrieval; it is a high-consequence decision process that requires integrating evidence, comparing possibilities, and reaching a final judgment. Thus, when AI not only provides background data but also directly analyzes patient information, proposes a diagnosis, explains the basis for its judgment, and contributes an opinion that enters the doctor’s final diagnosis, its contribution to the diagnostic process becomes more identifiable. Although patients generally believe that doctors should retain final discretion and bear supervisory responsibility, they also recognize that AI recommendations may change doctors’ judgments, affect treatment outcomes, and expose patients to real risks [18]. These beliefs are not contradictory. The doctor may remain the final decision maker institutionally and legally, while AI is perceived as an agentic participant in the diagnostic process. As AI becomes increasingly involved in medical decision-making, questions of responsibility have become an increasingly important concern in medical AI applications [43,44].

2.3. Diagnostic Mode and Perceived Doctor Responsibility

Building on the preceding analysis, patients and observers in medical settings do not necessarily understand AI as merely a passive tool invoked by a doctor. When AI can analyze patient information, generate diagnostic recommendations, and inform the doctor’s judgment, these observable task capabilities offer important cues for identifying AI as an actor with some degree of agency [40,45]. Compared with doctor-only diagnosis, doctor–AI collaborative diagnosis is therefore more likely to lead people to believe that AI does more than provide technical support: it actively participates in diagnostic judgment and in producing the outcome. This belief may increase perceived shared agency in the diagnosis.
More specifically, perceived shared agency is the subjective judgment that a task is not performed independently by a single actor but jointly by a human and AI, with both making substantive contributions to the outcome [27]. The construct captures how people identify the acting agents in human–AI collaboration. In a doctor-only diagnosis, the analysis of diagnostic information, formation of the diagnostic judgment, and final diagnosis are all associated primarily with the doctor, so people perceive little involvement by another actor. In doctor–AI collaborative diagnosis, by contrast, AI independently analyzes patient information and generates a recommendation that the doctor receives and considers. People are therefore more likely to perceive both the doctor and AI as participants in the diagnostic process, each contributing to the final outcome. Accordingly, doctor–AI collaborative diagnosis should produce greater perceived shared agency than doctor-only diagnosis.
Perceived shared agency should, in turn, shape judgments of doctor responsibility. Responsibility attribution generally depends on an actor’s causal contribution to an outcome and the role the actor played in producing it [13,14,26]. When perceived shared agency is low, people are more likely to view the doctor as the sole or primary source of the diagnostic judgment and to attribute the diagnosis and any associated error primarily to the doctor. In this single-agent structure, the doctor’s causal contribution to the outcome is highly salient, concentrating responsibility on the doctor.
When perceived shared agency is high, however, people understand the diagnosis as an outcome jointly produced by the doctor and AI. The doctor is no longer seen as the sole or exclusive source of the diagnosis; AI’s analysis of patient information and its diagnostic recommendation are also incorporated into people’s understanding of how the outcome was produced. Although the doctor retains final discretion and the corresponding professional duties of review and supervision, the doctor is no longer perceived as the sole contributor, thereby reducing responsibility that would otherwise be concentrated on the doctor. Research on multi-actor responsibility similarly shows that when multiple actors jointly produce an outcome, people reassess and allocate responsibility according to each actor’s causal contribution [43,44]. We therefore propose:
H1. 
Diagnostic mode affects perceived doctor responsibility. Specifically, when people perceive that a diagnostic error has occurred, doctor–AI collaborative diagnosis, compared with doctor-only diagnosis, reduces perceived doctor responsibility.
H2. 
Perceived shared agency mediates the effect of diagnostic mode on perceived doctor responsibility. Specifically, when people perceive that a diagnostic error has occurred, doctor–AI collaborative diagnosis, compared with doctor-only diagnosis, increases perceived shared agency in the diagnosis, which in turn reduces perceived doctor responsibility.

2.4. Perceived Diagnostic Error Type as a Boundary Condition

The addition of another diagnostic actor does not necessarily reduce perceived doctor responsibility. Whether doctor–AI collaborative diagnosis diffuses perceived doctor responsibility depends critically on whether AI is perceived to have substantively contributed to the final diagnosis and thus increased perceived shared agency. Consequently, when the perceived shared agency created by AI involvement is weakened, its responsibility-reducing effect should also diminish.
In doctor–AI collaborative diagnosis, doctors may respond to AI advice in different ways. On the one hand, a doctor may rely too heavily on AI and incorporate an incorrect recommendation into the final diagnosis. On the other hand, a doctor may underuse AI and reject correct advice that could have prevented the error. Research on human–AI collaboration in medicine shows that doctors may be misled by incorrect AI advice or fail to improve a diagnosis because they disregard correct AI advice [46,47]. We therefore distinguish between two perceived sources of diagnostic error: the doctor accepts incorrect AI advice versus the doctor rejects correct AI advice.
When a doctor accepts incorrect AI advice, AI first forms a diagnostic judgment based on patient information, after which the doctor accepts that recommendation and incorporates it into the final decision. AI’s diagnostic input and the doctor’s acceptance and decision are thus substantively linked within the same diagnostic process. The final error reflects both AI’s recommendation and the doctor’s acceptance, making people more likely to view the erroneous diagnosis as jointly produced by the doctor and AI. This should heighten perceived shared agency and reduce perceived responsibility assigned to the doctor alone. By contrast, when the doctor rejects correct AI advice, AI participates in the diagnostic process and provides an opinion, but that advice is neither accepted nor incorporated into the final erroneous diagnosis. The final diagnosis primarily reflects the doctor’s own judgment rather than a joint product of the doctor’s and AI’s inputs. Relative to accepting incorrect AI advice, rejecting correct AI advice should therefore reduce perceived shared agency and weaken the responsibility-reducing effect of AI involvement. What matters for responsibility judgment is not AI’s mere formal presence in the diagnostic process, but whether AI substantively contributes to producing the final erroneous diagnosis. We therefore propose:
H3. 
Relative to doctor-only diagnosis, the indirect effect of doctor–AI collaborative diagnosis on perceived doctor responsibility through perceived shared agency is weaker when the doctor rejects correct AI advice than when the doctor accepts incorrect AI advice.

3. Overview of the Studies

To test the hypotheses systematically, we conducted five studies from both the patient and observer perspectives. Patients directly bear the adverse medical outcome, and their perceptions of doctor responsibility may therefore be accompanied by stronger physiological responses. Study 1A used event-related potentials (ERPs) to provide complementary physiological evidence of how AI involvement changes patients’ perceived doctor responsibility when they perceive a diagnostic error. Study 1B used a scenario experiment to provide a further test of this effect from the patient perspective. Study 1C then tested the mediating role of perceived shared agency and examined whether the relative indirect effect differed between the two perceived diagnostic-error conditions from the patient perspective.
Observers do not directly bear the adverse medical outcome and are therefore less likely to exhibit strong physiological responses. Study 2A used a scenario experiment to examine how AI involvement changes observers’ perceived doctor responsibility when they perceive a diagnostic error, thereby testing whether the effect extends beyond patients who directly experience harm. Study 2B further tested the mediating role of perceived shared agency and compared the relative indirect effects across the two perceived diagnostic-error conditions from the observer perspective. Table 1 summarizes the design and primary findings of the five studies.

4. Study 1A: An ERP Test from the Patient Perspective

Study 1A used event-related potentials (ERPs) as a complementary method to examine early neural responses to different diagnostic modes, thereby providing physiological corroboration for the association between perceived shared agency and perceived doctor responsibility. Specifically, when a diagnostic error occurs, patients not only evaluate the negative outcome but also judge which actors causally contributed to it and whether those actors could have controlled or prevented the error. Appraisal theories of emotion propose that when a negative outcome and its causal agent can be identified, evaluations of causality and responsibility shape both the type of emotional response and the target toward which it is directed [48,49]. Responsibility attribution theory further suggests that when an actor is perceived as able to control or prevent a negative outcome and is therefore assigned greater responsibility, people typically experience stronger negative emotions, such as anger, toward that actor. When the actor is assigned less responsibility, actor-directed negative emotion is correspondingly weaker [14,50]. Research on service failures likewise shows that the more consumers perceive a negative outcome as controllable by the service provider, the more responsibility they assign to that provider and the greater the anger they experience [51].
The P2 component is generally considered an ERP component associated with early attention and negative emotional responses, and it can capture the intensity of emotional reactions to negative outcomes [52]. Prior ERP research indicates that responsibility judgments can influence the early neural processing of negative outcomes. Li et al. (2011), for example, found that greater subjective responsibility for an outcome elicited a larger feedback-related error negativity (fERN), whose amplitude was significantly correlated with self-reported responsibility [53]. More directly, Pu and Yu (2019) found that negative outcomes elicited smaller P2 amplitudes in a no-responsibility condition than in shared- and full-responsibility conditions, indicating that responsibility level modulates early emotional-attentional processing of negative outcomes [54].
We do not treat P2 as a neural marker specific to responsibility attribution, but rather as a complementary electrophysiological indicator associated with responsibility judgment. Thus, if AI involvement reduces perceived doctor responsibility when the doctor is the explicit target of evaluation, a diagnostic error in the doctor–AI collaboration condition should elicit a smaller P2 amplitude than one in the doctor-only condition.

4.1. Participants and Procedure

Participants. Study 1A recruited participants at a Chinese university for an in-person ERP experiment. The study used a two-condition within-subjects design (diagnostic mode: doctor only vs. doctor–AI collaboration), such that each participant completed trials in both diagnostic modes. An a priori power analysis was conducted in G*Power 3.1 [55]. For the focal within-subject effect, the power analysis was conducted assuming a large effect size (f = 0.40), α = 0.05, power (1 − β) = 0.90, one group, two measurements, a correlation of 0.50 among repeated measures, and a nonsphericity correction of ε = 1. The analysis indicated a minimum required sample size of 19 participants. To allow for artifacts and unusable EEG data, we recruited 26 participants. Two did not proceed to formal EEG recording because electrode impedance could not be reduced to the required range, yielding a final sample of 24 participants (Mage = 23.96, SD = 1.49; 50% women).
Procedure. Before the experiment, participants read and signed an informed consent form. The experimenter then fitted each participant with an EEG cap and prepared the electrodes, ensuring that impedance met the required threshold before formal recording. Before the task, participants read background information about an ophthalmic condition and the use of AI in ophthalmic diagnosis. The experimenter explained the material, and a comprehension check confirmed that participants understood the disease, the AI-diagnosis context, and the task requirements. Participants then began the ERP task. The instructions read: “Imagine that you recently developed an eye condition and went to a hospital for examination. The diagnosis indicated that you needed eye surgery. Your condition did not improve after the surgery, so you believe the diagnosis was incorrect. Please judge how much responsibility the doctor should bear.” During the task, the system randomly presented two types of diagnostic-mode stimuli: “This diagnosis was made by a doctor” and “This diagnosis was made collaboratively by a doctor and AI.” Each trial began with a 100 ms fixation point, followed by a blank screen of random duration (500–800 ms) and a diagnostic-mode display presented for up to 5000 ms. Participants were instructed to respond while the diagnostic-mode display remained on screen. Upon response, the display disappeared immediately, followed by another blank screen of random duration (500–800 ms) before the next trial began. All participants used the same response mapping: the “F” key indicated that the doctor should bear full responsibility, whereas the “J” key indicated that the doctor should not bear full responsibility. Keypresses and EEG data were recorded simultaneously. Each diagnostic-mode condition was presented 48 times, yielding 96 trials in total. The stimuli were presented in a random order using E-Prime. Doctor names varied across trials, whereas the AI was consistently named DeepSeek to ensure a consistent AI-agent cue. Figure 1 illustrates the sequence of a single trial.

4.2. Data Collection and Processing

EEG was recorded using a NeuroScan EEG acquisition system and a 64-channel Ag/AgCl electrode cap positioned according to the extended international 10–20 system. The ground electrode was located midway between FCz and Fz. Vertical and horizontal electrooculograms (VEOG and HEOG) were recorded simultaneously to monitor ocular activity. EEG signals were sampled at 1000 Hz with an online band-pass filter of 0.01–100 Hz, and electrode impedances were maintained below 5 kΩ throughout data acquisition.
Following data acquisition, the EEG data were processed offline using CURRY 7. Continuous EEG data were re-referenced offline to the algebraic average of the left and right mastoids (M1 and M2). Continuous EEG data were filtered offline using a 0.01 Hz high-pass filter and a 30 Hz low-pass filter, both with a slope of 48 dB/oct, together with a 50 Hz notch filter to attenuate line noise. Filtering was performed on the continuous data before epoching. VEOG deflections with an absolute amplitude of 100 μV or greater were identified as blink or vertical eye-movement artifacts and corrected using the Gratton and Coles method implemented in CURRY 7. The continuous EEG data were then segmented into epochs time-locked to stimulus onset, extending from −200 to 800 ms. Baseline correction was performed using the mean voltage during the −200 to 0 ms prestimulus interval. Epochs in which the absolute amplitude exceeded 100 μV at any scalp electrode were excluded. Following artifact rejection, at least 75% of the trials were retained for every participant. The retained artifact-free epochs were averaged separately for the doctor-only and doctor–AI collaboration conditions to obtain the corresponding ERP waveforms. The preprocessing pipeline followed established ERP procedures [56].
Based on the characteristic frontocentral scalp distribution of the P2 component and prior ERP research [57,58], P2 was quantified as the mean amplitude during the 120–200 ms poststimulus interval at Fz, FC1, FCz, FC2, C1, Cz, C2, and CPz. P2 mean amplitudes were submitted to a 2 (diagnostic mode: doctor only vs. doctor–AI collaboration) × 8 (electrode) repeated-measures analysis of variance, with both diagnostic mode and electrode treated as within-subject factors. Mauchly’s test of sphericity was conducted for effects involving the electrode factor, and Greenhouse–Geisser corrections were applied when the sphericity assumption was violated. Additionally, P2 mean amplitudes averaged across the eight electrodes were compared between the two diagnostic modes using a paired-samples t test.

4.3. Results

Behavioral results. For each condition, we calculated the proportion of trials in which participants selected “the doctor should bear full responsibility” by dividing the number of F-key responses by the 48 trials presented in that condition. This proportion served as the direct behavioral measure of full-responsibility judgments. A paired-samples t test showed that the proportion of full-responsibility judgments was significantly higher in the doctor-only condition (M = 0.65, SD = 0.15) than in the doctor–AI collaboration condition (M = 0.37, SD = 0.14), t(23) = 6.921, p < 0.001, Cohen’s dz = 1.41.
EEG results. Mean P2 amplitudes were analyzed using a 2 × 8 repeated-measures analysis of variance with diagnostic mode (doctor only vs. doctor–AI collaboration) and electrode (Fz, FC1, FCz, FC2, C1, Cz, C2, and CPz) as within-participant factors. The participant-level difference scores in P2 amplitude between the two diagnostic modes, averaged across the eight electrodes, did not significantly deviate from normality, Shapiro–Wilk W = 0.933, p = 0.114. Because the diagnostic mode had only two levels, the sphericity assumption was not applicable to its main effect. Mauchly’s test indicated that sphericity was violated for the main effect of electrode, W < 0.001, χ2(27) = 144.77, p < 0.001, and for the diagnostic mode × electrode interaction, W < 0.001, χ2(27) = 154.00, p < 0.001. Greenhouse–Geisser-corrected results are therefore reported for the main effect of electrode and the diagnostic mode × electrode interaction.
The main effect of diagnostic mode was significant, F (1, 23) = 58.054, p < 0.001, partial η2 = 0.716. Mean P2 amplitude was significantly lower in the doctor–AI collaboration condition (M = 1.72 μV, SD = 2.56) than in the doctor-only condition (M = 3.98 μV, SD = 2.47). The main effect of electrode was also significant, F (2.18, 50.23) = 4.401, p = 0.015, partial η2 = 0.161, Greenhouse–Geisser ε = 0.312. The diagnostic mode × electrode interaction was not significant, F (2.27, 52.29) = 0.352, p = 0.732, partial η2 = 0.015, Greenhouse–Geisser ε = 0.325. Thus, there was no evidence that the effect of diagnostic mode on P2 amplitude differed across the eight electrodes. Because the diagnostic mode had only two levels, its ANOVA main effect was algebraically equivalent to a paired-samples t test. We report the corresponding paired contrast to provide an interpretable effect estimate and confidence interval. P2 mean amplitude was 2.26 μV higher in the doctor-only condition than in the doctor–AI collaboration condition, 95% CI [1.65, 2.88], t(23) = 7.619, p < 0.001, Cohen’s dz = 1.56, 95% CI [0.95, 2.15]. Figure 2 presents the ERP waveforms by diagnostic mode.

4.4. Discussion

Using behavioral choices and physiological data, Study 1A provided initial evidence that when patients perceive a diagnostic error, doctor–AI collaborative diagnosis reduces perceived doctor responsibility relative to doctor-only diagnosis. However, although the physiological evidence from Study 1A captures patients’ real-time responses, and the observed difference in P2 suggests physiological processing associated with responsibility perceptions or judgments, P2 may also reflect early visual and attentional processing. It therefore cannot be interpreted as a direct measure of perceived responsibility. Accordingly, Study 1A provides only complementary physiological evidence, and the relationship between diagnostic mode and patients’ perceived doctor responsibility requires further validation.

5. Study 1B: Testing the Main Effect from the Patient Perspective

To provide a further test of H1, Study 1B used a scenario experiment to directly measure perceived doctor responsibility among participants adopting the patient perspective when they believed a diagnostic error had occurred. We adapted established measures from prior research to fit the scenario [17,28]. The final four-item measure of perceived doctor responsibility (α = 0.790, AVE = 0.616) comprised the following items: (1) “How responsible is this doctor for the diagnostic error?”; (2) “To what extent is this doctor at fault for the diagnostic error?”; (3) “This doctor could have completely avoided the diagnostic error”; and (4) “This doctor could have completely prevented the diagnostic error.” Participants responded on seven-point scales, with higher scores indicating greater perceived doctor responsibility.

5.1. Participants and Procedure

Manipulation pretest. Before the main experiment, we recruited 100 Chinese residents through Credamo (Mage = 33.54, SD = 10.52; 57% women) and randomly assigned them to the doctor-only condition or doctor–AI collaboration condition (n = 50 per condition). After reading the corresponding materials, participants answered the question, “Which actors or systems participated in this diagnosis?” The proportion selecting “doctor only” differed significantly between the doctor-only and doctor–AI collaboration conditions (96% vs. 0%), χ2(1) = 92.308, p < 0.001, φ = 0.961. Specifically, in the doctor-only condition, 48 participants selected “doctor only” and two selected “doctor and AI”; in the doctor–AI collaboration condition, all 50 participants selected “doctor and AI.” These results confirmed the effectiveness of the diagnostic-mode manipulation and supported the use of the materials in the main experiment.
Participants. The study used a two-condition between-subjects design (diagnostic mode: doctor only vs. doctor–AI collaboration). An a priori power analysis was conducted in G*Power 3.1 [55]. Assuming a two-tailed independent-samples test, a medium effect size (d = 0.50), α = 0.05, and power (1 − β) = 0.90, the analysis indicated a minimum sample of 172 participants. We therefore recruited 200 Chinese residents through the Credamo online platform (Mage = 28.47, SD = 7.45; 51.5% women). Participants were randomly assigned to the doctor-only condition or doctor–AI collaboration condition (n = 100 per condition).
Procedure. The procedure was largely identical to that of Study 1A, except that perceived responsibility was measured with a scale and the stimuli were not repeated. The instructions closely followed those used in Study 1A: “Imagine that you recently developed an eye condition and went to a hospital for examination. The diagnosis indicated that you needed eye surgery. Your condition did not improve after the surgery, so you believe the diagnosis was incorrect. Please judge how much responsibility the doctor should bear.” The final sentence stated either “The diagnosis was made by a doctor alone” or “The diagnosis was made collaboratively by a doctor and an AI system.”

5.2. Results

An independent-samples t test showed that patients reported significantly higher perceived doctor responsibility in the doctor-only condition (M = 6.34, SD = 0.41) than in the doctor–AI collaboration condition (M = 5.24, SD = 0.70), with a mean difference of 1.10, 95% CI [0.94, 1.27], t(198) = 13.607, p < 0.001, Cohen’s d = 1.92. Figure 3 presents the results.

5.3. Discussion

The results of Study 1B provide further support for H1. From the patient perspective, when a diagnostic error was perceived, doctor–AI collaborative diagnosis reduced perceived doctor responsibility relative to doctor-only diagnosis. Study 1B did not, however, examine the proposed mediating mechanism or boundary condition. Study 1C therefore tests both the mechanism through which diagnostic mode shapes patients’ perceived doctor responsibility and the boundary condition on this effect.

6. Study 1C: Perceived Diagnostic Error Type as a Boundary Condition

Study 1C tested H2 and H3. First, we examined whether diagnostic mode changes patients’ perceived doctor responsibility through perceived shared agency. Second, we tested whether the relative indirect effect differs across the two perceived diagnostic-error conditions, thereby providing a further test of the mediating role of perceived shared agency. We distinguished two forms of error in doctor–AI collaborative diagnosis. In the first condition, the doctor rejects correct AI advice: AI makes the correct judgment, but the doctor does not follow it. In the second condition, the doctor accepts incorrect AI advice: AI makes an incorrect judgment, which the doctor accepts in forming the final diagnosis. Similar distinctions have been used in research on responsibility in human–AI collaboration and autonomous driving [16,59].

6.1. Participants and Procedure

Participants. The study used a four-condition between-subjects design: doctor only, doctor–AI collaboration, doctor–AI collaboration in which the doctor accepted incorrect AI advice, and doctor–AI collaboration in which the doctor rejected correct AI advice. An a priori power analysis in G*Power 3.1 [55], assuming a medium effect size (f = 0.25), α = 0.05, and power (1 − β) = 0.90, indicated a minimum sample of 232 participants. We recruited 400 Chinese residents through Credamo (Mage = 32.56, SD = 9.91; 64% women). Participants were randomly assigned to the four conditions (n = 100 per condition).
Procedure. The procedure was largely identical to that of Study 1B. The stimuli differed in three main respects. First, the AI was no longer given a specific name but was described as a generative AI system deployed by the hospital, and the medical condition was changed from an eye disorder to a more readily understood thyroid condition. Second, we added the doctor-rejects-correct-AI-advice and doctor-accepts-incorrect-AI-advice conditions. In the latter condition, the doctor–AI collaboration stimulus was followed by: “The AI recommended surgery because surgery could effectively treat this thyroid condition. The doctor accepted this recommendation and concluded that you needed surgery.” In the former condition, it was followed by: “The AI recommended against surgery because surgery could not effectively treat this thyroid condition. However, the doctor rejected this recommendation and concluded that you needed surgery.” Third, after the measure of perceived doctor responsibility (α = 0.866, AVE = 0.737), participants completed a four-item measure of perceived shared agency: “The attending doctor and another diagnostic agent jointly produced this diagnosis”; “The attending doctor and another diagnostic agent both made substantive contributions to producing this diagnosis”; “This diagnosis resulted from the joint judgment of the attending doctor and another diagnostic agent”; and “In forming this diagnosis, the attending doctor was not the only agent that determined the outcome” (α = 0.946, AVE = 0.861). Participants responded on seven-point scales, with higher scores indicating greater perceived shared agency. Discriminant validity was also supported. The square roots of the AVEs for perceived doctor responsibility (0.858) and perceived shared agency (.928) both exceeded the absolute correlation between the two constructs (r = −0.699), and the HTMT ratio was 0.752. Participants then completed manipulation checks. All conditions included the question “Which actors or systems participated in this diagnosis?” with the response options “doctor only” or “doctor and AI.” The two conditions manipulating perceived diagnostic error type also included the question “During the diagnostic process, did the doctor accept or reject the AI’s opinion?” with the response options “accept” and “reject.” Finally, as a supplementary self-report check for potential demand characteristics, participants indicated whether, while answering the preceding questions, they had immersed themselves in the scenario and selected responses that genuinely reflected their views rather than attempting to infer the research purpose. The response options were “yes” and “no.”

6.2. Results

Manipulation and demand-characteristics checks. The proportion selecting “doctor only” in response to “Which actors or systems participated in this diagnosis?” differed significantly across the four conditions (doctor only: 97%; doctor–AI collaboration: 2%; doctor-accepts-incorrect-AI-advice condition: 1%; doctor-rejects-correct-AI-advice condition: 4%), χ2(3) = 349.584, p < 0.001, Cramer’s V = 0.935, indicating that the diagnostic-mode manipulation was successful. The error-type manipulation was also successful. The proportion indicating that the doctor accepted the AI’s opinion was 99% in the doctor-accepts-incorrect-AI-advice condition and 9% in the doctor-rejects-correct-AI-advice condition, χ2(1) = 163.043, p < 0.001, φ = 0.903. Finally, 398 of the 400 participants reported that they had responded based on their genuine impressions of the scenario rather than attempting to infer the research purpose. These responses provide supplementary, although not conclusive, evidence against widespread conscious hypothesis guessing.
Perceived doctor responsibility. A one-way analysis of variance revealed significant differences across the four conditions: doctor only (M = 6.35, SD = 0.60), doctor–AI collaboration (M = 4.87, SD = 0.89), doctor-accepts-incorrect-AI-advice condition (M = 5.03, SD = 0.98), and doctor-rejects-correct-AI-advice condition (M = 6.05, SD = 0.79), F (3, 396) = 78.747, p < 0.001, partial η2 = 0.374. Levene’s test indicated that the homogeneity-of-variance assumption was violated, F (3, 396) = 4.222, p = 0.006. We therefore conducted post hoc pairwise comparisons using Tamhane’s T2 procedure. Tamhane’s post hoc tests showed that perceived doctor responsibility was significantly higher in the doctor-only condition than in the doctor–AI collaboration condition (95% CI [1.1906, 1.7644], p < 0.001), confirming that AI collaboration reduced patients’ perceived doctor responsibility. The doctor-only condition was also significantly higher than both the doctor-accepts-incorrect-AI-advice condition (95% CI [1.0141, 1.6259], p < 0.001) and the doctor-rejects-correct-AI-advice condition (95% CI [0.0364, 0.5636], p = 0.017). Perceived doctor responsibility was significantly lower when the doctor accepted incorrect AI advice than when the doctor rejected correct AI advice (95% CI [−1.3537, −0.6863], p < 0.001). Thus, perceived diagnostic error type attenuated the responsibility-reducing effect of AI involvement. The omnibus difference remained significant after controlling for gender and age, F (3, 394) = 78.400, p < 0.001, partial η2 = 0.374. Figure 4 presents patients’ perceived doctor responsibility across the four diagnostic conditions.
Perceived shared agency. A one-way analysis of variance revealed significant differences across the four conditions: doctor only (M = 1.81, SD = 1.17), doctor–AI collaboration (M = 5.87, SD = 0.81), doctor-accepts-incorrect-AI-advice condition (M = 5.75, SD = 0.70), and doctor-rejects-correct-AI-advice condition (M = 2.87, SD = 1.44), F (3, 396) = 366.141, p < 0.001, partial η2 = 0.735. Levene’s test indicated that the homogeneity-of-variance assumption was violated, F (3, 396) = 15.235, p < 0.001. We therefore conducted post hoc pairwise comparisons using Tamhane’s T2 procedure. Tamhane’s post hoc tests showed that perceived shared agency was significantly lower in the doctor-only condition than in the doctor–AI collaboration condition (95% CI [−4.4418, −3.6832], p < 0.001), indicating that AI collaboration increased patients’ perceived shared agency. The doctor-only condition was also significantly lower than both the doctor-accepts-incorrect-AI-advice condition (95% CI [−4.3095, −3.5805], p < 0.001) and the doctor-rejects-correct-AI-advice condition (95% CI [−1.5539, −0.5661], p < 0.001). Perceived shared agency was significantly higher when the doctor accepted incorrect AI advice than when the doctor rejected correct AI advice (95% CI [2.4574, 3.3126], p < 0.001). Thus, perceived diagnostic error type attenuated the effect of AI involvement on perceived shared agency. The omnibus difference remained significant after controlling for gender and age, F (3, 394) = 355.609, p < 0.001, partial η2 = 0.730.
Mediation analysis. Using participants in the doctor-only and doctor–AI collaboration conditions (n = 200), we conducted a bootstrapping analysis using PROCESS Model 4 [60] with 5000 bootstrap samples to test whether perceived shared agency mediated the effect of diagnostic mode on perceived doctor responsibility. Diagnostic mode was coded 0 = doctor only and 1 = doctor–AI collaboration, perceived doctor responsibility was the dependent variable, and perceived shared agency was the mediator. The indirect effect through perceived shared agency was significant (b = −1.3862, SE = 0.2727, 95% CI [−1.9295, −0.8394]), whereas the direct effect was not (b = −0.0913, SE = 0.2178, 95% CI [−0.5208, 0.3383]). This result is consistent with an indirect pathway through perceived shared agency, providing support for H2.
Multicategorical mediation analysis and comparison of relative indirect effects. To test H3, we treated diagnostic condition as a multicategorical independent variable and used indicator coding in PROCESS Model 4 with 5000 bootstrap samples. We first used the doctor-only condition as the reference category. The three indicator variables represented doctor–AI collaboration (X1), doctor-accepts-incorrect-AI-advice condition (X2), and doctor-rejects-correct-AI-advice condition (X3). With perceived doctor responsibility as the dependent variable and perceived shared agency as the mediator, the relative indirect effect for the doctor-accepts-incorrect-AI-advice condition versus the doctor-only condition was significant (b = −1.3190, SE = 0.1541, 95% CI [−1.6362, −1.0145]), whereas the corresponding direct effect was not significant (b = −0.0010, SE = 0.1737, 95% CI [−0.3426, 0.3405]). The relative indirect effect for the doctor-rejects-correct-AI-advice condition versus the doctor-only condition was also significant but smaller (b = −0.3544, SE = 0.0751, 95% CI [−0.5105, −0.2116]), whereas the corresponding direct effect was not significant (b = 0.0544, SE = 0.1119, 95% CI [−0.1656, 0.2744]). To directly compare the relative indirect effects across the two perceived diagnostic-error conditions, we re-estimated the model using the doctor-rejects-correct-AI-advice condition as the reference category. The resulting indicator variables represented doctor only (D1), doctor–AI collaboration (D2), and doctor-accepts-incorrect-AI-advice condition (D3). The relative indirect effect associated with D3, which represents the contrast between the doctor-accepts-incorrect-AI-advice condition and the doctor-rejects-correct-AI-advice condition, was significant (b = −0.9646, SE = 0.1155, 95% CI [−1.1990, −0.7411]), whereas the corresponding direct effect was not significant (b = −0.0554, SE = 0.1460, 95% CI [−0.3425, 0.2317]). Thus, the responsibility-reducing indirect effect through perceived shared agency was significantly weaker when the doctor rejected correct AI advice than when the doctor accepted incorrect AI advice, supporting H3. Different types of perceived diagnostic error significantly altered patients’ levels of perceived shared agency, with the magnitude of the corresponding indirect effect varying accordingly. This result provides further evidence consistent with an indirect pathway through perceived shared agency.
Robustness analysis. As a robustness check, we repeated the key analyses using only the two items that directly assessed the doctor’s responsibility and fault (r = 0.829). The indirect effect of diagnostic mode on perceived doctor responsibility through perceived shared agency remained significant (b = −1.1004, SE = 0.2670, 95% CI [−1.6011, −0.5421]). Moreover, the difference in relative indirect effects between the doctor-accepts-incorrect-AI-advice and doctor-rejects-correct-AI-advice conditions remained significant (b = −0.8124, SE = 0.1218, 95% CI [−1.0547, −0.5730]). Thus, the substantive conclusions were unchanged when perceived doctor responsibility was operationalized using only the direct responsibility and fault items.

6.3. Discussion

Study 1C showed that, from the patient perspective, the effect of diagnostic mode on perceived doctor responsibility was consistent with an indirect pathway through perceived shared agency, and that perceived diagnostic error type moderated this indirect effect. Specifically, when AI participated in diagnosis, patients attributed some degree of agency to it: they perceived AI as capable of forming its own diagnostic judgment and making a substantive contribution to the final diagnosis, and thus regarded the doctor and AI as joint actors. Consistent with an indirect pathway through perceived shared agency, higher perceived shared agency under AI involvement was associated with lower perceived doctor responsibility. Furthermore, the study altered patients’ levels of perceived shared agency by manipulating the type of perceived diagnostic error. Compared with the condition in which the doctor accepted incorrect AI advice, perceived shared agency was lower when the doctor rejected correct AI advice, and both the corresponding indirect effect and the responsibility-reducing effect of AI involvement were weaker. These findings provide further evidence consistent with an indirect pathway through perceived shared agency. Study 1C nevertheless adopted the perspective of patients who directly experienced harm. Their responsibility judgments may have been influenced by personal loss, emotional involvement, and expectations of compensation. The subsequent studies therefore shift to the observer perspective to determine whether the effect extends beyond patients’ immediate self-interest.

7. Study 2A: Testing the Main Effect from the Observer Perspective

Study 2A retested H1 from the observer perspective. Unlike patients, observers do not directly experience the perceived diagnostic error. Their responsibility judgments should therefore involve less personal engagement and more closely resemble public opinion or third-party evaluation. If the main effect also emerges among observers, the diffusion of perceived doctor responsibility by AI cannot be explained solely as a self-protective cognition among patients; instead, it has broader implications for social judgment. Because observers are also less likely to exhibit strong physiological responses, Study 2A used a scenario experiment.

7.1. Participants and Procedure

Manipulation pretest. Before the main experiment, we recruited 100 Chinese residents through Credamo (Mage = 32.98, SD = 10.55; 57% women) and randomly assigned them to the doctor-only condition or doctor–AI collaboration condition (n = 50 per condition). After reading the corresponding observer scenario, participants answered the question, “Which actors or systems participated in this diagnosis?” The proportion selecting “doctor only” differed significantly between the doctor-only and doctor–AI collaboration conditions (94% vs. 4%), χ2(1) = 81.032, p < 0.001, φ = 0.900. Specifically, in the doctor-only condition, 47 participants selected “doctor only” and three selected “doctor and AI”; in the doctor–AI collaboration condition, two participants selected “doctor only” and 48 selected “doctor and AI.” These results confirmed the effectiveness of the diagnostic-mode manipulation and supported the use of the materials in the main experiment.
Participants. The study used a two-condition between-subjects design (diagnostic mode: doctor only vs. doctor–AI collaboration). Assuming a two-tailed independent-samples t test, a medium effect size (d = 0.50), α = 0.05, and power (1 − β) = 0.90, the analysis indicated a minimum sample of 172 participants. We recruited 200 Chinese residents through Credamo (Mage = 27.37, SD = 6.99; 68.5% women). Participants were randomly assigned to the doctor-only condition or doctor–AI collaboration condition (n = 100 per condition).
Procedure. The procedure was largely identical to that of Study 1B, except that the instructions adopted an observer perspective: “Imagine that you recently saw a news story on Weibo. A patient with an eye condition went to a hospital for examination, and the diagnosis indicated that the patient needed eye surgery. After the surgery, the patient’s eye condition did not improve. You believe that a diagnostic error occurred. Please judge how much responsibility the doctor should bear.” The final sentence stated either “The diagnosis was made by a doctor alone” or “The diagnosis was made collaboratively by a doctor and an AI system.” The perceived doctor responsibility scale demonstrated satisfactory reliability and convergent validity (α = 0.835, AVE = 0.670).

7.2. Results

An independent-samples t test showed that observers reported significantly higher perceived doctor responsibility in the doctor-only condition (M = 6.31, SD = 0.38) than in the doctor–AI collaboration condition (M = 5.08, SD = 0.72), with a mean difference of 1.23, 95% CI [1.07, 1.39], t(198) = 15.136, p < 0.001, Cohen’s d = 2.14. Thus, AI participation significantly reduced observers’ perceived doctor responsibility relative to doctor-only diagnosis. Figure 5 presents observers’ perceived doctor responsibility by diagnostic mode.

7.3. Discussion

Study 2A again supported H1 and showed that the effect extends to observers. In other words, not only patients but also third-party observers assign less responsibility to a doctor when AI is involved. This finding gives the results more direct implications for public discourse, social evaluation, and healthcare governance. The attenuation of perceived doctor responsibility is therefore unlikely to arise solely from patients’ psychological efforts to cope with their own victimization. Instead, it appears to reflect a more general social-attribution pattern: when third parties learn that a diagnosis was made collaboratively by a doctor and AI, they also incorporate AI into their responsibility-attribution framework. Like Study 1B, however, Study 2A tested only the overall main effect of collaborative diagnosis and did not examine the proposed mediating mechanism or boundary condition.

8. Study 2B: Perceived Diagnostic Error Type as a Boundary Condition from the Observer Perspective

Study 2B provided a further test of H2 and H3 from the observer perspective. We examined whether diagnostic mode changes observers’ perceived doctor responsibility through perceived shared agency and whether the relative indirect effect through perceived shared agency differed between the two perceived diagnostic-error conditions, thereby testing perceived diagnostic error type as a boundary condition.

8.1. Participants and Procedure

Participants. The study again used a four-condition between-subjects design: doctor only, doctor–AI collaboration, doctor–AI collaboration in which the doctor accepted incorrect AI advice, and doctor–AI collaboration in which the doctor rejected correct AI advice. Assuming a one-way analysis of variance with four groups, a medium effect size (f = 0.25), α = 0.05, and power (1 − β) = 0.90, the analysis indicated a minimum sample of 232 participants. Accordingly, we recruited 400 participants from countries other than China through CloudResearch. Of these participants, 82.25% were from the United States, and the remainder were from Canada, the United Kingdom, Australia, and other countries (Mage = 38.09, SD = 12.41; 54.3% women). Participants were randomly assigned to the four conditions (n = 100 per condition).
Procedure. The procedure was largely identical to that of Study 1C, except that the four scenarios adopted an observer perspective and involved a knee condition: “Imagine that you recently saw a news story. A patient with a knee condition went to a hospital for examination, and the diagnosis indicated that the patient needed knee arthroscopy. After the surgery, the patient’s knee condition did not improve. You believe that a diagnostic error occurred. Please judge how much responsibility the doctor should bear.” The final sentence stated either “The diagnosis was made by a doctor alone” or “The diagnosis was made collaboratively by a doctor and an AI system.” The stimuli in the two conditions manipulating perceived diagnostic error type followed those used in Study 1C. Both the perceived doctor responsibility scale (α = 0.969, AVE = 0.914) and the perceived shared agency scale (α = 0.916, AVE = 0.803) exhibited good reliability and convergent validity. Discriminant validity was also supported. The square roots of the AVEs for perceived doctor responsibility (0.956) and perceived shared agency (0.896) both exceeded the absolute correlation between the two constructs (r = −0.593), and the HTMT ratio was 0.630.

8.2. Results

Manipulation and demand-characteristics checks. The proportion selecting “doctor only” in response to “Which actors or systems participated in this diagnosis?” differed significantly across the four conditions (doctor only: 88%; doctor–AI collaboration: 6%; doctor-accepts-incorrect-AI-advice condition: 8%; doctor-rejects-correct-AI-advice condition: 15%), χ2(3) = 224.542, p < 0.001, Cramer’s V = 0.749, indicating that the diagnostic-mode manipulation was successful. The error-type manipulation was also successful. The proportion indicating that the doctor accepted the AI’s opinion was 98% in the doctor-accepts-incorrect-AI-advice condition and 19% in the doctor-rejects-correct-AI-advice condition, χ2(1) = 128.535, p < 0.001, φ = 0.802. Finally, 396 of the 400 participants reported that they had responded based on their genuine impressions of the scenario rather than attempting to infer the research purpose. These responses provide supplementary, although not conclusive, evidence against widespread conscious hypothesis guessing.
Perceived doctor responsibility. A one-way analysis of variance revealed significant differences across the four conditions: doctor only (M = 6.31, SD = 1.01), doctor–AI collaboration (M = 4.74, SD = 1.65), doctor-accepts-incorrect-AI-advice condition (M = 4.94, SD = 1.70), and doctor-rejects-correct-AI-advice condition (M = 5.93, SD = 1.44), F (3, 396) = 26.416, p < 0.001, partial η2 = 0.167. Levene’s test indicated that the homogeneity-of-variance assumption was violated, F (3, 396) = 15.372, p < 0.001. We therefore conducted post hoc pairwise comparisons using Tamhane’s T2 procedure. Tamhane’s post hoc tests showed that perceived doctor responsibility was significantly higher in the doctor-only condition than in the doctor–AI collaboration condition (95% CI [1.0560, 2.0890], p < 0.001), indicating that AI collaboration reduced observers’ perceived doctor responsibility. The doctor-only condition was also significantly higher than the doctor-accepts-incorrect-AI-advice condition (95% CI [0.8451, 1.8999], p < 0.001), but did not differ significantly from the doctor-rejects-correct-AI-advice condition (95% CI [−0.0820, 0.8570], p = 0.164). Perceived doctor responsibility was significantly lower when the doctor accepted incorrect AI advice than when the doctor rejected correct AI advice (95% CI [−1.5777, −0.3923], p < 0.001). Thus, perceived diagnostic error type attenuated the responsibility-reducing effect of AI involvement. The omnibus difference remained significant after controlling for gender and age, F (3, 394) = 26.591, p < 0.001, partial η2 = 0.168.
Perceived shared agency. A one-way analysis of variance revealed significant differences across the four conditions: doctor only (M = 2.60, SD = 1.81), doctor–AI collaboration (M = 5.44, SD = 1.24), doctor-accepts-incorrect-AI-advice condition (M = 5.29, SD = 1.54), and doctor-rejects-correct-AI-advice condition (M = 3.02, SD = 1.52), F (3, 396) = 93.400, p < 0.001, partial η2 = 0.414. Levene’s test indicated that the homogeneity-of-variance assumption was violated, F (3, 396) = 7.071, p < 0.001. We therefore conducted post hoc pairwise comparisons using Tamhane’s T2 procedure. Tamhane’s post hoc tests showed that perceived shared agency was significantly lower in the doctor-only condition than in the doctor–AI collaboration condition (95% CI [−3.4231, −2.2569], p < 0.001), indicating that AI collaboration increased observers’ perceived shared agency. The doctor-only condition was also significantly lower than the doctor-accepts-incorrect-AI-advice condition (95% CI [−3.3179, −2.0571], p < 0.001), but did not differ significantly from the doctor-rejects-correct-AI-advice condition (95% CI [−1.0406, 0.2156], p = 0.403). Perceived shared agency was significantly higher when the doctor accepted incorrect AI advice than when the doctor rejected correct AI advice (95% CI [1.7005, 2.8495], p < 0.001). Thus, perceived diagnostic error type attenuated the effect of AI involvement on perceived shared agency. The omnibus difference remained significant after controlling for gender and age, F (3, 394) = 94.225, p < 0.001, partial η2 = 0.418.
Mediation analysis. Using participants in the doctor-only and doctor–AI collaboration conditions (n = 200), we conducted a bootstrapping analysis using PROCESS Model 4 [60] with 5000 bootstrap samples to test whether perceived shared agency mediated the effect of diagnostic mode on perceived doctor responsibility. Diagnostic mode was coded 0 = doctor only and 1 = doctor–AI collaboration, perceived doctor responsibility was the dependent variable, and perceived shared agency was the mediator. The indirect effect through perceived shared agency was significant (b = −1.2013, SE = 0.2151, 95% CI [−1.6718, −0.8288]), whereas the direct effect was not (b = −0.3712, SE = 0.2322, 95% CI [−0.8292, 0.0867]). This result is consistent with an indirect pathway through perceived shared agency, providing support for H2.
Multicategorical mediation analysis and comparison of relative indirect effects. To test H3, we treated diagnostic condition as a multicategorical independent variable and used indicator coding in PROCESS Model 4 with 5000 bootstrap samples. We first used the doctor-only condition as the reference category. The three indicator variables represented doctor–AI collaboration (X1), doctor-accepts-incorrect-AI-advice condition (X2), and doctor-rejects-correct-AI-advice condition (X3). With perceived doctor responsibility as the dependent variable and perceived shared agency as the mediator, the relative indirect effect for the doctor-accepts-incorrect-AI-advice condition versus the doctor-only condition was significant (b = −1.2267, SE = 0.1762, 95% CI [−1.5974, −0.9143]), whereas the corresponding direct effect was not significant (b = −0.1458, SE = 0.2166, 95% CI [−0.5717, 0.2800]). The relative indirect effect for the doctor-rejects-correct-AI-advice condition versus the doctor-only condition was not significant (b = −0.1883, SE = 0.1108, 95% CI [−0.4234, 0.0158]), whereas the corresponding direct effect was also not significant (b = −0.1992, SE = 0.1849, 95% CI [−0.5627, 0.1643]). To directly compare the relative indirect effects across the two perceived diagnostic-error conditions, we re-estimated the model using the doctor-rejects-correct-AI-advice condition as the reference category. The resulting indicator variables represented doctor only (D1), doctor–AI collaboration (D2), and doctor-accepts-incorrect-AI-advice condition (D3). The relative indirect effect associated with D3, which represents the contrast between the doctor-accepts-incorrect-AI-advice condition and the doctor-rejects-correct-AI-advice condition, was significant (b = −1.0384, SE = 0.1514, 95% CI [−1.3587, −0.7609]), whereas the corresponding direct effect was not significant (b = 0.0534, SE = 0.2079, 95% CI [−0.3554, 0.4621]). Thus, the responsibility-reducing indirect effect through perceived shared agency was significantly weaker when the doctor rejected correct AI advice than when the doctor accepted incorrect AI advice, supporting H3. Different types of perceived diagnostic error significantly altered observers’ levels of perceived shared agency, with the magnitude of the corresponding indirect effect varying accordingly. This manipulation provides further evidence consistent with an indirect pathway through perceived shared agency.
Robustness analysis. As a robustness check, we repeated the key analyses using only the two items that directly assessed the doctor’s responsibility and fault (r = 0.952). The indirect effect of diagnostic mode on perceived doctor responsibility through perceived shared agency remained significant (b = −1.1800, SE = 0.2115, 95% CI [−1.6381, −0.8069]). Moreover, the difference in relative indirect effects between the doctor-accepts-incorrect-AI-advice and doctor-rejects-correct-AI-advice conditions remained significant (b = −1.0183, SE = 0.1558, 95% CI [−1.3372, −0.7344]). Thus, the substantive conclusions were unchanged when perceived doctor responsibility was operationalized using only the direct responsibility and fault items.

8.3. Discussion

The results of Study 2B were highly consistent with those of Study 1C. Specifically, Study 2B showed that, from the observer perspective, the indirect effect of diagnostic mode on perceived doctor responsibility through perceived shared agency was significant, and that perceived diagnostic error type moderated this indirect effect. When AI was involved in the diagnosis, observers attributed a certain degree of agency to the AI, perceiving it as capable of forming its own diagnostic judgments and making a substantive contribution to the final diagnosis, and consequently regarded the doctor and AI as joint agents involved in the diagnostic process. Consistent with this explanatory pathway, higher perceived shared agency under AI-involved diagnostic conditions was associated with lower perceived doctor responsibility. Furthermore, by manipulating perceived diagnostic error type, the study altered observers’ levels of perceived shared agency. Compared with the condition in which the doctor accepted incorrect AI advice, perceived shared agency was lower when the doctor rejected correct AI advice; correspondingly, both the indirect effect and the attenuating effect of AI involvement on perceived doctor responsibility were weaker. These findings provide further support for the indirect pathway through perceived shared agency. Overall, Study 2B extended the findings from the patient perspective to the observer perspective, providing a more robust empirical basis for the general discussion.

9. General Discussion

This research addresses a rapidly emerging but insufficiently answered question: When a diagnostic error is perceived, how does the involvement of AI, particularly generative AI, change patients’ and observers’ perceived doctor responsibility? Across five studies, we obtained two core findings.
First, in the context of perceived diagnostic errors, the presence of AI generally reduces perceived doctor responsibility. Both ERP evidence and scenario experiments showed that doctor–AI collaborative diagnosis, relative to doctor-only diagnosis, reduced perceived doctor responsibility among patients and observers. This effect emerged primarily through perceived shared agency: collaborative diagnosis led people to understand the outcome as jointly produced by the doctor and AI. Thus, people did not view AI merely as a tool that enhanced the doctor’s capabilities. To some extent, they treated AI as an agentic source that could jointly shape the diagnostic judgment, thereby diffusing responsibility that would otherwise be concentrated on the doctor.
Second, the responsibility-reducing effect of AI involvement has a clear boundary condition: perceived diagnostic error type. We manipulated perceived diagnostic error type to alter perceived shared agency. When the doctor rejected correct AI advice rather than accepting incorrect AI advice, the AI recommendation was not incorporated into the final erroneous diagnosis. Perceived shared agency was consequently lower, and the responsibility-reducing effect of AI involvement was therefore weaker.

9.1. Theoretical Contributions

This research makes four primary theoretical contributions.
First, by introducing perceived shared agency, this research advances the explanatory logic of responsibility attribution theory in human–AI collaboration. Existing work typically begins with a known set of potentially responsible actors and explains responsibility judgments in terms of causal contribution, control, intention, avoidability, and normative obligation [13,14,26]. In human–AI collaboration, however, people must first determine whether AI is merely a passive tool used by a human or a co-actor that participates in judgment and influences the outcome. Prior research shows that robot autonomy affects how people allocate credit and blame between robots and humans [61]. We use perceived shared agency to capture this antecedent psychological process. At its core, perceived shared agency reflects the belief that a task was not completed entirely by a human acting alone, but jointly by a human and AI, each of which made a substantive contribution to the outcome [27]. Our findings show that AI involvement leads people to interpret a diagnosis more as the joint product of the doctor and AI, expanding the perceived boundary of acting agents and reducing responsibility assigned to the doctor alone. Critically, in the doctor-rejects-correct-AI-advice condition, AI is present in the diagnostic process but does not substantively contribute to producing the erroneous outcome; the responsibility-reducing effect of AI involvement is therefore weaker. We thus move the starting point of responsibility attribution theory upstream, from how responsibility is divided among known actors to whom people first identify as co-actors, revealing two connected stages of agent identification and responsibility allocation in human–AI collaboration.
Second, this research shifts AI-responsibility scholarship from legal and normative attribution toward psychological attribution. Existing research has largely examined whether AI can qualify as a responsible entity from legal, ethical, and machine-morality perspectives, emphasizing normative issues such as responsibility gaps, responsibility allocation, and whether AI should be held accountable [19,23,24]. This stream provides an important foundation for AI governance but offers limited insight into how people perceive and judge responsibility in specific human–AI collaborative settings. We therefore distinguish normative or legal responsibility attribution from individual-level perceived responsibility. The former concerns whether AI qualifies as a responsible entity and who should bear responsibility institutionally; the latter concerns whom people perceive as participating in producing an outcome and how that perception shapes judgments of a human actor’s responsibility. We do not seek to determine whether AI should bear legal or ethical responsibility. Rather, drawing on responsibility attribution theory, we examine how AI involvement changes perceived doctor responsibility. The findings show that even when AI has not been recognized as a responsible entity in legal or ethical terms, people may psychologically identify it as a co-actor involved in producing a diagnostic outcome and accordingly adjust their judgments of the doctor’s responsibility. This perspective complements normative research on AI responsibility with a behavioral, micro-level account.
Third, this research extends medical AI research from technology acceptance to responsibility judgments concerning human professionals. Existing work has focused primarily on patients’ acceptance of or resistance to AI, the value of medical AI applications, and mechanisms of trust repair and forgiveness following AI service failures. Its central concern remains whether users will accept AI and how relationships can be restored after AI fails [20,21,22]. As generative AI becomes deeply embedded in medical diagnosis, however, the relevant question is no longer only whether AI will be adopted, but also how its entry into doctors’ decision processes changes people’s understanding of the doctor’s responsibility. We show that AI collaboration not only changes how diagnoses are made but also systematically alters patients’ and observers’ perceived doctor responsibility, with the direction and magnitude of that change depending on AI’s role in the causal chain that produced the error. We thereby move medical AI research beyond adoption, trust, and remediation to examine how AI reshapes responsibility judgments about human professionals and the boundaries of their accountability.
Fourth, this research connects individual responsibility attribution to the broader healthcare sociotechnical system by showing that the structure of human–AI participation itself can shape how responsibility is perceived. In healthcare AI systems, responsibility is not only formally allocated through professional, organizational, or regulatory rules. Patients and observers also reconstruct responsibility psychologically from how doctors and AI participate in producing a diagnostic outcome. Our findings show that this psychological allocation depends not simply on whether AI is present, but on whether AI is perceived to have substantively contributed to the erroneous outcome. When AI contributes to the final diagnosis, people are more likely to perceive the doctor and AI as joint agents and assign less responsibility to the doctor; when AI is present but does not substantively contribute to the error, this responsibility-reducing effect is substantially weaker. This finding adds an attributional layer to the understanding of healthcare sociotechnical systems by showing that the configuration of human–AI collaboration can alter not only decision processes but also stakeholders’ perceptions of responsibility. Within such systems, formal responsibility may be distributed across doctors, hospitals, AI developers, and regulators, whereas patients and observers may psychologically reconstruct responsibility according to the perceived contributions of human and AI actors. Because perceived responsibility is closely related to trust, complaints, accountability seeking, and acceptance of medical AI, these stakeholder responses may, in turn, feed back into the broader sociotechnical system by informing how hospitals, AI developers, and regulators approach AI adoption, oversight, disclosure, and responsibility governance [10,18,19].

9.2. Managerial Implications

For hospitals, the findings show that introducing AI changes not only clinical workflows but also perceived doctor responsibility when a diagnostic error is perceived. In the short term, AI’s presence may reduce the perceived burden of holding the doctor solely responsible. In the long term, however, unclear responsibility boundaries may encourage doctors to over-rely on the system or may create opportunities for strategic responsibility avoidance in disputes. Hospitals deploying AI should therefore attend not only to efficiency gains but also to clear human–AI collaboration protocols, outcome-review procedures, and responsibility-disclosure mechanisms [19,21].
For medical AI firms, the market value of AI should not be framed solely in terms of improving diagnostic accuracy; firms must also recognize that AI can change how people judge the doctor’s responsibility when they perceive a diagnostic error. In communicating with hospitals and designing product documentation, firms should explain AI’s specific roles in supporting diagnosis and treatment, flagging risks, and facilitating secondary review, while clearly presenting the scope of AI recommendations, the possibility of error, and the doctor’s ultimate review responsibility. The findings should not be interpreted as supporting the claim that “AI can share responsibility with doctors.” On the contrary, precisely because AI may reduce perceived doctor responsibility, interfaces, usage records, and risk disclosures should help hospitals and patients understand that AI is an assistive system participating in clinical judgment, not an entity that replaces the doctor’s ultimate responsibility for diagnosis [19,20,21].
For regulators, the findings suggest that policies encouraging medical AI adoption should also recognize that AI involvement may systematically change perceived doctor responsibility when a diagnostic error is perceived. In settings where AI’s presence may lead people to underestimate the doctor’s responsibility, regulation should specify doctors’ duties of care, review, and explanation with respect to final diagnostic decisions, preventing technology from being used to deflect responsibility [19,24].

9.3. Limitations and Future Research

This research has several limitations. First, although our studies included samples and contexts from China, the United States, Australia, and the United Kingdom, the evidence was drawn primarily from China and the United States and relied on hypothetical scenarios. The cross-cultural generalizability and ecological validity of the findings therefore remain to be established. Future research could include samples from a broader range of countries, particularly African countries, and incorporate high-fidelity doctor–AI interactions, dynamic diagnostic tasks, and clinical field data to examine the scope conditions of these findings. Second, although participants correctly understood the scenario materials, explicitly presenting the diagnostic error and the respective roles of the doctor and AI may have created demand characteristics that influenced responsibility attribution. Future research could use more indirect manipulations and conceal the study purpose. It could also incorporate behavioral measures, such as actual complaint behavior, follow-up-care choices, and AI adoption, and conduct experiments in high-fidelity clinical simulations or hospital field settings to reduce biases associated with scenario cues and self-report measures. Third, although Studies 1C and 2B identified significant indirect effects through perceived shared agency, these effects should be interpreted as statistical rather than conclusive evidence of causal mediation. Perceived shared agency was measured after perceived doctor responsibility rather than independently manipulated, and the two diagnostic-error conditions may also have differed in perceived doctor controllability and culpability. Future research should disentangle these alternative explanations and provide stronger causal tests of the proposed mechanism. Fourth, this research focuses on perceived doctor responsibility when people perceive a diagnostic error; it does not examine objective responsibility determinations or the allocation of legal responsibility. The findings are therefore most applicable to individual-level psychological attribution and social evaluation and should not be directly generalized to formal legal or institutional judgments of responsibility. Fifth, this research treats AI as a broad category without distinguishing among different types of AI systems. Conversational large language models, medical-image analysis systems, and clinical-record generation systems differ in their degree of autonomy, professional functions, and modes of participation in decision-making. They may therefore have different effects on perceived shared agency and perceived doctor responsibility. Future research could compare different AI technologies and degrees of AI involvement to identify the technological boundary conditions of the effects documented here.

10. Conclusions

As AI, particularly generative AI, becomes deeply integrated into healthcare, patients and observers no longer necessarily view the doctor as the sole diagnostic decision maker. This research shows that when a diagnostic error is perceived, doctor–AI collaborative diagnosis significantly reduces perceived doctor responsibility among both patients and observers, primarily by increasing perceived shared agency. This diffusion of responsibility nevertheless depends on perceived diagnostic error type. When the error is perceived to result from the doctor rejecting correct AI advice, the responsibility-reducing effect of AI involvement is substantially attenuated. Specifically, the difference in perceived doctor responsibility between the doctor-only and doctor-rejects-correct-AI-advice conditions was small but significant in Study 1C, whereas this pairwise difference was not significant in Study 2B. When the error is perceived to result from the doctor accepting incorrect AI advice, AI continues to diffuse perceived responsibility away from the doctor. AI is therefore changing not only medical decision-making but also the psychological models patients and observers use to determine who should be held responsible.

Author Contributions

Methodology and visualization, R.C.; writing—original draft, R.C. and W.T.; writing—review and editing, R.C. and W.T.; supervision, R.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research project was supported by the National Social Science Fund Later-funded Projects, PRC (Grant No. 21FGLB041), and Research Innovation Platform Project for Graduate Students under the Special Fund for Basic Scientific Research of Central Universities at Zhongnan University of Economics and Law, Grant No. 202611036.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki, and approved by the Ethics Committee of Huaqiao University (M2023009).

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

All data supporting the conclusions in the manuscript can be found at: https://osf.io/mxapw/overview?view_only=1d46ec11d4074c2086c7653f91a3cfbb (accessed on 30 July 2026).

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

For Pretest 1, we recruited 100 participants (Mage = 33.62, SD = 10.36; 54.0% women). After reading an adverse-medical-outcome scenario from the patient perspective, participants responded to one item measuring perceived diagnostic error: “I believe that this medical outcome was caused by a diagnostic error made by the doctor.” Responses were provided on a seven-point Likert scale ranging from 1 (strongly disagree) to 7 (strongly agree). Participants adopting the patient perspective reported a high level of perceived diagnostic error (M = 5.72, SD = 0.94), significantly above the scale midpoint of 4, t(99) = 18.23, p < 0.001.
Pretest 2 randomly recruited 100 participants (Mage = 32.77, SD = 9.60; 52.0% women). After reading an adverse-medical-outcome scenario from the observer perspective, participants responded to the same perceived-diagnostic-error item: “I believe that this medical outcome was caused by a diagnostic error made by the doctor.” Responses were provided on a seven-point Likert scale ranging from 1 (strongly disagree) to 7 (strongly agree). Participants adopting the observer perspective likewise reported a high level of perceived diagnostic error (M = 5.67, SD = 0.92), significantly above the scale midpoint of 4, t(99) = 18.12, p < 0.001.

References

  1. Sahni, N.R.; Carrus, B. Artificial Intelligence in U.S. Health Care Delivery. N. Engl. J. Med. 2023, 389, 348–358. [Google Scholar] [CrossRef] [Scilit]
  2. Silcox, C.; Zimlichmann, E.; Huber, K.; Rowen, N.; Saunders, R.; McClellan, M.; Kahn, C.N., III; Salzberg, C.A.; Bates, D.W. The Potential for Artificial Intelligence to Transform Healthcare: Perspectives from International Health Leaders. npj Digit. Med. 2024, 7, 88. [Google Scholar] [CrossRef] [Scilit]
  3. Topol, E. High-Performance Medicine: The Convergence of Human and Artificial Intelligence. Nat. Med. 2019, 25, 44–56. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S.S.; Wei, J.; Chung, H.W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S. Large Language Models Encode Clinical Knowledge. Nature 2023, 620, 172–180. [Google Scholar] [CrossRef] [Scilit]
  5. Thirunavukarasu, A.J.; Ting, D.S.J.; Elangovan, K.; Gutierrez, L.; Tan, T.F.; Ting, D.S.W. Large Language Models in Medicine. Nat. Med. 2023, 29, 1931–1940. [Google Scholar] [CrossRef] [Scilit]
  6. Ayers, J.W.; Poliak, A.; Dredze, M.; Leas, E.C.; Zhu, Z.; Kelley, J.B.; Faix, D.J.; Goodman, A.M.; Longhurst, C.A.; Hogarth, M.; et al. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Intern. Med. 2023, 183, 589–596. [Google Scholar] [CrossRef] [Scilit]
  7. Lee, P.; Bubeck, S.; Petro, J. Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. N. Engl. J. Med. 2023, 388, 1233–1239. [Google Scholar] [CrossRef] [Scilit]
  8. Vincent, C.; Young, M.; Phillips, A. Why Do People Sue Doctors? A Study of Patients and Relatives Taking Legal Action. Obstet. Gynecol. Surv. 1995, 50, 103–105. [Google Scholar] [CrossRef] [Scilit]
  9. Studdert, D.M.; Mello, M.M.; Gawande, A.A.; Gandhi, T.K.; Kachalia, A.; Yoon, C.; Puopolo, A.L.; Brennan, T.A. Claims, Errors, and Compensation Payments in Medical Malpractice Litigation. N. Engl. J. Med. 2006, 354, 2024–2033. [Google Scholar] [CrossRef] [Scilit]
  10. Kistler, C.E.; Walter, L.C.; Mitchell, C.M.; Sloane, P.D. Patient Perceptions of Mistakes in Ambulatory Care. Arch. Intern. Med. 2010, 170, 1480–1487. [Google Scholar] [CrossRef] [Scilit]
  11. Schlesinger, M.; Dhingra, I.; Fain, B.A.; Prentice, J.C.; Parkash, V. Adverse Events and Perceived Abandonment: Learning from Patients’ Accounts of Medical Mishaps. BMJ Open Qual. 2024, 13, e002848. [Google Scholar] [CrossRef] [Scilit]
  12. Weiner, B. Intrapersonal and Interpersonal Theories of Motivation from an Attributional Perspective. Educ. Psychol. Rev. 2000, 12, 1–14. [Google Scholar] [CrossRef] [Scilit]
  13. Alicke, M.D. Culpable Control and the Psychology of Blame. Psychol. Bull. 2000, 126, 556–574. [Google Scholar] [CrossRef]
  14. Weiner, B. Judgments of Responsibility: A Foundation for a Theory of Social Conduct; Guilford Press: New York, NY, USA, 1995. [Google Scholar]
  15. El Zein, M.; Bahrami, B.; Hertwig, R. Shared Responsibility in Collective Decisions. Nat. Hum. Behav. 2019, 3, 554–559. [Google Scholar] [CrossRef] [Scilit]
  16. Gill, T. Blame It on the Self-Driving Car: How Autonomous Vehicles Can Alter Consumer Morality. J. Consum. Res. 2020, 47, 272–291. [Google Scholar] [CrossRef] [Scilit]
  17. Huo, W.; Zheng, G.; Yan, J.; Sun, L.; Han, L. Interacting with Medical Artificial Intelligence: Integrating Self-Responsibility Attribution, Human–Computer Trust, and Personality. Comput. Hum. Behav. 2022, 132, 107253. [Google Scholar] [CrossRef] [Scilit]
  18. Richardson, J.P.; Smith, C.; Curtis, S.; Watson, S.; Zhu, X.; Barry, B.; Sharp, R.R. Patient Apprehensions about the Use of Artificial Intelligence in Healthcare. npj Digit. Med. 2021, 4, 140. [Google Scholar] [CrossRef] [Scilit]
  19. Sullivan, Y.W.; Fosso Wamba, S. Moral Judgments in the Age of Artificial Intelligence. J. Bus. Ethics 2022, 178, 917–943. [Google Scholar] [CrossRef] [Scilit]
  20. Longoni, C.; Bonezzi, A.; Morewedge, C.K. Resistance to Medical Artificial Intelligence. J. Consum. Res. 2019, 46, 629–650. [Google Scholar] [CrossRef] [Scilit]
  21. Cadario, R.; Longoni, C.; Morewedge, C.K. Understanding, Explaining, and Utilizing Medical Artificial Intelligence. Nat. Hum. Behav. 2021, 5, 1636–1642. [Google Scholar] [CrossRef] [Scilit]
  22. Chen, A.; Pan, Y.; Li, L.; Yu, Y. Are You Willing to Forgive AI? Service Recovery from Medical AI Service Failure. Ind. Manag. Data Syst. 2022, 122, 2540–2557. [Google Scholar] [CrossRef] [Scilit]
  23. Matthias, A. The Responsibility Gap: Ascribing Responsibility for the Actions of Learning Automata. Ethics Inf. Technol. 2004, 6, 175–183. [Google Scholar] [CrossRef] [Scilit]
  24. Bigman, Y.E.; Waytz, A.; Alterovitz, R.; Gray, K. Holding Robots Responsible: The Elements of Machine Morality. Trends Cogn. Sci. 2019, 23, 365–368. [Google Scholar] [CrossRef] [Scilit]
  25. Shaver, K.G. The Attribution of Blame: Causality, Responsibility, and Blameworthiness; Springer Science+Business Media: New York, NY, USA, 2012. [Google Scholar]
  26. Malle, B.F.; Guglielmo, S.; Monroe, A.E. A Theory of Blame. Psychol. Inq. 2014, 25, 147–186. [Google Scholar] [CrossRef] [Scilit]
  27. Voyer, B.G.; Sangle-Ferriere, M.; Sajtos, L.; Sung, B. The Measurement of Perceived Shared Agency in Customer–Artificial Intelligence Interactions. J. Serv. Theory Pract. 2025, 35, 632–658. [Google Scholar] [CrossRef] [Scilit]
  28. Gailey, J.A. Attribution of Responsibility for Organizational Wrongdoing: A Partial Test of an Integrated Model. J. Criminol. 2013, 2013, 1–10. [Google Scholar] [CrossRef] [Scilit]
  29. Varkey, B. Principles of Clinical Ethics and Their Application to Practice. Med. Princ. Pract. 2021, 30, 17–28. [Google Scholar] [CrossRef] [Scilit]
  30. Hamilton, V.L. Who Is Responsible? Toward a Social Psychology of Responsibility Attribution. Soc. Psychol. 1978, 41, 316–328. [Google Scholar] [CrossRef] [Scilit]
  31. Fincham, F.D.; Jaspars, J.M. Attribution of Responsibility: From Man the Scientist to Man as Lawyer. Adv. Exp. Soc. Psychol. 1980, 13, 81–138. [Google Scholar] [CrossRef] [Scilit]
  32. Gallagher, T.H.; Waterman, A.D.; Ebers, A.G.; Fraser, V.J.; Levinson, W. Patients’ and Physicians’ Attitudes Regarding the Disclosure of Medical Errors. JAMA 2003, 289, 1001–1007. [Google Scholar] [CrossRef] [Scilit]
  33. Hotvedt, R.; Førde, O.H. Doctors Are to Blame for Perceived Medical Adverse Events. A Cross Sectional Population Study. The Tromsø Study. BMC Health Serv. Res. 2013, 13, 46. [Google Scholar] [CrossRef] [Scilit]
  34. Robbennolt, J.K. Outcome Severity and Judgments of “Responsibility”: A Meta-Analytic Review. J. Appl. Soc. Psychol. 2000, 30, 2575–2609. [Google Scholar] [CrossRef] [Scilit]
  35. Rudolph, U.; Roesch, S.; Greitemeyer, T.; Weiner, B. A Meta-Analytic Review of Help Giving and Aggression from an Attributional Perspective: Contributions to a General Theory of Motivation. Cogn. Emot. 2004, 18, 815–848. [Google Scholar] [CrossRef] [Scilit]
  36. Bismark, M.; Dauer, E.; Paterson, R.; Studdert, D. Accountability Sought by Patients Following Adverse Events from Medical Care: The New Zealand Experience. Can. Med. Assoc. J. 2006, 175, 889–894. [Google Scholar] [CrossRef] [Scilit]
  37. Prentice, J.C.; Bell, S.K.; Thomas, E.J.; Schneider, E.C.; Weingart, S.N.; Weissman, J.S.; Schlesinger, M.J. Association of Open Communication and the Emotional and Behavioural Impact of Medical Error on Patients and Families: State-Wide Cross-Sectional Survey. BMJ Qual. Saf. 2020, 29, 883–894. [Google Scholar] [CrossRef] [Scilit]
  38. Rejeleene, R.; Mehta, N.B. Artificial Intelligence in Medicine: How It Works, How It Fails. Cleve. Clin. J. Med. 2026, 93, 113–120. [Google Scholar] [CrossRef] [Scilit]
  39. Shortliffe, E.H.; Sepúlveda, M.J. Clinical Decision Support in the Era of Artificial Intelligence. JAMA 2018, 320, 2199–2200. [Google Scholar] [CrossRef] [Scilit]
  40. Baird, A.; Maruping, L.M. The Next Generation of Research on IS Use: A Theoretical Framework of Delegation to and from Agentic IS Artifacts. MIS Q. 2021, 45, 315–341. [Google Scholar] [CrossRef] [Scilit]
  41. Epley, N.; Waytz, A.; Cacioppo, J.T. On Seeing Human: A Three-Factor Theory of Anthropomorphism. Psychol. Rev. 2007, 114, 864–886. [Google Scholar] [CrossRef] [Scilit]
  42. Shank, D.B.; DeSanti, A. Attributions of Morality and Mind to Artificial Intelligence after Real-World Moral Violations. Comput. Hum. Behav. 2018, 86, 401–411. [Google Scholar] [CrossRef] [Scilit]
  43. Grote, T.; Berens, P. On the Ethics of Algorithmic Decision-Making in Healthcare. J. Med. Ethics 2020, 46, 205–211. [Google Scholar] [CrossRef] [Scilit]
  44. Smith, H. Clinical AI: Opacity, Accountability, Responsibility and Liability. AI Soc. 2021, 36, 535–545. [Google Scholar] [CrossRef] [Scilit]
  45. Sundar, S.S. Rise of Machine Agency: A Framework for Studying the Psychology of Human–AI Interaction (HAII). J. Comput.-Mediat. Commun. 2020, 25, 74–88. [Google Scholar] [CrossRef] [Scilit]
  46. Gaube, S.; Suresh, H.; Raue, M.; Merritt, A.; Berkowitz, S.J.; Lermer, E.; Coughlin, J.F.; Guttag, J.V.; Colak, E.; Ghassemi, M. Do as AI Say: Susceptibility in Deployment of Clinical Decision-Aids. npj Digit. Med. 2021, 4, 31. [Google Scholar] [CrossRef] [Scilit]
  47. Reverberi, C.; Rigon, T.; Solari, A.; Hassan, C.; Cherubini, P.; Cherubini, A. Experimental Evidence of Effective Human–AI Collaboration in Medical Decision-Making. Sci. Rep. 2022, 12, 14952. [Google Scholar] [CrossRef] [Scilit]
  48. Lazarus, R.S. Emotion and Adaptation; Oxford University Press: New York, NY, USA, 1991. [Google Scholar]
  49. Roseman, I.J.; Smith, C.A. Appraisal Theory Overview, Assumptions, Varieties, Controversies. In Appraisal Processes in Emotion: Theory, Methods, Research; Scherer, K.R., Schorr, A., Johnstone, T., Eds.; Oxford University Press: New York, NY, USA, 2001; pp. 3–19. [Google Scholar]
  50. Schmidt, G.; Weiner, B. An Attribution-Affect-Action Theory of Behavior: Replications of Judgments of Help-Giving. Pers. Soc. Psychol. Bull. 1988, 14, 611–621. [Google Scholar] [CrossRef] [Scilit]
  51. Folkes, V.S. Consumer Reactions to Product Failure: An Attributional Approach. J. Consum. Res. 1984, 10, 398–409. [Google Scholar] [CrossRef] [Scilit]
  52. Carretié, L.; Mercado, F.; Tapia, M.; Hinojosa, J.A. Emotion, Attention, and the ‘Negativity Bias’, Studied through Event-Related Potentials. Int. J. Psychophysiol. 2001, 41, 75–85. [Google Scholar] [CrossRef] [Scilit]
  53. Li, P.; Han, C.; Lei, Y.; Holroyd, C.B.; Li, H. Responsibility Modulates Neural Mechanisms of Outcome Processing: An ERP Study. Psychophysiology 2011, 48, 1129–1133. [Google Scholar] [CrossRef] [Scilit]
  54. Pu, M.; Yu, R. Personal Responsibility Modulates Neural Representations of Anticipatory and Experienced Pain. Psychophysiology 2019, 56, e13294. [Google Scholar] [CrossRef] [Scilit]
  55. Faul, F.; Erdfelder, E.; Lang, A.-G.; Buchner, A. G*Power 3: A Flexible Statistical Power Analysis Program for the Social, Behavioral, and Biomedical Sciences. Behav. Res. Methods 2007, 39, 175–191. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  56. Fu, H.; Niu, J.; Wu, Z.; Cheng, B.; Guo, X.; Zuo, J. Exploration of Public Stereotypes of Supply-and-Demand Characteristics of Recycled Water Infrastructure: Evidence from an Event-Related Potential Experiment in Xi’an, China. J. Environ. Manag. 2022, 322, 116103. [Google Scholar] [CrossRef] [Scilit]
  57. Ho, H.T.; Schröger, E.; Kotz, S.A. Selective Attention Modulates Early Human Evoked Potentials during Emotional Face–Voice Processing. J. Cogn. Neurosci. 2015, 27, 798–818. [Google Scholar] [CrossRef] [Scilit]
  58. Qin, J.; Han, S. Neurocognitive Mechanisms Underlying Identification of Environmental Risks. Neuropsychologia 2009, 47, 397–405. [Google Scholar] [CrossRef] [Scilit]
  59. Awad, E.; Levine, S.; Kleiman-Weiner, M.; Dsouza, S.; Tenenbaum, J.B.; Shariff, A.; Bonnefon, J.-F.; Rahwan, I. Drivers Are Blamed More than Their Automated Cars When Both Make Mistakes. Nat. Hum. Behav. 2020, 4, 134–143. [Google Scholar] [CrossRef] [Scilit]
  60. Hayes, A.F. Introduction to Mediation, Moderation, and Conditional Process Analysis: A Regression-Based Approach; Guilford Publications: New York, NY, USA, 2017. [Google Scholar]
  61. Kim, T.; Hinds, P. Who Should I Blame? Effects of Autonomy and Transparency on Attributions in Human-Robot Interaction. In Proceedings of the 15th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN 2006), Hatfield, Hertfordshire, UK, 6–8 September 2006; IEEE: Piscataway, NJ, USA, 2006; pp. 80–85. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Sequence of a single trial.
Figure 1. Sequence of a single trial.
Systems 14 01081 g001
Figure 2. Event-related potential waveforms by diagnostic mode.
Figure 2. Event-related potential waveforms by diagnostic mode.
Systems 14 01081 g002
Figure 3. Patients’ perceived doctor responsibility by diagnostic mode. Note: Error bars represent ±1 SD. *** p < 0.001.
Figure 3. Patients’ perceived doctor responsibility by diagnostic mode. Note: Error bars represent ±1 SD. *** p < 0.001.
Systems 14 01081 g003
Figure 4. Patients’ perceived doctor responsibility across the four diagnostic conditions. Note: Error bars represent ±1 SD. * p < 0.05, *** p < 0.001.
Figure 4. Patients’ perceived doctor responsibility across the four diagnostic conditions. Note: Error bars represent ±1 SD. * p < 0.05, *** p < 0.001.
Systems 14 01081 g004
Figure 5. Observers’ perceived doctor responsibility by diagnostic mode. Note: Error bars represent ±1 SD. *** p < 0.001.
Figure 5. Observers’ perceived doctor responsibility by diagnostic mode. Note: Error bars represent ±1 SD. *** p < 0.001.
Systems 14 01081 g005
Table 1. Study Designs and Primary Findings.
Table 1. Study Designs and Primary Findings.
StudyMethodParticipantsScenarioPrimary Findings
Study 1AERP experimentFinal N = 24; recruited in person at a Chinese universityPatient perceives an error in the diagnosis of an eye conditionDoctor–AI collaboration reduced perceived doctor responsibility and elicited a smaller P2 amplitude than doctor-only diagnosis.
Study 1BScenario experiment200 Chinese residents; CredamoPatient perceives an error in the diagnosis of an eye conditionDoctor–AI collaboration reduced perceived doctor responsibility relative to doctor-only diagnosis.
Study 1CScenario experiment400 Chinese residents; CredamoPatient perceives an error in the diagnosis of a thyroid conditionA significant indirect effect through perceived shared agency was observed, consistent with the proposed mediating pathway. The indirect effect was weaker when the doctor rejected correct AI advice than when the doctor accepted incorrect AI advice.
Study 2AScenario experiment200 Chinese residents; CredamoObserver perceives an error in the diagnosis of a patient’s eye conditionDoctor–AI collaboration reduced perceived doctor responsibility relative to doctor-only diagnosis.
Study 2BScenario experiment400 participants from six English-speaking countries; CloudResearchObserver perceives an error in the diagnosis of a patient’s knee conditionA significant indirect effect through perceived shared agency was observed, consistent with the proposed mediating pathway. The indirect effect was weaker when the doctor rejected correct AI advice than when the doctor accepted incorrect AI advice.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Cheng, R.; Sun, R.; Tang, W. When AI Joins the Diagnosis: How Doctor–AI Collaboration Shapes Perceived Doctor Responsibility Under Perceived Diagnostic Errors. Systems 2026, 14, 1081. https://doi.org/10.3390/systems14091081

AMA Style

Cheng R, Sun R, Tang W. When AI Joins the Diagnosis: How Doctor–AI Collaboration Shapes Perceived Doctor Responsibility Under Perceived Diagnostic Errors. Systems. 2026; 14(9):1081. https://doi.org/10.3390/systems14091081

Chicago/Turabian Style

Cheng, Ruxia, Rui Sun, and Wenlong Tang. 2026. "When AI Joins the Diagnosis: How Doctor–AI Collaboration Shapes Perceived Doctor Responsibility Under Perceived Diagnostic Errors" Systems 14, no. 9: 1081. https://doi.org/10.3390/systems14091081

APA Style

Cheng, R., Sun, R., & Tang, W. (2026). When AI Joins the Diagnosis: How Doctor–AI Collaboration Shapes Perceived Doctor Responsibility Under Perceived Diagnostic Errors. Systems, 14(9), 1081. https://doi.org/10.3390/systems14091081

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop