Next Article in Journal
Alzheimer’s Disease Detection Based on Machine Learning and Deep Learning Frameworks: A Cross-Dataset Comparative Performance Analysis and Assessment of Clinical Readiness
Next Article in Special Issue
A Configurable Framework for Quantifying and Comparing Interpretability Across ML Models and Methods
Previous Article in Journal
A Hybrid CNN-MLP-DWD Framework for Robust Medical Image Classification Under High-Dimensional Low-Sample Size Conditions
Previous Article in Special Issue
From Explainable AI to Knowledge Extraction for Trustworthy Energy Forecasting Systems: A Systematic Review
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Cognitive Friction in Clinical Decision Support: A Comparative Study of Judicial and Adjunct Human–AI Interaction Protocols

1
Department of Electrical, Computer and Biomedical Engineering, University of Pavia, 27100 Pavia, Italy
2
Department of Global Public Health and Primary Care, Faculty of Medicine, University of Bergen, 5009 Bergen, Norway
3
Department of Mathematics, University of Bergen, 5007 Bergen, Norway
4
Department of Radiology, IRCCS Policlinic San Matteo Foundation, 27100 Pavia, Italy
5
Department of Clinical, Surgical, Diagnostic, and Pediatric Sciences, University of Pavia, 27100 Pavia, Italy
6
Infectious Diseases Unit, IRCCS Policlinic San Matteo Foundation, 27100 Pavia, Italy
7
Emergency Medicine Unit and Emergency Medicine Postgraduate Training Program, Department of Internal Medicine, IRCCS Policlinic San Matteo Foundation, 27100 Pavia, Italy
8
PhD Program in Experimental Medicine, University of Pavia, 27100 Pavia, Italy
9
Telfer School of Management, University of Ottawa, Ottawa, ON K1N 6N5, Canada
*
Authors to whom correspondence should be addressed.
Mach. Learn. Knowl. Extr. 2026, 8(7), 216; https://doi.org/10.3390/make8070216
Submission received: 22 June 2026 / Revised: 16 July 2026 / Accepted: 17 July 2026 / Published: 22 July 2026

Abstract

Artificial intelligence is increasingly used to support clinical decision making, yet concerns remain regarding algorithmic aversion, automation bias and the preservation of meaningful human oversight; while explainable AI aims to improve transparency, less attention has been devoted to the design of human–AI interaction protocols. This study investigates Frictional AI, an interaction paradigm that introduces cognitive friction to encourage critical engagement with AI recommendations. First, semi-structured interviews were conducted with a legal expert and a psychologist and analyzed through thematic analysis to identify legal, ethical, and cognitive requirements for AI-assisted decision support. Second, a user study involving 96 medical residents compared three interaction protocols: a conventional explainable AI-first design (XAI) and two friction-based protocols, namely a judicial protocol based on juxtaposed explanations (Judicial AI, JAI) and an adjunct protocol requiring an initial unsupported decision before AI exposure (AAI). Diagnostic accuracy and confidence, perceived usefulness, completion time, and reliance patterns were evaluated. The interviews highlighted the importance of human-centered explanations, contrastive reasoning, preservation of professional responsibility, and the role of user studies in evaluating human–AI interaction. The quantitative results showed that none of the AI-assisted conditions improved diagnostic accuracy relative to the no-support baseline. However, JAI achieved performance comparable to the baseline, outperforming XAI and AAI, and exhibited the lowest level of over-reliance. Overall findings suggest that the effectiveness of decision-support systems depends not only on model performance and explanation quality but also on interaction design. In conclusion, while preserving diagnostic performance, judicial protocols showed promise in mitigating automation bias and promoting active cognitive engagement in clinical decision support.

1. Introduction

Artificial intelligence (AI) has demonstrated remarkable potential in supporting clinical decision making [1,2], yet its successful integration into healthcare depends not only on predictive performance but also on how clinicians interact with algorithmic recommendations; while explainable AI (XAI) has emerged as a central strategy for improving transparency and trust, increasing evidence suggests that XAI alone may be insufficient to ensure appropriate human oversight and may even contribute to excessive reliance on automated systems. This section discusses the limitations of current AI-based decision support, introduces the role of interaction design as a mechanism for mitigating algorithmic aversion and automation bias, and presents the rationale, objectives, and research questions of the present study.

1.1. Limitations of AI-Based Decision Support

AI and AI-driven decision-support systems (DSS) are increasingly permeating nearly every aspect of modern life, including high-stakes domains such as justice and healthcare [3,4]. Despite their remarkable performance, AI systems reason in ways that fundamentally differ from human cognition. Rather than relying on causal understanding, machines primarily operate through pattern recognition and statistical inference [5]. Consequently, AI systems may learn non-causal, spurious, or otherwise misleading associations that can result in erroneous predictions. Identifying such issues is often challenging because the internal decision-making processes of many modern AI models remain opaque, a characteristic commonly referred to as the black-box problem. This lack of transparency has contributed to widespread skepticism toward AI-assisted decision making, a phenomenon known as algorithmic aversion [6].
A natural response to these concerns has been the growing adoption of explainable AI methods [7,8]. XAI aims to improve the transparency and interpretability of AI systems, enabling developers to identify biases, errors, and undesirable model behaviors while helping end-users understand the rationale underlying machine-generated predictions. Beyond its technical benefits, explainability has become one of the fundamental principles of trustworthy AI and an important requirement of emerging regulatory, ethical, and governance frameworks, which emphasize transparency, fairness, and accountability [3,4].
Specifically, the European Union AI Act [9], which builds on earlier European regulatory frameworks, including the General Data Protection Regulation (GDPR) [10] and the Medical Device Regulation (MDR) [11], as well as related initiatives such as the proposed AI Liability Directive [12]: collectively, these frameworks establish legal requirements concerning transparency, human oversight, data protection, risk management, and information provision for high-risk AI systems.
These legal requirements are complemented by governance frameworks, including the Assessment List for Trustworthy Artificial Intelligence (ALTAI) [13], the NIST AI Risk Management Framework (AI RMF 1.0) [14], and, in the healthcare domain, the FUTURE-AI guidelines [15]. These initiatives extend the concept of trustworthy AI beyond explainability alone by emphasizing robust data governance, technical robustness and safety, continuous risk management, institutional oversight, human agency, accountability, transparency, and the protection of fundamental human rights throughout the AI life cycle. Finally, from a broader ethical perspective, IEEE developed the Ethically Aligned Design framework [16], which provides guidance on embedding ethical considerations into AI development, ranging from accountability to the protection of fundamental human rights.
Although algorithmic aversion is commonly regarded as one of the primary barriers to the effective adoption of AI technologies, often framed as a major source of negative technology impact [17], an equally important yet frequently overlooked challenge is automation bias [18]. Continuous efforts to improve model performance, increase user trust through transparency, and enhance the persuasiveness of AI systems—for example, through natural-language interactions enabled by large language models (LLM)—may inadvertently encourage excessive reliance on automated recommendations [19,20]. Automation bias is typically defined as a detrimental over-reliance on DSSs, whereby users defer their judgment to the machine even when independent verification would be warranted [17,18]. Such behavior is particularly dangerous in healthcare, where incorrect decisions may have profound consequences for patient well-being.
Despite receiving relatively limited attention in the literature [21], automation bias has been widely documented in healthcare settings [22,23]. Moreover, it has been associated with several detrimental consequences. First, it may reduce clinicians’ sense of agency during decision making, leading healthcare professionals to follow AI recommendations without feeling empowered to challenge them [24]. Second, automation bias may promote automation complacency [25], whereby users progressively reduce their vigilance and critical monitoring of automated systems because they perceive them as sufficiently reliable. Third, automation bias may foster a sense of deresponsibilization [26], whereby individuals perceive themselves as less accountable for outcomes because responsibility is implicitly shifted to the AI system. In this respect, the phenomenon may also be related to defensive medicine [27], whereby clinicians may defer to AI recommendations or modify their clinical decisions primarily to reduce legal liability or professional criticism. Finally, excessive dependence on AI may promote deskilling [28], limiting both the maintenance and, in some cases, the development of professional expertise [29].
These considerations suggest that the challenge of trustworthy AI extends beyond predictive performance and transparency alone. As highlighted by both O’Neil [3] and Christian [4], algorithmic systems can become harmful when their decisions are scaled to large populations without adequate oversight, accountability, or opportunities for human contestation. Moreover, contemporary governance frameworks consistently recognize that explainability represents only one component of trustworthy AI, which must be complemented by appropriate data governance, institutional oversight, technical risk management, and other mechanisms that preserve meaningful human control over AI-assisted decisions. Consequently, the goal of AI in healthcare should not merely be to produce accurate recommendations but to support clinicians in making informed decisions while preserving professional autonomy, responsibility, and critical judgment. Achieving this balance requires approaches that promote meaningful human oversight and encourage users to fully engage with, rather than passively accept, algorithmic recommendations.

1.2. The Role of Interaction Design in Clinical Decision Making

As discussed by Cabitza et al. [22], when improvements in model performance and interpretability are insufficient to address the limitations of AI-assisted decision making, the key solution becomes interaction design. Building on established principles of user-centered design, which aim to minimize unnecessary cognitive friction [30,31], conventional AI-based decision-support systems typically adopt an extremely simple interaction paradigm: the AI system provides a prediction—usually a diagnostic label—in an “oracular” form, often without explanations and, in some cases, without even reporting a confidence score. Norman emphasizes that while simplicity is an important design goal [31], many domains such as healthcare are inherently complex and cannot be meaningfully simplified without loss of critical information [32]. Rather than attempting to eliminate complexity, Norman invites interaction designers to engage with it, enabling the development of systems that are enriching and empowering for users.
In line with these principles, according to Cabitza et al. [22], this overly simplistic interaction design is one of the main drivers of automation bias. Consequently, the authors argue that effective human–AI interaction protocols should actively shape how clinicians engage with AI-generated information rather than merely presenting recommendations. To address this issue, they introduced the concept of Frictional AI (FAI), which leverages cognitive friction to encourage users to critically evaluate the available information in the raw data along with the model’s suggestions.
This perspective is grounded in the idea that human cognition tends to favor fast and intuitive reasoning (System 1) over slower and more analytical reasoning (System 2), as pointed out by Kahneman [33]. Nevertheless, although System 1 enables efficient decision making, it is also prone to systematic biases. In a similar way, accepting an AI recommendation solely because the system is known to perform well, without critically examining the rationale behind the prediction, may represent a cognitive shortcut.
To improve decision quality, Kahneman advocates the use of nudges, interventions designed to influence behavior and promote more deliberate decision making. Frictional AI applies this principle to human–AI interaction by introducing obstacles into the decision-making process in the form of cognitive friction. These obstacles encourage users to engage System 2 reasoning, thereby reducing the risk of blindly accepting AI recommendations. To operationalize this concept, FAI proposes some distinct interaction protocols, including for example:
  • Adjunct protocols: the AI system acts as a second opinion rather than providing immediate guidance. Users are first required to take time to formulate an independent diagnosis before receiving AI support.
  • Judicial protocols: instead of explicitly presenting a prediction, the system provides arguments supporting competing diagnostic hypotheses. AI instruments (e.g., XAI or LLMs) are used to generate juxtaposed evidence for the most plausible hypotheses—typically two alternatives, which represent opposing positions in a judicial debate. The user, acting as a “judge”, is then required to evaluate the evidence and reach their own conclusion. The study and application of these protocols is also referred to as Judicial AI (JAI).
  • Analogical protocols: rather than providing direct recommendations, the system presents representative cases similar to the current instance. The most similar examples associated with different candidate hypotheses are retrieved from the dataset and displayed together with their ground-truth labels.
  • It is important to note that these protocols were not intended as entirely novel interaction designs. Rather, they represent a conceptual organization of interaction patterns that have previously appeared in the literature under different names and in different application domains. For example, Reingold et al. [34] introduced, outside the healthcare domain, an interaction design based on dissenting explanations, in which users are presented with evidence supporting alternative conclusions. This concept is closely aligned with the notion of juxtaposed evidence employed in judicial protocols.
Similarly, adjunct protocols build upon experimental paradigms that are widely used in the literature on decision support to study human–AI collaboration, in which users are first asked to make an independent decision before being exposed to algorithmic recommendations and given the opportunity to revise their judgment. Examples in healthcare include the works of Cao et al. [35], Bashkirova et al. [36], and Buçinca et al. [37].
In addition to this, some recent studies have investigated the effectiveness of frictional interaction designs in healthcare settings. For example, Cabitza et al. [22] implemented an analogical protocol for vertebral fracture detection from X-ray images. In this study, clinicians—particularly less experienced users—perceived the frictional support as beneficial. Diagnostic accuracy and clinician confidence improved with AI assistance, although not significantly, and the protocol overall demonstrated a positive technology impact.
In a subsequent study, the same research group [38] evaluated judicial protocols, in the same vertebral fracture diagnosis task. The results showed an overall improvement in diagnostic accuracy. Although, this time, the protocol appeared less beneficial for less experienced users, the increase in accuracy (effect size) attributable to AI support was statistically significant among experienced clinicians. For the most challenging cases, a positive effect size was also observed, although statistical significance was not reached. Consistent improvements in diagnostic confidence were likewise reported. Overall, the authors described a positive technology impact characterized by a leveler effect [39], whereby users with lower baseline performance benefited the most from AI support. In this study, JAI evidence was generated using XAI techniques. Specifically, the Grad-CAM algorithm [40] was employed as an example for juxtaposed evidence generation: class activation maps highlight image regions that contributed most strongly to a given prediction; by producing separate maps for both the fracture and non-fracture classes, the system was able to provide users with evidence supporting competing diagnostic hypotheses.
It is important to note, however, that both of the aforementioned studies required participants to first evaluate each case without AI assistance, before receiving algorithmic support. This design choice was necessary to enable the quantification of algorithm aversion, automation bias, and overall technology impact [17]. As a consequence, both protocols were partially confounded by elements characteristic of the adjunct protocol. Nevertheless, subsequent studies involving knee MRI interpretation and ECG classification [41] demonstrated that AI-/XAI-first designs generally yielded higher diagnostic performance than adjunct, human-first approaches. Beyond highlighting potential limitations of adjunct protocols, this result suggests—yet, without proving it—that the positive effects observed in analogical and judicial protocols are not driven by this necessary design overlap.
Among the proposed friction-based interaction protocols, judicial protocols are particularly noteworthy because they do not merely delay or restrict access to AI recommendations. Rather, they require users to actively evaluate competing pieces of evidence before reaching a decision. This design resembles the differential diagnosis framework [42] widely used in healthcare, and may therefore represent a promising strategy for mitigating automation bias while preserving users’ sense of agency, responsibility, and critical judgment.

1.3. Study Rationale and Research Questions

Despite growing recognition that interaction design plays a crucial role in determining the impact of AI-assisted decision making, the literature investigating how interaction protocols can promote positive technology impact remains limited. In particular, little is known about how different interaction designs can mitigate algorithm aversion and automation bias while satisfying user-centered, ethical, and legal requirements. To address this gap, our first contribution consists of a qualitative study in the form of semi-structured interviews, involving a legal expert and a psychologist. The study investigates the characteristics of effective decision support from legal and psychological perspectives, focusing on three themes: the level of explainability required for trustworthy AI, the relationship between reliance on AI and responsibility, and the adequacy of existing evaluation methodologies.
In addition, we have discussed the epistemological foundations of friction-based human–AI interaction protocols and illustrated their potential through examples from the literature. Nevertheless, the available evidence is predominantly produced by the same research group and it remains limited in scope. For example, the earliest studies about analogical [22] and judicial protocols [38] were conducted on the same task setting, namely vertebral fracture diagnosis from a dataset of only 18 X-ray images. Furthermore, the number of participating clinicians was relatively small: the former case involved 8 musculoskeletal radiologists and 8 specialized spine surgeons, whereas the latter involved 16 orthopedic clinicians, including 10 board-certified subspecialists and 6 residents.
Judicial protocols have emerged as a particularly promising interaction design. However, their evaluation has been partially limited by their entanglement with adjunct protocols, which acted as a confounding factor in the validation study conducted [38]. To further investigate the effectiveness of friction-based interaction designs in clinical decision support, our second contribution consists of a systematic comparison of the following interaction protocols:
  • XAI—an XAI-first protocol (state of the art), in which participants are shown the model prediction and corresponding XAI explanations;
  • JAI—a pure judicial protocol, in which participants are shown juxtaposed explanations for the two classes derived from the same XAI method;
  • AAI—a pure adjunct protocol, in which participants first provide an unsupported diagnosis and subsequently revise it after being presented with the model prediction and corresponding XAI explanation.
  • All groups will be compared also against a baseline, NoAI, obtained from the first decision in AAI group.
We emphasize that the proposed study extends the current evidence on friction-based human–AI interaction in two important ways. First, it includes a pure adjunct protocol that is conceptually equivalent to those previously investigated in the literature [41], thereby providing an independent validation of earlier findings in a different clinical setting. Second, and more importantly, it evaluates a judicial protocol within a JAI-first interaction framework. Unlike the study by Cabitza et al. [38], where the judicial protocol was embedded within an adjunct-like workflow requiring both an initial unsupported decision and a subsequent post-support revision, the present study evaluates JAI as an unentangled, standalone interaction protocol. Consequently, any observed effects of the judicial protocol can be attributed directly to the interaction design itself, without the confounding influence of the adjunct workflow.
This comparison is performed through a user study involving 96 clinicians from the IRCCS Policlinic San Matteo Foundation in Pavia (Italy), conducted within the ALFABETO project [43,44], which aims to develop an AI-based pipeline integrating diagnostic data into decision support for the triage of COVID-19 patients.
The objectives of the qualitative and quantitative components of our study are summarized by the following research questions (RQs):
  • RQ1—Design of clinical decision support from legal and psychological perspectives
  • RQ1a: What are the characteristics of sufficiently transparent and understandable explanations?
  • RQ1b: How can over-reliance be reduced, and how is it related to responsibility in human–AI decision making?
  • RQ1c: Are user studies an effective tool for assessing decision support quality and demonstrating compliance with legal requirements?
  • RQ2—Characterizing Frictional AI Systems
  • RQ2a Is the performance associated with FAI-based decision-support systems non-inferior to that achieved with either no support or conventional XAI-first approaches? We address this question by comparing diagnostic accuracy, diagnostic confidence, and perceived usefulness. In addition to statistical significance testing, we evaluate intervention impact through effect size analyses.
  • RQ2b: Do FAI-based decision-support systems slow down the decision-making process compared with conventional XAI-first support? We address this question by comparing survey completion times.
  • RQ2c: Is the risk of automation bias and algorithm aversion associated with FAI-based decision-support systems lower than that associated with XAI-first support? We address this question by comparing over-reliance (used as a proxy for automation bias [45]) and under-reliance (used here as a proxy for algorithm aversion).

2. Materials and Methods

This study adopted a mixed-methods design to investigate how different human–AI interaction protocols influence clinical decision making. The methodological framework consisted of two complementary components. First, interviews with legal and psychological experts were conducted to identify design requirements for AI-assisted decision-support systems. Second, these insights were used to motivate and contextualize a user study comparing alternative human–AI interaction protocols.
The background review presented in Section 1 served not only to position the present work within the broader context of explainable AI and human–AI interaction, but also to guide the methodological development of both study components. Specifically, the literature review informed the development of the interview themes, guides, and questions, while providing the conceptual basis for the design of the user study. In particular, it motivated the selection of the interaction protocols under comparison, the overall structure of the user study, and the user-centered outcomes assessed throughout the survey. The following sections describe the qualitative and quantitative components of the study.

2.1. Interviews with Domain Experts

To investigate the legal, ethical, and psychological requirements of AI-assisted decision-support systems, we conducted a qualitative study involving experts in law and psychology. The interviews were designed to explore themes related to explainability, automation bias, responsibility, human oversight, and the evaluation of AI-assisted decision support. The collected data were subsequently analyzed through thematic analysis.

2.1.1. Organization of the Interviews

To investigate how clinical decision-support systems should be designed to promote positive technology impact while satisfying legal, ethical, and psychological requirements, we conducted a qualitative study involving two domain experts: a full professor (male) of general psychology at the University of Bergen (Norway) with expertise in human decision-making and cognitive biases, with four years of experience working with AI; an associate professor (female) at the Faculty of Law of the University of Bergen, with approximately five years of experience in AI regulation and liability.
These disciplines were selected to provide complementary perspectives on the interaction between humans and AI systems in high-stakes environments, specifically healthcare. Psychology provides insights into cognitive processes, decision-making biases, and human reliance on AI, whereas law contributes perspectives on accountability, responsibility, transparency, and regulatory compliance. Experts were purposively selected based on their established research activity and prior experience with AI-related topics within their respective domains.
Although only two experts were interviewed, the sample size is consistent with the concept of information power [46], according to which smaller samples may be sufficient when the study aim is narrow, participants possess highly relevant expertise, and the collected data are rich and focused. Given the targeted nature of the study and the complementary expertise of the participants in law and psychology, the objective was not to achieve thematic saturation but rather to obtain in-depth interdisciplinary insights into human–AI interaction and decision-support design.
The interview guides were developed following a targeted literature review on explainable AI, trustworthy AI, human oversight, automation bias, liability, and AI regulation in healthcare (Section 1) and were tailored to the expertise of each participant while maintaining a common thematic structure. Based on this review, three overarching themes were identified and used to structure the interviews: (i) the characteristics of sufficiently transparent and understandable explanations, (ii) the relationship between reliance on AI systems, automation bias, responsibility, and accountability in clinical decision making, and (iii) the adequacy of current evaluation methodologies and user studies for assessing the quality and legal compliance of AI-assisted decision support.
Separate interview guides were prepared for the psychologist and the legal expert, containing domain-specific questions while preserving a common thematic structure. Questions were intentionally designed to be open-ended and exploratory, encouraging participants to reflect on their professional experience and discuss both current challenges and potential solutions. A detailed description of the interview guides and questions is provided in Appendix A.
Semi-structured interviews were conducted individually. This format was chosen to ensure consistency across participants while allowing experts to elaborate on topics within their area of expertise. During the interviews, participants were invited to comment on the role of explainability, human oversight, trust calibration, automation bias, responsibility attribution, and interaction design in AI-assisted decision making. Particular attention was devoted to identifying design strategies capable of supporting clinicians while preserving critical thinking, professional autonomy, and accountability.
To reduce preparation effects and encourage spontaneous reflections, only a brief overview of the study objectives was provided before the interviews. All recordings and identifiable information were stored on secure university servers and deleted after completion of the analysis.
The interviews were conducted in person by C.A.S., who had not previously participated in studies involving Frictional AI, Judicial AI, or other human–AI interaction protocols. The interviews were audio-recorded, transcribed verbatim, and analyzed using a thematic analysis approach described in the following section. The resulting insights were subsequently synthesized into a set of design requirements that informed the development of the interaction protocols evaluated in this study.
Immediately after each interview, analytic notes and first impressions were recorded to support familiarization with the data and reflexive memoing throughout the analysis process. Reflexive memoing was employed to document analytic decisions and reduce the influence of preconceived assumptions. To support reflexivity, analytic notes and first impressions were revisited during theme development for the Thematic Analysis.

2.1.2. The Thematic Analysis

Data from the transcriptions of the interviews were analyzed using thematic analysis following the six-phase framework proposed by Braun and Clarke [47]. The analysis was conducted with the support of NVivo (QSR International) [48], which was used to organize transcripts, manage codes, and facilitate the identification of recurring patterns across interviews.
The analysis proceeded iteratively through six stages. First, the transcripts were read repeatedly to achieve familiarity with the data and to record initial observations and analytic notes. Second, meaningful text segments related to the research questions were assigned initial codes. Coding was primarily inductive and data-driven, whereas theme interpretation and organization followed a deductive approach informed by the existing literature on explainability, automation bias, responsibility, and human–AI interaction.
Third, related codes were grouped into candidate themes representing broader concepts relevant to the design, evaluation, and regulation of AI-assisted decision-support systems. Fourth, candidate themes were reviewed and refined through repeated comparison with the original transcripts to ensure coherence within themes and clear distinctions between themes. Fifth, themes were defined and named by identifying their central meaning and contribution to the research questions. Finally, the findings were synthesized and organized according to the three interview themes investigated in the study: (i) characteristics of understandable and transparent explanations, (ii) automation bias, responsibility, and human oversight, and (iii) evaluation methodologies for AI-assisted decision support. Representative quotations were also selected to illustrate the themes and are reported in the Results Section (Section 3.1) along with the identified themes.

2.2. Comparison of Human–AI Interaction Protocols

A user study was designed to compare alternative human–AI interaction protocols in a clinical decision-support scenario. The study involved a COVID-19 hospitalization risk prediction task and evaluated the impact of the protocols on different aspects, including diagnostic performance, confidence, and decision-making behavior.

2.2.1. Hospitalization Risk Prediction Task

The core of the user study is a diagnostic task consisting of predicting hospitalization risk for COVID-19 patients, formulated as a binary classification problem (“hospitalize” vs. “home”). The dataset consists of 660 chest X-ray images collected from patients treated at the emergency department of the IRCCS Policlinic San Matteo Foundation upon their arrival.
In preceding studies [43,49], radiomic features were extracted from the images and associated with personal and clinical data obtained from each patient’s electronic health record. An example of the collected patient information is provided in Appendix B.1, which describes how data were shown during the user study. All patient data were securely stored within the dedicated IT infrastructure of the IRCCS Policlinico San Matteo Foundation and fully anonymized before being used for research purposes. Access to the data was restricted to authorized physicians and data analysts involved in the study, in accordance with the approved study protocol, institutional data protection policies, and applicable ethical requirements. Here, 23-dimensional feature vectors were extracted and an eXtreme Gradient Boosting (XGBoost) classifier was trained for the binary classification task as described by Bergomi et al. [49].
The SHAP algorithm [50] was selected for generating visual explanations due to its simplicity, widespread use in the field of XAI, model-agnostic nature (making it suitable for the explanation of a ML model), and because it was identified as the preferred method by clinicians in previous studies in the context of ALFABETO [49]. Considering the binary nature of the model, SHAP values were always computed with respect to the positive output class (hospitalize): positive SHAP values are thus supporting the positive classification, negative values support the classification for “home”.
To maximize understandability, SHAP values were represented using the standard tornado plot visualization, showing positive contributions in magenta (left side) and negative contributions in blue (right side) ordered for decreasing importance. Features are labeled by name and the horizontal axis represents the magnitude of the feature importance score.
For classical explanations (used in the XAI-first and adjunct protocols), the model confidence score associated with the prediction is reported along with the prediction itself, before presenting the SHAP explanation. In contrast, for judicial protocols, in accordance with the principles of Judicial AI [22,38], SHAP explanations are always presented as juxtaposed pairs with no clue to the model’s predicted class. To create such visualization, the information present in the previously-described tornado plot is split into two plots (one for each class). SHAP values are displayed without their sign (and colored in gray), and are normalized in magnitude. No predicted class or confidence score are shown beforehand. For a detailed illustration of the different ways in which model outputs and explanations were presented under the XAI-first, adjunct, and judicial protocols, refer to Appendix B.1.

2.2.2. User Study Design

The user study was conceived as a multi-arm study involving residents from the IRCCS Policlinic San Matteo Foundation in Pavia with varying levels of experience in chest X-ray interpretation and belonging to three different hospital departments (Radiology, Infectious Diseases, and Emergency Medicine). The study was conducted within the context of the ALFABETO project [43,44]. It was implemented and administered through the KoboToolbox platform [51] in the form of a survey and aimed to assess the impact of different human–AI interaction protocols in a simple clinical decision-support task (hospitalize vs. home).
A short profiling questionnaire was initially distributed to potential participants through their institutional email addresses. The purpose of this preliminary survey was twofold: first, to estimate the available sample size and, second, to collect participant characteristics, including gender, department, approximate number of COVID-19 patients treated during their career, familiarity with AI, and trust in AI. Details of the profiling questionnaire are reported in Appendix B.2.
Responding clinicians were then randomized into three groups, each corresponding to a specific human–AI interaction protocol:
  • XAI-first (XAI)—for each case, participants were shown the model suggested prediction, the corresponding confidence score, and the SHAP explanation; this can be considered as the state of the art in clinical decision support.
  • Judicial (JAI)—a pure judicial protocol in which participants were shown only the juxtaposed SHAP explanations associated with the two classes.
  • Adjunct (AAI)—a pure adjunct protocol in which participants first provided an unsupported diagnosis and subsequently revised it after being presented with the model prediction, confidence score, and SHAP explanation.
A total of ten survey versions was created by stratifying and distributing 100 patient cases across the surveys. Consequently, each questionnaire consisted of the evaluation of 10 cases and the case mix presented in each version can be considered balanced. In accordance with the approved study protocol, all patient images and accompanying tabular information were fully anonymized before being uploaded to the KoboToolbox platform, where they were displayed to participants exclusively in anonymized form during the survey administration.
Participants were not informed about the nominal performance of the underlying XGBoost model before or during the study. Moreover, the evaluated cases per questionnaire were intentionally balanced such that, approximately, half of the AI recommendations corresponded to correct model predictions and half to incorrect predictions, with the cases properly randomized. This design choice was adopted to prevent participants from assuming a priori that the decision-support system was highly reliable, thereby reducing potential biases in the evaluation of trust and reliance behaviors.
For each case, participants were shown the patient’s tabular information and chest X-ray image and were asked to provide a classification (hospitalize vs. home). In addition, participants were asked to rate, using a six-point Likert scale, their confidence in the decision (1 = not at all confident; 6 = extremely confident), the perceived complexity of the case (1 = extremely easy; 6 = extremely challenging), and, when applicable, the usefulness of the decision support (1 = not at all useful; 6 = extremely useful). Subjects were briefly introduced to the survey with instructional videos (Appendix B.3).
For the XAI and JAI protocols, decision support was provided simultaneously with the image and tabular information. In contrast, under the AAI protocol, participants were first required to provide an initial unsupported decision, then a decision support (equivalent to that used in the XAI condition) was displayed, after which participants were asked to provide a revised opinion. During the statistical analysis, the initial decision collected under the AAI protocol was used as a baseline representing a no-support condition (NoAI).
The overall design of the user study was informed by the qualitative findings of the expert interviews, which provided practical guidance for the development and evaluation of the proposed human–AI interaction protocols. As described in greater detail in Section 3.1, the interviews reinforced the importance of evaluating AI-assisted decision support beyond diagnostic performance alone, emphasizing the role of surveys in capturing user-centered dimensions. They also highlighted contrastive reasoning as a promising strategy for encouraging critical evaluation of AI recommendations. These considerations motivated both the selection of the interaction protocols compared in this study—namely, the adjunct and judicial protocols, which were explicitly regarded by the experts as promising approaches for promoting meaningful human oversight in the context of Frictional AI—and the inclusion of user-centered outcome measures, such as diagnostic confidence, perceived usefulness, over-reliance, and under-reliance, alongside conventional performance metrics. Finally, the interviews emphasized the importance of exposing clinicians to multiple cases (namely more than five) in order to obtain a realistic assessment of AI-assisted decision making, a recommendation that was incorporated into the study design.

2.2.3. Statistical Evaluation

G*Power (v3.1) [52] was used to perform sample size calculations. The analysis focused on the primary outcomes (in the scope of RQ2), considering a two-tailed t-test for two independent groups. The significance level ( α ) was set to 0.05 and statistical power ( 1 β ) to 0.80, while all other parameters were left at their default values. The analysis indicated a minimum required sample size of 64 participants per group. Meanwhile, repeating the procedure in the case of matched pairs, the analysis indicated a minimum of 45 subjects.
All statistical analyses were performed using R (v4.3.0) on data aggregated at the participant level. Data normality was assessed using the Shapiro–Wilk test. For independent-group analyses, homoscedasticity was evaluated using Levene’s test for normally distributed variables and the Brown–Forsythe test otherwise. Group means were compared using Student’s t-test when assumptions of normality and equal variances were met, Welch’s t-test when variances differed, and the Mann–Whitney U test for non-parametric data. For comparisons between paired groups (i.e., AAI vs. NoAI), paired t-tests were used for normally distributed differences, whereas the Wilcoxon signed-rank test was used otherwise.
Group comparisons were first performed to assess statistical differences between conditions (e.g., H 0 : JAI accuracy = NoAI accuracy). Whenever a significant difference was detected, a one-sided non-inferiority test was subsequently performed. Rejection of the non-inferiority hypothesis was interpreted as evidence of superiority. When no difference between groups was observed, the p-value from the initial comparison was reported; otherwise, the p-value from the non-inferiority test was reported.
Given the expected limited sample size, results are reported as mean ± standard deviation (SD). In addition, effect sizes were computed for diagnostic accuracy and confidence to quantify intervention impact in a clinically interpretable manner.
Effect sizes comparing the intervention conditions (XAI, JAI, and AAI) with the control condition (NoAI) were computed following established practice [38,53]. Accuracy and confidence were first aggregated at the participant level, after which Cohen’s d with Hedges’ correction for small samples was calculated. Although AAI and NoAI involved matched observations, the independent-group formulation was adopted for all analyses to ensure methodological consistency and facilitate comparisons across protocols. 95%-confidence intervals were estimated using bias-corrected and accelerated (BCa) bootstrap resampling [54]. Additional subgroup analyses stratified by participant characteristics, perceived case complexity, and perceived usefulness were performed; details about data stratification are described in Appendix C.
Effect sizes were interpreted according to Cohen’s guidelines [55]: absolute values below 0.2 were considered negligible, values between 0.2 and 0.5 small, between 0.5 and 0.8 medium, and above 0.8 large. Following previous work on intervention impact [38], ± 0.2 was adopted as the threshold for practical relevance.

3. Results

This section presents the findings from both components of the study. First, the results of the thematic analysis of interviews with domain experts are reported, highlighting key legal and psychological considerations for the design of AI-assisted decision-support systems. Subsequently, the outcomes of the user study are presented, including statistical comparisons across interaction protocols and effect size analyses of the intervention.

3.1. Themes Emerging from Expert Interviews

The interviews lasted approximately 65 min (legal expert) and 44 min (psychology expert), respectively, and the subsequent thematic analysis identified three overarching domains corresponding to the research questions: (i) characteristics of understandable explanations; (ii) over-reliance on AI and accountability; and (iii) evaluation of support quality and regulatory compliance. Figure 1 provides an overview of the resulting thematic structure and associated sub-themes. The findings for each theme are presented in the following sections.
Understandable explanations. As the psychology expert explained, “explanations should sort of connect to the mental model that the user already has of the phenomenon”, rather than the model’s internal mechanics. They stressed that explanations should connect to clinicians’ existing reasoning processes and prioritize the information they actually use when making decisions: explanation design should begin by understanding “what are the issues or the facts that this clinician is interested in”.
The legal expert similarly argued that explanations should be tailored to the recipient. The legal expert repeatedly emphasized the importance of identifying the recipient of the explanation, asking “who is actually receiving this explanation?”. The expert further commented on this idea of tailoring explanations that “for a clinician, this is definitely something that is manageable … however, for a patient, it might be very difficult”. A recurrent recommendation was to present information selectively, starting from the most influential factors and avoiding unnecessary technical detail.
Contrastive reasoning. Both interviews highlighted the value of “contrastive” explanations—the psychology expert suggested that explanations should present both arguments supporting and opposing the recommendation. This converged with the rationale behind Judicial AI, which exposes competing evidence rather than a single authoritative output. Furthermore, the legal expert explicitly mentioned contrast-based explanations as a promising approach, mentioning for example counterfactual explanations.
Simplicity and presentation format. The psychology expert raised concerns that numerical probabilities and complex SHAP plots may be difficult to integrate cognitively and could trigger biases such as confirmation bias [56] or base-rate neglect [33]. Practical suggestions included replacing detailed numerical displays with qualitative categories (e.g., low, medium, high confidence), using verbal summaries, and presenting a small number of strong and weak arguments rather than a full feature list.
The legal expert argued that understandability should take precedence over technical completeness, noting that “without making it understandable, you cannot really expect trust”.
Right to explanation. The legal expert clarified that current European regulations provide only limited and still-evolving rights regarding explanations in AI-supported healthcare. Under the GDPR, the relevant rights concern information about the general logic of automated processing rather than detailed explanations of individual decisions, and the applicability to decision-support systems remains uncertain. The AI Act was described as broadening transparency obligations, but without prescribing specific explainability techniques.
Reducing over-reliance on AI. Both experts agreed that clinicians must retain the competence and opportunity to disagree with AI recommendations. The legal expert stressed that meaningful human oversight requires the ability to review and modify AI-supported decisions, whereas the psychology expert recommended explicit warnings about model limitations and excluded factors.
The adjunct protocol, in which clinicians first make an unsupported decision before viewing AI output, was considered a potentially useful strategy for mitigating automation bias. Judicial AI was also viewed positively because the presentation of competing evidence may increase cognitive engagement and discourage passive acceptance of AI recommendations.
Responsibility and accountability. According to the legal expert, responsibility for clinical decisions remains primarily with clinicians and healthcare institutions because the system functions as decision support rather than a fully autonomous decision maker. However, the precise allocation of responsibility among clinicians, hospitals, and developers was described as an open and unresolved issue.
Evaluating explanation quality. Regarding the survey design, both experts considered survey-based assessment useful for studying understandability, trust, confidence, and perceived usefulness, although they noted that no established legal metric currently exists for explainability compliance.
The psychology expert suggested incorporating open-ended questions to investigate clinicians’ reasoning processes and mental models, for example asking “how much of the available information did they use to make up their mind?” and from this “try to classify how sophisticated knowledge or explanations they had gained from reading the report.”.
Regarding study design, the psychology expert considered repeated exposure to multiple cases necessary for users to develop a realistic understanding of the AI system, specifically advising that “you need at least more than five” cases. The expert also highlighted the risk of order effects and confirmation bias, observing that “some clinicians are going to make up their mind based on what they first see” and recommending randomization of information presentation when possible.

3.2. User Study Results

This section reports the findings of the quantitative evaluation of the three human–AI interaction protocols. First, descriptive statistics and statistical comparisons are presented for diagnostic accuracy, confidence, perceived usefulness, decision-making time, and reliance patterns. Subsequently, effect size analyses are reported to assess the magnitude of the intervention effects, both overall and across participant- and case-specific subgroups.

3.2.1. Quantitative Results and Statistical Analysis

A total of 147 residents responded to the profiling questionnaire and were randomized into the three study branches. However, only 96 participants completed the final survey. In total, 24 responses were collected for the XAI branch, 31 for the JAI branch, and 41 for the NoAI/AAI branch. Although the achieved sample size was smaller than that estimated by the a priori power analysis for both independent and matched groups, the study still allowed the identification of statistically significant differences between conditions. Nevertheless, the reduced sample size may have limited the ability to detect smaller effects.
Table 1 summarizes the baseline characteristics of the 96 participants who completed the study, reporting both the overall cohort and the distributions stratified by study group. These characteristics correspond to the information collected through the profiling questionnaire (Appendix B.2), namely sex, clinical specialty, residency year, previous experience managing COVID-19 patients, familiarity with AI, and skepticism toward AI.
Figure 2 presents a comparison of the study conditions through box-plots illustrating diagnostic accuracy and confidence, perceived usefulness of the support, survey completion time (minutes), over-reliance, and under-reliance. Before analysis, measurements were aggregated at the participant level. Descriptive statistics for each outcome and experimental condition are summarized in Table 2:
Accuracy and confidence. Concerning accuracy (Figure 2a), the NoAI condition achieved the highest mean performance, whereas the AI-assisted approaches yielded slightly lower accuracies, with JAI showing the closest performance to the baseline. Nevertheless, no statistically significant differences were observed, except for AAI, which performed significantly worse than both NoAI ( p = 2.7 × 10 4 ) and JAI ( p = 2.8 × 10 4 ).
A similar pattern emerged for diagnostic confidence (Figure 2b). Although no statistically significant differences were detected, AI-assisted conditions generally exhibited higher average confidence scores than the baseline. Notably, both accuracy and confidence showed reduced variability under AI-assisted conditions compared with the NoAI baseline.
Perceived usefulness. Figure 2c compares the usefulness of the support across the three intervention groups. The dashed horizontal line represents the random baseline ( μ = 3.5 ), corresponding to the midpoint of the six-level Likert scale. Neither JAI nor AAI achieved usefulness ratings significantly different from the random baseline, whereas XAI demonstrated evidence of perceived usefulness exceeding randomness ( p = 2.6 × 10 3 ). No statistically significant differences were observed among the three intervention groups.
Survey duration. Survey completion times are reported in Figure 2d. Although no statistically significant differences were observed among the groups, both friction-based protocols required approximately two additional minutes on average compared with the conventional XAI-first approach. Furthermore, JAI exhibited greater variability in completion times than either XAI or AAI. Although this difference is reflected in both the box-plots and descriptive statistics, it did not reach statistical significance: Levene’s test did not reject the assumption of homoscedasticity for JAI versus XAI ( p = 3.2 × 10 1 ) or JAI versus AAI ( p = 7.3 × 10 2 ). This pattern may suggest heterogeneous user engagement, with some participants investing substantially more effort than others when interacting with the judicial protocol.
Automation bias and algorithmic aversion. Reliance patterns were investigated through over-reliance (Figure 2e), used as a proxy for automation bias, and under-reliance (Figure 2f), used as a proxy for algorithm aversion. Interestingly, JAI exhibited the lowest mean over-reliance score, even lower than the NoAI baseline, where agreement with AI predictions is by definition random. Although the difference was not statistically significant ( p = 8.6 × 10 1 ), this finding is consistent with the intended objective of judicial protocols. In contrast, AAI showed the highest level of over-reliance and significantly exceeded all other conditions: NoAI ( p = 2.9 × 10 6 ), XAI ( p = 6.8 × 10 3 ), and JAI ( p = 1.6 × 10 6 ).
The findings for under-reliance were less conclusive. JAI exhibited a slightly higher level of under-reliance than the other conditions, reaching statistical significance only when compared with AAI ( p = 6.4 × 10 3 ).
For both over-reliance and under-reliance, XAI showed the greatest dispersion in participant responses. Nevertheless, homoscedasticity was rejected only for XAI versus JAI ( p = 2.9 × 10 2 ) and XAI versus AAI ( p = 2.9 × 10 2 ) in the case of over-reliance, and for XAI versus JAI ( p = 4.3 × 10 2 ) in the case of under-reliance. Moreover, on average, users in the XAI condition exhibited less favorable reliance behavior than those in the NoAI condition.

3.2.2. Evaluation of the Impact

Figure 3 reports the effect sizes for diagnostic accuracy and confidence associated with the different intervention protocols relative to the no-support condition (NoAI). Consistent with the findings presented in the previous section, the average effect on accuracy was negative across all AI-assisted conditions. However, comparing groups via the simulation-based approach—evaluating the BCa 95% confidence intervals—only AAI showed evidence of a detrimental effect (non-negligible negative effect size). XAI also showed a small negative effect on average but this conclusion is still limited by the overlapping between the its BCa confidence interval and the non-negligibility margin. Finally, the slightly-negative effect associated with JAI can be considered negligible, given the extensive overlap with the area of negligibility. Although none of the intervention protocols improved diagnostic accuracy, JAI yielded the least detrimental performance among the three approaches.
The results for diagnostic confidence were comparatively more positive. XAI exhibited a clear positive effect, with confidence intervals entirely exceeding the non-negligibility threshold. JAI and AAI provided also carried a positive effect but this conclusion is limited by the overlapping between the their confidence interval and the non-negligibility margin.
To gain further insights into the acceptance and effectiveness of friction-based interaction protocols, Figure 4 presents a stratified analysis of intervention effects across participant- and case-related subgroups. Specifically, the effect size was evaluated according to participants’ skepticism toward AI, familiarity with AI, clinical specialty (hospital department), level of professional experience, and perceived case complexity—see Appendix C.
Figure 4a,b report the results for JAI: the stratified analysis reveals substantial heterogeneity between subgroups, although most confidence intervals indicate negligibility. Comparatively, some groups appeared to benefit from judicial protocols more than others experiencing neutral or slightly-negative effects on average. For example, skeptical residents showed a more favorable average effect than non-skeptical residents. Likewise, residents with limited clinical experience exhibited a modest positive effect on diagnostic accuracy. Emergency medicine residents also showed a positive effect on average; however, the considerable width of the BCa confidence intervals suggests substantial uncertainty, preventing any definitive conclusions from being drawn.
Regarding diagnostic confidence, JAI generally produced positive effects across most subgroups. A meaningful negative effect was observed only in the subgroup of emergency medicine residents. In contrast, fully positive effects were observed among infectious disease residents and among participants with limited clinical experience.
The results for AAI (Figure 4c,d) largely mirrored those observed in the aggregate analysis. On average, most subgroups experienced either neutral or negative effects on diagnostic accuracy, with several BCa confidence intervals suggesting negative, non-negligible effects relative to the baseline condition. For diagnostic confidence, the majority of subgroups showed a negligible effect size.

4. Discussion

This study investigated how different human–AI interaction protocols influence clinical decision making and explored the legal and psychological requirements that should guide the design of AI-assisted decision-support systems. Through a mixed-methods approach combining expert interviews and a user study, we evaluated the potential of Frictional AI as an alternative to conventional XAI-first decision support. This section presents the discussion of the results described in the previous one and it is organized according to the research questions introduced in Section 1.3.

4.1. The Good Design of Clinical Decision Support

In the context of RQ1a, the interviews with the psychology and legal experts converged on a common conclusion: explainability should primarily support human understanding rather than faithfully expose the internal mechanics of AI models. Both experts emphasized that explanations should be simple and adapted to the recipient and should connect with the user’s existing mental models. These findings are consistent with established human-centered design principles [4,30,31] and with recent guidelines and governance frameworks for trustworthy AI [13,14,15], which emphasize usability and understandability as prerequisites for meaningful human oversight and the possibility for users to overrule AI-assisted recommendations [3].
Interestingly, both experts also highlighted the value of contrastive reasoning. Rather than presenting a single authoritative recommendation, explanations should expose competing arguments and encourage users to actively evaluate alternative interpretations. This observation provides qualitative support for the rationale underlying Judicial AI [17], which was specifically designed around “juxtaposed” evidence. The convergence between independent expert opinions and the theoretical foundations of JAI suggests that contrastivity may represent a promising direction for future explainable decision-support systems.
This concept is not entirely new within the field of AI. As discussed by Miller [57], contrastivity is a fundamental characteristic of human reasoning and human-generated explanations and constitutes the foundation of both the established field of counterfactual explanations [58,59,60] and the more recent field of contrastive explanations [60,61,62]. The former focuses on identifying the minimal changes required to obtain a different prediction, whereas the latter seeks to explain why one specific outcome (the fact) was predicted instead of a plausible alternative (the foil). Within the context of judicial protocols, both paradigms may provide valuable foundations for constructing juxtaposed evidence beyond—or in association with—traditional XAI approaches. In this respect, AlRegib and Prabhushankar [63] argue that classical, contrastive, and counterfactual explanations can contribute together to the generation of “abductive” evidence, a concept closely related to the differential diagnosis process routinely employed by clinicians in everyday practice. Differential diagnosis, indeed, consists of the use of symptoms and signs found on the patients to generate a set of hypotheses enabling abductive inference, ultimately leading to the identification of a diagnosis.
Regarding over-reliance and responsibility (RQ1b), the interviews revealed that meaningful human oversight requires more than merely providing explanations. Clinicians must retain both the ability and the responsibility to challenge AI recommendations. This finding is particularly relevant given the increasing attention devoted by the AI Act [9] and related regulations to human oversight mechanisms. Notably, the legal expert emphasized that responsibility remains primarily with clinicians and healthcare institutions, despite the growing use of AI support and the consequent tendency of automation complacency [23,25] and medical deresponsibilization [26], in which users tend to defer responsibility to machines. Consequently, decision-support systems should be designed not only to improve performance but also to preserve professional agency and accountability [15,23].
Finally, concerning RQ1c, both experts considered user studies a valuable instrument for evaluating explainability and human–AI interaction, although neither identified existing legal metrics capable of directly demonstrating regulatory compliance. This finding reinforces the importance of empirical validation studies and suggests that user-centered evaluation should complement technical assessments of the underlying model and explainability.
Several studies have highlighted the importance of evaluating the human dimension of AI-assisted decision making through user studies and surveys. Beyond conventional performance metrics such as diagnostic accuracy and confidence, user studies enable the assessment of human-centered aspects including understandability, perceived usefulness, cognitive support, and trust in the system. Furthermore, they provide a means to investigate human–AI interaction dynamics that may ultimately lead to undesirable consequences such as automation complacency, deskilling, reduced agency, and deresponsibilization [23,24,26,28,29]. In particular, automation bias [18] and algorithmic aversion [6] can be quantified by comparing clinician behavior and performance before and after exposure to AI assistance [17]. Besides, such methodologies make it possible to characterize different patterns of human–AI interaction and evaluate the overall technology impact [17] on decision making. When pre-support decisions are unavailable, over-reliance and under-reliance have been proposed as practical proxies for automation bias [34,45] and algorithmic aversion, respectively.
The user study conducted in this work was specifically designed to evaluate these dimensions, with particular emphasis on automation bias, algorithmic aversion, and the overall impact of different interaction protocols on clinical decision making.

4.2. Friction-Based Systems and User Performance

The quantitative findings provide a nuanced perspective on Frictional AI and help address RQ2a. Contrary to our initial hypothesis, neither friction-based nor conventional XAI-based protocols improved diagnostic accuracy relative to the no-support condition. However, these findings should be interpreted within the context of the selected decision-support task, which combined relatively simple input modalities (tabular patient information and chest X-ray images) with a clinically challenging prediction problem for less experienced clinicians.
Despite the absence of overall improvements, Judicial AI consistently achieved more favorable outcomes than the alternative intervention protocols. Its diagnostic accuracy was the closest to that observed in the no-support condition and thus higher than that achieved by the state-of-the-art XAI-first protocol. Furthermore, although the BCa confidence intervals remained wide, the average detrimental effect size associated with JAI was limited, corresponding to a “negligible” interpretation according to Cohen’s guidelines [55], while XAI exhibited on average a small negative effect on diagnostic accuracy.
These findings partially overlap with those reported by Cabitza et al. [38], where JAI was also associated with a negligible overall effect size despite producing improvements in accuracy. In that study, the most relevant findings emerged from subgroup analyses, which revealed “large” positive effect size among more experienced clinicians. Although our results also suggest heterogeneity across participant groups, none of the observed differences reached statistical significance. Interestingly, non-zero, less experienced residents appeared to derive greater benefits from JAI, whereas the effect among more experienced participants was close to neutral or slightly negative. However, direct comparisons between the two studies should be interpreted cautiously, since the previous investigation involved a smaller but substantially more experienced population, whereas our sample consisted exclusively of residents. Finally, it is noteworthy that skeptical participants appeared to derive greater benefits from judicial protocols than those who reported higher trust in AI, although the corresponding effect sizes remained “small”.
The findings for adjunct protocols were considerably less encouraging. Consistent with previous evidence [41], AAI demonstrated a detrimental effect on diagnostic accuracy. The corresponding effect size was negative and, in some subgroups, approached levels conventionally interpreted as “large”. This observation is further supported by the statistically significant reduction in performance relative to the no-support condition.
The confidence results provide a different perspective. All AI-assisted protocols tended to increase clinicians’ confidence, with XAI showing the strongest positive effect. This finding is consistent with previous studies reporting that explainability and AI recommendations can increase users’ perceived certainty both in general [22] and specifically to the assessment of judicial protocols [38]. Although no statistically significant differences emerged, on average, the estimated effect sizes ranged from small in the FAI conditions to medium for XAI-first. Notably, XAI was also the only protocol whose perceived usefulness was shown to be significantly greater than the random baseline, suggesting that participants appreciated receiving a direct prediction accompanied by a conventional explanation.
The subgroup analyses for confidence largely confirmed the overall findings, showing generally positive effect sizes for JAI. The magnitude of these effects ranged from small (approximately 0.2) to medium, with the largest positive effects observed for diagnostic confidence among infectious disease residents and participants with non-zero, limited clinical experience. Interestingly, among emergency medicine residents, the positive effect observed on diagnostic accuracy was accompanied by a markedly negative effect on diagnostic confidence, reaching a very large negative effect size according to Cohen’s guidelines. In contrast, the effects associated with AAI were generally negligible, with most effect sizes falling within the interval between −0.2 and 0.2.
However, increased confidence does not necessarily translate into improved diagnostic performance. From a human-factors perspective, this dissociation may be problematic because it can promote inappropriate trust calibration and potentially increase automation bias. As discussed in Section 1.1, automation complacency [23,25], namely an excessive confidence in automated systems, may negatively affect human–AI interaction and ultimately result in decision deference and, in its most extreme form, deresponsibilization [26]. On the other hand, elevated confidence may also reflect the presence of confirmation bias [56], whereby users selectively attend to information that supports their initial beliefs or decisions. This phenomenon is well documented in healthcare [64] and was explicitly highlighted both in the literature [33] and during the expert interviews (Section 3.1) as a potential risk associated with AI-assisted decision support. The analysis of automation bias and algorithmic aversion, discussed in the following section in relation to RQ2c, may provide additional insight into these dynamics.

4.3. The Benefits of Cognitive Friction

Concerning RQ2b, the analysis of completion times revealed that friction-based protocols required approximately two additional minutes compared with the conventional XAI-first approach. Although these differences did not reach statistical significance, they suggest that participants invested more effort when interacting with frictional systems. Overall, these findings are consistent with the core rationale of cognitive friction [22], namely introducing obstacles into the decision-making process to slow down judgments and encourage more deliberate reflection on the available evidence.
In the AAI condition, the increased completion time can be largely attributed to the requirement of making two separate decisions: one before and one after exposure to AI support. In contrast, the delay observed for JAI likely stems from the absence of a direct, oracular recommendation and the consequent need for users to actively interpret and compare the provided explanations. Conversely, the shorter completion times observed in the XAI condition may indicate that participants frequently relied on the model prediction while only partially engaging with the SHAP explanations. Nevertheless, the greater variability in completion times observed under JAI further suggests heterogeneous engagement strategies, with some participants carefully evaluating the competing evidence and others interacting with the protocol more superficially.
The most compelling evidence regarding the value of cognitive friction emerged from the analysis of reliance patterns (RQ2c). JAI exhibited the lowest level of over-reliance, used in this study as a proxy for automation bias [45], among all intervention conditions and even lower values than the baseline in which any trace of reliance is completely random. Although the difference did not reach statistical significance, the direction of the effect is highly consistent with the theoretical objective of Judicial AI and the points raised by the psychologist and legal expert in the course of the semi-structured interviews: by requiring users to evaluate competing explanations rather than providing a single recommendation, JAI appears capable of discouraging passive acceptance of AI-generated advice.
In contrast, AAI produced the highest level of over-reliance despite being explicitly designed to mitigate automation bias. One possible explanation is that requiring users to formulate an initial decision may increase their commitment to the subsequent interaction with AI support: once exposed to the recommendation, participants may interpret the AI output either as confirmation or correction of their initial judgment and consequently assign greater weight to it. Such behavior may be associated with well-known cognitive phenomena including confirmation bias [33,56,64] and automation complacency [23,25]. Although this interpretation remains speculative, it highlights the complexity of human–AI interactions and suggests that adjunct protocols do not necessarily mitigate automation bias in practice—despite their potential highlighted by Cabitza et al. [22] while introducing Frictional AI methods and by the experts in the interview phase.
Nevertheless, these findings are consistent with another previous observation by Cabitza et al. [41] as discussed in Section 1.2 and align with similar evidence reported in the literature. For example, Buçinca et al. [37] compared multiple interaction protocols, including an adjunct design in which users first made an independent decision and subsequently received a contrastive explanation [60,61,62] comparing the model prediction with the first user’s decision (as the foil). Although their results in terms of classification accuracy were broadly comparable to those obtained with XAI-first approaches, the authors reported higher levels of over-reliance on AI in the adjunct condition. Furthermore, their evaluation of subjective dimensions—including perceived competence and autonomy, relatedness to AI, interest and enjoyment, and mental demand—showed that adjunct protocols resulted in a significantly poorer user experience across most evaluated dimensions. Taken together, these findings suggest that exposing users to AI corrections after they have committed to an initial decision may inadvertently increase reliance on AI while simultaneously reducing the quality of the interaction experience.
The findings regarding under-reliance were less straightforward to interpret. Although JAI exhibited slightly higher levels of under-reliance than the other conditions, the magnitude of the effect remained relatively limited. From a theoretical perspective, this observation is not entirely unexpected: interaction protocols designed to reduce automation bias necessarily encourage users to critically evaluate AI recommendations and preserve their autonomy in the decision-making process. As a consequence, a moderate increase in under-reliance may represent an unavoidable trade-off associated with reducing passive acceptance of AI outputs.
Future studies should further investigate this balance between automation bias and algorithm aversion, as the optimal interaction protocol is likely one that promotes critical engagement without inducing systematic rejection of correct AI advice.
As a final remark, we note that participants were not informed of the nominal performance of the underlying AI model and were presented with both correctly and incorrectly classified cases. This design choice prevented participants from assuming a priori that the decision-support system was highly reliable or inherently trustworthy, thereby reducing potential user biases arising from assumptions of near-perfect system performance rather than from the interaction protocol itself.

4.4. Judicial AI vs. Adjunct Protocols

Although the quantitative evaluation was primarily designed to address RQ2 by comparing friction-based protocols against the no-support and conventional XAI-first conditions, the results revealed substantial differences between the two friction-based approaches themselves. Since both JAI and AAI were conceived as implementations of the Frictional AI paradigm—introducing cognitive friction to mitigate automation bias—yet exhibited markedly different effects on diagnostic performance and reliance patterns, a direct comparison between them provides additional insight into the mechanisms through which cognitive friction influences human–AI interaction. The following analysis therefore focuses on the relative strengths and limitations of the two protocols, with the aim of identifying which characteristics of friction-based interaction design are most likely to promote positive technology impact.
From the perspective of diagnostic accuracy, JAI consistently outperformed AAI. This difference was statistically significant in the aggregate analysis and was reflected in the effect size evaluation, where JAI exhibited a negligible overall effect relative to the no-support condition, whereas AAI produced a large negative effect. Importantly, several subgroup analyses suggested that the detrimental effect of AAI may become even more pronounced in specific populations, such as familiar residents and high-experience ones. These findings indicate that introducing friction alone is insufficient to improve decision quality and that the manner in which friction is operationalized is crucial.
The comparison becomes even more striking when considering reliance patterns. JAI achieved the lowest level of over-reliance among all evaluated protocols, whereas AAI exhibited the highest. The difference between the two approaches was statistically significant, representing one of the strongest findings of the study. This result suggests that forcing users to evaluate competing explanations may be more effective at preserving critical thinking than requiring an unsupported preliminary decision followed by exposure to an AI recommendation; while both protocols increase cognitive effort, only JAI appears capable of reducing the tendency to defer to AI advice.
A similar pattern emerged for under-reliance. Although JAI exhibited slightly higher levels of algorithm aversion than AAI, the magnitude of this increase remained relatively limited. Interestingly, the difference between the two protocols reached statistical significance, suggesting that the reduction in automation bias achieved by JAI may come at the cost of a modest increase in users’ willingness to challenge AI recommendations. However, unlike AAI, this increase in under-reliance was not accompanied by a deterioration in diagnostic performance, indicating that users were not simply rejecting AI advice indiscriminately. This finding suggests that encouraging users to question AI recommendations does not necessarily impair decision quality and may instead contribute to a healthier calibration of trust.
The comparison between JAI and AAI is particularly interesting when interpreted in light of the expert interviews. Both experts emphasized that clinicians should remain actively engaged in the decision-making process and should be encouraged to critically evaluate AI recommendations rather than passively accept them. Furthermore, both highlighted the importance of presenting alternative perspectives and supporting contrastive reasoning: from this perspective, JAI appears more closely aligned with the legal and psychological requirements identified during the qualitative phase of the study. In brief, Judicial AI behaves like an “honest broker” [65], presenting competing evidence and alternative interpretations and this fact allows the human decision maker to retain professional agency and forces them to actively compare alternative hypotheses—thus, operationalizing many of the principles advocated by the experts, including meaningful human oversight and active engagement with the decision process.
Conversely, the results obtained with AAI suggest that simply delaying exposure to AI recommendations may not be sufficient to achieve these objectives. Although adjunct protocols are often motivated by the intention of reducing automation bias, our findings indicate that the subsequent presentation of a direct recommendation may still favor or, even, encourage deference to AI outputs. This interpretation is consistent with previous works [37,41], which reported increased over-reliance on AI and poorer subjective user experiences under adjunct interaction designs.
These results are also consistent with the theoretical considerations discussed in Section 1.2; while previous studies have shown that adjunct-based interaction protocols may negatively affect decision quality [41], Cabitza et al. [38] demonstrated that combining adjunct protocols with judicial principles can produce more favorable outcomes. Our findings reinforce the interpretation that the benefits previously observed are primarily attributable to the judicial paradigm itself rather than to the adjunct administration strategy.
Taken together, these findings suggest that the benefits of cognitive friction depend less on the amount of friction introduced and more on its nature. Judicial protocols appear to channel cognitive effort toward the evaluation of competing evidence and alternative explanations, whereas adjunct protocols primarily delay access to the AI recommendation without fundamentally changing the way users interact with it. Consequently, among the two friction-based approaches evaluated in this study, JAI appears to offer the most promising balance between maintaining diagnostic performance, reducing automation bias, and preserving meaningful human oversight.

4.5. Limitations and Future Work

Several limitations should be acknowledged. First, the achieved sample size was substantially smaller than the target estimated through the a priori power analysis. Consequently, many comparisons were underpowered, and the findings should be interpreted with caution: while several statistically significant differences emerged despite the limited number of participants, the study may have lacked sufficient power to detect smaller effects. Consequently, the absence of statistical significance should not be interpreted as evidence of equivalence between conditions and warrants confirmation in larger cohorts.
Another limitation, directly related to the small sample size, is the heterogeneity of the participant population. Although participants were randomly assigned to the three study arms and the groups were comparable overall (Table 1), some differences in baseline characteristics remained. For example, the AAI group included a 20%-higher proportion of participants familiar with AI than the XAI and JAI groups. Likewise, the XAI branch contained a larger proportion of residents with no previous experience managing COVID-19 patients (54.2%) than the JAI (41.9%) and AAI (39.0%) branches, whereas participants with experience treating more than 100 COVID-19 patients were much more frequently represented in the AAI group (34.1%) than in the XAI group (16.7%). Differences were also observed in the sex distribution, with the XAI branch including a lower proportion of female participants and a correspondingly higher proportion of male participants than the two frictional-based branches. In contrast, specialty and residency year were generally well balanced across the three protocols.
These observations illustrate the variability that may arise when randomization is performed on relatively small samples and suggest that part of the observed differences between interaction protocols may reflect baseline participant characteristics in addition to the intervention itself. For example, differences in AI trust and prior clinical experience may have influenced user-centered outcomes such as diagnostic confidence and perceived usefulness, while differences in AI familiarity may have contributed to the observed variability in survey completion time. Consequently, these measures should not be interpreted as reflecting exclusively the effect of the evaluated friction-based interaction protocols. Future studies should therefore recruit larger cohorts and consider stratified randomization or covariate-adjusted analyses to further reduce the potential influence of baseline heterogeneity.
Similarly, the qualitative component of the study (the semi-structured interviews) was limited by the inclusion of only two domain experts. Moreover, although the experts provided complementary legal and psychological perspectives, both were affiliated with the University of Bergen and had comparable levels of experience with AI, potentially limiting the diversity of viewpoints captured by the thematic analysis.
Second, automation bias and algorithmic aversion were assessed indirectly through over-reliance and under-reliance metrics. Although these measures have been adopted in previous studies [34,45], they remain proxies and may not fully capture the complexity of human reliance behavior compared with dedicated technology-impact assessments [17]. Nevertheless, the practical applicability of these metrics highlights the need for future research aimed at developing robust methodologies capable of evaluating reliance behaviors without requiring a pre-support assessment phase.
Third, in line with the work of Buçinca et al. [37], we also advocate for studies that extend beyond performance and reliance metrics to investigate subjective user experience and long-term skill development. The interviews conducted in this study not only reinforced the importance of user-centered interaction designs but also highlighted the need to better understand how different interaction protocols influence responsibility attribution, accountability, and trust calibration; while several user-centered metrics are currently available to assess psychological and human-factor requirements, no established legal metrics currently exist for evaluating compliance with explainability-related regulatory requirements. We therefore encourage future work aimed at defining measurable and operational criteria for explainability from a legal perspective.
Fourth, beyond interaction design, the successful translation of friction-based decision support into clinical practice also depends on broader AI governance and risk management considerations. Although issues such as data governance, privacy preservation, digital sovereignty, and the management of systemic risks were outside the scope of the present work, they represent essential prerequisites for the responsible deployment of AI-assisted decision-support systems. Future work should therefore evaluate FAI-based interaction protocols within governance frameworks that ensure compliance with emerging regulations, robust data governance, continuous risk management, and meaningful human oversight throughout the AI life cycle, such as the European AI Act [9], the NIST AI Risk Management Framework [14], and the FUTURE-AI guidelines [15].
Finally, the user study focused on a single clinical task involving COVID-19 hospitalization decisions, and the experimental setting deliberately simplified the underlying clinical decision-making process. Although participants were provided with patient characteristics and chest X-ray images, real-world hospitalization decisions are typically informed by a broader set of information, including additional diagnostic examinations, the patient’s overall clinical condition, and resource availability. Moreover, clinicians must generally consider the potential consequences associated with incorrect decisions, a risk dimension that was necessarily absent from the experimental scenario.
In addition to focusing on a single clinical task, the evaluated decision-support systems employed a single explanation technique (SHAP). Different clinical tasks, explanation paradigms, or AI models may therefore produce different outcomes. Furthermore, the participant population consisted exclusively of residents from a single institution. Broader studies involving clinicians with different backgrounds, specialties, and levels of expertise are necessary to assess the generalizability of the findings and to facilitate comparisons with previous studies. For example, both the preliminary evaluations of adjunct and judicial protocols included a mixture of residents and board-certified clinicians.
In conclusion, future research should evaluate friction-based interaction protocols in larger and more diverse clinical populations, across different medical domains, data modalities, and explanation paradigms. In agreement with both previous studies [22,38] and the findings reported in Section 3.2.2, additional work is also needed to investigate how cognitive friction can be personalized according to user characteristics and contextual factors.
Given the promising results observed for Judicial AI, particularly in terms of reducing automation bias while preserving diagnostic performance, our future work will focus on evaluating JAI across different datasets, decision-support tasks, explanation techniques, and clinical settings and comparing it against alternative human–AI interaction protocols. The aim is to better characterize its strengths, limitations, and applicability in real-world healthcare environments.

5. Conclusions

To characterize human–AI interaction protocols based on the concept of cognitive friction (Frictional AI), we combined qualitative insights from semi-structured interviews with legal and psychological experts with a quantitative comparison of three decision-support designs through a user study.
The expert interviews highlighted that effective explanations should be simple, prioritize human understanding, be adapted to the recipient, and promote contrastive reasoning rather than merely exposing model internals (RQ1a). The experts further emphasized that reducing automation bias requires preserving clinician agency and responsibility (RQ1b), and identified user studies as a valuable instrument for evaluating explainability and human–AI interaction, despite the current lack of established legal metrics for assessing explainability compliance (RQ1c).
Friction-based protocols did not improve diagnostic accuracy relative to the no-support condition. Nevertheless, JAI achieved performances comparable to the baseline and generally superior to XAI and AAI, while maintaining similar levels of diagnostic confidence and perceived usefulness (RQ2a). Friction-based designs also increased decision-making time, suggesting greater cognitive engagement (RQ2b). Most importantly, JAI exhibited the lowest level of over-reliance among the evaluated protocols, supporting the hypothesis that appropriately designed cognitive friction can mitigate automation bias. Although a modest increase in under-reliance was observed, this was not associated with reduced diagnostic performance, suggesting a healthier calibration of trust rather than indiscriminate rejection of AI recommendations (RQ2c).
JAI consistently outperformed AAI across the most relevant outcomes, suggesting that the benefits of cognitive friction stem less from delaying access to AI advice, as in adjunct protocols, and more from encouraging users to engage in critical reasoning by evaluating competing evidence. Judicial AI showed the most encouraging results in reducing automation bias while maintaining diagnostic performance comparable to the no-support baseline and fostering active clinician engagement.
Taken together, these findings support the central premise of Frictional AI: the quality of human–AI collaboration depends not only on model performance and explanation quality, but also on the structure of the interaction itself. As AI systems become increasingly integrated into clinical workflows, interaction design may therefore become as important as model optimization for ensuring positive technology impact.

Author Contributions

Conceptualization, S.P. and L.B.; methodology, S.P. and L.B.; formal analysis, S.P. and C.A.S.; data curation, S.P.; writing—original draft preparation, S.P.; writing—review and editing, S.P., L.B., G.N., C.A.S., P.K., E.D., G.A., A.I.H.F., V.C., C.B., V.Z., F.S., L.P. and E.P.; supervision, G.N., P.K., E.D., G.A., A.I.H.F., V.C., C.B., V.Z., F.S., L.P. and E.P. All authors have read and agreed to the published version of the manuscript.

Funding

The APC was funded by the Italian Ministry of Research, under the complementary actions to the NRRP “Fit4MedRob–Fit for Medical Robotics” Grant (# PNC0000007).

Institutional Review Board Statement

The user study was conducted in accordance with the Declaration of Helsinki, and approved by the Ethics Committee of IRCCS Policlinic San Matteo Foundation (Protocol “P-20200072983”, approved on 30 September 2020). The expert interview study was approved through the University of Bergen’s System for Risk and Compliance Processing of Personal Data in Research and Student Projects (RETTE; project ID S3978).

Informed Consent Statement

Informed consent was obtained from all subjects involved either in the interview or in the user study.

Data Availability Statement

The data used in this study, including chest X-ray images, radiomic features, participants’ responses to the questionnaires, and the recordings and transcriptions form the interviews, are not publicly available due to ethical and privacy restrictions.

Acknowledgments

S.P. is enrolled in the National PhD program in Artificial Intelligence, XXXIX cycle, course on Health and life sciences, organized by Università Campus Bio-Medico di Roma. L.B. and E.P. acknowledge funding support provided by the Italian project PRIN PNRR 2022 InXAID—Interaction with eXplainable Artificial Intelligence in (medical) Decision Making. CUP: H53D23008090001 funded by the European Union—Next Generation EU. The authors acknowledge the support and guidance of Anna Oleynik during the planning of the project. Most of the preliminary work leading to the analyses reported in this paper was conducted at the Pandemic Centre, University of Bergen (Norway). ChatGPT-5.5 has been used for language editing and small adjustments in some of the figures presented, but all content has been reviewed and approved by all the authors.

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
BCaBias-Corrected and accelerated
CRPC-Reactive Protein
DSSDecision-Support System
ECGElectrocardiogram
FAIFrictional AI
GDPRGeneral Data Protection Regulation
JAIJudicial AI
LLMLarge Language Model
MDRMedical Device Regulation
MLMachine Learning
MRIMagnetic Resonance Imaging
RQResearch Question
SDStandard Deviation
XAIeXplainable Artificial Intelligence
WBCWhite Blood Cell
XGBoostExtreme Gradient Boosting

Appendix A. Interview Materials

This appendix provides a detailed report of the interview guides and questions used during the interviews with the legal and psychology experts described in Section 2.1.1. Information was organized and summarized in Table A1. The interviews were organized around three main themes: the characteristics of sufficiently transparent and understandable explanations; the relationship between over-reliance in AI and accountability in clinical decision making; and the adequacy of current evaluation methodologies and user studies for assessing human–AI interaction and compliance with legal and ethical requirements.
The responses to these questions constituted the basis for the thematic analysis presented in Section 2.1.2, the results of which are reported in Section 3.1.
Table A1. Semi-structured interview guide for domain experts.
Table A1. Semi-structured interview guide for domain experts.
ThemeQuestions
Preliminary questions1. What is your name?
2. What is your age?
3. What gender were you assigned at birth?
4. Do you have professional experience with AI or AI-based decision-support systems?
Legal domain—questions for the legal expert
Degree of explainability5. To what extent is explainability legally required under regulations such as the MDR [11], GDPR [10], and AI Act [9]?
6. Do current regulations prescribe specific techniques for ensuring AI transparency? If so, which ones, and how effective are they?
7. Considering Article 13 of the AI Act, how would you assess the clarity and comprehensibility of the explanations presented in the survey? Are they sufficiently understandable, and what changes would you suggest?
8. Should AI explanations prioritize understandability or faithfulness to the model’s internal workings? Why? Should a global explanation be included? Is Frictional AI [22] problematic?
9. Do patients have a “right to explanation” in AI-driven healthcare, and should research focus on whether explanations are truly understandable from the patient’s perspective?
Reliance on AI and automation bias10. How do current liability frameworks address disagreements between clinicians and AI recommendations, and who is responsible when clinical errors occur?
11. Could AI-based clinical decision-support systems encourage defensive medicine [27] or reduce clinicians’ acceptance of AI due to concerns about liability?
12. What role can explainable AI techniques play in liability frameworks such as the EU AI Liability Directive [12], and how might they affect accountability in healthcare?
13. What responsibilities should developers and manufacturers bear in healthcare litigation involving AI, and how should they be held accountable for clinical errors or harm?
Adequacy of surveys14. What metrics could demonstrate compliance with legal requirements such as the MDR or AI Act, and what factors should guide their definition?
15. Can a survey-based approach adequately demonstrate regulatory compliance? If not, how should studies assessing transparency requirements be designed?
Psychological domain—questions for the psychology expert
Degree of explainability5. What are the key characteristics of a “good explanation”, and how should they guide the design of AI explanations?
6. How could SHAP [50] be made more intuitive and user-friendly for non-technical users?
7. What criteria should guide feature selection in explanations? Should fewer, more impactful features be prioritized?
8. How should probabilities and model confidence be presented in AI-driven healthcare applications to maximize user understanding? Do users generally consider such information important?
Reliance on AI and automation bias9. What measures could reduce clinicians’ over-reliance on AI and mitigate automation bias [18]?
10. How can simple explanations be balanced with the complexity of high-risk healthcare decisions? Could simplified explanations increase over-reliance, and could model confidence help calibrate trust?
11. What are your views on Frictional AI, and could it help reduce automation bias in healthcare?
12. Beyond automation bias, what other cognitive biases might AI explanations trigger, and how could they be mitigated?
Adequacy of surveys14. Do the current survey questions accurately measure cognitive constructs such as trust, confidence, and usefulness? How could they be improved?
15. Is the “Clinical Explanation Satisfaction Scale” [66] clear and easy to understand? What improvements would enhance clarity and response quality?

Appendix B. User Study Materials

This appendix provides additional materials related to the user study described in Section 2.2.1. Specifically, it includes representative examples of the data presented to participants during the survey, details of the profiling questionnaire used to characterize the study population, and information about the instructional videos shown to participants before administering the survey.

Appendix B.1. Survey Inputs

Figure A1 illustrates the information presented to participants during the user study. Figure A1a shows a representative chest X-ray image together with the corresponding patient information. Each patient use case was condensed into an input table containing the following patient information:
  • Radiomic data—consolidation, infiltration, edema, effusion, and lung opacity;
  • Generalities—age and gender;
  • Respiratory symptoms—presence of respiratory issues, cough, breathing difficulties, chronic obstructive pulmonary disease, and respiratory failure;
  • Laboratory data—white blood cell (WBC) count and C-reactive protein (CRP) test result;
  • Comorbidities—hypertension, type 2 diabetes mellitus, cardiovascular disease, chronic renal failure, stroke, ischemic heart disease, atrial fibrillation, heart failure, dementia, and active cancer in the last 5 years;
Figure A1b presents the decision support provided under the XAI-first and adjunct protocols as described in Section 2.2.1. Specifically, participants were shown the model prediction, the associated confidence score, and a SHAP explanation plot corresponding to the prediction of the class “hospitalize”: positive SHAP values indicate features supporting hospitalization and are displayed in magenta, whereas negative SHAP values, displayed in blue, indicate features supporting discharge to home. Features are ordered according to their relative contribution to the model prediction—a visualization usually referred to as tornado plot as a consequence of its appearance.
Finally, Figure A1c illustrates the support provided under the judicial protocol. In this case, participants were not shown either the model prediction or the confidence score. Instead, two juxtaposed SHAP explanations corresponding to the competing classes (hospitalize vs. home) were presented simultaneously. To avoid revealing the model’s preferred prediction, SHAP values were displayed in absolute value, normalized in magnitude, and represented using a neutral color scheme. The bar plot on the left represents the evidence supporting the “hospitalize” class and is obtained from the positive (magenta) component of the original SHAP tornado plot. Conversely, the bar plot on the right represents the evidence supporting the “home” class and is derived from the negative (blue) component. This presentation aims to encourage users to compare competing hypotheses and independently determine which interpretation is better supported by the available evidence.
Figure A1. Survey input data, including (a) patient information and the corresponding chest X-ray image; (b) model prediction and SHAP explanation for the XAI-first and adjunct protocols; (c) juxtaposed SHAP explanations for the judicial protocol.
Figure A1. Survey input data, including (a) patient information and the corresponding chest X-ray image; (b) model prediction and SHAP explanation for the XAI-first and adjunct protocols; (c) juxtaposed SHAP explanations for the judicial protocol.
Make 08 00216 g0a1

Appendix B.2. The Profiling Questionnaire

Table A2 summarizes the information collected through the profiling questionnaire administered prior to the user study. As described in Section 2.2.2, the questionnaire was designed to characterize the participant population and support the subsequent randomization process.
Specifically, it collected demographic information, including sex, years of professional experience, and clinical specialty, as well as information regarding participants’ familiarity with AI technologies and their attitudes toward the use of AI in healthcare. A summary of the resulting participant characteristics is reported in Table 1 (Section 3.2.1).
These variables were subsequently used to perform the stratified analyses reported in Section 3.2.2 and described in Appendix C.
Table A2. Summary of the profiling questionnaire.
Table A2. Summary of the profiling questionnaire.
QuestionPossible Answers
Demographics and background
E-mail addressOpen-ended
SexFemale/Male/Prefer not to answer
How many years of experience do you have as a medical specialist?Open-ended (years)
What is your team?Radiology/Infectious Diseases/Emergency Medicine
Familiarity with AI
I have a good knowledge of Artificial Intelligence.Yes/No
I have worked with and/or used Artificial Intelligence systems in my job.Yes/No
Trust in AI
I believe that Artificial Intelligence can help me answer questions more accurately and quickly when I am uncertain about the answer.Yes/No
I believe that using Artificial Intelligence (e.g., a virtual assistant) to support my work or study can increase my productivity.Yes/No
I believe that Artificial Intelligence can improve the effectiveness of my work.Yes/No

Appendix B.3. Instructional Videos

To ensure that participants fully understood the task and the study procedures, they were provided with instructional material before survey administration. A total of three instructional videos were prepared and distributed to participants after assigning them into the user study groups. All videos began with a common introduction presenting the ALFABETO project [43,44], the diagnostic task and the classification model (Section 2.2.1), and examples of chest X-ray images and extracted radiomic features (Figure A1a).
The second part of each video was tailored to the assigned study branch and described the corresponding interaction protocol and the interpretation of its specific SHAP-based explanation (Figure A1b,c).
Participants were, finally, told about the information that would be presented during the survey and the evaluation procedure: for each case, they reviewed the patient’s tabular information and chest X-ray image, provided a hospitalization decision, and rated their confidence, perceived case complexity, and, when applicable, the usefulness of the decision support using six-point Likert scales.

Appendix C. Stratified Effect Size Analysis Details

In line with previous works [17,22,38], to investigate whether the impact of the evaluated interaction protocols varied across participant characteristics and case-specific factors, the effect size analysis—described in Section 2.2.3—was stratified. This stratification happened according to the responses collected through the profiling questionnaire (Appendix B.2), as well as participants’ ratings of case complexity and decision-support perceived usefulness.
The following subgroup analyses were performed:
  • Level of clinical experience—zero-experience (no COVID-19 patients treated), little-experience (≤100 patients treated), and high-experience residents (>100 patients treated);
  • Medical specialty—radiology, infectious diseases, and emergency medicine residents;
  • Familiarity with AI—residents reporting familiarity (at least one positive response) or unfamiliarity with AI technologies;
  • Trust in AI—skeptic (at least two negative responses) and unskeptic residents;
  • Perceived case complexity—simple (complexity score   3 ) and complex cases (complexity score >   3 );
  • Perceived usefulness of the decision support—support perceived as not useful (usefulness score   3 ) and useful (usefulness score >   3 ).
For each subgroup, intervention effects on diagnostic accuracy and diagnostic confidence were estimated using the same methodology described in Section 2.2.3, including Hedges-corrected Cohen’s d and BCa bootstrap confidence intervals. The results of this stratified analysis are presented in Section 3.2.2.

References

  1. Tripathy, N.; Nayak, S.K.; Addula, S.R.; Panigrahi, A.; Pati, A.; Rout, A. A Comprehensive Analysis and Prediction of Diabetes Diseases Using Deep Learning Techniques. In Advances in Healthcare using Machine Learning; CRC Press: Boca Raton, FL, USA, 2025. [Google Scholar] [CrossRef] [Scilit]
  2. Nicora, G.; Pe, S.; Santangelo, G.; Billeci, L.; Aprile, I.G.; Germanotta, M.; Bellazzi, R.; Parimbelli, E.; Quaglini, S. Systematic review of AI/ML applications in multi-domain robotic rehabilitation: Trends, gaps, and future directions. J. NeuroEng. Rehabil. 2025, 22, 79. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. O’Neil, C. Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy; Crown Publishing Group: New York, NY, USA, 2016. [Google Scholar]
  4. Christian, B. The Alignment Problem: Machine Learning and Human Values; W. W. Norton & Company: New York, NY, USA, 2020. [Google Scholar]
  5. Cristianini, N. La Scorciatoia: Come le Macchine Sono Diventate Intelligenti Senza Pensare in Modo Umano; Contemporanea, Il Mulino: Bologna, Italy, 2023. [Google Scholar]
  6. Dietvorst, B.J.; Simmons, J.P.; Massey, C. Algorithm aversion: People erroneously avoid algorithms after seeing them err. J. Exp. Psychol. Gen. 2015, 144, 114–126. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Guidotti, R.; Monreale, A.; Ruggieri, S.; Turini, F.; Giannotti, F.; Pedreschi, D. A Survey of Methods for Explaining Black Box Models. ACM Comput. Surv. (CSUR) 2018, 51, 1–42. [Google Scholar] [CrossRef] [Scilit]
  8. Gohel, P.; Singh, P.; Mohanty, M. Explainable AI: Current status and future directions. arXiv 2021, arXiv:2107.07045. [Google Scholar] [CrossRef] [Scilit]
  9. European Union. Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act); European Union: Brussels, Belgium, 2024. [Google Scholar]
  10. European Parliament; Council of the European Union. Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the Protection of Natural Persons with Regard to the Processing of Personal Data and on the Free Movement of Such Data, and Repealing Directive 95/46/EC (General Data Protection Regulation); European Union: Brussels, Belgium, 2016. [Google Scholar]
  11. European Parliament; Council of the European Union. Regulation (EU) 2017/745 of the European Parliament and of the Council of 5 April 2017 on Medical Devices, Amending Directive 2001/83/EC, Regulation (EC) No 178/2002 and Regulation (EC) No 1223/2009 and Repealing Council Directives 90/385/EEC and 93/42/EEC; European Union: Brussels, Belgium, 2017. [Google Scholar]
  12. European Commission. Proposal for a Directive of the European Parliament and of the Council on Adapting Non-Contractual Civil Liability Rules to Artificial Intelligence (AI Liability Directive); COM(2022) 496 final; European Commission: Brussels, Belgium, 2022. [Google Scholar]
  13. European Commission; AI HLEG. The Assessment List for Trustworthy Artificial Intelligence (ALTAI) for Self-Assessment; European Commission: Brussels, Belgium, 2020. [Google Scholar]
  14. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0); National Institute of Standards and Technology: Gaithersburg, MD, USA, 2023. [CrossRef] [Scilit]
  15. Lekadir, K.; Frangi, A.F.; Porras, A.R.; Glocker, B.; Cintas, C.; Langlotz, C.P.; Weicken, E.; Asselbergs, F.W.; Prior, F.; Collins, G.S.; et al. FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ 2025, 388, e081554. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. IEEE Global Initiative on Ethics of Autonomous and Intelligent Systems. Ethically Aligned Design: A Vision for Prioritizing Human Well-Being with Autonomous and Intelligent Systems; IEEE Standards Association: Piscataway, NJ, USA, 2019. [Google Scholar]
  17. Cabitza, F.; Campagner, A.; Angius, R.; Natali, C.; Reverberi, C. AI Shall Have No Dominion: On How to Measure Technology Dominance in AI-supported Human decision-making. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems; CHI ’23; Association for Computing Machinery: New York, NY, USA, 2023; pp. 1–20. [Google Scholar] [CrossRef] [Scilit]
  18. Lyell, D.; Coiera, E. Automation bias and verification complexity: A systematic review. J. Am. Med. Inform. Assoc. JAMIA 2017, 24, 423–431. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Sergeeva, E.; Sergeeva, A.; Tang, H.; Bongard-Blanchy, K.; Szolovits, P. Right, No Matter Why: AI Fact-checking and AI Authority in Health-related Inquiry Settings. arXiv 2023, arXiv:2310.14358. [Google Scholar] [CrossRef] [Scilit]
  20. Natali, C. Per Aspera ad Astra, or Flourishing via Friction: Stimulating Cognitive Activation by Design through Frictional Decision Support Systems. In CHItaly-DC 2023 CHItaly 2023 Doctoral Consortium CHItaly 2023 Proceedings of the Doctoral Consortium of the 15th Biannual Conference of the Italian SIGCHI Chapter (CHItaly 2023); CEUR-WS: Aachen, Germany, 2023. [Google Scholar]
  21. Ud Din, S.; Kemna, R.; Ket, J.C.F.; Iqbal, M.; Bohoudi, O.; Hoogendoorn, M.; Beretta, E.; Lisowska, A. Which explainable AI methods in medical imaging are clinically impactful? A systematic literature review addressing the clinician’s perspective. Front. Artif. Intell. 2026, 9, 1819422. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Cabitza, F.; Natali, C.; Famiglini, L.; Campagner, A.; Caccavella, V.; Gallazzi, E. Never tell me the odds: Investigating pro-hoc explanations in medical decision making. Artif. Intell. Med. 2024, 150, 102819. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Carr, N. The Glass Cage: Automation and Us; W. W. Norton & Company: New York, NY, USA, 2014. [Google Scholar]
  24. Moore, J.W. What Is the Sense of Agency and Why Does it Matter? Front. Psychol. 2016, 7, 1272. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Parasuraman, R.; Manzey, D.H. Complacency and bias in human use of automation: An attentional integration. Hum. Factors 2010, 52, 381–410. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Sureau, C. Medical deresponsibilization. J. Assist. Reprod. Genet. 1995, 12, 552–558. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Grote, T.; Berens, P. How competitors become collaborators—Bridging the gap(s) between machine learning algorithms and clinicians. Bioethics 2022, 36, 134–142. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Aquino, Y.S.J.; Rogers, W.A.; Braunack-Mayer, A.; Frazer, H.; Win, K.T.; Houssami, N.; Degeling, C.; Semsarian, C.; Carter, S.M. Utopia versus dystopia: Professional perspectives on the impact of healthcare artificial intelligence on clinical roles and skills. Int. J. Med. Inform. 2023, 169, 104903. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Shen, J.H.; Tamkin, A. How AI Impacts Skill Formation. arXiv 2026, arXiv:2601.20245. [Google Scholar] [CrossRef] [Scilit]
  30. Cooper, A. The Inmates Are Running the Asylum: Why High-Tech Products Drive Us Crazy and How to Restore the Sanity; Sams Publishing: Carmel, IN, USA, 2004. [Google Scholar]
  31. Norman, D.A. The Design of Everyday Things; Basic Books: New York, NY, USA, 2013. [Google Scholar]
  32. Norman, D.A. Living with Complexity; MIT Press: Cambridge, MA, USA, 2011. [Google Scholar]
  33. Kahneman, D. Thinking, Fast and Slow; Farrar, Straus and Giroux: New York, NY, USA, 2011. [Google Scholar]
  34. Reingold, O.; Shen, J.H.; Talati, A. Dissenting Explanations: Leveraging Disagreement to Reduce Model Overreliance. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence; AAAI Press: Washington, DC, USA, 2024. [Google Scholar] [CrossRef] [Scilit]
  35. Cao, S.; Liu, A.; Huang, C.M. Designing for Appropriate Reliance: The Roles of AI Uncertainty Presentation, Initial User Decision, and User Demographics in AI-Assisted Decision-Making. In Proceedings of the ACM on Human–Computer Interaction; Association for Computing Machinery: New York, NY, USA, 2024. [Google Scholar] [CrossRef] [Scilit]
  36. Bashkirova, A.; Krpan, D. Confirmation bias in AI-assisted decision-making: AI triage recommendations congruent with expert judgments increase psychologist trust and recommendation acceptance. Comput. Hum. Behav. Artif. Hum. 2024, 2, 100066. [Google Scholar] [CrossRef] [Scilit]
  37. Buçinca, Z.; Swaroop, S.; Paluch, A.E.; Doshi-Velez, F.; Gajos, K.Z. Contrastive explanations that anticipate human misconceptions can improve human decision-making skills. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Yokohama, Japan, 26 April–1 May 2025; pp. 1–25. [Google Scholar] [CrossRef] [Scilit]
  38. Cabitza, F.; Famiglini, L.; Fregosi, C.; Pe, S.; Parimbelli, E.; La Maida, G.A.; Gallazzi, E. From Oracular to Judicial: Enhancing Clinical Decision Making through Contrasting Explanations and a Novel Interaction Protocol. In Proceedings of the 30th International Conference on Intelligent User Interfaces, Cagliari, Italy, 24–27 March 2025; pp. 745–754. [Google Scholar] [CrossRef] [Scilit]
  39. Mollick, E. Everyone Is Above Average. 2023. Available online: https://www.oneusefulthing.org/p/everyone-is-above-average (accessed on 7 June 2026).
  40. Pe, S.; Famiglini, L.; Gallazzi, E.; Bortolotto, C.; Carone, L.; Cisarri, A.; Salina, A.; Preda, L.; Bellazzi, R.; Cabitza, F.; et al. Alternative Strategies to Generate Class Activation Maps Supporting AI-based Advice in Vertebral Fracture Detection in X-ray Images. Methods Inf. Med. 2024, 63, 122–136. [Google Scholar] [CrossRef] [Scilit]
  41. Cabitza, F.; Campagner, A.; Ronzio, L.; Cameli, M.; Mandoli, G.E.; Pastore, M.C.; Sconfienza, L.M.; Folgado, D.; Barandas, M.; Gamboa, H. Rams, hounds and white boxes: Investigating human–AI collaboration protocols in medical diagnosis. Artif. Intell. Med. 2023, 138, 102506. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Jain, B. The key role of differential diagnosis in diagnosis. Diagnosis 2017, 4, 239–240. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  43. Nicora, G.; Tito, A.L.; Donatelli, A.; Callea, G.; Biasibetti, C.; Galli, M.V.; Comotto, F.; Bortolotto, C.; Perlini, S.; Preda, L.; et al. ALFABETO: Supporting COVID-19 hospital admissions with Bayesian Networks. In Proceedings of the Workshop on Towards Smarter Health Care: Can Artificial Intelligence Help? co-located with 20th International Conference of the Italian Association for Artificial Intelligence (AIxIA2021), Online Event, 29 November 2021; Volume 3060, pp. 79–84. [Google Scholar]
  44. Catalano, M.; Bortolotto, C.; Nicora, G.; Achilli, M.F.; Consonni, A.; Ruongo, L.; Callea, G.; Lo Tito, A.; Biasibetti, C.; Donatelli, A.; et al. Performance of an AI algorithm during the different phases of the COVID pandemics: What can we learn from the AI and vice versa. Eur. J. Radiol. Open 2023, 11, 100497. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  45. Vasconcelos, H.; Jörke, M.; Grunde-McLaughlin, M.; Gerstenberg, T.; Bernstein, M.; Krishna, R. Explanations Can Reduce Overreliance on AI Systems During Decision-Making. arXiv 2023, arXiv:2212.06823. [Google Scholar] [CrossRef] [Scilit]
  46. Malterud, K.; Siersma, V.D.; Guassora, A.D. Sample Size in Qualitative Interview Studies: Guided by Information Power. Qual. Health Res. 2016, 26, 1753–1760. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Ahmed, S.K.; Mohammed, R.A.; Nashwan, A.J.; Ibrahim, R.H.; Abdalla, A.Q.; M. Ameen, B.M.; Khdhir, R.M. Using thematic analysis in qualitative research. J. Med. Surg. Public Health 2025, 6, 100198. [Google Scholar] [CrossRef] [Scilit]
  48. NVIVO: Software per l’analisi dei Dati Qualitativi–QSR International. Available online: https://yosigo.ugr.es/it/recurso/nvivo-qualitative-data-analysis-software-qsr-internazionale/ (accessed on 8 June 2026).
  49. Bergomi, L.; Nicora, G.; Orlowska, M.A.; Podrecca, C.; Bellazzi, R.; Fregosi, C.; Salinaro, F.; Bonzano, M.; Crescenzi, G.; Speciale, F.; et al. Which explanations do clinicians prefer? A comparative evaluation of XAI understandability and actionability in predicting the need for hospitalization. BMC Med. Inf. Decis. Mak. 2025, 25, 269. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  50. Lundberg, S.; Lee, S.I. A Unified Approach to Interpreting Model Predictions. arXiv 2017, arXiv:1705.07874. [Google Scholar] [CrossRef] [Scilit]
  51. KoboToolbox. Available online: https://www.kobotoolbox.org/ (accessed on 7 June 2026).
  52. Erdfelder, E.; Faul, F.; Buchner, A. GPOWER: A general power analysis program. Behav. Res. Methods Instrum. Comput. 1996, 28, 1–11. [Google Scholar] [CrossRef] [Scilit]
  53. Marfo, P.; Okyere, G.A. The accuracy of effect-size estimates under normals and contaminated normals in meta-analysis. Heliyon 2019, 5, e01838. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  54. Chen, D.; Fritz, M.S. Comparing Alternative Corrections for Bias in the Bias-Corrected Bootstrap Test of Mediation. Eval. Health Prof. 2021, 44, 416–427. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Cohen, J. Statistical Power Analysis for the Behavioral Sciences, 2nd ed.; Lawrence Erlbaum Associates: Hillsdale, NJ, USA, 1988. [Google Scholar]
  56. Nickerson, R.S. Confirmation Bias: A Ubiquitous Phenomenon in Many Guises. Rev. Gen. Psychol. 1998, 2, 175–220. [Google Scholar] [CrossRef]
  57. Miller, T. Explanation in artificial intelligence: Insights from the social sciences. Artif. Intell. 2019, 267, 1–38. [Google Scholar] [CrossRef] [Scilit]
  58. Chou, Y.L.; Moreira, C.; Bruza, P.; Ouyang, C.; Jorge, J. Counterfactuals and Causability in Explainable Artificial Intelligence: Theory, Algorithms, and Applications. arXiv 2021, arXiv:2103.04244. [Google Scholar] [CrossRef] [Scilit]
  59. Guidotti, R. Counterfactual explanations and how to find them: Literature review and benchmarking. Data Min. Knowl. Discov. 2024, 38, 2770–2824. [Google Scholar] [CrossRef] [Scilit]
  60. Stepin, I.; Alonso, J.M.; Catala, A.; Pereira-Fariña, M. A Survey of Contrastive and Counterfactual Explanation Generation Methods for Explainable Artificial Intelligence. IEEE Access 2021, 9, 11974–12001. [Google Scholar] [CrossRef] [Scilit]
  61. Prabhushankar, M.; Kwon, G.; Temel, D.; AlRegib, G. Contrastive Explanations In Neural Networks. In Proceedings of the 2020 IEEE International Conference on Image Processing (ICIP), Abu Dhabi, United Arab Emirates, 25–28 October 2020; pp. 3289–3293. [Google Scholar] [CrossRef] [Scilit]
  62. Jacovi, A.; Swayamdipta, S.; Ravfogel, S.; Elazar, Y.; Choi, Y.; Goldberg, Y. Contrastive Explanations for Model Interpretability. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing; Moens, M.F., Huang, X., Specia, L., Yih, S.W.T., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 1597–1611. [Google Scholar] [CrossRef] [Scilit]
  63. AlRegib, G.; Prabhushankar, M. Explanatory paradigms in neural networks: Towards relevant and contextual explanations. IEEE Signal Process. Mag. 2022, 39, 59–72. [Google Scholar] [CrossRef] [Scilit]
  64. Loncharich, M.F.; Robbins, R.C.; Durning, S.J.; Soh, M.; Merkebu, J. Cognitive biases in internal medicine: A scoping review. Diagnosis 2023, 10, 205–214. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  65. Pielke, R.A.J. The Honest Broker: Making Sense of Science in Policy and Politics; Cambridge University Press: Cambridge, UK, 2007. [Google Scholar]
  66. Lesley, U.; Kuratomi Hernández, A. Improving XAI Explanations for Clinical Decision-Making–Physicians’ Perspective on Local Explanations in Healthcare. In Proceedings of the Artificial Intelligence in Medicine: 22nd International Conference, AIME 2024, Salt Lake City, UT, USA, 9–12 July 2024; Proceedings, Part II; Springer: Berlin/Heidelberg, Germany, 2024; pp. 296–312. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Thematic map of expert interview findings. The three overarching themes correspond to the interview research questions (RQ1) and are organized into the sub-themes identified through thematic analysis.
Figure 1. Thematic map of expert interview findings. The three overarching themes correspond to the interview research questions (RQ1) and are organized into the sub-themes identified through thematic analysis.
Make 08 00216 g001
Figure 2. Comparison between user study branches in terms of (a) diagnostic accuracy, (b) diagnostic confidence, (c) perceived usefulness of the support, (d) survey completion time in minutes, (e) user over-reliance on AI when AI is wrong, and (f) user under-reliance on AI when AI is right. Each metric is obtained by aggregating user-averaged metrics.
Figure 2. Comparison between user study branches in terms of (a) diagnostic accuracy, (b) diagnostic confidence, (c) perceived usefulness of the support, (d) survey completion time in minutes, (e) user over-reliance on AI when AI is wrong, and (f) user under-reliance on AI when AI is right. Each metric is obtained by aggregating user-averaged metrics.
Make 08 00216 g002
Figure 3. Comparison of the user study branches against the baseline no-support (NoAI) condition in terms of effect size for (a) diagnostic accuracy and (b) confidence. Effect sizes are reported as Cohen’s d with Hedges’ correction, together with 95% BCa bootstrap confidence intervals. The gray shaded area represents the region of negligible effect sizes, bounded by the red dashed vertical lines at the conventional Cohen’s d thresholds.
Figure 3. Comparison of the user study branches against the baseline no-support (NoAI) condition in terms of effect size for (a) diagnostic accuracy and (b) confidence. Effect sizes are reported as Cohen’s d with Hedges’ correction, together with 95% BCa bootstrap confidence intervals. The gray shaded area represents the region of negligible effect sizes, bounded by the red dashed vertical lines at the conventional Cohen’s d thresholds.
Make 08 00216 g003
Figure 4. Stratified comparison of the JAI and AAI protocols against the baseline no-support (NoAI) condition in terms of effect size. The figure shows (a) diagnostic accuracy for JAI, (b) diagnostic confidence for JAI, (c) diagnostic accuracy for AAI, and (d) diagnostic confidence for AAI. Effect sizes are reported as Cohen’s d with Hedges’ correction and corresponding 95% BCa bootstrap confidence intervals. The gray shaded area represents the region of negligible effect sizes, bounded by the red dashed vertical lines at the conventional Cohen’s d thresholds. The analyzed subgroups are reported on the vertical axis and represented using distinct colors.
Figure 4. Stratified comparison of the JAI and AAI protocols against the baseline no-support (NoAI) condition in terms of effect size. The figure shows (a) diagnostic accuracy for JAI, (b) diagnostic confidence for JAI, (c) diagnostic accuracy for AAI, and (d) diagnostic confidence for AAI. Effect sizes are reported as Cohen’s d with Hedges’ correction and corresponding 95% BCa bootstrap confidence intervals. The gray shaded area represents the region of negligible effect sizes, bounded by the red dashed vertical lines at the conventional Cohen’s d thresholds. The analyzed subgroups are reported on the vertical axis and represented using distinct colors.
Make 08 00216 g004
Table 1. Characteristics of participants included in the quantitative evaluation, overall and stratified by human–AI interaction protocol. Values are reported as n (%).
Table 1. Characteristics of participants included in the quantitative evaluation, overall and stratified by human–AI interaction protocol. Values are reported as n (%).
CharacteristicOverallXAIJAIAAI
Sex
Female54 (56.2%)11 (45.8%)18 (58.1%)25 (61.0%)
Male39 (40.6%)13 (54.2%)11 (35.5%)15 (36.6%)
Prefer not to answer3 (3.1%)0 (0.0%)2 (6.5%)1 (2.4%)
Specialty
Radiology57 (59.4%)15 (62.5%)17 (54.8%)25 (61.0%)
Infectious Diseases30 (31.2%)7 (29.2%)11 (35.5%)12 (29.3%)
Emergency Medicine9 (9.4%)2 (8.3%)3 (9.7%)4 (9.8%)
Residency year
Year I17 (17.7%)3 (12.5%)6 (19.4%)8 (19.5%)
Year II17 (17.7%)5 (20.8%)5 (16.1%)7 (17.1%)
Year III29 (30.2%)9 (37.5%)9 (29.0%)11 (26.8%)
Year IV30 (31.2%)5 (20.8%)10 (32.3%)15 (36.6%)
Year V3 (3.1%)2 (8.3%)1 (3.2%)0 (0.0%)
COVID-19 patients treated
0 patients42 (43.8%)13 (54.2%)13 (41.9%)16 (39.0%)
≤100 patients28 (29.2%)7 (29.2%)10 (32.3%)11 (26.8%)
>100 patients26 (27.1%)4 (16.7%)8 (25.8%)14 (34.1%)
AI familiarity
Familiar with AI52 (54.2%)11 (45.8%)14 (45.2%)27 (65.9%)
Not familiar with AI44 (45.8%)13 (54.2%)17 (54.8%)14 (34.1%)
AI skepticism
Not skeptical toward AI71 (74.0%)16 (66.7%)24 (77.4%)31 (75.6%)
Skeptical toward AI25 (26.0%)8 (33.3%)7 (22.6%)10 (24.4%)
Table 2. Summary of study outcomes across the no-support scenario and the three experimental conditions. Values are reported as mean ± SD. The best available average value for each metric is highlighted in bold.
Table 2. Summary of study outcomes across the no-support scenario and the three experimental conditions. Values are reported as mean ± SD. The best available average value for each metric is highlighted in bold.
MetricNoAIXAIJAIAAI
Accuracy 0.68 ± 0.17 0.63 ± 0.14 0.67 ± 0.11 0.57 ± 0.12
Diagnostic confidence 3.39 ± 0.71 3.69 ± 0.76 3.56 ± 0.69 3.45 ± 0.67
Perceived usefulnessNot available 3.92 ± 0.66 3.63 ± 0.98 3.63 ± 0.91
Completion time (min)Not available 9.53 ± 5.32 11.65 ± 6.53 11.27 ± 4.32
Over-reliance 0.47 ± 0.31 0.55 ± 0.30 0.44 ± 0.26 0.75 ± 0.22
Under-reliance 0.19 ± 0.18 0.20 ± 0.25 0.24 ± 0.15 0.15 ± 0.16
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Pe, S.; Bergomi, L.; Nicora, G.; Simonelli, C.A.; Kour, P.; Diaz, E.; Alendal, G.; Hernáiz Ferrer, A.I.; Corso, V.; Bortolotto, C.; et al. Cognitive Friction in Clinical Decision Support: A Comparative Study of Judicial and Adjunct Human–AI Interaction Protocols. Mach. Learn. Knowl. Extr. 2026, 8, 216. https://doi.org/10.3390/make8070216

AMA Style

Pe S, Bergomi L, Nicora G, Simonelli CA, Kour P, Diaz E, Alendal G, Hernáiz Ferrer AI, Corso V, Bortolotto C, et al. Cognitive Friction in Clinical Decision Support: A Comparative Study of Judicial and Adjunct Human–AI Interaction Protocols. Machine Learning and Knowledge Extraction. 2026; 8(7):216. https://doi.org/10.3390/make8070216

Chicago/Turabian Style

Pe, Samuele, Laura Bergomi, Giovanna Nicora, Camilla A. Simonelli, Prabhjot Kour, Esperanza Diaz, Guttorm Alendal, Ana I. Hernáiz Ferrer, Valeria Corso, Chandra Bortolotto, and et al. 2026. "Cognitive Friction in Clinical Decision Support: A Comparative Study of Judicial and Adjunct Human–AI Interaction Protocols" Machine Learning and Knowledge Extraction 8, no. 7: 216. https://doi.org/10.3390/make8070216

APA Style

Pe, S., Bergomi, L., Nicora, G., Simonelli, C. A., Kour, P., Diaz, E., Alendal, G., Hernáiz Ferrer, A. I., Corso, V., Bortolotto, C., Zuccaro, V., Salinaro, F., Preda, L., & Parimbelli, E. (2026). Cognitive Friction in Clinical Decision Support: A Comparative Study of Judicial and Adjunct Human–AI Interaction Protocols. Machine Learning and Knowledge Extraction, 8(7), 216. https://doi.org/10.3390/make8070216

Article Metrics

Back to TopTop