1. Introduction
With the advancement of information technology, the human–computer interaction process has become more complex [
1], leading to a higher susceptibility to human error [
2]. Human error in interactive systems is an inherent aspect of system performance [
3]. Interaction design for human error can effectively improve the usability of interactive systems [
4,
5]. Human reliability analysis of interactive systems can reduce human error by providing recommendations for design improvements, thereby minimizing the negative impact of these errors and improving system usability [
6,
7,
8].
Human reliability analysis (HRA) provides insights into human–computer interaction behaviors during tasks and quantifies failures [
9]. The first generation of HRA methods is recognized for its intuitiveness and practicality [
10]. However, the increasing demand for a deeper understanding of human–computer interaction has driven the development of second-generation HRA methods [
11]. The cognitive reliability and error analysis method (CREAM) is a representative second-generation HRA method [
12]. It is known for identifying cognitive tasks and determining the causes of cognitive reliability degradation [
13]. Existing studies have applied the CREAM to analyze human reliability in system operations [
14]. For example, Chen et al. [
15] used the CREAM to assess human error in the operation of high-speed railway systems. Additionally, the CREAM has been employed in accident risk analysis. For instance, Elidolu et al. [
16] used the CREAM to quantify the causes of failures in a cruise ship explosion at sea. However, to the best of the authors’ knowledge, limited research has been conducted on applying the CREAM in interaction design to analyze cognitive function failures.
The CREAM defines the factors influencing cognitive, technological, environmental, and organizational performance as common performance conditions (CPCs) [
17]. However, two key limitations encountered in the practical application of the traditional CREAM include the objective assessment of the level of CPCs and the detailed analysis of the interaction among CPCs [
18]. The traditional CREAM lacks refined evaluation factors for CPCs [
19], and the evaluation process relies heavily on expert knowledge and experience, which are often subjective and inconsistent [
20]. Fuzzy comprehensive evaluation (FCE) is a widely used uncertainty assessment method that improves the objectivity of assessment results by representing expert judgments using fuzzy linguistic terms [
21]. By simplifying complex decision-making processes, FCE provides a reliable and adaptable tool for addressing ambiguity in evaluation scenarios [
22]. It has been demonstrated to be useful in optimizing evaluation effectiveness [
23,
24]. Therefore, to mitigate the impact of subjective expert assessments on the quantification of human reliability, this paper integrates FCE into the traditional CREAM framework to evaluate the levels of CPCs in interactive systems.
The traditional CREAM primarily examines the effects of CPCs on cognitive reliability in the cognitive dimension, providing only a simplified analysis of the interactions between factors [
25]. However, the interrelationships among CPCs can affect their impact on interactive systems, influencing the results of human reliability analysis [
26]. Therefore, a correlation impact analysis of CPCs is crucial for interaction systems. The decision-making trial and evaluation laboratory (DEMATEL) method effectively evaluates complex systems by analyzing factor relationships [
27]. It has been extensively applied to investigate the interrelationships among factors that influence system security [
28]. For example, Shi et al. [
29] used the DEMATEL method to analyze the interdependence and interactions of the factors influencing the cargo loading process. DEMATEL improves the reliability of failure analysis in interactive systems by identifying the causal relationships and interactions among factors [
30]. Therefore, to quantify the impact of human–computer interaction factors on the cognitive reliability assessment of interactive systems, DEMATEL is incorporated in this paper.
Human intrinsic factors (HIFs) influence the user’s comprehension and execution of tasks, which is crucial for determining the human reliability of interactions with the system [
31]. Considering the impact of HIFs on user behavior in interaction design can enhance interactive quality [
32]. The traditional CREAM, which does not explicitly incorporate the influence of HIFs on cognitive function failures, may have limited accuracy in predicting human error in interactive systems [
19]. Therefore, to provide a more comprehensive assessment of cognitive function failures, this paper integrates the impact of CPCs and HIFs.
Building on the above analysis of DEMATEL’s role and the integration of CPCs and HIFs, this study offers distinct research novelties. The novelty of this study lies in the synergistic effect of four dimensions: first, a human–computer interaction-specific CPC/HIF analysis system, which has a wide range of applications and is highly suitable for human reliability analysis; second, a normalized product-based weight integration mechanism, which solves the problem of low fusion accuracy; third, a closed-loop design improvement cycle, which realizes iterative optimization of human–computer interaction design; and finally, an end-to-end workflow, which covers the entire link from indicator establishment to effect verification.
This paper introduces a method for human reliability analysis in interaction design, integrating the CREAM, FCE, and DEMATEL. This research includes three main contributions that enhance the applicability and accuracy of human reliability analysis in interactive systems. First, to reduce uncertainty in evaluating CPC levels, FCE is incorporated. Second, to comprehensively explore the impact of human error in interactive systems, the impact of CPCs and HIFs is integrated into the traditional CREAM framework. Third, to enhance the accuracy of cognitive reliability analysis, the interrelationships among CPCs are analyzed, and the factor centrality weights of CPCs and HIFs are calculated using the DEMATEL method.
The structure of the paper is as follows:
Section 2 outlines the research background, while
Section 3 explains the proposed methodology.
Section 4 presents a case study, followed by a discussion of its results in
Section 5. Finally,
Section 6 provides the conclusions.
3. Proposed Methodology
This paper proposes a novel method integrating the CREAM, FCE, and DEMATEL to analyze human error in interaction design and reduce cognitive load on users interacting with the system. The framework of the proposed methodology is shown in
Figure 1 and consists of five stages. In stage 1, cognitive function failures are evaluated. In stage 2, CPC levels are assessed. In stage 3, the HIF and CPC weights are calculated. In stage 4, the cognitive failure probability is determined. In stage 5, recommendations for design improvements are proposed.
3.1. Analysis of Cognitive Function Failures
3.1.1. Conducting Hierarchical Task Analysis and Establishing a Cognitive Demand Profile
To analyze the tasks of the target interactive system, hierarchical task analysis (HTA) was performed. HTA provides a clear, structured breakdown of complex tasks [
48]. It is a widely used method for decomposing tasks into subtasks and represents the result of a hierarchical tree diagram [
49].
A cognitive demand profile is constructed based on the user’s performance during the completion of subtasks from the hierarchical tree diagram. The key cognitive activities and associated cognitive functions involved in interacting with the system are identified through the observation and recording of task completion. The fifteen cognitive activity types and four cognitive functions used to construct the cognitive demand profile are presented in
Table 1.
3.1.2. Obtaining the Nominal Cognitive Failure Probability
The CREAM categorizes four types of cognitive functions into various cognitive function failures and assigns the corresponding nominal CFPs, as shown in
Table 2 [
19]. Human error is analyzed based on the cognitive demand profile, and the cognitive function failures occurring during system interaction are identified.
3.2. Assessment of CPC Levels
Interaction contexts are established by analyzing CPCs defined based on the target interaction system [
37]. By determining CPC levels, interaction contexts where human error occurs during human–computer interaction can be identified. This paper evaluates the impact level of each CPC using FCE. Each CPC is further divided into sub-factors. The system’s performance level is then derived through FCE, which involves the following five steps:
Step 1: Establish the factor set.
The factor set is a collection of factors that affect the level of each CPC, represented by set C, , and represents the refined sub-factors for each CPC.
Step 2: Establish the weight set.
The weight set reflects the importance of each sub-factor in the CPCs by assigning weight to sub-factor ci. The weights of all sub-factors form set A, , where , with .
Step 3: Establish the evaluation set.
The evaluation set is a collection of evaluation results of the CPCs of the interactive system. The evaluation set is set as L, , and denotes the impact level possessed by the sub-factors of the CPCs.
Step 4: Perform a fuzzy comprehensive evaluation.
The evaluation is conducted on the
ith sub-factor
ci of the CPCs, with the affiliation degree of the CPCs to the
jth level
lj in the evaluation set denoted as
rij. The evaluation results of the sub-factors for each CPC are then aggregated into an evaluation matrix
R, as shown in Equation (1).
Based on the weight set
A and evaluation matrix
R, the fuzzy synthesized evaluation set
B is calculated using Equations (2) and (3).
where
is the fuzzy synthesized evaluation index and “
” represents the fuzzy synthesis process.
Step 5: The evaluation indicators are processed using the maximum affiliation method to determine CPC levels.
3.3. Calculation of HIF and CPC Weights
In this paper, to calculate the factor centrality weights of CPCs and HIFs, DEMATEL is employed. This method quantifies the relationships between factors and their respective degrees of influence with precision. DEMATEL quantitatively analyzes the relationships between factors to identify their causality and centrality, typically following five steps:
Step 1: Define and evaluate factors.
Identify the factors influencing the system and assess the relationships between factors using expert judgment. These relationships are categorized into four levels of influence: 0 represents no influence, 1 represents slight influence, 2 represents moderate influence, and 3 represents strong influence.
Step 2: Obtain the direct relationship matrix X.
Based on the evaluations from Step 1, the factors are compared pairwise to assess their influence. The arithmetic mean of the scores is used to construct an n × n direct relationship matrix, denoted as , where xij represents the degree to which factor i influences factor j, with all diagonal elements set to 0.
Step 3: Calculate the normalized direct relation matrix D.
The normalized direct relation matrix
D can be calculated using Equations (4) and (5).
Step 4: Calculate the total impact relationship matrix T.
Due to
, the total matrix
can be calculated using Equation (6).
where
I is the identity matrix, 0 is the null matrix, and
tij represents the degree to which factor
i affects factor
j.
Step 5: Conduct an analysis to get the sum of rows and columns.
Let
h and
g denote the sum of the rows and columns in matrix
T, which can be calculated using Equations (7) and (8).
Let
hi represent the sum of the
ith row and
gi represent the sum of the
ith column. In this paper,
hi indicates the degree to which CPC
i influences other CPCs, while
gi represents the degree to which it is influenced by others. The value of
hi +
gi represents the degree of centrality of CPC
i, reflecting its relative importance. Therefore, the value of
hi +
gi can then be normalized to calculate the factor centrality weight of CPC
i using Equation (9).
where
wi is the factor centrality weight of CPC
i on subtasks.
In this paper,
hj indicates the degree to which HIF
j influences other HIFs, while
gj represents the degree to which it is influenced by others. Based on the analysis above, the value of
hj +
gj can then be normalized to calculate the factor centrality weight of HIF
j using Equation (10).
where
vj is the impact weight of HIF
j on subtasks.
Based on the DEMATEL analysis results, we identify the CPCs influenced by others as CPCs with values higher than the average. Then, the relevance adjustment rules are applied to analyze the interactions among CPCs. The cognitive impact weights are determined based on CPC interactions and levels.
3.4. Calculation of the Cognitive Failure Probability
The factor centrality weight of CPCs and the impact weight of HIFs are integrated via Equation (11) to yield the comprehensive weight.
where
pi is the comprehensive weight,
ui is the cognitive impact weight of the
ith CPC,
wi is the factor centrality weight of CPC
i on subtasks, and
vj is the impact weight of HIF
j on subtasks.
Based on the derived comprehensive weight, the CFP for each subtask is calculated via Equation (12).
where
CFP0 denotes the nominal CFP of each subtask.
Note that Equation (11) combines the node centrality and the factor impact degree of cognitive risk factors within the DEMATEL causal network. The normalized product form is adopted to eliminate dimensional differences and highlight their synergistic effect, which is more reasonable than additive or unnormalized product forms, and this combination is a fusion operator rather than a weight update or risk amplification operator. Equation (12), which is closely related to Equation (11) in the cognitive reliability analysis process, uses multiplicative aggregation for CFP. This method assumes the conditional independence of cognitive failure events across different factors within the human–computer interaction scenario. To reconcile DEMATEL-based interaction modeling with multiplicative aggregation that may implicitly assume separability, we adopt a two-stage strategy: first using DEMATEL to screen key cognitive factors, and then applying multiplicative aggregation to calculate CFP. The key cognitive factors screened by DEMATEL inherently have independent characteristics in terms of their impact on cognitive failure, which ensures the rationality of the multiplicative aggregation method and effectively resolves potential conflicts between the two methods.
3.5. Recommendations for Design Improvements
To improve the design, we recommend the following: Identify human-error-prone subtasks in the interactive system based on the CFPs. Analyze the causes of human error using CREAM’s cause-effect matrix proposed by Hollnagel [
33]. Develop optimization strategies and implement design improvements for the target interactive system [
50].
5. Discussion
This paper introduces a hybrid approach that combines the CREAM, FCE, and DEMATEL for human reliability analysis, focusing on cognitive demands in interaction design. By integrating FCE with the traditional CREAM, the subjectivity of expert evaluations in assessing CPC levels was reduced, enhancing the accuracy of the assessment. Additionally, the impact of HIFs and CPCs on cognitive function failures was comprehensively considered. To quantify the factor centrality weights of HIFs and CPCs, DEMATEL was employed, allowing for the analysis of the correlations between CPCs. Furthermore, to provide valuable insights for optimizing interaction design, the underlying causes of human error were explored. A case study of an IVIS design was conducted to demonstrate the proposed approach.
Based on the analysis of drivers using the in-vehicle navigation functions of the original IVIS, cognitive errors can be categorized into misinterpretation of functional information and misperception of visual symbols. The physiological and psychological characteristics of drivers contribute to these errors, with causes such as cognitive biases and workload leading to prioritization errors and misinterpretation during interaction. Additionally, failures in IVIS interface design and excessive task demands result in information overload and low visibility. Design improvements are proposed to mitigate human error in the IVIS.
The main page and the navigation page of the revised IVIS are presented in
Figure 6. The research team reanalyzed the revised IVIS, and post-improvement CFPs were reassessed using the same method as the pre-improvement CFPs. To compare the CFPs before and after the improvement, A/B testing was employed. As a widely used method in human–computer interaction research, A/B testing allows for evaluating two versions of a design to identify the statistically more effective one [
64]. The results showed that CFPs after improvement (M = 0.002, SD = 0.002) were significantly lower than CFPs before improvement (M = 0.092, SD = 0.078),
t(10) = 3.79,
p = 0.004. The box plot is shown in
Figure 7, and the CFPs before and after improvement are marked using two * symbols, indicating a significant difference between the two groups. These results suggest a significant difference in CFPs during task completion with the revised IVIS, demonstrating the effectiveness of the proposed improvements and confirming the human reliability of the improved interaction system, thus validating the proposed methodology.
To further elaborate on the design improvement measures and their correlation with CPCs and HIFs, four optimization strategies were implemented to achieve a significant reduction in CFP, with their specific mapping relationships clarified as follows: (1) UI layout optimization, which effectively improved the visual attention allocation of CPCs; (2) feedback timing optimization, which reduced the response delay of HIFs; (3) icon redesign, which alleviated the cognitive ambiguity of HIFs; (4) hierarchy adjustment, which enhanced the information processing efficiency of CPCs. In addition, the reliability of the post-improvement CFP values was verified by two complementary methods: recalculation based on updated CPCs assessments, and repeated user experiments involving 30 subjects, which further confirmed the robustness of the improvement effects.
In order to consider the impact of CPC correlations on cognitive function failures, the traditional CREAM proposes basic adjustment rules. However, in the practical application of human reliability analysis for interactive systems, CPCs should be adjusted based on their interaction. To analyze the interrelationships among CPCs, the proposed method used DEMATEL. In the IVIS case study, the effects of CPCs were adjusted for interrelationships and the adjustment rules. For example, the original effect of CPC2 was not significant, and the adjusted effect was reduced. The proposed method provides a more accurate understanding of the effects of CPCs on cognitive function failures.
HIFs play a crucial role in user-centered interaction design, especially among IVIS users. However, the traditional CREAM does not focus on the impact of HIFs on human error in interactive systems during human reliability analysis. The proposed method identifies nine HIFs relevant to interactive systems operations and uses DEMATEL to examine their influence on human error. It also integrates the combined weights of CPCs and HIFs when calculating the CFPs. As a result, this approach offers more accurate outcomes, improving effectiveness and adaptability in practice.
To clarify the differences between the proposed method and related methods (including the traditional CREAM, single FCE, and single DEMATEL) in terms of output results and uncertainty handling, the details are presented in
Table 19. Compared with these related methods, the proposed method has the following innovation points: it constructs a human–computer interaction-specific CPC/HIF system, adopts DEMATEL for causal structure analysis of factors, uses FCE for quantitative assessment of factor levels, and integrates these three methods through a normalized product integration formula to form a systematic cognitive reliability analysis pipeline for interaction design.
This study was primarily conducted with drivers as the representative sample. Future research could include a broader range of participants to improve the applicability of the results. To resolve the limitations of the traditional CREAM in analyzing the interrelationships among CPCs, this study used DEMATEL to explore the interaction effects within interactive systems. This paper focused on CPC analysis and excluded the analysis of interaction among HIFs. Future research will investigate the mutual influences of HIFs to enhance human reliability analysis. Furthermore, this study mainly focused on identifying human errors and the causes of cognitive function failures. The interrelationship between cognitive function failures has not been explored in depth. Bayesian network analysis can be employed to explore the interrelationships among factors [
26,
37]. Bayesian network analysis effectively handles conditional dependencies, providing more accurate insights for optimizing human–computer interaction [
65,
66]. To comprehensively analyze human error in interactive systems, future research could integrate Bayesian networks with the proposed method to analyze the interrelationships among CPCs and the relationships between cognitive function failures. To ensure participant safety, the tests were conducted in static scenarios; however, future studies are planned for dynamic scenarios.
The proposed approach is relatively complex and involves multiple computational steps. Nevertheless, it is implemented systematically. More importantly, the approach integrates five well-defined simplification steps: analysis of cognitive function failures, assessment of CPC levels, calculation of HIF and CPC weights, calculation of the CFP, and recommendations for design improvements. With these detailed and standardized steps, the proposed approach is sufficiently reproducible. Although an IVIS was used as a case study, the approach is generalizable to other interactive systems for human error reduction. Future work will further validate and extend the approach by incorporating additional application cases.