2.1. Materials and Methods
2.1.1. Participants
Sixty-six volunteers participated in the study (52 women). The mean age of participants was 21.51 years (SD = 3.58, range 18–40). Participants provided informed consent and received the equivalent of 5.56 EUR. This study was approved by the Research Ethics Committee of the Institute of Psychology of Jagiellonian University (opinion no. KE/5_2023; 16 February 2023).
Since no prior study testing the effect of executive processes load on confidence was available to guide sample size determination, we estimated the minimal sample size required to observe the main effects of the executive component processes on response time (RT) based on Rietbergen et al. [
39]. To ensure sufficient statistical power, we adopted a convenience sample of sixty-six participants, which exceeds the minimum required
N = 22 calculated with G*Power (version 3.1.9.7) [
42] (α = 0.05, Power = 0.95), with the expected effect size estimated based on the study reported by Rietbergen et al. [
39].
2.1.2. Materials
We used a modified version of the classic Stroop task [
43], based on the design introduced by Rietbergen et al. [
39]. Stimuli were presented as either simple or compound displays. Simple stimuli consisted only of a target colour word. Compound stimuli consisted of the target word together with a dashed line, presented either above or below the word.
The word stimulus was one of four Polish colour words corresponding to red, blue, green, and brown. Each word was displayed in one of four font colours: red, blue, green, or brown. All words were presented in lowercase and were approximately 2 cm high and 6 cm wide.
The task manipulated three factors: conflict, complexity, and sequence. Conflict referred to the relation between word meaning and font colour. In congruent trials, the word meaning matched the font colour; in conflict trials, the word meaning and font colour differed. Complexity referred to the amount of information present in the stimulus. In simple trials, participants were presented with only the colour word. In compound trials, they were presented with both the colour word and the additional dashed line. Sequence referred to whether the stimulus complexity repeated or changed from the preceding trial. Thus, a trial could be repeated, as in simple trials following simple trials or compound trials following compound trials, or switched, as in simple trials following compound trials or compound trials following simple trials.
These manipulations allowed us to estimate three behavioural effects. The conflict effect indexed the cost of resolving interference between word meaning and font colour, corresponding to the classic Stroop effect. The complexity effect indexed the additional demands associated with processing the compound stimulus and determining the presence and position of the dashed line. The sequence effect indexed the cost of switching between simple and compound trial types.
To reduce low-level stimulus repetitions, we used two alternating stimulus lists. Blue and green colour-word stimuli appeared on odd trials, whereas red and brown colour-word stimuli appeared on even trials. This ensured that stimulus features, such as word meaning and font colour, did not repeat across consecutive trials, allowing us to focus on changes related to higher-level cognitive adaptation. Each word colour appeared equally often in congruent and incongruent displays. For example, the list contained the same number of green words printed in green font and green words printed in blue font. This balancing was used to prevent participants from learning accidental stimulus-frequency regularities.
The stimulus lists were balanced across the main task conditions. The task included 448 trials in total: 112 simple-congruent trials, in which word meaning and font colour matched and no additional line was present; 112 compound-congruent trials, in which word meaning and font colour matched and an additional line was present; 112 simple-conflict trials, in which word meaning and font colour differed and no additional line was present; and 112 compound-conflict trials, in which word meaning and font colour differed and an additional line appeared either above or below the word.
2.1.3. Procedure
The experiment was programmed in PsychoPy (version 2021.2.3) [
44] and conducted in the Laboratory of the Department of Experimental Psychology at the Institute of Psychology, Jagiellonian University. Before the study, all participants provided written informed consent.
Participants were instructed to ignore the meaning of the word and respond to its font colour. After the colour response, they indicated whether the dashed line was absent, below the word, or above the word. At the end of each trial, they rated their confidence in the correctness of their response.
The experimental task was preceded by two training sessions. The first training session consisted of 8 neutral trials with letter strings, such as “OOOOO” or “XXXXX.” Its purpose was to familiarize participants with the colour-response procedure. Each trial began with a black screen presented for 300 ms, followed by a fixation point for 500 ms and another black screen for 300 ms. The target stimulus was then presented for 700 ms. After the target disappeared, a fixation point was shown, and 500 ms later, the response options appeared on the left and right sides of the screen, together with the fixation point in the centre. The response options were two colour words displayed in white, for example, “brown” on the left and “red” on the right. They remained on the screen until the participant responded, but no longer than 1.8 s.
In the second training session, participants responded first to the font colour and then to the line component. They indicated whether the dashed line was absent, below the word, or above the word using three keys: “SPACE” for no line, “K” for a line below the word, and “O” for a line above the word. The line-position response was made without time pressure. Participants did not provide confidence ratings during the training sessions.
The experimental session consisted of 14 blocks of 32 trials. Stimuli were presented in a random order. Each trial began with a black screen for 300 ms, followed by a fixation point for 500 ms and another black screen for 300 ms. The target colour-word stimulus was then presented for 500 ms. After the target disappeared, a fixation point was shown alone for 300 ms. Next, the current colour-response options appeared on the left and right sides of the screen, with the fixation point remaining in the centre. The response screen remained visible until the participant responded, but no longer than 1.6 s.
Participants responded to the font colour using two keys. The “Z” key was assigned to the left response and was pressed with the left index finger. The “M” key was assigned to the right response and was pressed with the right index finger. The mapping between the target colour and response side was balanced, such that the correct response appeared on the left side in 50% of trials and on the right side in the remaining 50%. If no colour response was given within the response window, the trial ended with the feedback “Too slow.”
After the colour response, the word “Line” and a visual reminder of the line-response keys appeared on the screen. Participants then indicated whether the dashed line was absent, presented below the word, or presented above the word. They pressed the space bar with the right thumb to indicate no line, the “K” key to indicate a line below the word, and the “O” key to indicate a line above the word. This response was made without time pressure. After the line-position response, a black screen was presented for 600 ms. Participants then rated their confidence in the accuracy of their response on a 13-point scale. Importantly, participants evaluated their confidence in the entire response sequence, first identifying the colour and then the position of the line. The scale ranged from “certainly wrong” at point 1 to “certainly correct” at point 13. The midpoint, point 7, was labelled “I am guessing”; the remaining points were unlabelled. Participants moved along the scale using the “Z” and “M” keys and confirmed their selected rating with the space bar. To reduce accidental or impulsive confirmation, the selected point could be confirmed only after 500 ms. Confidence ratings were made without time pressure. After completing the experiment, participants were debriefed and paid for their participation. The entire procedure lasted approximately 60 min. The sequence of events within each experimental trial is illustrated in
Figure 1.
2.2. Results
2.2.1. Data Preparation
Accuracy was defined such that a trial was coded as correct only when both word colour and line-position responses were accurate. Accuracy was analyzed on all valid trials, whereas response times and confidence ratings were analyzed on correct trials only. For the response-time and confidence analyses, RTs more than 3 SD above the overall median (>1228.1 ms) or shorter than 100 ms were excluded. The first trial of the task and the first trial after each scheduled break were also excluded from these analyses. Higher confidence ratings indicated greater confidence that the response was correct, whereas lower ratings indicated greater confidence that an error had been made. We used an alpha level of 0.05 for all statistical tests. When applicable, post hoc comparisons were adjusted using the Tukey correction.
2.2.2. Data Analysis
To examine whether conflict, complexity, and sequence affected performance and confidence in the modified Stroop task, we analyzed accuracy, response times, and confidence ratings as a function of three within-participant factors: conflict, complexity, and sequence. Across the mixed-effects analyses in both experiments, categorical predictors were coded using simple contrasts.
Accuracy was analyzed using a generalized mixed-effects logistic model with a binomial distribution and logit link function. The model included conflict, complexity, sequence, all two-way interactions, and the three-way interaction as fixed effects, with random intercepts for participants:
Accuracy ~ 1 + Conflict + Complexity + Sequence + Conflict × Complexity + Conflict × Sequence + Complexity × Sequence + Conflict × Complexity × Sequence + (1 | participant).
Response times and confidence ratings were analyzed using linear mixed-effects models with the same fixed-effects structure and random intercepts for participants. The response-time model was specified as follows:
RT ~ 1 + Conflict + Sequence + Complexity + Conflict × Sequence + Conflict × Complexity + Sequence × Complexity + Conflict × Sequence × Complexity + (1 | participant).
The confidence-rating model was specified as follows:
Confidence rating ~ 1 + Conflict + Sequence + Complexity + Conflict × Sequence + Conflict × Complexity + Sequence × Complexity + Conflict × Sequence × Complexity + (1 | participant).
Mixed-effects analyses were conducted in Jamovi (version 2.6.26) [
45] using the GAMLj module (version 3.6.5) [
46]. Models were estimated with the bobyqa optimizer and degrees of freedom for linear mixed-effects models were obtained using Satterthwaite’s approximation. All models achieved successful convergence. The inclusion of random intercepts allowed us to account for between-participant variability in baseline accuracy, response times, and confidence ratings.
For the exploratory signal-detection analyses, the colour decision provided a binary left- versus right-hand response structure, with the required response hand balanced equally across trials. Because confidence concerned the correctness of the complete colour-plus-line response sequence, first-order correctness was defined using the same criterion: a trial was considered correct only when both responses were correct. On all trials except those containing an isolated line-position error, the binary first-order response corresponded directly to the participant’s observed left- or right-hand colour response. When the colour response was correct but the subsequent line-position response was incorrect (3.5% of the primary SDT trials), the complete response sequence was nevertheless incorrect and was therefore represented as a first-order error. This approach aligned the first-order correctness criterion with the outcome evaluated by the confidence judgement while retaining the binary response structure required for the SDT analysis.
To assess the robustness of the SDT findings to the binary coding used in the primary analysis, we repeated the complete analysis after excluding all trials with an incorrect line-position response. This additional restriction removed 3.9% of the primary SDT trials. In the sensitivity analysis, performance on the additional response component was held constant as correct; first-order correctness was therefore determined entirely by the colour response, and the binary first-order response corresponded directly to the observed left- or right-hand choice on every included trial. All other preprocessing steps, confidence ratings, condition definitions, and estimation procedures were identical to those used in the primary analysis.
Confidence was entered on the original 13-point scale used in the task. First trials and trials with omitted responses were excluded before SDT estimation. First-order sensitivity (d′), metacognitive sensitivity (meta-d′), and metacognitive efficiency (M-ratio = meta-d′/d′) were estimated separately for the four Conflict × Complexity conditions using maximum-likelihood estimation in metadPy 0.1.2 [
41,
47,
48]. Sequence was not included because further subdivision resulted in sparse participant-level cells. All 66 participants contributed estimates to both the primary and sensitivity analyses. The resulting estimates were analyzed using 2 × 2 repeated-measures ANOVAs with conflict and complexity as within-participant factors. Detailed model reports for all analyses in Experiments 1 and 2 are presented in the
Supplementary Materials (Sections S1.1–S2.7).
2.2.3. Accuracy
Firstly, we examined whether accuracy in the modified Stroop task was decreased by conflict, stimulus complexity, and sequence. Accuracy was analyzed using a generalized linear mixed-effects model with a binomial distribution and a logit link function. The model included conflict, complexity, sequence, and their interactions as fixed effects, with a random intercept for participant. The model was fitted to 29,502 trials from 66 participants. The fixed-effects structure was significant, χ2(7) = 426.56, p < 0.001. The marginal R2 was 0.08, whereas the conditional R2 was 0.23, suggesting additional between-participant variability in baseline accuracy beyond that captured by the fixed effects.
There were significant main effects of conflict, χ2(1) = 142.86, p < 0.001, complexity, χ2(1) = 212.45, p < 0.001, and sequence, χ2(1) = 48.73, p < 0.001. Overall, participants were more accurate in congruent than in conflict trials (M = 0.97 vs. 0.94, OR = 1.93, 95% CI [1.74, 2.15], p < 0.001), in simple than in compound trials (M = 0.97 vs. 0.94, OR = 2.23, 95% CI [2.01, 2.49], p < 0.001), and in repeated than in switched trials (M = 0.96 vs. 0.95, OR = 1.47, 95% CI [1.32, 1.64], p < 0.001). Accuracy was therefore affected by all three experimental manipulations.
There was a significant Conflict × Complexity interaction, χ2(1) = 8.48, p = 0.004. The conflict effect was significant for both simple and compound trials, but it was larger for simple trials (OR = 2.27, 95% CI [1.80, 2.87] vs. OR = 1.65, 95% CI [1.40, 1.93], respectively; p < 0.001 for both comparisons).
There was also a significant Complexity × Sequence interaction, χ2(1) = 39.83, p < 0.001. The sequence effect was significant for simple trials, whereas the corresponding comparison for compound trials did not reach significance. Participants were more accurate in simple-repeated than in simple-switched trials (M = 0.98 vs. 0.96, OR = 2.08, 95% CI [1.65, 2.63], p < 0.001). In contrast, the estimated difference between compound-repeated and compound-switched trials did not reach significance (both Ms = 0.94, OR = 1.04, 95% CI [0.88, 1.22], Tukey-adjusted p = 0.933). Thus, the sequence effect was larger for simple than compound trials. The Conflict × Sequence interaction, χ2(1) = 0.10, p = 0.753, and the three-way Conflict × Complexity × Sequence interaction, χ2(1) = 0.77, p = 0.381, did not reach significance.
In summary, participants were less accurate in conflict than in congruent trials, in compound than in simple trials, and in switched than in repeated trials. However, these effects were qualified by two interactions. First, the conflict effect was larger for simple than compound trials. Second, the sequence effect was larger for simple than compound trials: accuracy was lower in simple-switched than in simple-repeated trials, whereas the estimated difference between compound-switched and compound-repeated trials did not reach significance. The model-estimated accuracy pattern across conditions is presented in
Table 1 and
Figure 2.
2.2.4. Response Times
Next, we examined whether response times in the modified Stroop task were modulated by conflict, stimulus complexity, and sequence. Response times were analyzed using a linear mixed-effects model with conflict, complexity, sequence, all two-way interactions, and the three-way interaction as fixed effects, with random intercepts for participants. The model was fitted to 26,787 observations from 66 participants. The fixed-effects structure was significant, χ2(7) = 1750.30, p < 0.001. Marginal and conditional R2 were 0.05 and 0.21, respectively.
There were significant main effects of conflict, F(1, 26,714.42) = 83.50, p < 0.001, and complexity, F(1, 26,715.31) = 1507.79, p < 0.001. Participants responded more slowly in conflict than in congruent trials (estimated difference = 20 ms, 95% CI [10, 20]) and in compound than in simple trials (estimated difference = 80 ms, 95% CI [70, 80]). The estimated overall difference between repeated and switched trials did not reach significance, F(1, 26,715.22) = 0.29, p = 0.592.
There were significant Conflict × Sequence, F(1, 26,714.85) = 88.52, p < 0.001, Conflict × Complexity, F(1, 26,714.24) = 42.10, p < 0.001, and Sequence × Complexity interactions, F(1, 26,714.28) = 54.26, p < 0.001. These two-way interactions were further qualified by a significant Conflict × Sequence × Complexity interaction, F(1, 26,715.07) = 24.05, p < 0.001. We therefore decomposed the three-way interaction by examining conflict and complexity effects separately in repeated and switched trials.
In repeated trials, the conflict effect depended on stimulus complexity. Conflict slowed responses for simple stimuli by 20 ms, p < 0.001, whereas for compound stimuli, the direction of the conflict effect was reversed: congruent–compound-repeated responses were 20 ms slower than conflict–compound-repeated responses, p < 0.001. The complexity effect was significant at both levels of conflict, but it was larger in congruent than conflict trials (110 vs. 70 ms; p < 0.001 for both comparisons).
In switched trials, conflict increased response times for both simple and compound stimuli by 40 and 30 ms, respectively (p < 0.001 for both comparisons). The complexity effect was also significant in switched trials and was estimated at 60 ms under both congruent and conflict conditions (p < 0.001 for both comparisons).
Finally, we examined the sequence effect within each Conflict × Complexity combination. In congruent–simple trials, the estimated difference between simple-switched and simple-repeated trials did not reach significance (Tukey-adjusted p = 0.660). In conflict–simple trials, participants responded more slowly in simple-switched than simple-repeated trials by 20 ms, p < 0.001. In congruent–compound trials, participants responded 40 ms faster in compound-switched than compound-repeated trials, p < 0.001. In conflict–compound trials, participants responded 10 ms more slowly in compound-switched than compound-repeated trials, Tukey-adjusted p = 0.006.
In summary, response times were longer in conflict than in congruent trials and in compound than in simple trials, whereas the estimated main effect of sequence did not reach significance. However, the significant three-way interaction showed that sequence effects depended jointly on conflict and stimulus complexity. In repeated trials, conflict slowed responses for simple stimuli but showed the opposite pattern for compound stimuli. In switched trials, conflict slowed responses for both simple and compound stimuli. Complexity increased response times in all conditions, but this effect was largest in congruent-repeated trials and smaller in conflict-repeated and switched trials. The condition-specific response-time pattern underlying the three-way interaction is shown in
Table 2 and
Figure 3.
2.2.5. Confidence Judgments
Finally, we examined whether confidence ratings in the modified Stroop task were modulated by conflict, stimulus complexity, and sequence. Confidence ratings were analyzed using a linear mixed-effects model with conflict, complexity, sequence, all two-way interactions, and the three-way interaction as fixed effects, with random intercepts for participants. The model was fitted to 26,787 observations from 66 participants. The fixed-effects structure was significant, χ2(7) = 102.28, p < 0.001. Marginal and conditional R2 were 0.003 and 0.23, respectively.
There was a significant main effect of conflict, F(1, 26,714.21) = 19.85, p < 0.001. Participants reported lower confidence in conflict than in congruent trials (M = 12.50 vs. 12.56, b = −0.06, 95% CI [−0.09, −0.03]).
The main effects of sequence, F(1, 26,714.75) = 4.37, p = 0.037, and complexity, F(1, 26,714.81) = 11.65, p < 0.001, were also significant. When averaged across conditions, confidence was lower in switched than in repeated trials and in compound than in simple trials. Importantly, however, both marginal effects were qualified by the Sequence × Complexity interaction reported below.
There was a significant Sequence × Complexity interaction, F(1, 26,714.11) = 65.14, p < 0.001. The sequence effect was significant for both simple and compound trials, but it differed in direction and was larger for simple trials. Participants reported higher confidence in simple-repeated than simple-switched trials (M = 12.62 vs. 12.48, difference = 0.14, p < 0.001). For compound trials, the sequence effect was reversed: confidence was higher in compound-switched than compound-repeated trials (M = 12.54 vs. 12.46, difference = 0.08, p < 0.001). The Conflict × Sequence, Conflict × Complexity, and three-way interactions did not reach significance (all p values ≥ 0.755).
In summary, confidence was lower in conflict than in congruent trials, consistent with the view that conflict reduced subjective certainty, although this difference was small in absolute terms. Sequence and complexity also affected confidence, but their marginal effects were qualified by a crossover Sequence × Complexity interaction. Specifically, confidence was higher in simple-repeated than simple-switched trials, whereas the reverse pattern was observed for compound trials. The Sequence × Complexity interaction in confidence is shown in
Table 3 and
Figure 4.
2.2.6. Exploratory Analyses of Metacognitive Sensitivity and Efficiency
We next examined first-order sensitivity and the sensitivity and efficiency of the corresponding confidence judgments across conflict and complexity conditions. First-order sensitivity was indexed by d′, metacognitive sensitivity by meta-d′, and metacognitive efficiency by M-ratio. Estimates were obtained separately for the four Conflict × Complexity conditions.
The analysis of d′ revealed significant main effects of conflict, F(1, 65) = 85.37, p < 0.001, ηp2 = 0.57, and complexity, F(1, 65) = 41.59, p < 0.001, ηp2 = 0.39. First-order sensitivity was lower in conflict than in congruent trials and in compound than in simple trials. The Conflict × Complexity interaction did not reach significance, F(1, 65) = 0.03, p = 0.868.
Meta-d′ showed a significant main effect of complexity, F(1, 65) = 22.71,
p < 0.001, ηp
2 = 0.26. Metacognitive sensitivity was higher in simple than compound trials. The main effect of conflict did not reach significance, F(1, 65) = 2.42,
p = 0.125, ηp
2 = 0.04. However, we observed a significant Complexity × Conflict interaction, F(1, 65) = 4.54,
p = 0.037, ηp
2 = 0.07. The interaction was most clearly reflected in the complexity effect: meta-d′ was substantially higher in simple than in compound trials under conflict (mean difference = 0.69, Tukey-adjusted
p < 0.001), whereas the corresponding comparison under congruence did not reach significance after adjustment (mean difference = 0.33, Tukey-adjusted
p = 0.083). The reduction in metacognitive sensitivity from simple to compound trials was larger under conflict than under congruence. The Conflict × Complexity pattern in meta-d′ is shown in
Table 4 and
Figure 5.
We next quantified metacognitive efficiency with the M-ratio, which adjusts meta-d′ for primary-task sensitivity d′. Conflict had a significant main effect, F(1, 65) = 15.54,
p < 0.001, ηp
2 = 0.19: M-ratio was higher in conflict than in congruent trials. The effects of complexity, F(1, 65) = 0.28,
p = 0.597, and Conflict × Complexity, F(1, 65) = 0.00,
p = 0.971, did not reach significance. Because M-ratio expresses meta-d′ relative to d′, this result should not be interpreted as an absolute enhancement of metacognitive sensitivity.
Table 5 and
Figure 6 show the corresponding M-ratio estimates across conflict and complexity conditions.
To determine whether the SDT findings depended on the first-order coding used in the primary analysis, we repeated the analyses after excluding trials with an incorrect line-position response. Conflict continued to reduce d′, F(1, 65) = 91.68, p < 0.001, ηp2 = 0.59, whereas the complexity effect did not reach significance, F(1, 65) = 3.31, p = 0.073. Complexity continued to reduce meta-d′, F(1, 65) = 25.47, p < 0.001, ηp2 = 0.28, and the Conflict × Complexity interaction was again significant, F(1, 65) = 9.46, p = 0.003, ηp2 = 0.13. The simple–compound difference was significant under conflict (Tukey-adjusted p < 0.001), whereas the corresponding adjusted comparison under congruence did not reach significance (p = 0.052).
The conflict-related increase in M-ratio was also retained, F(1, 65) = 22.77, p < 0.001, ηp2 = 0.26, but was qualified by complexity: the Conflict × Complexity interaction was significant, F(1, 65) = 9.74, p = 0.003, ηp2 = 0.13. M-ratio was higher under conflict in simple trials (Tukey-adjusted p < 0.001), whereas the corresponding comparison in compound trials did not reach significance (Tukey-adjusted p = 0.383). M-ratio was also lower in compound than in simple trials under congruent and conflict conditions.
Taken together, excluding trials with incorrect line-position responses preserved the conflict-related reduction in d′, the complexity-related reduction in meta-d′, and the stronger expression of this effect under conflict, as well as the overall increase in M-ratio under conflict. However, the effect of complexity on d′ was no longer significant. In addition, M-ratio exhibited both a main effect of complexity and a Conflict × Complexity interaction.
2.3. Interim Discussion of Experiment 1
In this experiment, we examined whether cognitive-control demands influence performance, reported confidence, and metacognitive monitoring in a modified Stroop task. We observed the effect of all three manipulations on response accuracy: it was lower in conflict compared to congruent trials, in compound compared to simple trials, and in switched compared to repeated trials. These findings suggest that experimental manipulations increased demands on the intended control processes: inhibition, updating, and shifting [
38]. We assumed that the presence of conflicting stimuli would increase inhibition demands by requiring the cognitive system to resolve competition between the relevant stimulus dimension and the irrelevant word meaning [
43]. Compound trials were assumed to increase updating demands by the requirement to maintain and process an additional line-position component, and switched trials required the cognitive system to shift between simple and compound stimuli.
We also observed the effect of conflict and complexity on response times. Responses were slower in conflict compared to congruent trials, and in compound compared to simple trials. Complexity increased response times across conditions, although the size of this effect varied depending on conflict and sequence. Thus, the response-time results confirmed the effects of conflict and stimulus complexity but also showed that the temporal pattern of performance depended on the specific combination of task conditions. These results closely resemble the pattern reported by Rietbergen et al. [
39].
Surprisingly, in repeated trials, conflict increased response times to simple stimuli but reduced response times to complex stimuli. In switched trials, conflict slowed responses for both simple and compound stimuli. This pattern may reflect visual-working-memory dynamics. The complexity manipulation was intended to increase the need to update visual working memory by adding the line-position component to the colour-response requirement [
38,
39,
40]. When a compound trial followed a simple trial, the line component may have been relatively salient because it introduced new task-relevant visual information. Novelty and visual salience can facilitate encoding into visual working memory [
49,
50]. By contrast, when compound trials appeared consecutively, participants had to update the current line-position representation in the context of recently processed line information from the preceding trial. This may have increased interference between adjacent visual representations, consistent with evidence that previously stored items can interfere with current visual-working-memory contents [
51,
52]. Thus, switching between simple and compound trials may have affected memory-demanding trials differently than trials requiring conflict resolution. The novelty effect may have been more pronounced in congruent trials, where word meaning did not activate a competing response [
43,
53]. Participants may therefore have been more able to postpone the colour response and accumulate information about line position. This could explain why congruent compound-repeated trials produced the longest RTs.
The main question of Experiment 1 was whether increased cognitive-control demands influenced confidence in response correctness. The confidence analysis was restricted to objectively correct trials. Participants reported lower confidence in conflict than in congruent trials even when the final task response was correct. The conflict-related difference was statistically reliable but small in absolute magnitude (0.06 points on the 13-point scale), and the Conflict × Complexity, Conflict × Sequence, and three-way interactions did not reach significance. Confidence was also affected by complexity and sequence, but these effects were qualified by a crossover interaction. Confidence was higher in simple-repeated than in simple-switched trials, whereas confidence was higher in compound-switched than in compound-repeated trials. Thus, there was no uniform confidence cost of switching. Instead, confidence depended on the relation between the complexity of the current trial and that of the preceding trial. Because this interaction was not predicted, its interpretation should remain cautious.
The exploratory SDT analyses showed that conflict and complexity affected first-order and metacognitive performance differently. Conflict substantially reduced first-order sensitivity but did not produce an overall reduction in meta-
d′. Correspondingly, M-ratio was higher under conflict than under congruent conditions. Because M-ratio expresses meta-
d′ relative to
d′, this pattern indicates that first-order sensitivity declined more strongly than metacognitive sensitivity under conflict; it should not be interpreted as an absolute improvement in metacognitive monitoring [
41,
48]. Complexity showed a different pattern. Meta-
d′ was lower in compound than in simple trials, and this reduction was strongest when compound processing occurred under conflict. Importantly, this effect was also observed when trials with an incorrect line-position response were excluded, whereas the corresponding complexity effect on
d′ was no longer detected. The sensitivity analysis therefore strengthens the conclusion that the complexity-related reduction in metacognitive sensitivity does not depend on the particular treatment of trials containing line-position errors in the primary SDT representation. Together, these findings suggest that increased stimulus complexity compromises metacognitive sensitivity, particularly when additional processing demands coincide with response conflict. This pattern is consistent with evidence that active manipulation of working-memory contents can impair metacognitive monitoring [
35].