1. Introduction
Artificial intelligence is increasingly embedded in routine organizational decision-making rather than confined to specialist technical functions. Employees now encounter AI-generated recommendations, forecasts, rankings, summaries, and explanations across logistics, procurement, operations, marketing, human resources, finance, and risk management. In many of these settings, AI does not make the final decision autonomously; instead, employees must interpret machine-generated advice and decide whether to accept, modify, verify, or reject it. Recent work on human–AI collaboration therefore increasingly emphasizes the quality of joint decision-making rather than technology adoption alone (
Do Khac & Leyer, 2026;
Gonzalez & Heidari, 2025;
Latto et al., 2026;
Li & Tian, 2026).
This shift creates a practical and psychological challenge. AI-generated advice can be fluent, confident, and professionally presented even when it is inaccurate, incomplete, or based on weak assumptions. Hallucinations and reliability failures can therefore create risks when users infer correctness from presentation quality or apparent certainty (
Cheng et al., 2026). At the same time, people do not respond to AI advice only on the basis of objective quality. Reliance is also shaped by trust, confidence, source information, and prior beliefs about the system (
Ding et al., 2026;
Hao et al., 2026;
Meincke et al., 2026;
Pearson et al., 2026). The quality of the AI output and the quality of the human response to that output are therefore distinct components of AI-supported decision performance.
A central concern is whether reliance is appropriately calibrated. Automation-bias research shows that users may give excessive weight to automated recommendations despite contradictory evidence, whereas algorithm-aversion research shows that users may also reject useful algorithmic advice after observing errors (
Dietvorst et al., 2015;
Lee & See, 2004;
Romeo & Conti, 2026). More recent work similarly emphasizes that appropriate reliance depends on aligning trust and acceptance with the actual quality of the recommendation rather than maximizing trust in AI (
Dogru & Krämer, 2025;
Liebherr et al., 2026). Cognitive offloading adds another layer: AI can reduce effort, but reliance may become problematic when users disengage from information that still requires independent judgment (
Gerlich, 2025;
Zhu et al., 2026). These risks are especially relevant under time pressure, interruptions, competing information, and salient confidence cues.
Existing measures provide important but incomplete perspectives on this problem. AI-literacy scales assess knowledge, competencies, and beliefs about artificial intelligence (
Carolus et al., 2023;
Koch et al., 2024;
Ng et al., 2024), AI self-efficacy measures assess perceived capability to use AI effectively (
Morales-García et al., 2024;
Wang & Chuang, 2024), and trust measures assess users’ evaluations of automated systems (
Jian et al., 2000;
Lee & See, 2004). These constructs are valuable, but they do not directly establish whether a person can identify an inaccurate recommendation, integrate conflicting evidence, revise an initial judgment, or resist a confidently expressed but unsupported conclusion in a specific decision. Performance-based human–AI research already examines advice taking, reliance, confidence, and decision accuracy; the measurement gap addressed here is therefore not the absence of behavioral assessment itself. Rather, it concerns the limited integration of repeated workplace decision scenarios into a psychometric individual-differences framework that can examine structure, cognitive correlates, process indicators, and criterion-related evidence within the same assessment.
Traditional cognitive measures address a different part of this gap. Tests of reasoning, working memory, processing speed, and attention provide established indicators of cognitive functioning, but they generally do not place these abilities inside AI-supported organizational decisions. The Cattell–Horn–Carroll (CHC) framework offers a useful organizing basis because it conceptualizes cognition as a hierarchy of broad and narrower abilities that contribute differently across tasks (
Carroll, 1993;
McGrew, 2009;
Schneider & McGrew, 2018). In the present study, CHC is used as an interpretive framework rather than as the basis for proposing a new broad ability. Fluid reasoning may support evaluation of novel or inconsistent recommendations; working memory may support the maintenance and comparison of multiple constraints; processing speed may contribute to efficient evaluation under time limits; and attentional control may help maintain focus on diagnostic information when distracting or misleading cues are present.
Accordingly, the study conceptualizes cognitive performance under AI advice as an applied performance domain in which established cognitive resources are expressed while individuals evaluate machine-generated recommendations. The assessment was designed around three theoretically specified and closely related content/performance dimensions—AI error detection, evidence integration, and cognitive control and adaptive reliance—while confidence, response time, and reliance behavior were treated as complementary process indicators. This framing does not assume that the three dimensions are independent abilities or that working with AI constitutes a new form of intelligence. Instead, it treats them as related manifestations of performance within a common AI-supported decision context.
The present study therefore develops and initially validates a CHC-informed, performance-based assessment of cognitive performance under AI advice in organizational decision-making. Its contribution is primarily psychometric and integrative. First, it organizes repeated AI-supported workplace scenarios within an individual-differences assessment framework. Second, it examines whether performance shows a coherent psychometric structure while remaining related to established cognitive measures. Third, it combines final accuracy with confidence, response time, reliance behavior, and decision revision to characterize performance more fully than a total score alone. Fourth, it evaluates concurrent criterion-related and incremental validity beyond demographic and work characteristics, AI experience, conventional cognitive-performance measures, and AI-related self-reports. Finally, it examines preliminary group comparability through measurement-invariance and differential-item-functioning analyses. Scenario characteristics in the 24-item fractional design were fixed properties of particular scenarios rather than independently randomized manipulations; consequently, scenario-level coefficients are interpreted as adjusted associations and not as isolated causal effects.
3. Materials and Methods
3.1. Assessment Development
The assessment was developed through construct specification, scenario design, expert review, cognitive interviewing, pilot testing, and item refinement. The conceptual framework was CHC-informed rather than intended to define a new CHC ability. Fluid reasoning, working memory, and processing speed were treated as established cognitive resources, while attentional control was treated as a companion individual-difference resource. The 24 scenarios were theoretically organized into three closely related content/performance dimensions—AI error detection, evidence integration, and cognitive control and adaptive reliance—with confidence calibration and processing efficiency treated as cross-cutting process indicators.
An initial pool of 30 organizational scenarios was developed across broadly applicable contexts, including supplier selection, logistics disruption, inventory planning, employee scheduling, project prioritization, customer complaints, budget allocation, risk evaluation, marketing, recruitment, quality control, and operational planning. Specialized occupational terminology was minimized to reduce dependence on job-specific knowledge. Each scenario presented a concise organizational problem, task-relevant evidence, an AI-generated recommendation, and a required decision. Participants made an initial judgment, evaluated the AI advice, made a final decision, reported confidence, and could retain or revise their original response. Response-time variables were derived from the timing information recorded for the pre- and post-advice response stages.
The complete 24-scenario assessment administered in the study, including scenario evidence, AI-generated recommendations, response options, scoring keys, and scenario-characteristic coding, is provided in
Supplementary S2 of the Supporting Information. The complete scenario-level blueprint is reported in
Table S1.
Table 2 summarizes the assessment blueprint by linking each content/performance dimension to its organizational task context, principal cognitive demand, relevant scenario characteristics, required response, and process indicator.
Scenario presentation order was randomized, but the scenario characteristics themselves were fixed properties of particular scenarios in a fractional design rather than independently randomized within otherwise identical content. The final assessment contained 12 scenarios with correct and 12 with incorrect AI advice; 12 high- and 12 low-confidence AI cues; 6 scenarios with time pressure and 18 without; 4 scenarios with an interruption and 20 without; 12 high- and 12 low-information-conflict scenarios; and 8 low-, 8 moderate-, and 8 high-complexity scenarios. Consequently, coefficients for these characteristics are interpreted as adjusted associations within the observed scenario set rather than as isolated causal effects.
Table 3 summarizes the scenario characteristics embedded in the final assessment and the principal cognitive demands and outcomes associated with each characteristic.
The original 30 scenarios were reviewed by 10 experts representing psychometrics, cognitive assessment, supply chain decision-making, human–AI interaction, and organizational decision-making. Their professional experience averaged 15.9 years. Experts evaluated relevance, clarity, realism, cognitive demand, occupational neutrality, and ambiguity. Across the initial pool, the scale-level average CVI was 0.800. Twenty-four scenarios achieved I-CVI values of at least 0.80; within this group, I-CVI values ranged from 0.80 to 1.00, CVRs from 0.60 to 1.00, and modified kappa values from approximately 0.79 to 1.00. Their S-CVI/Ave was 0.908. Complete item-level results and the expert evaluation form are provided in
Table S2 and Supplementary S3.
Cognitive interviews were then conducted with 20 employed participants representing procurement, customer service, planning, operations, warehousing, and transport functions. Mean ratings were 4.40/5 for clarity, 4.20/5 for realism, and 3.05/5 for perceived difficulty. Eleven participants identified no substantive problems; the remaining interviews identified calculation burden, ambiguous evidence, instruction wording, or terminology that required refinement. These findings were used to simplify unnecessary calculations and clarify wording without reducing the intended cognitive demands. Full cognitive-interview findings and the interview protocol are provided in
Table S3 and Supplementary S4.
Pilot testing involved 120 employees. Across the 30 candidate scenarios, mean difficulty was 0.512, mean discrimination was 0.376, and mean response time was 47.53 s. Twenty-three scenarios were retained without major change. Scenario I06 showed acceptable content validity but weaker discrimination and was refined rather than removed. Scenarios I25–I30 showed a less favorable combination of content validity, discrimination, response burden, and interpretability and were removed, producing the final 24-scenario assessment. Complete pilot statistics are reported in
Table S4.
Table 4 summarizes the content-validation and pilot evidence used in the item-selection decisions.
Item-retention decisions considered content validity, cognitive-interview evidence, item difficulty, discrimination, response burden, ambiguity, and occupational neutrality jointly rather than relying on a single statistical threshold.
Figure 2 summarizes the resulting development sequence from construct definition through psychometric evaluation and validity testing.
3.2. Participants and Procedure
Data were collected in Türkiye from 13 May to 10 June 2026 using a nonprobability purposive employee-recruitment strategy. Potential participants were reached through company email invitations and LinkedIn. The study was administered online through Google Forms in Turkish, and no financial or other participation compensation was provided. Eligibility required participants to be at least 18 years old, currently employed at least part-time, use digital technologies at work, and have basic familiarity with AI-supported tools or recommendations; advanced AI expertise was not required.
The dataset initially contained 900 participant records; 780 were retained after the prespecified eligibility and data-quality screening procedures described below. The final sample ranged from 20 to 64 years of age (
M = 38.82,
SD = 9.31) and averaged 10.50 years of work experience (
SD = 5.29). Participants represented transport operations, procurement, warehousing, inventory planning, customer operations, supply chain analytics, and management. Detailed demographic, occupational, and AI-use characteristics are reported in
Table 5. No single-effect a priori power analysis was used because the primary aim was psychometric development and validation across multiple models rather than testing one focal effect. The final
N = 780 yielded 18,720 participant-by-scenario observations and substantial subgroup sizes for the planned psychometric and group-comparability analyses, while inference involving scenario clustering was treated conservatively because only 24 scenarios were available.
Participants first completed demographic, occupational, and AI-experience questions, followed by the 24-scenario organizational AI-advice assessment. They then completed the companion cognitive-performance measures, AI-related self-report measures, and the organizational decision-quality criterion task. Scenario presentation order was randomized; the scenario characteristics described in
Table 3 were not independently randomized within scenario content. The administration captured initial and final decisions, AI-evaluation responses, confidence, acceptance or rejection of AI advice, decision revision, and the response-time variables available for the pre- and post-advice stages. Full focal assessment instructions are provided in
Supplementary S1.
Table 5 illustrates the characteristics of the final analytic sample.
Participation was voluntary and anonymous, and informed consent was obtained before data collection. Ethical approval was granted before recruitment by the Istanbul Atlas University Social and Humanities Research and Publication Ethics Committee on 13 April 2026 (Decision No. 03/02; Official Document No. E-22686390-050.99-95101). Participants could discontinue participation at any time.
3.3. Measures and Task Design
The focal measure was the 24-scenario performance-based organizational AI-advice assessment developed in
Section 3.1. Each scenario presented an organizational problem and relevant evidence followed by an AI-generated recommendation. Participants made an initial judgment, evaluated the AI advice, made a final decision, and rated confidence on a 0–10 scale. The principal 24-scenario indicators were final decision accuracy (0–24), AI-evaluation accuracy (0–24), appropriate reliance (0–24), confidence-calibration indices, response-time measures, and decision revision. AI-evaluation accuracy refers to the correct evaluation of the AI recommendation across all 24 scenarios and is distinct from the eight-scenario AI-error-detection content/performance dimension. The three dimension scores were based on the eight scenarios assigned to each theoretically specified content/performance grouping.
Supplementary S6 and Table S14 document the study-specific variables, response coding, and scoring rules. Published measures are identified by their original sources. All study materials were administered in Turkish, and scenario wording and task comprehensibility were evaluated through the cognitive-interview and pilot stages described above. The present study does not claim a separate translation-validation study for the published measures.
Conventional cognitive-performance measures were included to examine convergent relationships. Fluid/general reasoning was measured with the 16-item International Cognitive Ability Resource (ICAR-16) (
Condon & Revelle, 2014;
Dworak et al., 2021), scored 0–16 (α = 0.809). Cognitive reflection was assessed with four binary-scored items based on established cognitive-reflection tasks (
Frederick, 2005;
Thomson & Oppenheimer, 2016), scored 0–4 (α = 0.542). Working memory was represented by a brief aggregate performance score informed by working-memory span principles (
Conway et al., 2005;
Unsworth et al., 2005); observed scores ranged from 3 to 14. Processing speed was represented by a brief timed performance score; observed scores ranged from 30 to 80. The retained dataset contains aggregate working-memory and processing-speed scores rather than trial-level responses, so item-level internal-consistency estimates are not reported for these two companion measures. Attentional control was represented by eight study-specific dichotomously scored indicators (ATT1–ATT8), coded 0/1 and summed to a 0–8 total (α = 0.718). Because these indicators were study-specific rather than a published standardized attention task, attentional control was treated as a companion individual-difference measure rather than as a separately validated CHC ability in the focal assessment.
AI-related measures included MAILS-10 (
Carolus et al., 2023;
Koch et al., 2024), scored as the mean of 10 six-point items (α = 0.914), and a 12-item trust-in-AI measure (
Jian et al., 2000), scored on a six-point scale (α = 0.921). AI self-efficacy was measured with six items scored from 1 to 6 and averaged across items (α = 0.873). AI-use frequency was coded ordinally from 1 (rarely) to 5 (several times daily). Prior generative-AI experience was recorded as Limited, Moderate, or Extensive and coded 1, 2, and 3, respectively. These variables were used as related self-report constructs or experience controls rather than as substitutes for demonstrated performance.
Concurrent criterion-related validity was evaluated with a separately scored six-item organizational decision-quality task. Each item was scored correct or incorrect and summed from 0 to 6 (α = 0.748). Because the criterion task and focal assessment were administered in the same study session, the resulting evidence is interpreted as concurrent rather than predictive validity. The criterion was scored independently from the focal 24-scenario assessment.
Table 6 shows the measures, scoring procedures, reliability information, and analytic roles.
3.4. Data Analysis
Analyses proceeded from data screening and item diagnostics to dimensionality, item-response modelling, process analysis, validity testing, and preliminary group-comparability assessment. Analyses were implemented in R version 4.6.1 and Python version 3.14.7; the complete software, package, version, coding, and reproducibility specification is provided in
Supplementary S5 of the Supporting Information. Screening considered missingness, incomplete assessments, technical failures, repeated response patterns, implausibly short completion times, and extreme response-time observations. Very slow responses were not removed solely because of their duration because long response times may reflect genuine processing difficulty.
Preliminary item analysis examined difficulty, discrimination, corrected item–total correlations, response-time intensity, confidence distributions, and potential floor or ceiling effects. For binary performance items, difficulty was the proportion correct. Dimensionality was evaluated using tetrachoric correlations, parallel analysis, and the minimum average partial procedure. Confirmatory models compared one-factor, correlated three-factor, higher-order, and bifactor specifications. CFI, TLI, RMSEA, SRMR, AIC, and BIC were considered jointly, and the correlated dimensions were not assumed to represent independent abilities.
Item-level measurement properties were evaluated with two-parameter logistic (2PL) item-response models. With i denoting participants and j denoting items, the probability of a correct response was represented as follows:
Item-level measurement properties were then examined using two-parameter logistic item response theory models. For participant
responding to item
, the probability of a correct response was represented as
Here,
aj denotes item discrimination,
bj item difficulty, and
θi the participant’s latent performance level. A unidimensional 2PL model was compared with a correlated multidimensional 2PL model and a bifactor specification. Item fit, information, conditional standard errors, and model fit were examined. Conditional measurement precision was summarized as:
where
I(
θ) is the test information available at a given latent-performance level.
Accuracy and response time were also examined jointly using a log-normal response-time component. With
i denoting participants and
j scenarios, the timing component was represented as:
where
Tij is response time,
λj is scenario time intensity,
νi is participant processing speed, and
εij is the residual term. Response time was interpreted jointly with accuracy; faster responding was not assumed to indicate better performance by itself.
The primary scenario-level analysis modeled binary final decision accuracy using logistic regression with two-way participant- and scenario-clustered covariance estimation. The prespecified primary model included AI-advice correctness, AI confidence, time pressure, interruption, information conflict, a high-complexity contrast, ICAR-16, working memory, processing speed, attentional control, AI literacy, correctness × confidence, correctness × time pressure, and interruption × working memory. Continuous person-level predictors were standardized before estimation. The model estimated an ordinary logistic mean structure; participant and scenario random intercepts were not included in the primary model.
In Equation (4),
xj denotes the vector of scenario characteristics,
zi the standardized person-level predictors, and
wij the prespecified interaction terms. Because the smaller clustering dimension contained only 24 scenarios, cluster-robust inference was reported conservatively using a t distribution with 23 degrees of freedom. Robustness was also examined through leave-one-scenario-out refits and an alternative scenario-random-intercept model. A fully crossed participant- and scenario-random-intercept model was explored but was not used for inference because it did not converge reliably. The previously considered remaining-23-scenario performance score was not included in the primary model because it is a within-assessment rest score rather than an independent cognitive measure; models using the rest score, balanced split-half scores, other-dimension scores, and an independent cognitive composite are reported as sensitivity analyses in the
Supporting Information. H5 was evaluated directly among incorrect-AI scenarios by testing whether participant performance moderated the association between high AI confidence and decision accuracy/reliance.
Confidence ratings were rescaled from 0–10 to 0–1 and evaluated with calibration bias, absolute calibration error, and the Brier score (
Brier, 1950). Calibration bias was calculated as:
where
ck is reported confidence and
yk is observed correctness coded 0 or 1. Positive values indicate overconfidence and negative values underconfidence. Absolute calibration error was calculated as:
and probabilistic calibration was additionally summarized using the Brier score:
Lower absolute calibration error and Brier scores indicate better calibration. Confidence discrimination was examined by comparing confidence for correct and incorrect decisions and by inspecting calibration patterns across relevant scenario characteristics.
Validity analyses distinguished convergent, discriminant, concurrent criterion-related, and incremental evidence. Convergent analyses examined relationships with ICAR-16, cognitive reflection, working memory, processing speed, and attentional control; H4 was additionally evaluated with dimension-aligned correlations linking ICAR-16 to AI error detection and working memory to evidence integration. AI literacy, trust in AI, AI self-efficacy, and AI experience were treated as related but conceptually distinct constructs. Incremental concurrent validity was evaluated hierarchically: Model 1 entered demographic variables; Model 2 added work experience; Model 3 added AI-use frequency and prior generative-AI experience; Model 4 added ICAR-16, cognitive reflection, working memory, and processing speed; Model 5 added AI literacy, trust in AI, and AI self-efficacy; and Model 6 added focal 24-scenario final accuracy. Attentional control was evaluated separately in a sensitivity model because its status differs from the conventional cognitive performance block. Ten-fold cross-validation was used as an internal assessment of statistical stability and is not interpreted as prospective external validation. Incremental contribution was summarized as:
Measurement invariance and differential item functioning (DIF) were examined as preliminary group-comparability tests. Two-group invariance comparisons used male (n = 427) versus female (n = 348), age <40 years (n = 415) versus ≥40 years (n = 365), bachelor’s degree or below (n = 498) versus postgraduate education (n = 282), nonmanager (n = 546) versus manager (n = 234), daily-or-more AI use (n = 433) versus weekly-or-less use (n = 347), and operations-oriented roles (n = 374) versus planning/analytical roles (n = 406). The five participants identifying as nonbinary/another gender were retained in overall analyses but were not included in the two-group gender invariance comparison. Invariance progressed from configural to loading and threshold models, with residual constraints examined where appropriate. Logistic DIF models used the remaining 23 items as the matching score and tested uniform group and non-uniform group-by-matching-score terms. Benjamini–Hochberg false-discovery-rate adjustment was applied separately to uniform- and non-uniform-DIF test families within each group comparison, and statistical significance was interpreted together with ΔMcFadden R2 and adjusted probability differences rather than as proof of fairness.
4. Results
4.1. Sample Characteristics, Data Quality, and Preliminary Item Results
The dataset contained 900 participant records. Following eligibility and data-quality screening, 780 cases were retained for analysis and 120 were excluded. Exclusions were due to implausibly short completion times (n = 33), incomplete assessments (n = 28), technical failures (n = 19), external AI assistance (n = 17), repeated response patterns (n = 15), or failure to meet the eligibility criteria (n = 8). Among the 780 retained cases, no missing values were observed in the demographic, scenario-level, or scale-item variables used in the analyses.
Participants ranged from 20 to 64 years of age (M = 38.82, SD = 9.31) and reported an average of 10.50 years of work experience (SD = 5.29). The sample included 427 men (54.7%), 348 women (44.6%), and 5 participants identifying as nonbinary or another gender (0.6%). Occupational representation included transport operations (20.9%), procurement (16.4%), warehousing (16.0%), inventory planning (15.5%), customer operations (11.0%), supply chain analytics (10.6%), and management (9.5%). AI use was relatively frequent: 15.1% used AI several times daily, 40.4% daily, 25.9% weekly, 11.7% monthly, and 6.9% rarely.
Response times were screened across the 18,720 participant–scenario observations. Raw pre-AI and post-AI response times showed moderate positive skewness of 0.77 and 0.78, respectively. Log transformation reduced the skewness to −0.04 and −0.01. Using an absolute standardized log-response-time threshold of 3.29, 19 pre-AI responses and 9 post-AI responses were flagged as extreme, representing less than 0.1% of observations. These responses were not automatically excluded because unusually slow responses could reflect genuine processing difficulty. Mean participant-level response time was 46.96 s before AI advice (SD = 5.13) and 35.84 s after AI advice (SD = 3.99).
Participants answered an average of 11.40 of the 24 scenarios correctly (SD = 5.66), corresponding to an overall accuracy rate of 0.475. Only one participant obtained a score of zero and seven achieved the maximum score of 24, indicating minimal floor and ceiling effects. The full 24-item accuracy score showed good internal consistency (α = 0.877). Reliability was also acceptable for the three content/performance groupings: α = 0.726 for AI error detection, α = 0.734 for evidence integration, and α = 0.778 for cognitive control and adaptive reliance. Correlations among these grouping scores ranged from 0.571 to 0.594, indicating substantial overlap alongside content-specific variation.
Table 7 summarizes the principal study variables. Final accuracy was positively related to AI-evaluation accuracy (r = 0.594), appropriate reliance (r = 0.646), ICAR-16 performance (r = 0.461), working memory (r = 0.343), cognitive reflection (r = 0.339), processing speed (r = 0.259), and attentional control (r = 0.440), all
p < .001. Associations with AI literacy (r = 0.246) and AI self-efficacy (r = 0.161) were smaller, while trust in AI was essentially unrelated to performance (r = −0.007,
p = .841). Frequency of AI use also showed only a weak relationship with final accuracy, Spearman’s ρ = 0.063,
p = .077. Final accuracy was moderately associated with the concurrently administered organizational decision-quality criterion, r = 0.408,
p < .001.
Preliminary item analysis indicated satisfactory variation across the 24 retained scenarios. Item difficulty, defined as the proportion of correct final decisions, ranged from 0.181 to 0.824 (
M = 0.475). Upper–lower 27% discrimination indices ranged from 0.351 to 0.791 (
M = 0.589), while corrected item–total correlations ranged from 0.301 to 0.566 (
M = 0.446). Mean post-AI response times ranged from 26.86 to 46.70 s. S24 was the easiest scenario (
p = 0.824), whereas S09 was the most difficult (
p = 0.181). S22 showed the strongest discrimination (
D = 0.791) and corrected item–total relationship (r = 0.566). Even the weakest retained item, S24, maintained a corrected item–total correlation above 0.30. Accordingly, all 24 scenarios were retained for the subsequent dimensionality and IRT analyses. Complete item-level descriptive statistics are provided in
Table S5.
Table 8 demonstrates the preliminary item-level performance of the 24 retained scenarios, including difficulty, discrimination, item-total relationships, response time, and retention decisions.
4.2. Dimensionality and Multidimensional IRT Results
Dimensionality was examined using the 24 binary final-decision items from the 780 retained cases. Parallel analysis of the tetrachoric correlation matrix revealed a strong first dimension. The first observed eigenvalue was 10.08, well above the corresponding 95th-percentile random eigenvalue of 1.66. The second observed eigenvalue was 1.55 and was essentially equal to, but slightly below, its random criterion of 1.55, while the third observed eigenvalue of 1.32 was below the corresponding random value of 1.51. Thus, parallel analysis primarily identified a strong general component, with only borderline evidence for a second component. In contrast, the minimum average partial procedure reached its minimum after three components, suggesting that meaningful multidimensional structure remained after accounting for the dominant general component.
Confirmatory model comparisons clarified this pattern. The one-factor model showed inadequate fit, χ2(252) = 2106.06, CFI = 0.799, TLI = 0.780, RMSEA = 0.097, and SRMR = 0.063. Fit improved substantially for the theoretically specified correlated three-factor model, χ2(249) = 1207.78, CFI = 0.896, TLI = 0.885, RMSEA = 0.070, and SRMR = 0.043. Standardized loadings were positive and ranged from 0.472 to 0.811. The three factors were strongly correlated (0.756–0.798), indicating that the content/performance dimensions were distinguishable in specification but shared considerable common variance. Because CFI and TLI remained below commonly preferred levels, the three-factor CFA is interpreted as an improvement over the one-factor model rather than as unequivocal evidence of three sharply distinct latent abilities.
A higher-order model produced essentially the same fit as the correlated three-factor model because only three first-order factors were specified. Higher-order loadings ranged from 0.863 to 0.911, indicating a substantial common cognitive-performance component. The bifactor model yielded CFI = 0.904, TLI = 0.883, RMSEA = 0.071, and SRMR = 0.043; approximately 74.7% of common loading variance was attributable to the general factor. Taken together, the CFA results therefore indicate a dominant general cognitive-performance component together with additional structure corresponding to the three theoretically specified content/performance dimensions. The correlated three-factor representation was retained for multidimensional IRT because it preserved the prespecified content organization without requiring the additional complexity of a bifactor parameterization.
Table 9 compares the fit of the competing confirmatory factor and item-response models used to evaluate the dimensional structure of the assessment.
The IRT model comparisons favored the correlated three-dimensional specification over a unidimensional 2PL model. Allowing the three content/performance dimensions to correlate improved fit, Δ−2LL = 227.53, Δdf = 3, p < .001, and reduced both AIC and BIC. Estimated latent correlations ranged from 0.734 to 0.789, again indicating substantial shared performance variance. Adding bifactor-specific slopes improved −2LL by only 6.03 despite 21 additional parameters (p = .999) and increased AIC and BIC. The correlated multidimensional 2PL model was therefore retained as a useful content-domain representation, not as evidence that the three dimensions are independent cognitive abilities.
Item discrimination parameters in the selected model ranged from 0.889 to 2.497. Most scenarios showed moderate-to-high discrimination, indicating that they differentiated meaningfully among employees at different points on the relevant ability continuum. Item difficulty parameters ranged from −1.685 to 1.479, providing broad coverage of ability. S24 was the easiest item (
b = −1.685), whereas S09 was the most difficult (
b = 1.479). Approximate infit statistics ranged from 0.769 to .0946, with no item falling outside commonly acceptable ranges. Complete multidimensional IRT item parameters are reported in
Table S7, with item characteristic and item information curves shown in
Figures S1 and S2.
Table 10 presents the multidimensional IRT item parameters and related measurement-information statistics for the 24 retained scenarios.
Conditional-information analyses showed that measurement precision differed somewhat across the three dimensions. Cognitive control and adaptive reliance provided the greatest information, reaching a maximum of 5.22 around θ = −0.18, corresponding to a conditional standard error of 0.44. Evidence integration reached maximum information of 3.64 around θ = 0.03 (SE = 0.52), while AI error detection peaked at information = 3.42 around θ = 0.74 (SE = 0.54). Precision declined toward the tails of the ability distribution, particularly below θ = −2 and above θ = 2. Model-based marginal reliability was 0.787 for AI error detection, 0.801 for evidence integration, and 0.818 for cognitive control and adaptive reliance.
Measurement precision was strongest through the low-to-moderately high portion of the performance continuum and declined toward the extremes. Broad item difficulties, generally strong discrimination, acceptable item fit, and model-based marginal reliabilities of 0.787–0.818 supported use of the correlated multidimensional 2PL model for item-level description while preserving the qualification that a strong general component underlies the three content/performance dimensions.
Figure 3 presents the test information functions and corresponding conditional standard errors across the latent ability continuum for the three closely related content/performance dimensions. The figure illustrates where measurement precision is strongest and where it declines, providing a visual complement to the multidimensional IRT results reported above.
4.3. Accuracy, Response Time, Confidence, and Scenario-Level Associations
Accuracy and response time were examined jointly to determine whether higher performance reflected more efficient processing or merely greater time investment. At the person level, cognitive accuracy was positively associated with the item-adjusted processing-efficiency index, r = 0.249, p < .001. Thus, participants who performed more accurately also tended, on average, to respond somewhat more efficiently rather than requiring systematically longer processing time. At the item level, difficulty was strongly related to time intensity: more difficult scenarios produced longer response times, B = 0.646, SE = 0.095, 95% CI [0.448, 0.844], p < .001. Within participants and scenarios, correct responses were approximately 5.8% faster than incorrect responses, B = −0.060, SE = 0.004, 95% CI [−0.067, −0.053], p < .001.
The response-time pattern provided little evidence that rapid responding represented superficial engagement. Participants in the highest quartile of the processing-efficiency distribution achieved a mean accuracy of 0.544, compared with 0.452 among the remaining participants,
p < .001. Faster responding therefore tended to accompany, rather than undermine, successful performance. These findings suggest that the assessment captured meaningful variation in cognitive efficiency: difficult items required more processing time, but higher-performing individuals were generally able to reach correct decisions with greater efficiency. Response-time screening and sensitivity analyses are reported in
Table S8.
Table 11 summarizes the joint accuracy-response-time estimates used to evaluate processing efficiency across participants and scenarios.
Figure 4 illustrates the relationship between accuracy and processing efficiency. The median divisions are used only to make the distribution easier to visualize and do not represent empirically derived cognitive profiles. Of the 780 participants, 242 showed relatively high accuracy and high efficiency, 166 showed high accuracy with lower efficiency, 148 showed lower accuracy despite relatively higher processing efficiency, and 224 showed both lower accuracy and lower processing efficiency.
Final decision accuracy was examined using logistic regression with two-way participant- and scenario-clustered covariance estimation. Because only 24 scenario clusters were available, the primary inferential summary used a conservative t23 reference distribution. The primary model did not use the within-assessment 23-scenario rest score as a person-level ability predictor; instead, it included ICAR-16, working memory, processing speed, attentional control, and AI literacy together with the prespecified scenario characteristics and interactions. Scenario characteristics were fixed properties of the fractional design, so coefficients are interpreted as adjusted associations within the observed scenario set rather than isolated causal effects.
Correct AI advice was associated with substantially greater odds of a correct final decision, B = 1.844, OR = 6.32, 95% CI [3.22, 12.43], p < .001. High AI confidence was not independently associated with final accuracy, OR = 1.85, p = .213. The AI correctness × high-confidence interaction was negative but did not reach the 0.05 threshold under the conservative small-cluster inference, B = −1.171, OR = 0.31, 95% CI [0.09, 1.09], p = .066. Descriptively, accuracy under incorrect AI advice was lower for high-confidence than low-confidence advice (0.404 vs. 0.451), but this pattern is not treated as confirmatory support for H1. Accordingly, H1 was not supported.
Time pressure was positively associated with final accuracy as a main term, OR = 3.94, p = .009, but it also showed a strong negative interaction with AI correctness, B = −3.090, OR = 0.046, p < .001. Because time-pressure status was attached to particular scenarios, this crossover cannot be interpreted as an isolated effect of time pressure. More importantly, the hypothesis-specific analysis of inappropriate acceptance among incorrect-AI scenarios did not support H2: observed inappropriate acceptance was 35.1% without time pressure and 30.8% under time pressure, and the adjusted time-pressure association was essentially null (B ≈ −0.003, OR ≈ 1.00, p = .967). H2 was therefore not supported.
Interruption was not associated with final accuracy, OR = 1.05, p = .833, and the interruption × working-memory interaction was also nonsignificant, OR = 1.07, p = .288. In addition, only four scenarios contained an interruption and only one of those scenarios belonged to the evidence-integration grouping, limiting identification of the specific dimension-contingent claim in H3. High information conflict and high task complexity also showed nonsignificant adjusted main associations (p = .359 and p = .892, respectively). H3 was therefore not supported.
Among person-level predictors in the revised primary model, ICAR-16 was positively associated with final accuracy, B = 0.311, OR = 1.36, p < .001. Working memory was also positive, B = 0.100, OR = 1.10, p = .011, as were attentional control, B = 0.342, OR = 1.41, p < .001, and AI literacy, B = 0.155, OR = 1.17, p < .001. Processing speed was positive but did not reach the 0.05 threshold under t23 inference, OR = 1.06, p = .097. The original 23-scenario rest score remained strongly associated with focal-scenario accuracy in sensitivity analyses, but because it is derived from the same assessment it is interpreted as within-assessment cross-scenario consistency rather than independent cognitive validity.
Table 12 reports the revised primary two-way cluster-robust logistic model of scenario-level final decision accuracy.
Confidence calibration varied across AI-advice subsets. Overall mean confidence was 0.615 whereas observed accuracy was 0.475, producing a positive calibration bias of 0.141. Under correct AI advice, mean accuracy was 0.526 and mean confidence was 0.632 (Brier score = 0.198). Under incorrect low-confidence advice, accuracy was 0.451 and confidence was 0.559. The largest descriptive miscalibration occurred under incorrect high-confidence advice, where confidence remained 0.628 while accuracy was 0.404, yielding a calibration bias of 0.224 and a Brier score of 0.230.
Time-pressure scenarios showed mean accuracy of 0.460 and mean confidence of 0.569, corresponding to a calibration bias of 0.109. These calibration summaries are descriptive comparisons across fixed scenario subsets and should not be interpreted as causal effects of the scenario characteristics.
Table 13 summarizes confidence-calibration indicators across the principal AI-advice scenario subsets.
Figure 5 displays the confidence-calibration curves. Incorrect high-confidence advice showed the clearest descriptive tendency toward overconfidence. Because the corresponding correctness × confidence interaction in the revised primary accuracy model was not significant at α = 0.05 (
p = .066), the calibration pattern is treated as descriptive evidence of misalignment between confidence and correctness rather than evidence that high AI confidence independently changed decision accuracy.
The scenario-level hypothesis tests therefore produced mixed evidence. H1 was not supported because the correctness × confidence interaction did not reach the 0.05 threshold. H2 was not supported because time pressure did not increase inappropriate acceptance of incorrect AI advice. H3 was not supported because interruption and the interruption × working-memory term were nonsignificant and the evidence-integration-specific claim was weakly identified by the fractional design. H5 was also not supported: among incorrect-AI scenarios, the high-confidence × overall-performance interaction was negative and nonsignificant, B = −0.269, OR = 0.764, p = .111, opposite to the predicted protective moderation. These results do not alter the separate psychometric and validity questions addressed below.
4.4. Construct, Concurrent Criterion-Related, and Incremental Validity
Convergent evidence was evaluated using both overall assessment performance and hypothesis-aligned content/performance scores. Final assessment accuracy correlated with ICAR-16 at r = 0.461, 95% CI [0.403, 0.514], working memory at r = 0.343, 95% CI [0.279, 0.403], processing speed at r = 0.259, 95% CI [0.193, 0.324], and attentional control at r = 0.440, 95% CI [0.382, 0.495], all p < .001. More directly aligned with H4, ICAR-16 correlated with the AI-error-detection score at r = 0.417, 95% CI [0.357, 0.473], p < .001, while working memory correlated with the evidence-integration score at r = 0.286, 95% CI [0.221, 0.350], p < .001. Cross-domain correlations were also positive, indicating that these mappings are theoretical emphases rather than exclusive one-to-one relationships. H4 was supported.
Relationships with AI-related self-reports were weaker. AI literacy was positively related to final performance, r = 0.246, 95% CI [0.179, 0.311], p < .001, and AI self-efficacy showed a smaller association, r = 0.161, 95% CI [0.092, 0.229], p < .001. Trust in AI was essentially unrelated to performance, r = −0.007, 95% CI [−0.077, 0.063], p = .841. Frequency of AI use was also weak and nonsignificant, Spearman’s ρ = 0.063, 95% bootstrap CI [−0.008, 0.134], p = .077. This pattern supports the distinction between demonstrated performance and AI-related self-perceptions or exposure.
Concurrent criterion-related validity was evaluated against the separately scored six-item organizational decision-quality task administered in the same session. Focal assessment performance correlated moderately with criterion performance, r = 0.408, 95% CI [0.348, 0.465],
p < .001. Because the criterion was concurrent rather than prospective, this result is interpreted as concurrent criterion-related evidence and not as predictive validity.
Table 14 summarizes convergent, discriminant, and concurrent criterion-related validity evidence.
Incremental concurrent validity was examined using a corrected hierarchical regression specification with organizational decision quality as the dependent variable. Model 1 contained demographic controls and explained R2 = 0.008 (adjusted R2 = −0.002). Adding work experience in Model 2 produced essentially no change, ΔR2 < 0.001, ΔF(1, 770) = 0.06, p = .808. Model 3 added AI-use frequency and prior generative-AI experience, increasing explained variance by 0.007, ΔF(2, 768) = 2.70, p = .068.
Model 4 added the conventional cognitive-performance block—ICAR-16, cognitive reflection, working memory, and processing speed—and increased explained variance by ΔR2 = 0.075, ΔF(4, 764) = 15.68, p < .001, yielding R2 = 0.090. Model 5 then added the AI-related self-report block comprising AI literacy, trust in AI, and AI self-efficacy. This block contributed little additional variance, ΔR2 = 0.003, ΔF(3, 761) = 0.92, p = .430.
Adding the focal 24-scenario assessment in Model 6 produced the largest incremental gain, Δ
R2 = 0.103, ΔF(1, 760) = 97.04,
p < .001. The final model explained
R2 = 0.196 of concurrent organizational decision-quality variance (adjusted
R2 = 0.176). Each additional correctly solved focal scenario was associated with a 0.129-point increase in criterion performance,
B = 0.129,
SE = 0.013, 95% CI [0.103, 0.155],
p < .001; the standardized coefficient was β = 0.379. A sensitivity model that entered attentional control separately before the self-report block still showed a substantial focal-assessment increment, Δ
R2 = 0.087,
p < .001.
Table 15 summarizes the corrected hierarchical regression models for incremental concurrent validity.
Internal 10-fold cross-validation showed the same general pattern. Models containing only demographic and experience variables had essentially no held-out explanatory value. The conventional cognitive block increased cross-validated R2 to 0.059, the AI-related self-report block did not improve it (CV R2 = 0.055), and adding the focal assessment increased CV R2 to 0.160 while reducing RMSE to 1.764. These values support internal statistical stability but do not constitute prospective prediction in a new population.
Overall, the validity analyses were stronger than the scenario-specific hypothesis tests. H4 was supported by dimension-aligned cognitive associations, and H6 was supported by a focal-assessment increment of ΔR2 = 0.103 in the concurrent criterion model. In contrast, H1, H2, H3, and H5 were not supported under the revised hypothesis-aligned tests. The final hypothesis pattern was therefore H1: not supported; H2: not supported; H3: not supported; H4: supported; H5: not supported; H6: supported.
4.5. Measurement Invariance and Differential Item Functioning
Measurement invariance was examined across gender, age, education, managerial status, AI-use frequency, and occupational grouping using the correlated three-factor representation as the reference measurement model. For binary responses, configural and loading invariance were evaluated before threshold invariance. Residual invariance was not separately estimated because residual variances are not freely identified in the same manner under the categorical specification. These analyses are interpreted as preliminary group-comparability evidence, not as proof of fairness for consequential assessment use.
Configural and loading invariance results were broadly favorable across the six comparisons. Configural CFI values ranged from 0.975 to 0.992 and RMSEA values from 0.009 to 0.015. Equality constraints on factor loadings produced small changes in fit, with absolute ΔCFI values from 0.000 to 0.005 and ΔRMSEA values from 0.000 to 0.002.
Threshold invariance was supported for gender, age, education, managerial status, and AI-use frequency. The occupational comparison was the exception: constraining thresholds reduced CFI from 0.974 to 0.960 (ΔCFI = −0.014) and increased RMSEA from 0.015 to 0.019 (ΔRMSEA = 0.003), and the joint threshold-difference test was significant, χ2(24) = 75.66, p < .001. Full occupational threshold invariance was therefore not supported. Cross-occupation score comparisons should consequently be interpreted cautiously until this finding is replicated.
Table 16 presents the measurement-invariance results across the examined groups.
Differential item functioning was evaluated by logistic regression using performance on the remaining 23 items as the matching score. Uniform DIF was represented by the group term and non-uniform DIF by the group × matching-score interaction. Benjamini–Hochberg false-discovery-rate adjustment was applied separately to the uniform- and non-uniform-DIF families within each comparison. Practical importance was evaluated using ΔMcFadden R2 and the adjusted probability difference; values below approximately 0.02 were treated as small.
No item survived FDR correction for gender, age, education, managerial status, or occupational group. The largest observed item-level effects in these comparisons were small. Importantly, the absence of an individual occupational item surviving FDR correction does not negate the aggregate threshold-invariance result reported above; the two procedures address different aspects of group comparability.
One item, S01, showed significant uniform DIF across AI-use-frequency groups after FDR correction (q = 0.021, ΔMcFadden R2 = 0.011). At comparable matching-score levels, participants using AI weekly or less had an adjusted probability of correct performance approximately 0.101 lower than participants using AI daily or more frequently. The magnitude was small, so S01 was retained but should be monitored in future samples.
Table 17 summarizes the differential item functioning results across the examined groups.
Planning/analytical employees had a higher mean total score than operations-oriented employees (11.96 vs. 10.79, p = .004). This mean difference does not explain away the unfavorable occupational threshold invariance result. Although no individual occupational item survived FDR correction, the aggregate threshold test remained unfavorable; raw cross-occupation comparisons should therefore remain provisional. At the item level, the detected DIF effects were generally small and no item required deletion or group-specific scoring in this sample. However, item-level DIF and measurement-invariance results should be considered jointly. The occupational threshold finding prevents a broad claim of equivalent functioning across all examined groups. Overall, the analyses provide preliminary evidence of broadly similar measurement functioning across gender, age, education, managerial status, and AI-use-frequency groups, with an important occupational qualification. Full threshold invariance was not supported across occupational groups, and S01 showed a small AI-use-related DIF effect. These findings support continued psychometric evaluation and replication rather than claims that group comparability or fairness has been definitively established.
5. Discussion
5.1. Interpretation of the Main Findings
The purpose of this study was to examine whether cognitive performance under AI advice can be assessed as a meaningful applied performance domain in organizational decision-making. The results provide stronger support for the study’s psychometric and validity objectives than for several of its scenario-specific hypotheses. H4 and H6 were supported, whereas H1, H2, H3, and H5 were not supported under the revised hypothesis-aligned tests. This pattern does not imply that the assessment was unsuccessful. Reliability, item functioning, dimensionality, cognitive convergence, concurrent criterion-related validity, and incremental validity address different questions from the scenario-level hypotheses. The findings therefore support an initial measurement framework while indicating that several proposed contextual mechanisms require stronger experimental designs before they can be treated as established features of performance under AI advice.
The dimensionality results support a qualified interpretation. The one-factor CFA fit poorly, whereas the correlated three-factor model represented the data substantially better; however, its CFI (0.896) and TLI (0.885) remained below commonly preferred levels, and the three factors correlated strongly. Higher-order and bifactor analyses likewise indicated a dominant general cognitive-performance component. The most defensible interpretation is therefore not that AI error detection, evidence integration, and cognitive control and adaptive reliance are independent abilities, but that they are closely related content/performance dimensions embedded within a strong general performance structure. The multidimensional IRT results add useful content-level information, yet they should not be treated as proof of sharply distinct latent traits.
The convergent-validity pattern is consistent with the CHC-informed rationale. Overall performance was positively associated with ICAR-16, working memory, processing speed, and attentional control, while the dimension-aligned tests showed that ICAR-16 was associated with AI error detection (r = 0.417) and working memory with evidence integration (r = 0.286), both p < .001. These results support H4, while the positive cross-domain correlations also show that the theoretical mappings are not exclusive. Fluid reasoning and working memory appear to contribute across multiple task demands rather than map one-to-one onto a single content grouping. Attentional control should likewise be interpreted as a companion individual-difference resource rather than as evidence for a newly established CHC dimension of the assessment.
The process indicators provided complementary information. Higher-performing participants tended to respond somewhat more efficiently, and more difficult scenarios required more time, suggesting that response speed is informative only when interpreted jointly with accuracy. Confidence calibration showed a different pattern: confidence exceeded observed accuracy overall, and the largest descriptive miscalibration occurred when incorrect AI advice was expressed with high confidence. This pattern is consistent with concerns about reliance calibration in human–AI decision-making (
Dogru & Krämer, 2025;
Lee & See, 2004;
Meincke et al., 2026;
Pearson et al., 2026), but it should not be overstated. The correctness × confidence interaction did not reach the 0.05 threshold under conservative small-cluster inference (
p = .066), so H1 was not supported and the calibration pattern is best treated as descriptive evidence of confidence–accuracy misalignment.
The remaining scenario-specific hypotheses also require restraint. Time pressure did not increase inappropriate acceptance of incorrect AI advice as predicted in H2, and interruption was not significantly associated with accuracy or with the interruption × working-memory interaction, so H3 was not supported. H5 was also not supported: higher overall performance did not significantly reduce susceptibility to confidently expressed inaccurate AI recommendations in the hypothesis-specific moderation analysis. The strong correctness × time-pressure crossover in the accuracy model is therefore interpreted as an association within the observed scenario set rather than confirmation of H2. Because confidence, time pressure, interruptions, information conflict, and complexity were fixed characteristics of particular scenarios in the fractional design rather than independently randomized within otherwise identical content, these coefficients cannot isolate causal effects from scenario content.
The clearest evidence for the assessment’s added value comes from the validity analyses. Performance correlated moderately with the concurrently administered organizational decision-quality criterion (r = 0.408) and explained an additional 10.3% of criterion variance after demographic and work characteristics, AI experience, conventional cognitive-performance measures, and AI-related self-reports were entered (ΔR2 = 0.103, p < .001). H6 was therefore supported. The self-report block itself added little unique variance (ΔR2 = 0.003, p = .430). This pattern suggests that the focal assessment captures applied performance information not reducible to conventional cognition or AI-related self-perceptions. At the same time, the criterion was measured in the same session, so the evidence is concurrent rather than predictive. Likewise, the 23-scenario rest score is best viewed as a within-assessment cross-scenario consistency indicator and is retained only as a sensitivity analysis, not as independent validity evidence.
5.2. Theoretical and Measurement Contributions
The principal contribution is psychometric and integrative. Behavioral research on automation, advice taking, and human–AI reliance already examines whether people accept, reject, or revise machine advice; the present study does not claim to introduce behavioral measurement itself. Instead, it organizes repeated workplace-relevant AI-advice decisions into an individual-differences assessment framework that connects scenario performance with cognitive measures, process indicators, criterion-related evidence, item-response modelling, and group-comparability analyses. This framing sharpens the contribution relative to prior reliance paradigms by treating repeated behavioral performance as an assessment problem rather than only as a situational outcome.
The CHC contribution should also be understood as application rather than theoretical extension. CHC theory offers a well-established vocabulary for reasoning, working memory, and processing efficiency, allowing the cognitive demands of AI-supported decisions to be described without proposing a new broad intelligence. The results are compatible with that position: performance shared meaningful variance with established cognitive measures, yet the focal assessment also captured additional applied variance in a decision environment that requires evaluating machine-generated advice. Accordingly, cognitive performance under AI advice is best described as a CHC-informed applied performance domain rather than a new CHC ability or a distinct form of “AI intelligence.”
A second contribution concerns the separation of demonstrated performance from AI-related self-perceptions. AI literacy and AI self-efficacy were positively but more weakly associated with performance, trust in AI was essentially unrelated to final accuracy, and AI-use frequency showed only a weak nonsignificant relationship. These constructs remain important for understanding adoption, confidence, and engagement, but they do not substitute for observing whether a person detects an error, integrates evidence, or relies appropriately on a recommendation. The assessment therefore complements, rather than replaces, established AI-literacy, trust, and self-efficacy measures (
Carolus et al., 2023;
Jian et al., 2000;
Koch et al., 2024;
Wang & Chuang, 2024).
A third contribution is the combination of outcome and process information. Final accuracy provides the core performance score, while AI-evaluation accuracy, reliance behavior, confidence, decision revision, and response time help characterize how decisions were reached. Multidimensional IRT describes where scenarios provide information across the performance continuum, and internal cross-validation assesses statistical stability of the incremental models. The two-way cluster-robust scenario analysis additionally recognizes dependence within participants and scenarios, with conservative t23 inference used because only 24 scenario clusters were available. These methods provide complementary evidence, although none by itself establishes construct validity or causal mechanisms.
The group-comparability analyses provide an important qualification to the measurement contribution. Loading and threshold invariance were broadly favorable for gender, age, education, managerial status, and AI-use frequency, and item-level DIF effects were generally small. However, full occupational threshold invariance was not supported (ΔCFI = −0.014), even though no single occupational item survived FDR correction. The aggregate threshold result should not be dismissed because of the item-level DIF pattern. Cross-occupation score comparisons therefore require caution until the finding is replicated, and the present analyses should be described as preliminary group-comparability evidence rather than proof of fairness.
More broadly, the study provides initial validation evidence from one sample and one administration. The psychometric results are sufficiently coherent to justify continued development and independent replication, but they do not establish a finalized high-stakes instrument. The combination of a dominant general component, useful content-specific structure, convergent cognitive relationships, concurrent criterion-related evidence, and incremental validity is promising precisely because the conclusions can be stated without requiring every scenario-specific hypothesis to be confirmed.
5.3. Organizational and Practical Implications
The most defensible near-term applications are research, formative assessment, employee development, and training evaluation rather than personnel selection. Organizations often evaluate AI training through completion, satisfaction, or self-reported confidence. A performance-based assessment could complement those indicators by examining whether employees become better at evaluating inaccurate recommendations, integrating task evidence, regulating reliance, and calibrating confidence. Because the present study provides initial rather than definitive validation, such uses should focus on learning and development rather than consequential classification of employees.
The results also suggest that AI-literacy development may benefit from separating knowledge from performance. Employees can understand general AI limitations yet still struggle with verification in a concrete decision, while frequent AI use does not necessarily imply better judgment. Future training studies could use the assessment to test whether targeted verification routines, evidence-integration exercises, or calibration feedback improve performance. Such intervention effects were not tested here, so these possibilities should be treated as directions for applied research rather than established prescriptions.
For decision-support design, the calibration results highlight the importance of how confidence and uncertainty are communicated. Incorrect high-confidence advice was associated descriptively with the greatest confidence–accuracy mismatch, but the corresponding interaction with final accuracy was not statistically significant at α = 0.05. The practical implication is therefore not that confidence displays have been shown to cause overreliance, but that interface designs should be evaluated for whether they support verification and calibrated reliance. Explanations that expose evidence, assumptions, uncertainty, and omitted information may be more useful than confidence cues presented without context.
At the workflow level, organizations should evaluate the interaction between system characteristics, human capabilities, and decision context rather than focus only on AI accuracy (
Do Khac & Leyer, 2026;
Gonzalez & Heidari, 2025;
Li & Tian, 2026). The current fractional design does not establish that time pressure or interruptions causally reduce or improve performance, so workflow recommendations about these factors should await stronger within-scenario experiments. Nevertheless, the framework offers a way to study where oversight becomes cognitively demanding and where employees may need clearer evidence, more verification support, or better-calibrated system communication.
The assessment should not currently be used as a stand-alone basis for hiring, promotion, compensation, discipline, termination, or other high-stakes personnel decisions. Such applications would require substantially stronger evidence, including independent replication, prospective criterion validity, test–retest reliability, alternate forms, accessibility evaluation, adverse-impact analysis, cross-cultural and occupational comparability, and validation across different AI systems. The occupational threshold non-invariance observed here is an additional reason to avoid consequential cross-group interpretations at this stage.
5.4. Limitations and Future Research
Several limitations define the boundaries of the findings. First, the 24-scenario design was fractional: scenario characteristics were fixed to particular scenarios rather than independently crossed within otherwise identical content. Randomization applied to presentation order, not to independent assignment of confidence, time pressure, interruption, information conflict, or complexity. Scenario-level coefficients therefore represent adjusted associations within this scenario set and cannot establish causal effects. The small scenario-cluster dimension (24 scenarios) further limits precision, despite the conservative t23 inference and leave-one-scenario-out sensitivity analyses. Future work should use larger scenario banks and fully or partially crossed experimental manipulations when the goal is to identify contextual mechanisms.
Second, the study provides initial validation evidence from a single Turkish sample recruited through purposive nonprobability procedures and assessed once. This design limits claims about population generalizability, temporal stability, and cross-cultural equivalence. Test–retest studies, alternate forms, independent calibration samples, probability-based or more broadly recruited samples, and replications in other languages and countries are needed before the assessment can be treated as a stable individual-differences instrument.
Third, online administration introduces variability in device, screen size, connectivity, distraction, and testing environment. Response-time interpretation is especially sensitive to these factors. Timing information was derived from the study’s supplementary timestamping procedure rather than from a standardized laboratory platform, and very slow responses were not automatically discarded because they could reflect genuine processing difficulty. Future studies should archive detailed timing metadata and compare supervised with unsupervised administration to separate cognitive processing time from technical delay more precisely.
Fourth, the companion cognitive measures were intentionally brief to limit participant burden. The retained dataset contains aggregate working-memory and processing-speed scores rather than complete trial-level records, and attentional control was represented by eight study-specific binary indicators rather than a standardized CHC test. These measures were useful as correlates and sensitivity predictors, but future replications should employ fully documented standardized tasks, preserve trial-level data, and evaluate their reliability and validity independently. This is particularly important for reproducing the exact cognitive-measurement component of the study.
Fifth, the criterion evidence is concurrent and limited in scope. The six-item organizational decision-quality task was administered in the same session as the focal assessment, so the observed r = 0.408 and incremental ΔR2 = 0.103 do not demonstrate prediction of future workplace performance. Future studies should examine supervisor ratings, objective errors, forecasting performance, quality-control outcomes, and real AI-assisted decisions collected prospectively. Similarly, internal 10-fold cross-validation evaluates statistical stability within the present dataset and is not a substitute for external validation in a new sample.
Sixth, group comparability requires further study. Full occupational threshold invariance was not supported, and S01 showed a small AI-use-related DIF effect. Although the detected item-level effects were generally small, these results are sufficient to caution against strong claims of fairness or interchangeable score meaning across occupations. Larger subgroup samples, replication across industries, accessibility studies, and analyses of practical adverse impact are needed before cross-group comparisons can support consequential decisions.
Seventh, the within-assessment rest score used in earlier scenario-level analyses is not an independent measure of cognitive ability because it is derived from performance on the same 24-scenario assessment. Excluding the focal item prevents direct arithmetic overlap but does not create an external validity criterion. The revised primary model therefore relies on independently measured person-level predictors, while the rest score is retained only as an internal cross-scenario consistency sensitivity analysis. Future work should prioritize independent cognitive and behavioral predictors when testing person-level moderation.
Finally, both AI systems and workplace uses of AI are evolving rapidly. Scenario realism, expectations of system reliability, and the meaning of appropriate reliance may change as interfaces and models develop. Periodic content review will therefore be necessary. Future research should also test different AI systems, industry-specific scenarios, longitudinal training responsiveness, and real workplace deployment, while examining whether the same general performance structure emerges across contexts.
Further psychometric development could expand the item bank and evaluate computerized adaptive testing, particularly because item information varied across the performance continuum. Such work should follow, rather than precede, independent replication of the factor structure, stronger documentation of companion cognitive tasks, prospective criterion validation, and additional group-comparability evidence.
Overall, the study supports a restrained interpretation. Cognitive performance under AI advice is best understood as a CHC-informed applied performance domain with a dominant general component and meaningful content-specific structure. The central psychometric and validity objectives received stronger support than the proposed scenario-specific mechanisms. That distinction is informative: the assessment shows promise as a research and developmental instrument for observing how people evaluate and respond to AI advice, while the contextual moderators, temporal stability, external validity, and cross-group comparability require further independent study.
6. Conclusions
This study developed and initially validated a CHC-informed, performance-based assessment of cognitive performance under AI advice using repeated organizational decision scenarios. The findings do not support interpreting performance under AI advice as a new form of intelligence. Instead, they indicate a dominant general cognitive-performance component together with additional structure corresponding to three closely related content/performance dimensions: AI error detection, evidence integration, and cognitive control and adaptive reliance. This pattern is consistent with the view that established cognitive resources are expressed within a distinctive decision environment in which employees must evaluate the quality and usefulness of machine-generated advice.
The strongest evidence concerned the assessment’s psychometric and validity objectives. Performance showed meaningful relationships with established cognitive measures, and the dimension-aligned analyses supported H4: ICAR-16 was associated with AI error detection and working memory with evidence integration. Focal assessment performance was also moderately related to the concurrently administered organizational decision-quality task (r = 0.408) and explained additional concurrent criterion variance beyond demographic and work characteristics, AI experience, conventional cognitive-performance measures, and AI-related self-reports (ΔR2 = 0.103), supporting H6. These findings suggest that the assessment captures applied performance that is related to, but not reducible to, conventional cognitive measures or self-reported AI competence.
The scenario-specific hypotheses produced more limited evidence. H1, H2, H3, and H5 were not supported under the revised hypothesis-aligned analyses. In particular, the confidence interaction did not reach the conventional significance threshold, time pressure did not increase inappropriate acceptance of incorrect AI advice, the interruption hypothesis was not supported, and higher overall assessment performance did not provide the predicted protection against confidently expressed inaccurate advice. Because scenario characteristics were fixed properties of a 24-scenario fractional design rather than independently crossed manipulations, these results should be treated as adjusted associations within the observed scenario set. Stronger fully or partially crossed designs are needed before these contextual mechanisms can be established.
Accordingly, the contribution of the study is primarily psychometric and integrative rather than the introduction of behavioral AI-reliance measurement itself. The assessment combines final accuracy with AI-evaluation accuracy, reliance behavior, confidence, response time, and decision revision, while item-response and process analyses provide information about measurement precision and how performance unfolds. The present evidence should nevertheless be considered initial. It comes from one purposively recruited Turkish sample and one administration; the criterion evidence is concurrent rather than prospective; the scenario design contains only 24 fixed scenario clusters; and full occupational threshold invariance was not supported. The assessment is therefore best suited at present to research, formative assessment, and developmental or training contexts rather than hiring, promotion, or other consequential personnel decisions. Future work should establish test–retest reliability, prospective criterion validity, independent replication, cross-cultural comparability, stronger documentation of companion cognitive tasks, accessibility, and group comparability before higher-stakes applications are considered.