Next Article in Journal
Are Prisoners Getting the Right Education? Exploring Teaching and Learning of Functional Skills Mathematics and English in England and Wales
Previous Article in Journal
AR-GenSTEAM: Generation of STEAM-Based Serious Games with Augmented Reality
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

AI Literacy and Self-Perceived Cognitive Learning Outcomes Among University Students in AI-Integrated Courses: Associations with Instructor Feedback and AI Use Indicators

by
Yu Eun Lee
and
Jin Sook Kan
*
Center for Educational Innovation, Hallym University, Chuncheon 24252, Republic of Korea
*
Author to whom correspondence should be addressed.
Educ. Sci. 2026, 16(9), 1384; https://doi.org/10.3390/educsci16091384
Submission received: 12 August 2026 / Revised: 21 August 2026 / Accepted: 25 August 2026 / Published: 27 August 2026

Abstract

Artificial intelligence (AI) is rapidly being integrated into university curricula, yet quantitative indicators of AI use reveal little about how learners use AI as a learning resource or what educational outcomes follow. This cross-sectional survey study of 212 university students enrolled in AI-integrated courses examined the associations of AI literacy, instructor feedback, and two single-item AI use indicators—the proportion of in-class AI use and total weekly AI use time—with self-perceived cognitive learning outcomes, measured across the six cognitive processes of the revised Bloom’s taxonomy. Confirmatory factor analyses supported multidimensional and higher-order structures, but the cognitive domains overlapped substantially (interfactor correlations up to 0.943; HTMT up to 0.946), so domain-level distinctions should be interpreted with caution. A regression model with the four predictors explained 46.3% of the variance in overall self-perceived cognitive learning outcomes (R2 = 0.463, adjusted R2 = 0.452). When all predictors were considered simultaneously, only AI literacy showed a significant positive association (B = 0.615, β = 0.593, 95% CI [0.475, 0.754], p < 0.001); the data did not provide evidence for independent associations of instructor feedback or the two AI use indicators, whose weaker associations may partly reflect their single-item measurement. AI literacy remained significantly associated with all six cognitive domains after Benjamini–Hochberg correction. These findings suggest—within the limits of a cross-sectional, self-report design—that the quantity of AI use and learners’ competency to understand, evaluate, and self-regulate AI use are empirically distinct indicators that universities should measure separately.

1. Introduction

1.1. Background and Rationale

Artificial intelligence (AI), and generative AI (GenAI) in particular, is rapidly being integrated into university students’ learning processes, including information seeking, writing, problem solving, idea generation, feedback use, and self-regulated learning. Recent reviews indicate that AI serves diverse functions across learning, teaching, assessment, and educational administration, and that university students’ GenAI use is not a single behavior but a complex set of learning practices encompassing translation, content elaboration, information seeking, evaluation, conversation, content creation, and support for self-regulated learning (An et al., 2026; Chiu et al., 2023). This expansion of GenAI use is shifting the debate in higher education from whether to adopt and use AI toward which learning processes AI facilitates or transforms—and, in some cases, which cognitive performances it replaces.
How the outcomes of AI-integrated education should be evaluated remains an important challenge. Quantitative indicators of AI use, such as whether AI is used, the proportion of in-class AI use, and total weekly AI use time, are easy to measure but poorly capture the learning processes in which students critically examine AI outputs, check sources, detect and correct errors, and reconstruct AI-generated content by integrating it with their own thinking. Recent work suggests that AI use should not be framed as inherently beneficial or harmful; what matters is how learners use AI. Passively delegating thinking to AI may reduce opportunities for independent reasoning and problem solving, whereas deliberate use involving structured questioning, verification of results, revision, and reflection may foster higher-order thinking and productive human–AI collaboration (Abbosh et al., 2025; Girma, 2025; Zhai et al., 2024). Indeed, recent systematic analyses of AI governance in higher education identify over-reliance on automated systems and the erosion of students’ critical thinking capacity among the most frequently reported ethical risks, and argue that institutional policy frameworks remain underdeveloped relative to the pace of adoption (Kaşarcı et al., 2025). Understanding the outcomes of AI-integrated education therefore requires attention not only to the quantity of AI use but also to learners’ capacity to judge AI outputs critically and to integrate and regulate them appropriately within their own learning.
AI literacy is a key construct for explaining such differences in how AI is used. It has been conceptualized as a multidimensional competency that involves understanding the basic principles and characteristics of AI, using AI purposefully, critically evaluating AI-generated outputs, and using AI ethically and responsibly (Long & Magerko, 2020; Ng et al., 2021), and recent measurement research has extended it to psychological and self-management dimensions such as self-efficacy, self-regulation, and meta-competencies (Carolus et al., 2023; Ng et al., 2024). Advancing this line of research requires clearly defined constructs and psychometrically validated measures (Lintner, 2024); the intervention and process-oriented evidence linking AI literacy to learning is reviewed in Section 1.2.2.
The instructor’s role remains important even in AI-based learning environments. Although AI can provide learners with immediate explanations and suggestions, instructor feedback serves the instructional function of helping learners compare their current performance with learning goals and evaluation criteria, examine the validity and evidential basis of AI outputs, and revise and regulate their learning strategies accordingly. Research on formative assessment has shown that effective feedback goes beyond confirming right and wrong answers or conveying results and can promote improvement at the levels of task performance, learning processes, and self-regulation (Black & Wiliam, 1998; Hattie & Timperley, 2007; Sadler, 1989). Understanding learning outcomes in AI-integrated courses therefore requires considering not only learners’ AI literacy but also instructional conditions such as instructor feedback that support and regulate the learning process.
This study focuses on self-perceived cognitive learning outcomes—learners’ perceptions of their own level of cognitive performance—rather than intelligence measured by objective tests or actual cognitive performance itself. Specifically, students evaluated their own levels of remembering, understanding, applying, analyzing, evaluating, and creating; such self-assessments cannot be equated with cognitive achievement measured through performance tests. Self-assessment may also involve measurement error, as students may judge their performance to be higher or lower than it actually is. Nevertheless, perceptions of one’s cognitive performance provide educationally meaningful information, because they are involved in how learners monitor their current understanding and performance, form expectations for future learning, and select and adjust learning strategies. Meta-analyses have found significant associations between learners’ expectations and self-evaluations and their actual academic performance, while also documenting persistent discrepancies between self-evaluations and actual performance, pointing to the imperfection of self-knowledge (Pinquart & Ebeling, 2020; Z. Yan et al., 2022, 2023; Zell & Krizan, 2014). Recent performance-based GenAI literacy research likewise emphasizes that self-perceived competence and competence verified through actual performance should be measured and interpreted as distinct indicators rather than treated as interchangeable (Jin et al., 2025). Because such perceptions are self-reported, their associations with other self-reported competencies may also be inflated by common method variance or by stable dispositions such as academic self-efficacy or a generally favorable self-evaluation tendency; this rival explanation is addressed explicitly in the Discussion (Section 4).
AI research in higher education has accumulated rapidly around topics such as the conceptualization and measurement of AI literacy, instructional interventions, the educational potential and risks of GenAI use, cognitive dependence, and human agency. However, relatively few studies have examined, within a single analytic model, learners’ AI literacy, instructor feedback as an instructional condition, and two AI use indicators—the proportion of in-class AI use and total weekly AI use time—among university students from diverse majors enrolled in AI-integrated courses within the regular curriculum. Distinguishing quantitative indicators of AI use from AI literacy as a learner competency clarifies how these two types of indicators are differentially related to learning outcomes, and it provides a practical basis for delineating what quantity-centered indicators can and cannot capture in universities’ AI education performance management.

1.2. Theoretical Background and Prior Research

1.2.1. Self-Perceived Cognitive Learning Outcomes and the Revised Bloom’s Taxonomy

Learning outcomes can be assessed through objective performance data and through self-perception data in which learners judge their own understanding and capability. From a self-regulated learning perspective, such self-judgments are part of the self-monitoring process through which learners compare their performance with their goals and adjust their strategies and effort (Zimmerman, 2008). Meta-analyses show that expected achievement and self-evaluations are positively related to actual performance, but imperfectly and with individual variation (Pinquart & Ebeling, 2020; Z. Yan et al., 2022, 2023; Zell & Krizan, 2014); self-perceived outcomes are therefore complementary information about how learners perceive their learning, not substitutes for objective performance.
The self-perceived cognitive learning outcomes in this study were based on the cognitive process dimension of the revised Bloom’s taxonomy. Anderson and Krathwohl (2001) and Krathwohl (2002) classified cognitive processes into six categories: remember, understand, apply, analyze, evaluate, and create. These processes constitute an educational framework that systematically describes the diverse cognitive performances involved in learning, from retrieving learned information and constructing meaning to applying knowledge to new situations, analyzing relations among elements, making criterion-based judgments, and generating new products (Anderson & Krathwohl, 2001; Krathwohl, 2002). In this study, the taxonomy was used not as a criterion for directly measuring objective cognitive ability, but as a conceptual framework for measuring the extent to which students perceived themselves as able to carry out each cognitive performance in their AI-integrated courses. The self-perceived cognitive learning outcomes measured here are therefore a construct distinct from cognitive functioning measured by standardized intelligence tests or objective performance tests.
This distinction is particularly important in AI-supported learning environments. Because AI can externalize parts of cognitive tasks—information retrieval, summarizing, explaining, and drafting—the quality of products created with AI support may not correspond to the cognitive processing that occurred within the learner. Self-perceived cognitive learning outcomes do not directly identify such discrepancies, but they provide separate information about how learners perceive their cognitive engagement in AI-supported learning; accordingly, they are interpreted here strictly as subjective learning outcomes rather than as proxies for objective cognitive performance or actual achievement.

1.2.2. AI Literacy as a Learner Competency

AI literacy is being conceptualized as a multidimensional competency that goes beyond operating AI tools to understanding, using, critically evaluating, and responsibly employing AI. Long and Magerko (2020) defined AI literacy as a set of competencies that enables individuals to critically evaluate AI technologies and to communicate and collaborate effectively with AI. Ng et al. (2021) systematized AI literacy into four dimensions—know and understand, use and apply, evaluate and create, and understand ethical issues—presenting it as a concept encompassing not only knowledge about AI but also practical use, critical judgment, and ethical consideration. Subsequent measurement research has expanded AI literacy into a multidimensional construct that includes affective, behavioral, and ethical aspects as well as competencies related to self-regulation and self-management (Carolus et al., 2023; Ng et al., 2024). Because AI literacy scales differ in construct coverage, subdimensions, and level of psychometric validation, it is important to delimit the construct in line with the purpose of the research and to use instruments appropriate to that definition (Lintner, 2024).
Educational intervention studies with university students show that AI literacy is not a fixed individual trait but a competency that can be developed through instruction and learning experience. Kong et al. (2021, 2022) reported that AI literacy courses for students from diverse majors improved conceptual understanding of AI and self-perceived AI competence. Recent GenAI-based higher education research has also found significant associations between AI literacy and self-regulated learning and writing performance (Shi et al., 2025). A longitudinal qualitative study further reported that learners progressively developed prompting strategies, metacognitive monitoring, critical evaluation, ethical judgment, and integrated use of multiple AI tools as they formed GenAI literacy (W. Yan et al., 2025). These studies provide grounds for conceptualizing AI literacy not as an indicator of the level or amount of AI use, but as a learner competency for critically evaluating AI and using and regulating it in line with one’s learning goals.
This conceptual distinction also matters for measurement. Performance-based instruments such as the GLAT can predict actual GenAI-supported task performance better than self-reported proficiency scales in some contexts (Jin et al., 2025), so self-reported AI literacy should not be read as a direct indicator of actual performance capability. The AI literacy measured in this study refers to learners’ self-perceived AI-related knowledge and usage competence, critical evaluation of AI outputs, and ethical and self-regulated use; multimethod validation combining self-reported and performance-based measures remains an important task for future research.

1.2.3. Instructor Feedback as Scaffolding in AI-Integrated Learning

Instructor feedback is a core mechanism of formative assessment that helps learners identify the gap between their current performance and learning goals or evaluation criteria and adjust their subsequent strategies and behavior accordingly. Sadler (1989) emphasized that effective formative assessment requires learners to understand the expected standards, judge their current performance against them, and act to close the gap. Hattie and Timperley (2007) proposed that feedback can operate at the levels of the task, the processing of the task, and self-regulation. These feedback functions may be even more important in AI-integrated learning environments: the fluency and apparent completeness of AI-generated outputs can make them look like valid answers, but surface polish does not guarantee the accuracy or sufficiency of the evidence or alignment with learning goals. Instructor feedback can thus function as an instructional mechanism that supports learners in examining the evidence and validity of AI outputs rather than accepting them as given, and in monitoring and regulating their performance against learning goals.
Instructors can promote critical use of AI by asking whether an AI response meets the task demands and learning goals, whether its evidence is trustworthy, what needs to be revised, and what the learner can explain independently without AI assistance—redirecting learners’ attention from accepting AI outputs toward examining their validity and monitoring their own reasoning. Drawing on formative feedback theory and research on feedback levels, this study measured instructor feedback as clear communication of learning goals and performance standards, timely provision of feedback, specific guidance for improving performance, and support for learners’ self-monitoring and self-regulation (Compagna et al., 2025; Hattie & Timperley, 2007; Sadler, 1989).

1.2.4. AI Use Indicators, Cognitive Offloading, and Quality of Use

The educational meaning of AI use time varies with the learning environment and with how it is measured. In structured intelligent tutoring systems (ITSs), log-based learning time can reflect learning activity directly tied to a specific curriculum; in an ALEKS-based environment, for example, prior knowledge and log-based learning time significantly predicted achievement, alongside learner variables such as self-efficacy and resource management (Lee & Doo, 2025). Self-reported total weekly AI use time, which includes personal as well as course-related use, is not conceptually equivalent to such log-based time. This study therefore treated the proportion of in-class AI use and total weekly AI use time as AI use indicators with distinct characteristics and did not assume that quantitative increases in use imply equivalent educational engagement or learning outcomes.
Cognitive offloading refers to the use of external tools or resources to reduce the internal cognitive demands of task performance (Risko & Gilbert, 2016). Offloading is not inherently negative—it can free limited cognitive resources for more complex problem solving and higher-order thinking—but educational concern arises when external support replaces, rather than supports, learners’ verification, reflection, and independent judgment. Recent discussions similarly suggest that AI may extend learners’ cognitive performance while warning that sustained passive delegation of reasoning and judgment may foster cognitive dependence and reduced engagement (Abbosh et al., 2025; Girma, 2025; Zhai et al., 2024); Westerbeek (2026) extends this debate to questions of ‘cognitive authorship’ and ‘epistemic sovereignty.’ This study draws on these discussions as a theoretical lens without extending cross-sectional self-report associations into evidence of long-term cognitive or neurological change.

1.2.5. Human Agency, Self-Regulation, and the Shift from Technology Effects to Learning Design

The educational effects of AI may depend not only on the capabilities of AI systems but also on learners’ ability to judge and regulate when and how to use them. Ehlers (2026) emphasized that in a society where AI is ubiquitous, the role of education is not merely to increase the efficiency of information access but to develop the capacity to judge autonomously under uncertainty and to sustain human agency. From this perspective, AI literacy extends beyond tool-use skills to a self-regulatory competency for judging when to draw on AI assistance and when to think independently, and by what criteria to accept, reject, or revise AI-generated suggestions. Process-oriented research showing that GenAI literacy develops gradually through prompting, metacognitive monitoring, critical evaluation, and personalized adjustment of use also supports the importance of such self-regulation and learner agency (W. Yan et al., 2025).
From an educational technology perspective, learners’ AI use is shaped not only by individual competencies but also by instructional design; introducing AI systems into educational settings does not automatically guarantee particular learning outcomes, which also depend on learner characteristics and instructional conditions such as instructor feedback and task structure (Lee & Doo, 2025). The history of the field carries the same lesson. Analyzing four decades of research in Computers & Education, Zawacki-Richter and Latchem (2018) showed that research attention has repeatedly shifted from the effects of the technology itself toward learners, learning processes, collaboration, instructional design, and conditions such as assessment and feedback. Although that analysis predates the spread of GenAI, evaluating AI-integrated education solely through quantitative use indicators would risk repeating a technology-centered approach that neglects educational context and learning processes. Accordingly, this study does not test causal pathways or a comprehensive structural model; it examines the relative associations of learner-level AI literacy, instructor feedback as an instructional condition, and two AI use indicators with self-perceived cognitive learning outcomes, distinguishing the sheer amount of AI use from learners’ qualitative competency in using AI.

1.3. Purpose and Research Questions

The purpose of this study was to describe the levels of self-perceived cognitive learning outcomes among university students enrolled in AI-integrated courses and to analyze how AI literacy, instructor feedback, the proportion of in-class AI use, and total weekly AI use time are related to those outcomes. In particular, the study jointly examined the relative associations of two quantitative AI use indicators, AI literacy as a qualitative learner competency, and instructor feedback as an instructional condition. As a cross-sectional correlational study based on data collected at a single time point, it does not aim to estimate causal effects of enrollment in AI-integrated courses or of any individual predictor on learning outcomes. The specific research questions were as follows. Given the exploratory, cross-sectional nature of the design, the research questions are stated as questions rather than directional hypotheses, and the study was not preregistered.
  • Research Question 1. What are the levels of self-perceived cognitive learning outcomes, AI literacy, and instructor feedback among students in AI-integrated courses?
  • Research Question 2. How are self-perceived cognitive learning outcomes related to AI literacy, instructor feedback, the proportion of in-class AI use, and total weekly AI use time?
  • Research Question 3. When AI literacy, instructor feedback, the proportion of in-class AI use, and total weekly AI use time are considered simultaneously, which variables show independent associations with self-perceived cognitive learning outcomes?
  • Research Question 4. Are the associations between the predictors and self-perceived cognitive learning outcomes consistent across the six cognitive domains of remember, understand, apply, analyze, evaluate, and create after correction for multiple testing?

2. Materials and Methods

2.1. Study Design, Recruitment, and Participants

This cross-sectional online survey study was conducted from 26 May to 22 June 2026 with students enrolled at a private university in the Republic of Korea. A total of 439 students participated in the survey, of whom 215 reported taking one or more AI-integrated courses in the spring semester of 2026. Among these, three respondents were excluded because the required course-name item was missing or no valid course name could be identified, yielding a final analytic sample of 212. A duplicate check of student identifiers across all 439 respondents identified no duplicate cases. In addition, there were no item-level missing values in the final sample of 212 for the 52 measurement items or for the two AI use indicators (the proportion of in-class AI use and total weekly AI use time).
The final sample consisted of 109 men (51.4%) and 103 women (48.6%); 195 participants (92.0%) were undergraduates and 17 (8.0%) were graduate students (Table 1). Participants were distributed across diverse fields, including computing and AI, media, nursing, law, business, medicine, and the humanities and social sciences. Because respondents were asked to list all AI-integrated courses they were taking, a single respondent could report multiple courses; 39 respondents (18.4%) listed more than one course (29 listed two, 8 listed three, and 2 listed four). Individual respondents therefore could not be uniquely assigned to a single course, and sample sizes across courses were uneven. Accordingly, rather than estimating course-level effects or fitting course-level multilevel models, this study used the student as the unit of analysis and examined the associations of self-perceived cognitive learning outcomes with AI literacy, instructor feedback, and the AI use indicators.

2.2. Instrument Development and Content Validity

Three multi-item constructs—self-perceived cognitive learning outcomes, AI literacy, and instructor feedback—were measured on 5-point Likert-type scales. Items were selected on the basis of relevant theory and prior research and then revised and adapted to the university learning context and the purpose of this study. Two rounds of Delphi review with three experts in education were conducted to establish content validity. In the first round, the scale-level content validity index (S-CVI/Ave) was 1.00; in the second round, the standard deviation of expert ratings was 0.00 and the degree of consensus was 1.00, indicating full agreement among the experts on the final items. Through this process, the final instrument comprised 52 items. The full text of all 52 items is provided in Appendix A (Table A1).
The self-perceived cognitive learning outcome items were not adapted from any standardized cognitive ability test; they were self-report statements constructed on the basis of the cognitive process categories of the revised Bloom’s taxonomy to measure how students perceived their own level of performance. Likewise, the AI literacy and instructor feedback items did not reproduce any existing scale verbatim; they were constructed by revising and adapting the theoretical components and conceptual content of prior research to the AI-integrated course context of this study. Accordingly, the results are interpreted not as equivalent to objective cognitive ability, actual AI performance capability, or scores on any original scale, but within the limited scope of self-perceived cognitive learning outcomes, AI literacy, and instructor feedback as operationally defined and measured in this study.

2.2.1. Self-Perceived Cognitive Learning Outcomes

Self-perceived cognitive learning outcomes were measured with 24 items constructed on the basis of the cognitive process dimension of the revised Bloom’s taxonomy (Anderson & Krathwohl, 2001; Krathwohl, 2002). The subdomains comprised three items for remember, five for understand, four for apply, four for analyze, four for evaluate, and four for create. The items covered cognitive performances such as recalling key terms and content learned in class, explaining concepts in one’s own words, applying learned knowledge or procedures to new situations, analyzing relations among pieces of information or concepts, evaluating claims or results against explicit criteria, and generating new solutions or products. Each item asked students to rate the extent to which they perceived themselves as able to carry out the given cognitive performance, with higher scores indicating higher self-perceived performance. In this study, Cronbach’s α was 0.968 for the total scale and ranged from 0.849 to 0.903 across the six subdomains.

2.2.2. AI Literacy

AI literacy was measured with 20 items in five predefined content domains derived from relevant theory and prior research: conceptual understanding (4 items), use and application (4 items), critical evaluation (4 items), ethics and responsibility (4 items), and self-regulated use (4 items). Item construction drew on the conceptualization and measurement studies of Long and Magerko (2020), Ng et al. (2021, 2024), Carolus et al. (2023), and Lintner (2024). Item content included understanding the differences between AI and conventional software, selecting AI tools appropriate to learning purposes and constructing prompts, checking the sources of and errors in AI outputs, using AI in accordance with copyright and ethical principles, and monitoring one’s own learning process and level of dependence on AI. Higher item scores indicate higher self-perceived understanding, use, critical evaluation, ethics and responsibility, and self-regulated use of AI. Cronbach’s α for the total AI literacy scale was 0.947 in this study.

2.2.3. Instructor Feedback

Instructor feedback was measured with eight items assessing students’ perceptions of the formative feedback provided by their instructors during the learning process. The items covered the clarity and specificity of feedback, timeliness, revision guidance for improving performance, prompting re-examination of the problem-solving process, support for learner self-monitoring, alignment with learning goals and evaluation criteria, and direction for improving subsequent performance. Item construction was based on theoretical accounts of formative feedback and prior research on feedback levels (Compagna et al., 2025; Hattie & Timperley, 2007; Sadler, 1989). Higher scores indicate that students perceived their instructor’s feedback as clearer, more timely, and more helpful for monitoring and improving their performance. Cronbach’s α for the instructor feedback scale was 0.960 in this study.

2.2.4. AI Use Indicators

Two self-report AI use indicators with distinct characteristics were measured. First, the proportion of in-class AI use was measured with a single item: ‘How much of the course involved the use of AI?’ Responses were coded in increasing order of AI use as 1 = very low, 2 = low, 3 = moderate, 4 = high, and 5 = very high. In the final sample, 11 respondents chose very low, 20 low, 68 moderate, 67 high, and 46 very high. Second, total weekly AI use time was measured with a single item: ‘Over the past month, on average, how many hours per week did you use AI, including both class time and personal use?’ Responses were coded in increasing order as 1 = less than 3 h, 2 = 3 to less than 5 h, 3 = 5 to less than 10 h, and 4 = 10 h or more, with 85, 63, 35, and 29 respondents in each category, respectively. This indicator does not capture only the time invested in the learning activities of a specific course; it is a self-reported measure encompassing both class-related and personal use. Total weekly AI use time was therefore interpreted not as time-on-task in a specific course but as an indicator of students’ overall level of AI use.

2.3. Construct Validity Analyses

Because the measurement items were constructed by adapting prior research to the study context, confirmatory factor analysis (CFA) was conducted to further examine construct validity. CFAs were estimated using normal-theory maximum likelihood on standardized item data. Given the limited sample size (N = 212) relative to the 52 items, the measurement structure of each scale was examined first, and the 12-factor correlated model including self-perceived cognitive learning outcomes, AI literacy, and instructor feedback simultaneously was used as a diagnostic model for inspecting the overall structure among constructs rather than as a definitive measurement model test. For self-perceived cognitive learning outcomes, a one-factor model, a six-factor correlated model, and a second-order model in which the six cognitive domains loaded on a single higher-order factor were compared; for AI literacy, a one-factor model, a five-factor correlated model, and a second-order model were compared. Instructor feedback was examined primarily as a one-factor model in line with its theoretical structure. Model fit was evaluated using χ2, the comparative fit index (CFI), the Tucker–Lewis index (TLI), the root mean square error of approximation (RMSEA), and the standardized root mean square residual (SRMR), with composite reliability (CR), average variance extracted (AVE), interfactor correlations, and the heterotrait–monotrait ratio (HTMT) used as supplementary evidence of convergent and discriminant validity.
The CFAs were conducted on the same sample of 212 used in the main regression analyses. They therefore do not constitute cross-validation on an independent sample and are interpreted as providing additional evidence about the internal measurement structure of the instruments constructed in this study. To examine whether treating the 5-point items as continuous affected the conclusions, all measurement models were additionally re-estimated treating the items as ordinal, using diagonally weighted least squares (DWLS) on polychoric correlations. Because mean-and-variance-adjusted (WLSMV-type) test statistics were not available in the open-source environment used, fit indices are not reported for the ordinal solutions; their parameter estimates are summarized in Section 3.2.

2.4. Statistical Analysis

All statistical analyses were conducted in Python 3.13.5 using NumPy 2.3.5, SciPy 1.17.0, and statsmodels 0.14.6. Descriptive statistics and internal consistency reliability (Cronbach’s α) were computed for the main variables, and Pearson product–moment correlations were calculated to examine bivariate relations. In the main regression model, the total self-perceived cognitive learning outcome score was regressed simultaneously on AI literacy, instructor feedback, the proportion of in-class AI use, and total weekly AI use time. Coefficients were estimated by ordinary least squares (OLS), and the explanatory power and overall significance of the model were evaluated by R2 and the F test, respectively. Statistical inference for individual coefficients used heteroskedasticity-robust HC3 standard errors, with 95% confidence intervals and two-tailed p values (MacKinnon & White, 1985). Multicollinearity among predictors was checked using variance inflation factors (VIFs).
Additional regressions were estimated with each of the six cognitive process domains—remember, understand, apply, analyze, evaluate, and create—as the dependent variable. Testing four predictors across six domains yielded inferences on 24 regression coefficients; to control the accumulation of Type I error from multiple testing, the Benjamini–Hochberg false discovery rate (FDR) procedure was applied (Benjamini & Hochberg, 1995). In the domain-specific regressions, FDR-corrected q < 0.05 was used as the criterion for statistical significance.
Several further analyses addressed the robustness and interpretability of the results. Because the two AI use indicators are ordinal, the main regression model was re-estimated with each indicator entered as a set of categorical dummy variables. To quantify incremental contributions, a hierarchical regression was estimated (Block 1: the two AI use indicators; Block 2: instructor feedback; Block 3: AI literacy), and the change in R2 attributable to each predictor when entered last was computed. To address the asymmetry in measurement precision between the 20-item AI literacy scale and the single-item AI use indicators, the main model was re-estimated 20 times with the AI literacy composite replaced by every single item in turn, and once with a five-item short form (one item per content domain). To avoid interpreting statistical nonsignificance as evidence of no effect, equivalence was examined for the nonsignificant predictors using two one-sided tests (TOST) against a smallest effect size of interest of |β| = 0.20, based on the coefficient estimates and HC3 standard errors of the main model, complemented by BIC-approximated Bayes factors (BF01) comparing the full model with models omitting each predictor; a sensitivity power analysis determined the minimum detectable effect size for individual coefficients. Finally, Harman’s single-factor test was conducted as a coarse probe of common method variance.

2.5. Research Ethics

This study was conducted as a non-interventional educational study using an online survey and was confirmed by the Institutional Review Board of Hallym University to be exempt from formal review. Before participating, respondents received a verbal explanation of the purpose of the study and the participation procedures, and completed the online survey after giving consent. The raw survey file contained some participant-level information collected for research administration; no directly identifying information was used as an analysis variable, and such information was excluded from analysis and reporting. The survey data were used solely for the purposes of this study, and participant-level data were managed in accordance with institutional research ethics and privacy standards.

3. Results

3.1. Descriptive Statistics and Internal Consistency

Descriptive statistics and internal consistency reliabilities for the main variables are presented in Table 2. The mean of the total self-perceived cognitive learning outcome score was 3.63 (SD = 0.68). Among the six subdomains, evaluate was highest (M = 3.75, SD = 0.72), followed by analyze (M = 3.66, SD = 0.72), understand (M = 3.61, SD = 0.75), apply (M = 3.61, SD = 0.81), create (M = 3.58, SD = 0.80), and remember (M = 3.55, SD = 0.78). The mean was 3.90 (SD = 0.65) for AI literacy and 3.89 (SD = 0.83) for instructor feedback. Skewness and kurtosis of the main composite scores ranged from −0.69 to 0.05 and from −0.51 to 0.89, respectively, indicating no extreme departures from normality in skewness or kurtosis. Cronbach’s α for the scales and subdomains ranged from 0.849 to 0.968, indicating high internal consistency. High α values were not themselves interpreted as evidence of unidimensionality; the measurement structure was examined separately through confirmatory factor analysis.

3.2. Confirmatory Factor Analysis

Model fit indices for the CFAs are presented in Table 3. For self-perceived cognitive learning outcomes, the six-factor correlated model was clearly superior to the one-factor model and showed generally acceptable fit (CFI = 0.945, TLI = 0.936, RMSEA = 0.067, SRMR = 0.042). The second-order model, in which the six cognitive domains loaded on a single higher-order factor, also showed acceptable fit (CFI = 0.931, TLI = 0.922, RMSEA = 0.074, SRMR = 0.049), with standardized second-order loadings of 0.838–0.976. In the six-factor model, standardized item loadings were 0.671–0.881, composite reliabilities (CR) were 0.859–0.904, and average variance extracted (AVE) was 0.605–0.703. Interfactor correlations, however, were relatively high (0.712–0.943), and the maximum HTMT value was 0.946, indicating limited discriminant validity between some cognitive domains. These results suggest that the six domains have a distinguishable measurement structure while sharing substantial common variance. Accordingly, the domain scores were analyzed separately, and the total self-perceived cognitive learning outcome score reflecting the common higher-order construct was used alongside them. Practically, correlations of this magnitude mean that the six domains are close to empirically indistinguishable in these data; domain-level distinctions should therefore be read as descriptive rather than as evidence of separable constructs.
For AI literacy, the five-factor correlated model (CFI = 0.939, TLI = 0.928, RMSEA = 0.071, SRMR = 0.049) and the second-order model (CFI = 0.930, TLI = 0.920, RMSEA = 0.074, SRMR = 0.055) both fit better than the one-factor model. In the second-order model, standardized second-order loadings were 0.688–0.964, indicating that the five content domains were related to a common higher-order AI literacy construct. CRs were 0.822–0.877 and AVEs were 0.538–0.642, providing evidence of internal consistency and convergent validity. Some factors, however, were highly related, with a maximum interfactor correlation of 0.935 and a maximum HTMT of 0.924, indicating limited discriminant validity between some content domains. Taken together, the five domains showed a distinguishable measurement structure while sharing substantial common variance. The total AI literacy composite spanning the five domains was therefore used in the main analyses, and the results were interpreted with the potential overlap among content domains in mind. As with the outcome measure, these values indicate that the five content domains carry limited interpretive weight as distinct constructs in these data.
The one-factor model for instructor feedback showed high CFI and TLI and a low SRMR, but a relatively high RMSEA (CFI = 0.977, TLI = 0.968, RMSEA = 0.096, SRMR = 0.027). An exploratory two-factor model separating items 1–4 and items 5–8 improved overall fit relative to the one-factor model; however, the correlation between the two latent factors was very high (0.955), providing little basis for interpreting them as substantively distinct constructs. Instructor feedback was therefore retained as a single composite score in subsequent analyses rather than split into subfactors.
The 12-factor correlated model including all 52 items, examined diagnostically, showed acceptable absolute fit (RMSEA = 0.063, SRMR = 0.050) but relatively low incremental fit (CFI = 0.894, TLI = 0.883). The full measurement model therefore cannot be regarded as fully confirmed in this sample, and the CFA results are interpreted as supplementary internal evidence about the measurement structure. Future research should re-examine the full measurement structure and the discriminant validity among constructs in independent samples.
When the measurement models were re-estimated treating the items as ordinal (DWLS on polychoric correlations), the parameter estimates were consistent with the maximum likelihood solutions and, if anything, stronger: for the six-domain model, standardized item loadings were 0.737–0.924, interfactor correlations 0.740–0.949, and second-order loadings 0.870–0.972; for the five-domain AI literacy model, item loadings were 0.692–0.934, interfactor correlations reached 0.944, and second-order loadings were 0.714–0.966. The ordinal estimation therefore confirms both the strong measurement structure and the limited discriminant validity observed under maximum likelihood.

3.3. Correlations Between Self-Perceived Cognitive Learning Outcomes and the Main Variables

Pearson correlations between self-perceived cognitive learning outcomes and the main variables are presented in Table 4. AI literacy was significantly and positively correlated with the total score (r = 0.664, p < 0.001) and with all six subdomains. Instructor feedback was also significantly and positively correlated with the total score and all subdomains. In contrast, the proportion of in-class AI use and total weekly AI use time showed weak positive correlations with the total score (r = 0.109 and r = 0.075, respectively). Descriptively, AI literacy and instructor feedback showed larger correlations overall than the two AI use indicators. As these are cross-sectional bivariate correlations, they do not imply causal relations or causal direction.

3.4. Multiple Regression on Total Self-Perceived Cognitive Learning Outcomes

The OLS regression model including the four predictors simultaneously was statistically significant, F(4, 207) = 44.58, p < 0.001 (Table 5). The model explained 46.3% of the variance in the total self-perceived cognitive learning outcome score (R2 = 0.463, adjusted R2 = 0.452). With the other predictors held constant, AI literacy remained significantly and positively associated with self-perceived cognitive learning outcomes (B = 0.615, HC3 SE = 0.071, 95% CI [0.475, 0.754], p < 0.001, β = 0.593). In contrast, instructor feedback (B = 0.111, p = 0.091), the proportion of in-class AI use (B = 0.037, p = 0.231), and total weekly AI use time (B = 0.033, p = 0.331) did not show statistically significant coefficients in the simultaneous model; that is, the data did not provide evidence for independent associations of these predictors. VIFs ranged from 1.008 to 1.315, indicating no multicollinearity problems.
Robustness analyses supported and qualified these results. In the hierarchical regression, the two AI use indicators alone did not significantly explain the outcome (Block 1 R2 = 0.016, p = 0.177); adding instructor feedback increased R2 by 0.179 (p < 0.001), and adding AI literacy increased it by a further 0.268 (p < 0.001). When entered last, AI literacy uniquely accounted for ΔR2 = 0.268, against 0.014 for instructor feedback, 0.004 for the proportion of in-class AI use, and 0.003 for weekly use time. When the two AI use indicators were entered as categorical dummy variables, the pattern was unchanged (AI literacy B = 0.616, p < 0.001; R2 = 0.476; no dummy term significant). The association of AI literacy was also robust to its measurement length: replacing the 20-item composite with every single item in turn yielded standardized coefficients of 0.186–0.513 (median 0.353), significant in all 20 models, and a five-item short form (α = 0.768) yielded β = 0.503 (p < 0.001, R2 = 0.392) with the other predictors remaining nonsignificant. Equivalence tests indicated that the associations of the two AI use indicators were statistically equivalent to zero within |β| < 0.20 (TOST p = 0.003 for both; BIC-approximated BF01 = 7.2 and 8.6, moderate evidence for the null), whereas the test for instructor feedback was inconclusive (TOST p = 0.220; BF01 = 0.9): the data neither support an independent association nor rule out an educationally meaningful effect, as the upper bound of its confidence interval (B = 0.240) indicates. A sensitivity power analysis showed that, with N = 212 and four predictors, the design had 80% power (α = 0.05, two-tailed) to detect standardized coefficients of approximately |β| ≥ 0.19 (f2 ≈ 0.038). Finally, Harman’s single-factor test indicated that a single unrotated component accounted for 43.0% of the variance in the 52 items, below the conventional 50% threshold, although this test is insensitive and common method variance cannot be ruled out.

3.5. Exploratory Domain-Specific Regressions and FDR Correction

Given the limited discriminant validity among the six cognitive domains reported in Section 3.2, the domain-specific analyses in this section are exploratory. Results of the domain-specific regressions with FDR correction are presented in Table 6. After applying the Benjamini–Hochberg FDR procedure to the 24 predictor coefficient tests across the six cognitive domains, AI literacy remained significantly and positively associated with all six domains (standardized β = 0.412–0.615, all q < 0.001). Instructor feedback showed positive associations at p < 0.05 with the remember and apply domains before correction, but neither remained significant after FDR correction (q = 0.111 and q = 0.105, respectively). The proportion of in-class AI use was not significantly associated with any domain after correction. Total weekly AI use time retained a small positive association only with the remember domain after FDR correction (B = 0.116, 95% CI [0.035, 0.196], β = 0.157, p = 0.0048, q = 0.0164). However, total weekly AI use time is a broad self-report indicator encompassing both class-related and personal use, and the same pattern was not observed in the other five domains; this single result for the remember domain was therefore interpreted with caution.

4. Discussion

4.1. Interpretation of the Main Findings

This study examined the associations of AI literacy, instructor feedback, the proportion of in-class AI use, and total weekly AI use time with self-perceived cognitive learning outcomes among university students in AI-integrated courses. The central finding, consistent across bivariate correlations, multiple regression, and domain-specific analyses, was that AI literacy showed stronger and more consistent associations with self-perceived cognitive learning outcomes than the two AI use indicators. This does not constitute causal evidence that AI literacy improves cognitive learning outcomes. It does indicate that in this sample, a competency-centered construct—understanding, evaluating, and regulating the use of AI—was more closely related to self-perceived cognitive learning outcomes than quantitative indicators of how much AI was used. An important qualification is that this comparison partly confounds constructs with measurement quality: AI literacy was measured with a 20-item scale (α = 0.947), whereas the two AI use indicators were single ordinal items, so attenuation due to measurement error alone could account for part of the difference in associations. Illustratively, under an assumed single-item reliability of 0.60, the disattenuated correlation with the total outcome score would rise from 0.109 to approximately 0.14 for the proportion of in-class AI use and from 0.075 to approximately 0.10 for weekly use time, compared with a disattenuated correlation of approximately 0.69 for AI literacy. The ordering of the associations is therefore unlikely to be fully explained by attenuation, but the size of the differences should not be interpreted at face value. Consistent with this, the single-item sensitivity analyses in Section 3.4 show that the association of AI literacy is not an artifact of aggregation—every single-item model remained significant—while its magnitude attenuated exactly as measurement theory predicts.
The descriptive statistics for self-perceived cognitive learning outcomes did not follow a simple lower-to-higher-order sequence across the revised Bloom categories: evaluate had the highest mean and remember the lowest. The revised taxonomy, however, is a theoretical framework for classifying types of cognitive processes, not a scale ordering their psychometric difficulty. It is therefore not appropriate to interpret these mean differences as evidence that students developed higher-order processes more than lower-order ones. The results should be read as describing how students perceived their class-related cognitive performance at the end of the semester. The CFA likewise showed that the six domains form a multidimensional structure that is distinguishable yet highly interrelated. Accordingly, domain scores were analyzed separately alongside the total score reflecting the common higher-order construct, and caution is needed in interpreting individual domains as fully independent constructs.
AI literacy showed the highest bivariate correlation with the total score (r = 0.664) and retained a significant independent association in the regression model (B = 0.615, β = 0.593). It also remained significantly associated with all six cognitive domains after FDR correction. This pattern is consistent with conceptualizations of AI literacy not as mere AI use experience but as a learner competency encompassing understanding, critical evaluation, ethical judgment, and self-regulated use (Carolus et al., 2023; Long & Magerko, 2020; Ng et al., 2021, 2024). It also aligns with studies linking AI literacy to self-regulation and academic performance among university students (Kong et al., 2021, 2022; Shi et al., 2025) and with research suggesting that mature GenAI literacy involves process characteristics such as metacognitive monitoring, critical evaluation, and strategic use (W. Yan et al., 2025). Although the present data cannot directly distinguish active thinking with AI from passive dependence on it, the findings are consistent with theoretical accounts holding that the educational meaning of AI use depends less on its sheer amount than on how critically and deliberately learners handle AI outputs (Abbosh et al., 2025). An equally plausible alternative to any reading that treats AI literacy as the antecedent is the reverse: students who appraise their cognitive performance favorably may also appraise their AI-related competencies favorably, so the observed association is equally consistent with perceived competence shaping self-reported AI literacy, or with a common disposition such as academic self-efficacy shaping both.
In contrast, the proportion of in-class AI use and total weekly AI use time did not show independent associations with the total score when the other predictors were considered; in the hierarchical analysis, the two indicators alone explained only 1.6% of the variance (p = 0.177). These results should not be interpreted to mean that the amount of AI use is educationally unimportant. The proportion of in-class AI use is an ordinal self-report of how much of the course involved AI, and total weekly AI use time is a broad indicator including personal as well as academic use. Neither captures what tasks learners performed during that time nor how they examined, revised, and integrated AI outputs—that is, the quality of the actual learning process. By contrast, Lee and Doo (2025) reported that log-based learning time was related to objective achievement in a structured intelligent tutoring environment in which learning activities were directly tied to the curriculum. This contrast suggests that time-on-task invested in specific learning activities and self-reported AI use time spanning diverse purposes should not be interpreted as the same indicator.
Instructor feedback was significantly and positively related to the total score and all six subdomains at the bivariate level, but did not retain an independent association in the model including AI literacy and the two AI use indicators. The associations observed for the remember and apply domains before correction were also nonsignificant after FDR correction. The results therefore do not support claims that an independent association of instructor feedback was confirmed or that instructor feedback is selectively related to particular cognitive domains. One possible interpretation is that even if instructor feedback plays an educationally important role, its association may be partly shared with learners’ evaluative judgment, self-regulation, or AI literacy. As this explanation was not directly tested here, longitudinal or mediation studies are needed to examine whether instructor feedback is linked to later learning outcomes through AI literacy or self-regulation. A further measurement-based explanation is that respondents taking more than one AI-integrated course (18.4% of the sample; Section 2.1) may have answered the instructor feedback items with different course referents in mind, introducing measurement error that would attenuate the observed association. Together with the inconclusive equivalence test reported in Section 3.4, these considerations argue against reading the nonsignificant coefficient as evidence of no effect.
Finally, the single association that survived FDR correction—between total weekly AI use time and the remember domain (β = 0.157)—should not be over-interpreted: it was one of 24 tests, was not observed in the other five domains, and rests on domain distinctions whose discriminant validity is limited; it is best treated as an exploratory observation awaiting replication.

4.2. Theoretical Contributions

First, this study conceptually and empirically distinguished the quantity of AI use from AI literacy. Educational technology research has moved away from attributing learning outcomes to particular media or technologies and toward jointly considering the learner activities, interactions, and instructional designs that give technology its educational meaning (Zawacki-Richter & Latchem, 2018). Applying this perspective to AI-integrated higher education, the study found that AI literacy—the learner competency to understand, critically evaluate, and self-regulate the use of AI—was more consistently associated with self-perceived cognitive learning outcomes than the two AI use indicators. This does not mean that AI literacy is more important than the quantity of AI use in all situations. Rather, because the two reflect different aspects, the findings provide empirical grounds for measuring and interpreting them separately when analyzing the outcomes of AI-integrated education.
Second, the study shows that the meaning of total weekly AI use time can vary with context and measurement. Self-reported weekly AI use time encompassing personal and course-related use cannot be interpreted as equivalent to system-logged learning time tied directly to curricular activities. This distinction is particularly important given that technologies with very different functions and learning structures—from open-ended GenAI assistants to adaptive tutoring systems—are all entering AI education research. Subsuming such different environments under a single ‘AI use’ category can obscure differences in task structure, the cognitive activities required of learners, feedback loops, and actual time-on-task. Future research should therefore measure not only total weekly AI use time but also in which learning tasks, in what ways, and with which cognitive activities AI is used.
Third, the study positions self-perceived cognitive learning outcomes as a self-report educational outcome indicator distinct from objective cognitive performance. The construct measured here does not directly assess objective cognitive functioning or actual task performance; it reflects how learners perceive their class-related cognitive performance. This distinction is especially important in AI-supported learning environments, because task products created with AI may result from learner–AI interaction, and product quality alone does not license inferences about learners’ internal cognitive processing or independent capability. Self-perception data provide information about how learners judge their own understanding and performance, whereas performance-based assessment provides a different kind of evidence about actual task performance. Recent work on performance-based AI literacy assessment similarly argues for using the two types of measurement complementarily rather than equating self-reported competence with actual performance (Jin et al., 2025). Future research on AI education outcomes should combine self-reported perceptions, objective performance, and actual learning process data in multimethod designs.
Fourth, the study offers a more restrained, process-oriented interpretation of instructor feedback in AI-supported learning. Instructor feedback was related to self-perceived cognitive learning outcomes at the bivariate level but did not retain an independent association once AI literacy and the AI use indicators were considered. The results alone therefore cannot support claims about direct effects of instructor feedback. At the same time, this pattern suggests examining instructor feedback less as an independent predictor directly tied to outcomes and more as an instructional process supporting learners’ evaluative judgment, self-regulation, and AI literacy development. For example, instructor feedback may function as scaffolding that helps learners examine the validity of AI outputs, revise their judgments, and regulate how they use AI. This process view connects with recent arguments that in AI-rich learning environments, what matters is less the efficiency of production than learners’ retaining responsibility and agency for judgment and knowledge construction (Ehlers, 2026; Westerbeek, 2026). As these pathways were not directly tested here, they should be examined empirically in future longitudinal and mediation or process models.

4.3. Practical Implications

Two kinds of implications should be distinguished: the recommendation to measure the quantity of AI use, AI literacy, and learning outcomes as separate indicators follows directly from the present findings, whereas the more specific suggestions on task design and instructor development that follow are informed by the broader literature and should be read as theory-based extensions rather than as direct implications of our data. Universities should not use quantitative indicators such as AI adoption, the number of accounts, or use time as stand-alone evidence of educational effectiveness. Such indicators are useful for monitoring adoption and operation, but they do not by themselves guarantee the quality of learning or learning outcomes. University AI education performance management systems therefore need to measure the quantity of AI use, AI literacy, and actual learning outcomes as conceptually distinct. In particular, AI literacy—understanding the functions and limits of AI, checking sources and evidence, comparing multiple AI outputs, identifying errors and bias, making ethical decisions, and regulating one’s level of dependence on AI—should be included as an explicit learning objective.
AI-integrated assignments should likewise be designed so that learners retain authorship as agents of judgment and knowledge construction, rather than rewarding the mere production of unexamined AI outputs. For example, students can be asked to compare responses generated by multiple AI systems, to verify AI claims against original data and sources and record errors, to explain the rationale for their revisions of AI outputs, or to construct an independent solution first and then re-examine it using AI critique. Such task structures can help distinguish, at the level of actual learning activities, between delegating the entire problem-solving process to AI and using AI as a cognitive tool that supports reflection, verification, and higher-order problem solving (Abbosh et al., 2025).
Instructor development also needs to go beyond basic AI tool use and prompt-writing techniques. In AI-supported learning, instructor feedback can be designed to examine students’ reasoning processes: whether they constructed appropriate questions, verified the evidence AI provided, recognized the uncertainty and limitations of information, justified their final judgments, and retained responsibility for their submissions. Because instructor feedback did not retain a significant independent association with self-perceived cognitive learning outcomes in the simultaneous model, direct effects cannot be claimed from this study. Nevertheless, formative feedback theory and prior research on AI-supported learning suggest designing instructor support to promote learners’ evaluative judgment and self-regulation rather than the mere correction of products (Hattie & Timperley, 2007; Lee & Doo, 2025).
Finally, universities’ evaluation of AI education outcomes should use self-report perception data, objective performance data, and actual learning process data in a complementary manner. Institution-level performance management systems could include self-reported AI literacy, performance-based AI literacy assessment, course achievement, and rubric-based authentic task performance, together with process indicators such as source checking, revision histories of AI outputs, and AI use logs. This approach moves evaluation beyond asking ‘How much was AI used?’ toward asking ‘What did students understand, judge, verify, and ultimately produce in the process of using AI?’

5. Limitations and Future Research

Several limitations should be considered when interpreting the results. First, the study is cross-sectional, with all key variables measured at the same time point; temporal precedence and causal effects therefore cannot be established. For example, students with higher AI literacy may report higher self-perceived cognitive learning outcomes, but it is equally possible that students with higher prior achievement, confidence, or self-regulation rated both constructs higher. Longitudinal designs with repeated measures, cross-lagged models, and quasi-experimental designs are needed to test change and temporal precedence.
Second, the study did not include a comparison group of students not enrolled in AI-integrated courses, so the educational effects of enrollment itself cannot be estimated. The analyses examined associations among variables within students who reported taking AI-integrated courses. Future research should combine pre–post measurement with appropriate comparison conditions and account for learner characteristics such as prior achievement, discipline, prior AI experience, and learning time.
Third, because self-perceived cognitive learning outcomes, AI literacy, and instructor feedback were all measured by self-report at the same time point, common method variance and social desirability bias are possible (Podsakoff et al., 2003). In particular, the self-perceived cognitive learning outcomes measured here do not represent objective cognitive performance, and self-reported AI literacy is not equivalent to AI competence verified through performance. Recent research on performance-based GenAI literacy assessment shows that direct performance assessment can provide information distinct from self-report (Jin et al., 2025). Future research should combine self-report data with standardized performance tests or authentic tasks, course assessments, rubric-based scores, instructor ratings, and actual learning process data. No procedural separation of measurement occasions or sources was possible within the single online survey, and no marker variable was included. Harman’s single-factor test (Section 3.4) did not indicate a single dominant factor, but this test is insensitive, and marker-variable or CFA-based probes remain for future research with data collected for that purpose.
Fourth, the supplementary CFAs provide supporting but not definitive evidence of construct validity. The second-order models for self-perceived cognitive learning outcomes and AI literacy showed acceptable fit, but some first-order factor correlations were very high and some HTMT values exceeded conservative discriminant validity thresholds. The 12-factor diagnostic model including all 52 items showed acceptable RMSEA and SRMR but relatively low CFI and TLI. Moreover, the CFAs used the same sample of 212 as the main regressions, and 5-point Likert-type items were analyzed as approximately continuous using normal-theory maximum likelihood. Future research should secure larger independent samples, apply estimation methods appropriate for ordinal data such as weighted least squares mean and variance adjusted (WLSMV), and conduct cross-validation, along with measurement invariance tests across gender, educational level, discipline, and institution.
Fifth, because respondents could list multiple AI-integrated courses, individual responses could not be uniquely attributed to a single course, and course-level characteristics such as instructor, course content, assessment methods, and instructional design could not be related to outcomes. Future research should collect survey records anchored to a designated focal course or construct samples prospectively by course and section, enabling multilevel models that consider the student and course levels simultaneously.
Sixth, both AI use indicators were relatively broad single-item self-reports. The proportion of in-class AI use did not measure which learning activities actually involved AI, and total weekly AI use time included both class-related and personal use. In addition, AI literacy was measured with a 20-item scale whereas the two AI use indicators were single ordinal items, so differences in measurement precision may have contributed to differences in the observed associations. The results should therefore not be generalized as evidence that the quality of AI use inherently matters more than its quantity. Future research should collect finer-grained process indicators such as task-specific AI use logs, prompts and AI responses, fact-checking behavior, revision histories, source verification behavior, and purpose-specific time-on-task.
Seventh, the sample was drawn from a single university and consisted mostly of undergraduates with a small number of graduate students. Generalization to other institutions, disciplines, countries, and educational levels is therefore limited. Multi-institution, multi-group research is needed to test whether the observed relations replicate across educational and cultural contexts.
Finally, the study did not adequately include potential confounders and mechanisms such as prior achievement, prior AI experience, self-efficacy, self-regulated learning, motivation, and prior knowledge. Prior research on structured AI learning environments suggests that such learner characteristics can be meaningfully related to achievement (Lee & Doo, 2025). Future research should test longitudinal process models incorporating temporal relations among instructor feedback, AI literacy, self-regulation, cognitive offloading, self-perceived cognitive learning outcomes, and objective performance data. Moreover, as this study measured neither neural activity nor long-term change in cognitive functioning, claims about long-term cognitive decline, cognitive atrophy, or neurological change associated with AI use should not be inferred from its results.

6. Conclusions

As AI becomes an everyday learning tool in higher education, the central educational question is expanding beyond whether students use AI toward how they understand, evaluate, and regulate it and retain responsibility for its use. In this cross-sectional study of 212 university students in AI-integrated courses, AI literacy showed consistent positive associations with self-perceived cognitive learning outcomes and was the only predictor that retained a significant independent association with the total score after instructor feedback, the proportion of in-class AI use, and total weekly AI use time were considered simultaneously. The two AI use indicators showed no significant independent associations with the total score.
These findings must be interpreted within clear boundaries. The study does not demonstrate that taking AI-integrated courses improves objective cognitive performance, nor does it license the conclusion that AI literacy causally improves self-perceived cognitive learning outcomes. What it does show is that the quantity of AI use and the competency to understand, critically evaluate, and self-regulate the use of AI are empirically distinct indicators. In particular, the more consistent associations of AI literacy relative to the two AI use indicators suggest that the design and evaluation of AI-integrated courses should consider not only how much AI is used but also how learners use it.
AI-integrated courses should therefore be designed so that students question and examine AI outputs rather than accepting them as given—checking evidence and errors, applying ethical judgment, and continuously monitoring their own learning and their level of dependence on AI. Learning activities and instructional support that enable learners to retain cognitive authorship as the final agents of judgment and knowledge construction, even while drawing on AI assistance, are essential. Future evaluation of AI-integrated education should likewise draw on objective performance indicators, actual AI use process data, and learning behavior data rather than relying solely on self-reported perceptions. This perspective aligns with a learner-centered educational technology approach that judges outcomes by learners’ competencies, learning processes, and the quality of instructional design rather than by the mere adoption or amount of educational technology.

Author Contributions

Conceptualization, Y.E.L.; methodology, Y.E.L.; software, Y.E.L.; validation, Y.E.L.; formal analysis, Y.E.L.; investigation, Y.E.L.; resources, Y.E.L.; data curation, Y.E.L.; writing—original draft preparation, Y.E.L.; writing—review and editing, Y.E.L. and J.S.K.; visualization, Y.E.L.; project administration, Y.E.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the ANCHOR program (Glocal University 30) through the Gangwon ANCHOR Center, funded by the Ministry of Education (MOE) and the Gangwon State (G.S.), Republic of Korea, grant number 2026-ANCHOR-10-009. The APC was funded by the ANCHOR program (Glocal University 30), grant number 2026-ANCHOR-10-009.

Institutional Review Board Statement

Ethical review and approval were waived for this study by the Institutional Review Board of Hallym University, because it was a non-interventional survey study; the raw survey file contained participant-level information collected for research administration, no directly identifying information was used as an analysis variable, and such information was excluded from analysis and reporting. In accordance with the institution’s practice, no separate exemption number or confirmation date is issued for studies confirmed as exempt.

Informed Consent Statement

All participants received a verbal explanation of the study before the online survey link was provided, and the survey was administered after verbal informed consent was obtained.

Data Availability Statement

The raw data are not publicly available because they contain participant-level information. De-identified analysis data may be made available through the corresponding author upon reasonable request, within the scope permitted by institutional and ethical regulations.

Acknowledgments

During the preparation of this manuscript, the corresponding author used GenAI tools (OpenAI ChatGPT (GPT-5.6 Sol) and Anthropic Claude (Claude Fable 5); accessed in August 2026) for the purposes of reviewing the statistical values, references, interpretations of the results, and final wording reported in the manuscript. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Table A1 presents the full text of the 52 measurement items in the original Korean together with English translations. All items were administered on a 5-point Likert-type scale.
Table A1. Measurement items (Korean originals and English translations).
Table A1. Measurement items (Korean originals and English translations).
No.DomainItem (Korean)Item (English Translation)
Self-perceived cognitive learning outcomes (24 items)
1Remember나는 수업의 핵심 용어나 정의를 정확하게 떠올려 말할 수 있다.I can accurately recall and state the key terms and definitions from the course.
2Remember나는 수업 시간에 다룬 주요 내용을 다시 말할 수 있다.I can restate the main content covered in class.
3Remember나는 시험이나 과제를 할 때 배운 사실이나 정보를 잘 기억해 낸다.I recall learned facts and information well when taking exams or doing assignments.
4Understand나는 수업과 관련된 기본 개념, 정의, 절차를 정확히 이해하고 설명할 수 있다.I accurately understand and can explain the basic concepts, definitions, and procedures related to the course.
5Understand나는 원리에 근거하여 수업에서 배운 개념을 나만의 언어로 재구성하여 설명할 수 있다.I can reconstruct and explain concepts learned in class in my own words, based on their underlying principles.
6Understand나는 수업에서 배운 개념을 내 말로 설명할 수 있다.I can explain the concepts I learned in class in my own words.
7Understand나는 서로 비슷하거나 다른 개념의 차이를 이해하고 구분할 수 있다.I can understand and distinguish between similar or different concepts.
8Understand나는 수업 자료나 교수자의 설명에서 핵심 의미를 파악할 수 있다.I can grasp the key meaning in course materials or the instructor’s explanations.
9Apply나는 수업에서 배운 내용을 실제 문제 상황에 적용할 수 있다.I can apply what I learned in class to real problem situations.
10Apply나는 새로운 사례가 주어져도 배운 원리나 방법을 사용할 수 있다.I can use the principles or methods I have learned even when a new case is presented.
11Apply나는 과제를 수행할 때 수업에서 익힌 지식과 절차를 활용할 수 있다.I can use the knowledge and procedures acquired in class when carrying out assignments.
12Apply나는 배운 내용을 단순히 아는 수준을 넘어 실제 행동으로 옮길 수 있다.I can put what I have learned into practice, beyond simply knowing it.
13Analyze나는 복잡한 내용을 여러 요소로 나누어 이해할 수 있다.I can understand complex content by breaking it down into its components.
14Analyze나는 문제의 원인과 결과를 구분하여 파악할 수 있다.I can identify and distinguish the causes and effects of a problem.
15Analyze나는 자료나 주장 속에서 중요한 내용과 덜 중요한 내용을 구분하고, 정보를 논리적으로 분류할 수 있다.I can distinguish more important from less important content in materials or arguments and classify information logically.
16Analyze나는 여러 정보 사이의 관계나 구조를 논리적으로 파악할 수 있다.I can logically grasp the relationships or structure among multiple pieces of information.
17Evaluate나는 여러 의견이나 해결 방안 중 더 타당한 것을 판단할 수 있다.I can judge which of several opinions or solutions is more valid.
18Evaluate나는 객관적 근거를 바탕으로 제시된 정보의 논리적 오류나 편향성을 비판적으로 검토할 수 있다.I can critically examine logical errors or bias in presented information on the basis of objective evidence.
19Evaluate나는 명확한 기준과 근거를 바탕으로 판단하고 결정할 수 있다.I can judge and decide on the basis of clear criteria and evidence.
20Evaluate나는 수업에서 제시된 내용이나 결과를 그대로 받아들이지 않고 타당성을 따져본다.I do not accept content or results presented in class at face value but examine their validity.
21Create나는 배운 내용을 바탕으로 새로운 아이디어를 떠올릴 수 있다.I can come up with new ideas based on what I have learned.
22Create나는 기존의 방법을 변형하거나 결합하여 새로운 해결 방안을 만들 수 있다.I can create new solutions by modifying or combining existing methods.
23Create나는 과제나 문제를 해결할 때 독창적인 접근을 시도하는 편이다.I tend to attempt original approaches when solving assignments or problems.
24Create나는 수업에서 배운 여러 개념을 통합 및 재구성하여 기존에 없던 결과물(보고서, 작품, 모델 등)을 산출할 수 있다.I can integrate and reorganize the concepts learned in class to produce novel outputs (e.g., reports, works, models).
AI literacy (20 items)
1Conceptual understanding나는 AI와 일반적인 컴퓨터 프로그램의 차이를 설명할 수 있다.I can explain the difference between AI and ordinary computer programs.
2Conceptual understanding나는 AI가 어떤 방식으로 데이터를 바탕으로 작동하는지 기본 원리를 알고 있다.I know the basic principles of how AI operates on the basis of data.
3Conceptual understanding나는 생성형 AI와 다른 유형의 AI(예: 추천시스템, 음성인식)의 차이를 이해하고 있다.I understand the difference between generative AI and other types of AI (e.g., recommender systems, speech recognition).
4Conceptual understanding나는 AI의 가능성과 한계를 구분해서 이해하고 있다.I understand the possibilities and limitations of AI and can tell them apart.
5Use and application나는 학습 목적에 맞는 AI 도구를 선택할 수 있다.I can select AI tools appropriate to my learning purpose.
6Use and application나는 AI에게 원하는 결과를 얻기 위해 질문이나 지시문을 적절히 구성할 수 있다.I can appropriately construct questions or prompts to obtain the results I want from AI.
7Use and application나는 AI를 활용하여 과제 아이디어 정리, 정보 탐색, 초안 작성을 수행할 수 있다.I can use AI to organize ideas for assignments, search for information, and write drafts.
8Use and application나는 과제의 목적에 따라 AI를 활용하는 방식을 적절히 조절하며, AI가 원하는 답변을 주지 않을 때 질문을 구체화하여 다시 시도한다.I adjust how I use AI according to the purpose of the task, and when AI does not give the answer I want, I make my question more specific and try again.
9Critical evaluation나는 AI가 제시한 정보의 출처와 근거를 확인한다.I check the sources and evidence of the information AI provides.
10Critical evaluation나는 AI의 답변에 오류나 왜곡이 있을 수 있다는 점을 염두에 둔다.I keep in mind that AI answers may contain errors or distortions.
11Critical evaluation나는 AI가 만든 결과물을 그대로 사용하기보다 검토·수정한 뒤 활용한다.Rather than using AI-generated outputs as they are, I review and revise them before use.
12Critical evaluation나는 AI가 제시한 여러 답변 중 어떤 것이 더 타당한지 비교·판단할 수 있다.I can compare and judge which of several answers provided by AI is more valid.
13Ethics and responsibility나는 AI 활용 시 저작권 가이드라인 및 인용 원칙을 준수하여 결과물을 활용한다.When using AI, I comply with copyright guidelines and citation principles in using the outputs.
14Ethics and responsibility나는 AI 결과물에 편향이나 차별이 포함될 수 있다는 점을 알고 있다.I know that AI outputs may contain bias or discrimination.
15Ethics and responsibility나는 AI 활용이 사회와 직업 세계에 미칠 영향을 생각해 본다.I think about the impact that AI use will have on society and the world of work.
16Ethics and responsibility나는 학습에서 AI를 사용할 때 윤리적 기준과 학습 목적에 맞게 활용할 수 있다.When using AI in learning, I can use it in line with ethical standards and my learning goals.
17Self-regulated use나는 학습 목표에 따라 AI를 스스로 계획적으로 활용할 수 있다.I can plan and use AI on my own according to my learning goals.
18Self-regulated use나는 AI를 사용할 때 내가 무엇을 배우고 있는지 스스로 점검할 수 있다.When using AI, I can monitor for myself what I am learning.
19Self-regulated use나는 AI의 답변을 비판적으로 수용하며, 나만의 논리로 결과물을 완성한다.I accept AI answers critically and complete my work with my own reasoning.
20Self-regulated use나는 AI를 활용하여 학습 효율을 높이면서도, 결과에 과도하게 의존하지 않으려고 한다.I try to improve my learning efficiency with AI without becoming overly dependent on its outputs.
Instructor feedback (8 items)
1Instructor feedback교수자의 피드백은 명확하였다.The instructor’s feedback was clear.
2Instructor feedback교수자의 피드백은 내가 잘한 부분과 부족한 부분을 명확히 구분하여 구체적으로 설명해 주었다.The instructor’s feedback specifically explained what I did well and what I lacked, clearly distinguishing the two.
3Instructor feedback교수자의 피드백은 적절한 시점에 제공되었다.The instructor’s feedback was provided at appropriate times.
4Instructor feedback교수자의 피드백은 내가 무엇을 수정해야 하는지 알려주었다.The instructor’s feedback told me what I needed to revise.
5Instructor feedback교수자의 피드백은 문제해결 과정을 다시 생각하게 만들었다.The instructor’s feedback made me rethink my problem-solving process.
6Instructor feedback교수자의 피드백은 내가 학습과정을 스스로 점검하는 데 도움이 되었다.The instructor’s feedback helped me monitor my own learning process.
7Instructor feedback교수자는 학습 목표와 평가 기준에 부합하는 일관된 피드백을 제공하였다.The instructor provided consistent feedback aligned with the learning goals and evaluation criteria.
8Instructor feedback교수자의 피드백은 다음 학습에서 개선하고 보완해야 할 방향을 설정하는 데 도움이 되었다.The instructor’s feedback helped me set directions for improvement in my subsequent learning.
Note. Bold rows indicate the scale to which the items in the following rows belong.

References

  1. Abbosh, A., Al-Anbuky, A., Xue, F., & Mahmoud, S. S. (2025). Perspective on the role of AI in shaping human cognitive development. Information, 16(11), 1011. [Google Scholar] [CrossRef] [Scilit]
  2. An, Q., Koh, J. H. L., & Liu, Q. (2026). Generative artificial intelligence in higher education: A systematic review of student use and learning outcomes. Australasian Journal of Educational Technology. Advanced online publication. [Google Scholar] [CrossRef] [Scilit]
  3. Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001). A taxonomy for learning, teaching, and assessing: A revision of Bloom’s taxonomy of educational objectives. Longman. [Google Scholar]
  4. Benjamini, Y., & Hochberg, Y. (1995). Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B (Methodological), 57(1), 289–300. [Google Scholar] [CrossRef] [Scilit]
  5. Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7–74. [Google Scholar] [CrossRef] [Scilit]
  6. Carolus, A., Koch, M. J., Straka, S., Latoschik, M. E., & Wienrich, C. (2023). MAILS—Meta AI literacy scale: Development and testing of an AI literacy questionnaire based on well-founded competency models and psychological change- and meta-competencies. Computers in Human Behavior: Artificial Humans, 1(2), 100014. [Google Scholar] [CrossRef] [Scilit]
  7. Chiu, T. K. F., Xia, Q., Zhou, X., Chai, C. S., & Cheng, M. (2023). Systematic literature review on opportunities, challenges, and future research recommendations of artificial intelligence in education. Computers and Education: Artificial Intelligence, 4, 100118. [Google Scholar] [CrossRef] [Scilit]
  8. Compagna, K., Ross, S., & Lee, A. S. O. (2025). An exploration of feedback using Hattie and Timperley’s feedback levels. Family Medicine, 57(7), 508–512. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Ehlers, U.-D. (2026). How artificial intelligence is shaping the future of learning: Rethinking education, competence, and human agency. Journal of Innovative Business and Management, 18(1), 1–11. [Google Scholar] [CrossRef] [Scilit]
  10. Girma, A. H. (2025). The role of artificial intelligence in shaping human interaction and cognitive function. Kotebe Journal of Education, 3(1), 69–88. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81–112. [Google Scholar] [CrossRef] [Scilit]
  12. Jin, Y., Martinez-Maldonado, R., Gašević, D., & Yan, L. (2025). GLAT: The generative AI literacy assessment test. Computers and Education: Artificial Intelligence, 9, 100436. [Google Scholar] [CrossRef] [Scilit]
  13. Kaşarcı, İ., Akın Demircan, Z., Çeliker Ercan, G., & İnci, T. (2025). Managing artificial intelligence ethics in higher education: A systematic framework for issues and policy recommendations. International Journal of Current Educational Studies, 4(2), 112–137. [Google Scholar] [CrossRef] [Scilit]
  14. Kong, S.-C., Cheung, W. M.-Y., & Zhang, G. (2021). Evaluation of an artificial intelligence literacy course for university students with diverse study backgrounds. Computers and Education: Artificial Intelligence, 2, 100026. [Google Scholar] [CrossRef] [Scilit]
  15. Kong, S.-C., Cheung, W. M.-Y., & Zhang, G. (2022). Evaluating artificial intelligence literacy courses for fostering conceptual learning, literacy and empowerment in university students: Refocusing to conceptual building. Computers in Human Behavior Reports, 7, 100223. [Google Scholar] [CrossRef] [Scilit]
  16. Krathwohl, D. R. (2002). A revision of Bloom’s taxonomy: An overview. Theory Into Practice, 41(4), 212–218. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Lee, Y. E., & Doo, M. Y. (2025). What makes ALEKS learning successful?: Influences of prior knowledge, learning time, self-efficacy, and resource management on learning achievement. Journal of Computing in Higher Education. Advanced online publication. [Google Scholar] [CrossRef] [Scilit]
  18. Lintner, T. (2024). A systematic review of AI literacy scales. npj Science of Learning, 9, 50. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Long, D., & Magerko, B. (2020). What is AI literacy? Competencies and design considerations. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (pp. 1–16). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
  20. MacKinnon, J. G., & White, H. (1985). Some heteroskedasticity-consistent covariance matrix estimators with improved finite sample properties. Journal of Econometrics, 29(3), 305–325. [Google Scholar] [CrossRef] [Scilit]
  21. Ng, D. T. K., Leung, J. K. L., Chu, S. K. W., & Qiao, M. S. (2021). Conceptualizing AI literacy: An exploratory review. Computers and Education: Artificial Intelligence, 2, 100041. [Google Scholar] [CrossRef] [Scilit]
  22. Ng, D. T. K., Wu, W., Leung, J. K. L., Chiu, T. K. F., & Chu, S. K. W. (2024). Design and validation of the AI literacy questionnaire: The affective, behavioural, cognitive and ethical approach. British Journal of Educational Technology, 55(3), 1082–1104. [Google Scholar] [CrossRef] [Scilit]
  23. Pinquart, M., & Ebeling, M. (2020). Students’ expected and actual academic achievement—A meta-analysis. International Journal of Educational Research, 100, 101524. [Google Scholar] [CrossRef] [Scilit]
  24. Podsakoff, P. M., MacKenzie, S. B., Lee, J.-Y., & Podsakoff, N. P. (2003). Common method biases in behavioral research: A critical review of the literature and recommended remedies. Journal of Applied Psychology, 88(5), 879–903. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Sadler, D. R. (1989). Formative assessment and the design of instructional systems. Instructional Science, 18(2), 119–144. [Google Scholar] [CrossRef] [Scilit]
  27. Shi, J., Liu, W., & Hu, K. (2025). Exploring how AI literacy and self-regulated learning relate to student writing performance and well-being in generative AI-supported higher education. Behavioral Sciences, 15(5), 705. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Westerbeek, H. (2026). How AI is rewiring the human brain: The generational transformation of cognition and knowing. AI & Society, 41, 5327–5337. [Google Scholar] [CrossRef] [Scilit]
  29. Yan, W., Nakajima, T., & Sawada, R. (2025). Beyond tool use: Tracking the evolution of generative AI literacy among university students through a process-oriented investigation. Computers and Education: Artificial Intelligence, 9, 100465. [Google Scholar] [CrossRef] [Scilit]
  30. Yan, Z., Lao, H., Panadero, E., Fernández-Castilla, B., Yang, L., & Yang, M. (2022). Effects of self-assessment and peer-assessment interventions on academic performance: A meta-analysis. Educational Research Review, 37, 100484. [Google Scholar] [CrossRef] [Scilit]
  31. Yan, Z., Wang, X., Boud, D., & Lao, H. (2023). The effect of self-assessment on academic performance and the role of explicitness: A meta-analysis. Assessment & Evaluation in Higher Education, 48(1), 1–15. [Google Scholar] [CrossRef] [Scilit]
  32. Zawacki-Richter, O., & Latchem, C. (2018). Exploring four decades of research in Computers & Education. Computers & Education, 122, 136–152. [Google Scholar] [CrossRef] [Scilit]
  33. Zell, E., & Krizan, Z. (2014). Do people have insight into their abilities? A metasynthesis. Perspectives on Psychological Science, 9(2), 111–125. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Zhai, C., Wibowo, S., & Li, L. D. (2024). The effects of over-reliance on AI dialogue systems on students’ cognitive abilities: A systematic review. Smart Learning Environments, 11, 28. [Google Scholar] [CrossRef] [Scilit]
  35. Zimmerman, B. J. (2008). Investigating self-regulation and motivation: Historical background, methodological developments, and future prospects. American Educational Research Journal, 45(1), 166–183. [Google Scholar] [CrossRef] [Scilit]
Table 1. Characteristics of the study participants (N = 212).
Table 1. Characteristics of the study participants (N = 212).
Characteristicn%
Gender: male10951.4
Gender: female10348.6
Educational level: undergraduate19592.0
Educational level: graduate178.0
Table 2. Descriptive statistics and internal consistency of the main variables (N = 212).
Table 2. Descriptive statistics and internal consistency of the main variables (N = 212).
VariableMSDSkewnessKurtosisCronbach’s α
Self-perceived cognitive learning outcomes (total)3.630.680.050.060.968
Remember3.550.78−0.090.120.870
Understand3.610.75−0.200.470.898
Apply3.610.81−0.230.340.903
Analyze3.660.720.030.090.890
Evaluate3.750.720.02−0.460.849
Create3.580.80−0.11−0.220.859
AI literacy3.900.65−0.30−0.510.947
Instructor feedback3.890.83−0.690.890.960
Table 3. Model fit of the confirmatory factor analyses.
Table 3. Model fit of the confirmatory factor analyses.
Scale/Modelχ2dfCFITLIRMSEASRMR
Self-perceived cognitive learning outcomes: one-factor817.682520.8620.8490.1030.057
Self-perceived cognitive learning outcomes: six-factor correlated464.202370.9450.9360.0670.042
Self-perceived cognitive learning outcomes: second-order530.742460.9310.9220.0740.049
AI literacy: one-factor614.331700.8400.8210.1110.074
AI literacy: five-factor correlated329.351600.9390.9280.0710.049
AI literacy: second-order357.961650.9300.9200.0740.055
Instructor feedback: one-factor59.18200.9770.9680.0960.027
Overall 12-factor diagnostic model2208.6712080.8940.8830.0630.050
Table 4. Pearson correlations between self-perceived cognitive learning outcomes and the main variables.
Table 4. Pearson correlations between self-perceived cognitive learning outcomes and the main variables.
Outcome VariableAI LiteracyInstructor FeedbackProportion of In-Class AI UseTotal Weekly AI Use Time
Remember0.511 ***0.399 ***0.148 *0.181 **
Understand0.606 ***0.387 ***0.0880.090
Apply0.582 ***0.430 ***0.142 *0.098
Analyze0.611 ***0.359 ***0.0680.005
Evaluate0.636 ***0.339 ***0.0850.002
Create0.578 ***0.384 ***0.0630.038
Total0.664 ***0.432 ***0.1090.075
Note. * p < 0.05, ** p < 0.01, *** p < 0.001. The coefficients are uncorrected bivariate correlations; FDR correction was applied to the coefficient tests in the domain-specific regressions, not to this descriptive correlation table.
Table 5. Multiple regression on total self-perceived cognitive learning outcomes.
Table 5. Multiple regression on total self-perceived cognitive learning outcomes.
PredictorBHC3 SEStandardized β95% CIpVIF
Intercept0.6010.245[0.120, 1.081]0.014
AI literacy0.6150.0710.593[0.475, 0.754]<0.0011.312
Instructor feedback0.1110.0660.137[−0.018, 0.240]0.0911.315
Proportion of in-class AI use0.0370.0310.060[−0.024, 0.098]0.2311.011
Total weekly AI use time0.0330.0340.052[−0.034, 0.101]0.3311.008
Note. —, not applicable.
Table 6. Domain-specific regressions with Benjamini–Hochberg FDR correction.
Table 6. Domain-specific regressions with Benjamini–Hochberg FDR correction.
DomainAdj. R2AI Literacy β (q)Instructor Feedback β (q)Proportion of In-Class AI Use β (q)Total Weekly AI Use Time β (q)
Remember0.3150.412 (<0.001)0.185 (0.111)0.099 (0.258)0.157 (0.016)
Understand0.3730.546 (<0.001)0.115 (0.364)0.041 (0.601)0.071 (0.364)
Apply0.3710.485 (<0.001)0.185 (0.105)0.094 (0.201)0.072 (0.364)
Analyze0.3670.571 (<0.001)0.080 (0.572)0.029 (0.737)−0.013 (0.825)
Evaluate0.3960.615 (<0.001)0.038 (0.737)0.046 (0.601)−0.016 (0.800)
Create0.3370.512 (<0.001)0.132 (0.278)0.021 (0.782)0.020 (0.782)
Note. q values reflect Benjamini–Hochberg FDR correction applied to the 24 predictor × domain coefficient tests.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Lee, Y.E.; Kan, J.S. AI Literacy and Self-Perceived Cognitive Learning Outcomes Among University Students in AI-Integrated Courses: Associations with Instructor Feedback and AI Use Indicators. Educ. Sci. 2026, 16, 1384. https://doi.org/10.3390/educsci16091384

AMA Style

Lee YE, Kan JS. AI Literacy and Self-Perceived Cognitive Learning Outcomes Among University Students in AI-Integrated Courses: Associations with Instructor Feedback and AI Use Indicators. Education Sciences. 2026; 16(9):1384. https://doi.org/10.3390/educsci16091384

Chicago/Turabian Style

Lee, Yu Eun, and Jin Sook Kan. 2026. "AI Literacy and Self-Perceived Cognitive Learning Outcomes Among University Students in AI-Integrated Courses: Associations with Instructor Feedback and AI Use Indicators" Education Sciences 16, no. 9: 1384. https://doi.org/10.3390/educsci16091384

APA Style

Lee, Y. E., & Kan, J. S. (2026). AI Literacy and Self-Perceived Cognitive Learning Outcomes Among University Students in AI-Integrated Courses: Associations with Instructor Feedback and AI Use Indicators. Education Sciences, 16(9), 1384. https://doi.org/10.3390/educsci16091384

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop