1. Introduction
In this study, first-generation artificial intelligence refers to the current stage of AI development prior to fully autonomous artificial general intelligence (AGI). It encompasses GenAI, which creates novel content from learned data patterns, and general-purpose AI systems that perform diverse tasks but lack the adaptability, autonomous learning, and cross-domain reasoning associated with AGI. AI is changing higher education through automated feedback, learning analytics, intelligent tutoring, content generation, knowledge representation, and virtual simulation.
Crompton and Burke (
2023) describe a rapidly expanding field in which AI supports both teaching efficiency and new forms of learning.
Chan and Hu (
2023) likewise show that students perceive substantial benefits from GenAI while also raising concerns about accuracy, dependence, fairness, and academic integrity. UNESCO (
Miao & Holmes, 2023) therefore calls for human-centered governance, data protection, transparency, and the development of learners’ capacity to evaluate AI outputs. These developments imply that curriculum reform should align technology with educational purposes rather than treat AI as an optional add-on.
This alignment problem is especially important in professional engineering courses.
Chiu et al. (
2023),
Bearman et al. (
2023), and
Xia et al. (
2024) indicate that AI-enabled activities can support literature retrieval, data interpretation, visualization, formative feedback, and design comparison, but their educational value depends on the relationship among disciplinary content, pedagogy, assessment, and ethical use. A chatbot may produce a plausible explanation of cement hydration, for example, yet a student must still verify the explanation against material-science principles, test data, standards, and the conditions of a specific engineering application. The educational goal is therefore not more AI use in itself, but more capable and responsible problem solving.
The Chinese New Agricultural Science initiative asks agricultural higher education to serve new agriculture, new rural development, new farmers, and new ecological priorities through interdisciplinary and practice-oriented talent cultivation (
Ministry of Education of the People’s Republic of China, 2019;
General Office of the Ministry of Education of the People’s Republic of China, 2022). Within this context, Building Materials is a foundational course for civil engineering, agricultural engineering, and urban–rural construction. Conventional syllabi commonly organize the course by material categories such as cement, concrete, steel, masonry, asphalt, and polymers. This structure remains necessary, but it can underrepresent rural infrastructure, green construction, agricultural solid-waste valorization, life-cycle thinking, and the uncertainty of engineering decisions.
Fadele and Otieno (
2022) and
Endale et al. (
2023) show how agricultural-waste-derived supplementary cementitious materials, including rice-husk ash, can connect material chemistry, performance, environmental implications, and local agricultural resources.
Four gaps motivate this study. First,
Chiu et al. (
2023) and
Bearman et al. (
2023) show that AI-in-higher-education research often emphasizes opportunities and tools while leaving important questions about pedagogy, values, and integrated curriculum design unresolved. Second,
Healey (
2005),
Healey and Jenkins (
2009), and
Uaciquete and Valcke (
2022) emphasize inquiry and the research–teaching nexus, but provide limited course-specific guidance for converting research datasets and frontier problems into appropriately scaffolded undergraduate tasks. Third,
Campbell and Stanley (
1963),
Holland (
1986),
Shadish et al. (
2002), and
Vickers and Altman (
2001) underscore the need for defensible comparisons and appropriate treatment of baseline information when comparative or nonrandomized outcome evidence is interpreted. Fourth, curriculum reform reports frequently provide positive outcome claims without sufficiently transparent intervention descriptions, instrument evidence, comparison conditions, or safeguards against threats to validity. These concerns are particularly consequential when nonrandomized designs are used to support causal language.
Accordingly, this study had two linked aims: to develop a theory-informed instructional framework for integrating AI, research-informed teaching, and New Agricultural Science within an undergraduate Building Materials course, and to conduct a preliminary retrospective evaluation of outcomes observed under the resulting intervention. The conceptual and empirical components are treated separately: RQ1 and RQ2 concern the design logic and operationalization of the framework, whereas RQ3 concerns between-cohort outcome differences in the available archival data. The study addresses the following research questions:
RQ1. What theoretical and contextual principles can inform the integration of AI, research-informed teaching, New Agricultural Science, and intended learning outcomes in an undergraduate Building Materials course?
RQ2. How can these principles be operationalized in course content, learning activities, assessment, governance, and the broader educational environment?
RQ3. What between-cohort differences are observed in knowledge performance, practical performance, continuous-assessment performance, course satisfaction, and project innovation, and how should these differences be interpreted given the nonrandomized design and the unequal measurement strength of the available outcomes?
3. The One Core, Two Channels, and Four-Dimensional Drivers Framework
3.1. One Core: Responsible AI-Supported Engineering Problem Solving
The core outcome is the development of students who can solve Building Materials problems through disciplinary knowledge, engineering evidence, research awareness, and responsible AI use. The phrase ‘AI-literate innovative professional’ is therefore operationalized rather than treated as a general aspiration. A successful student should be able to explain material behavior, select or interpret tests, compare alternatives, identify uncertainty, use AI for an appropriate subtask, verify generated content, and justify a decision in relation to a rural or agricultural engineering scenario.
The core does not position AI as an autonomous decision maker. Teachers establish learning outcomes and task boundaries; students remain accountable for claims; and disciplinary evidence has priority over generated fluency. This human-in-the-loop principle is consistent with UNESCO’s human-centered guidance (
Miao & Holmes, 2023). The conceptual framework of the curriculum reform model empowered by artificial intelligence is shown in
Figure 1.
3.2. Two Channels: A Governed Research–Teaching Cycle
Channel 1 converts research outputs into teaching resources. Candidate papers, datasets, test images, standard-based problems, and engineering cases are screened for relevance and risk. The teaching team then reduces unnecessary complexity, retains the essential uncertainty, supplies prerequisite explanations, and designs questions or tasks that require evidence. AI may support classification, summarization, alternative representation, and preliminary question generation, but every resource is verified by a subject specialist before release.
Channel 2 converts aggregated learning evidence into teaching-improvement and possible research questions.
Susnjak et al. (
2022) show how learning-analytics dashboards can be designed to provide actionable, interpretable insights; accordingly, evidence in this course may include rubric patterns, anonymized misconceptions, commonly selected project topics, student explanations, and interpretable learning-analytics summaries. It is not a license to reuse identifiable prompts or personal data. The teaching team determines whether a pattern indicates a content gap, an instructional-design problem, an assessment problem, or a potentially researchable engineering issue. The data-flow architecture of the dual-channel research–teaching mechanism is shown in
Figure 2.
3.3. “Four-Dimensional Drivers”: Coordinated Reform of Content, Methods, Assessment, and Ecosystem
Content restructuring. Core material-science concepts are connected to green construction, rural infrastructure, agricultural-waste-derived materials, and life-cycle considerations through a knowledge map and case sequence.
Pedagogical innovation. Pre-class diagnosis, case inquiry, collaborative explanation, virtual and physical experiments, project work, and reflection are combined. AI use is attached to specific cognitive tasks rather than permitted without purpose.
Assessment transformation. Evidence includes knowledge tests, practical performances, project artifacts, oral defense, AI-use documentation, and reflective verification. Platform activity is supplementary evidence, not a proxy for learning quality.
Ecosystem development. The course links classroom, laboratory, research team, field or practice base, digital platform, and relevant industry or rural-construction problems while defining data responsibilities and access boundaries. The four-dimensional evaluation system for Building Materials is shown in
Figure 3.
3.4. Design Propositions
Alignment proposition: AI-enabled activities will be educationally coherent only when each activity is linked to an intended performance and an assessment criterion.
Authenticity proposition: Research cases and rural engineering scenarios will support transfer when they preserve meaningful constraints, evidence, and uncertainty.
Bidirectionality proposition: The research–teaching nexus will be sustainable when research outputs inform teaching and aggregated learning evidence informs iterative redesign.
Verification proposition: AI literacy will develop when students must document, test, and revise AI-supported work rather than merely use a tool.
Evidence proposition: Claims about effectiveness require validated measures, a specified comparison condition, baseline adjustment, mechanism evidence, and transparent reporting.
4. Materials and Methods
4.1. Study Design, Participants, and Instructional Conditions
Anonymized institutional records of students enrolled in the course were used in this study. A total of 123 student records were included in the analysis, including 68 in the intervention group and 55 in the comparison group. Students’ names and identification numbers were not used in the analysis. The Building Materials course comprised 32 teaching hours, with each teaching hour defined as 45 min of classroom time. Early course-orientation and diagnostic activities were used for instructional planning, but no participant-level pre-intervention measure from those activities was retained in the archival dataset for research analysis; consequently, no baseline-adjusted or pre–post analysis was possible. The content of the 32 h quasi-experimental implementation plan is shown in
Figure 4. The course-reform model described above guided the intervention, while the empirical evaluation relied on archived end-of-course records. The same course title and lead instructor were retained across the two cohorts. Intervention and comparison conditions, together with data-provenance controls, are summarized in
Table 2.
All participants were third-year undergraduate students majoring in Civil Engineering at the same Chinese university. Students were randomly allocated to classes by the university as part of routine administration; given this administrative allocation and the absence of sex-based placement, sex was not treated as a planned covariate, although the exact male/female balance cannot be verified from the archival dataset. Exact ages and individual prior academic achievement measures relevant to the course were not retained. However, all participants were at the same third-year stage within the same university and major, so age and prior disciplinary preparation were expected to be broadly comparable at the cohort level, while acknowledging that this comparability was not directly measured. Because Building Materials is a discipline-specific course offered in the third year, students had relatively limited prior exposure to AI applications specifically involving Building Materials concepts and problems before taking the course. Accordingly, the relevant course-entry limitation concerns domain-specific AI exposure rather than general familiarity with AI. No participant-level baseline measure of AI literacy or AI-supported Building Materials problem-solving capability was retained.
The records were drawn from autumn-semester cohorts in 2024 and 2025. The analytic dataset contained 68 records designated as intervention and 55 designated as comparison. To avoid conflating calendar year with instructional condition, the analysis is reported primarily by condition; calendar-year labels are retained only where the source institutional report itself used them. For the knowledge-test outcome, the same final-examination paper was administered to the two cohorts, with identical content, question structure, score allocation, and a total score of 100 points. This removes examination-form variation as an explanation for the score difference, although it does not remove cohort effects or other cross-year threats to validity. Five student-level variables were available: final-examination scores, experimental grades, usual grades, course-satisfaction scores, and project innovation scores. The project innovation score was an additional research indicator and was not part of the formal course grade.
4.2. Methods
A combined approach of curriculum-model development and retrospective cohort comparison was adopted. The course-reform design was informed by DBR principles and implemented in the authentic teaching context, while de-identified end-of-course records were used for outcome evaluation. AI was embedded within disciplinary learning and assessment rather than introduced as an independent technological component. Because the archived research materials did not include a complete design log, systematic version history, or independent fidelity dataset, the study does not claim to have empirically evaluated a full DBR cycle.
Course development drew on needs analysis, theoretical integration, prototype design, pilot enactment, and course-level adjustment. Attention was given to limited links with rural construction and green materials, insufficient use of research outputs as learning resources, and weak coordination among AI use, course objectives, assessment, and ethical requirements. Constructive alignment, TPACK, AI literacy research, and the research–teaching nexus were used to align learning outcomes, teaching activities, assessment tasks, disciplinary content, and technology use. AI-supported activities were permitted only when a disciplinary purpose was defined, and students were expected to verify information, judge physical plausibility, identify uncertainty, protect privacy, and disclose AI use where appropriate.
Implementation followed the One Core, Two Channels, and Four-Dimensional Drivers framework described in
Section 3. Research outputs were converted into scaffolded teaching resources, selected AI-supported tasks were attached to specific disciplinary purposes, and classroom, laboratory, project, and reflective activities were coordinated with assessment. Aggregated and anonymized learning evidence was also intended to inform course improvement. These elements describe the intervention architecture; the current administrative dataset does not independently verify the fidelity or mechanism of every framework component.
The model was preliminarily evaluated using records from 123 students, including 68 in the intervention group and 55 in the comparison group. Although students were randomly allocated to classes by the university, assignment to the intervention versus comparison instructional condition was cohort/year-based rather than individually randomized. Because random assignment was not used, the comparison group was treated as an observational reference, and no pre–post effect was estimated.
Five outcome domains were analyzed. Knowledge performance was represented by a 100-point final examination. Practical skills were represented by a 100-point experimental grade covering preparation, operation, data collection, analysis, engineering judgment, low-carbon material application, and reflection. Learning-process and AI-supported problem-solving performance was measured by the 100-point usual grade, comprising routine learning (20 points), course-task completion (30 points), and AI literacy problem solving (50 points). The AI component assessed five areas: problem identification, AI-tool use and prompt refinement, output verification, independent reasoning, and ethical AI use. As the largest and directly scored component, it provides curriculum-specific evidence of students’ AI-related capability and AI-supported problem solving. Course acceptability was represented by the 70-point student teaching-evaluation score and is interpreted as an affective/implementation outcome rather than a direct learning outcome. Project innovation was represented by an independent 100-point indicator covering topic relevance, professional knowledge integration, innovative thinking, research process, outcome transformation, and academic norms. Because non-submission could be recorded as zero, innovation scores were interpreted primarily in terms of documented participation rather than as a stable continuous cohort outcome. Descriptive statistics were reported for all outcomes. Welch’s t tests, 95% confidence intervals, and Hedges’ g were used for the four broadly observed continuous outcomes; the project innovation indicator was treated as sparse descriptive evidence.
4.3. Statistical Analysis and Computing Environment
Student-level records were summarized using sample size, mean, standard deviation, median, range, and, where informative, score-band counts and percentages. For the four broadly observed continuous outcomes, two-sided Welch independent-sample t tests were used to compare cohort means because group sizes differed and Welch’s procedure does not require equal variances. Mean differences are reported with 95% confidence intervals, and standardized between-cohort differences are reported as Hedges’ g. Statistical significance was evaluated at α = 0.05, while interpretation emphasizes effect magnitude and confidence intervals rather than p values alone. The highly zero-inflated project innovation indicator was summarized descriptively and was not treated as a stable continuous inferential outcome.
Data management and numerical analysis were conducted on the local Java-based teaching-analysis platform. The runtime environment used Java JDK 1.8, with Eclipse as the development environment; Apache Tomcat 7.0 and MySQL 5.7/8.0 provided application-server and database support. The platform was operated on Windows 7/8/10 or macOS systems.
5. Results
5.1. Knowledge-Test Results: Final-Examination Scores
Final-examination scores were used as the direct indicator of course knowledge performance. The intervention group (n = 68) achieved a mean final-examination score of 93.60 ± 4.78, with a median of 95 and a range of 82–100. The comparison group (n = 55) achieved 85.60 ± 8.43, with a median of 88 and a range of 63–98. The observed mean difference was 8.00 points (95% CI, 5.46–10.54). Welch’s t test indicated a statistically significant between-cohort difference, t (81.30) = 6.27, p < 0.001, with a large standardized difference (Hedges’ g = 1.19).
The distribution of scores provides additional context. All students in both cohorts obtained a passing final-examination score (≥60). However, 53 of 68 students in the intervention group (77.9%) scored 90 or above, compared with 22 of 55 students in the comparison group (40.0%). The remaining 15 intervention-group students were in the 80–89 range. In the comparison group, 22 students (40.0%) were in the 80–89 range, seven (12.7%) in the 70–79 range, and four (7.3%) in the 60–69 range. Thus, the observed difference was concentrated not in the basic pass rate, which was 100% in both groups, but in the proportion reaching high levels of examination performance. Because the two cohorts answered the same paper, this redistribution toward higher score bands represents a direct difference in performance on the same assessed knowledge content.
5.2. Practical Skills: Experimental Grades
Experimental grades were used as the direct indicator of practical skills. The intervention group achieved a mean experimental grade of 82.93 ± 2.10, with a median of 83 and a range of 81–86. The comparison group achieved 78.25 ± 4.88, with a median of 77 and a range of 69–89. The observed mean difference was 4.67 points (95% CI, 3.26–6.08). Welch’s t test showed a statistically significant between-cohort difference, t (70.05) = 6.62, p < 0.001, with a large standardized difference (Hedges’ g = 1.28).
The score distribution further differentiates the cohorts. All 68 intervention-group students (100%) were in the 80–89 ‘good’ band. In the comparison group, 21 of 55 students (38.2%) were in the 80–89 band, 33 (60.0%) were in the 70–79 ‘medium’ band, and one (1.8%) was in the 60–69 passing band. Neither cohort contained a failing experimental score, and neither contained a score of 90 or above. Therefore, as with the final examination, the practical-skill difference reflects the level and concentration of performance rather than a difference in basic pass/fail attainment.
The practical-skill rubric clarifies what the experimental grade represents. The largest single component is engineering-application judgment, and the rubric explicitly requires students to move through a sequence of experimental testing, data acquisition, performance analysis, environmental/application linkage, and engineering judgment. The higher experimental grades in the intervention cohort are therefore consistent with stronger observed performance in an integrated laboratory-to-engineering reasoning process, although the nonrandomized design prevents a causal interpretation.
5.3. AI Literacy and Routine Participation: Usual Grades
The 50-point AI component was not based on the frequency of AI use. It scored five observable capabilities—problem identification and description, effective AI-tool use and prompt adjustment, analysis and verification of AI outputs, independent thinking and problem solving, and appropriate/ethical AI use—at 10 points each. Students were expected to demonstrate the complete process of identifying a problem, formulating an appropriate AI query, using AI as an aid, checking the output against disciplinary evidence, and independently forming a conclusion or solution. Thus, one-half of the usual-grade score directly represented demonstrated AI-related capability in disciplinary learning tasks. Accordingly, the measure is treated as direct, curriculum-specific evidence of AI-supported problem-solving capability.
The intervention cohort obtained a mean continuous-assessment score of 85.38 ± 4.45 (median 84.5; range 78–99), compared with 49.82 ± 15.06 (median 46; range 22–95) in the comparison cohort. The mean difference was 35.56 points (95% CI, 31.36–39.76), Welch’s t (61.63) = 16.93, p < 0.001, with Hedges’ g = 3.34. No intervention student scored below 70, whereas 43 of 55 comparison students (78.2%) scored below 60.
The results show that, under the same scoring framework, there are significant differences in learning-process performance between the two groups. More importantly for the present study, the usual-grade result provides the clearest curriculum-specific signal of AI-capability performance in the available dataset. The intervention cohort exceeded the comparison cohort by 35.56 points, with a very large standardized difference (Hedges’ g = 3.34), and the score distributions were strongly separated. Because AI literacy and AI-assisted problem solving accounted for 50% of the total usual grade—and because that 50% explicitly assessed five observable AI competencies rather than simple frequency of tool use—the between-cohort difference is directly relevant to the targeted AI-capability outcomes of the reform. The same instructors and the same formal scoring criteria were used across the two semesters, reducing differences attributable to scorers or nominal grading standards. Taken together, these features support interpreting the result as strong evidence that the intervention cohort demonstrated better performance in the intended AI-supported problem-solving process, together with better routine participation and task completion. The archived dataset does not retain the five AI-domain subscores separately, so the exact proportion of the 35.56-point total-score gap attributable only to the AI component cannot be numerically isolated; however, this limitation does not reduce the usual grade to a generic engagement measure, because AI capability was explicitly and heavily weighted in the scoring design.
5.4. Course Satisfaction: Student Teaching-Evaluation Results
Course satisfaction was assessed using the institution’s student teaching-evaluation records. In this evaluation, students rated the quality of the teaching and their classroom experience; the resulting scores therefore represent students’ evaluations of the course and instructor, rather than assessments of the students themselves or of their academic performance. The maximum possible questionnaire score was 70. The official institutional evaluation summary reported a mean score of 69.5/70 for the 2024 autumn cohort and 68.0/70 for the 2025 autumn cohort, corresponding to 99.3% and 97.1% of the maximum possible score, respectively.
The student-level course-satisfaction scores recorded in the grade sheets closely corroborated these official summaries. Students in the intervention group gave the course a mean satisfaction rating of 69.49 ± 1.60 (median 70, range 60–70), whereas students in the comparison group gave a mean rating of 67.85 ± 3.76 (median 70, range 50–70). The mean difference in students’ course ratings between the two groups was 1.63 points (95% CI, 0.55–2.71), Welch’s t (69.73) = 3.01, p = 0.004, Hedges’ g = 0.58. Therefore, this result reflects the students’ opinions on the teaching methods of the teachers and their overall course experience. It also indirectly indicates that the students in the intervention group prefer the AI-assisted classroom approach more.
The institutional teaching-evaluation report also provides category-level information about students’ evaluations of the instructor’s teaching. Both calendar-year cohorts reported very high mean ratings for teaching content and methods, routine teaching management, and teaching effectiveness, with some difference in students’ overall impression of the instructor. These category-level statistics are not recombined here because the institutional questionnaire applies its own aggregation procedure, and they are not used for additional condition-level inference. Overall, the student evaluation data indicate that the course and teaching were rated highly by students in both cohorts, with a modest descriptive advantage for the intervention condition and a pronounced ceiling effect.
5.5. Project Innovation Performance: Independent Innovation Scores
Project innovation performance was assessed using an additional 100-point indicator that was separated from the formal course grade. The rubric covered six dimensions: topic innovation and linkage to real problems (20 points), integrated application of professional knowledge (20 points), innovative thinking and viewpoint formation (25 points), research process and problem solving (15 points), innovation-outcome transformation (15 points), and academic norms (5 points). The indicator was intended to evaluate students’ ability to transform Building Materials knowledge into a paper or comparable project outcome. A published article by
Yue and Zhang (
2025) in
Minerals was used as artifact-level evidence for the recorded excellent innovation outcome.
At the cohort level, the data were highly sparse. In the intervention group, only one of 68 students received a nonzero innovation score, which was 100, while the remaining 67 records were 0. All 55 records in the comparison group were 0. When zero entries were included, the intervention-group mean was 1.47 ± 12.13, with a median of 0 and a range of 0–100, compared with 0.00 ± 0.00 in the comparison group. The mean difference was 1.47 points, with Welch’s t (67.00) = 1.00, p = 0.321, and Hedges’ g = 0.16. These statistics were reported for completeness but were not considered evidence of a stable group-level innovation effect.
The published article (
Yue & Zhang, 2025) nevertheless provides concrete evidence for the quality of the single 100-point outcome. It addresses the sustainable use of industrial solid waste in low-carbon cement and demonstrates professional knowledge integration, research analysis, problem solving, viewpoint formation, and scholarly outcome transformation. Because the rubric allows non-submission to be recorded as non-participation, zero scores should not be interpreted as equivalent to weak project performance.
5.6. Integrated Outcome Findings for RQ3
Taken together, the five-domain results provide a more complete picture of the course-reform evidence. The intervention cohort showed clear and broadly distributed advantages in knowledge-test performance, practical skills, and the usual-grade composite of AI literacy and routine participation, together with a smaller advantage in course satisfaction. Project innovation performance differed in a qualitatively different way: the intervention cohort contained one documented excellent innovation outcome, whereas the comparison cohort contained none, but the indicator was too sparse for a stable statistical comparison. The five domains therefore represent complementary levels of evidence-knowledge mastery, practical application, learning-process and AI-related performance, affective evaluation, and knowledge-to-innovation transformation—rather than interchangeable measures of a single learning outcome.
To present these differences more clearly, the principal cohort results were visualized in
Figure 5.
Figure 5a compares the four continuous outcomes on a common percentage-of-scale basis. Higher values were observed for the intervention cohort in all four domains, although the degree of separation varied. The largest raw difference was found in the usual-grade composite, whereas course satisfaction remained close to the upper limit of the scale in both cohorts.
The standardized differences shown in
Figure 5b confirm this pattern. The largest separation was observed for the usual grade (Hedges’
g = 3.34), followed by practical skills (
g = 1.28), knowledge-test performance (
g = 1.19), and course satisfaction (
g = 0.58). These values indicate the magnitude of between-cohort score differences but should not be interpreted as causal treatment effects because random assignment and baseline adjustment were not available.
Figure 5c further shows that the differences were reflected in score distributions. For the final examination, most intervention students were concentrated in the 90–100 band, whereas comparison students were distributed across the 60–100 range. Experimental grades were concentrated in the 80–89 band for the intervention cohort, while most comparison students were in the 70–79 band. The largest distributional contrast was observed for the usual grade: no intervention student scored below 70, whereas 78.2% of comparison students scored below 60.
Because the project innovation data were highly zero-inflated,
Figure 5d presents documented nonzero participation rather than the arithmetic mean. Only one of 68 intervention students (1.5%) had a nonzero innovation outcome, compared with none in the comparison group. This result demonstrates the presence of one high-level knowledge-to-project transformation outcome but does not support a stable cohort-level innovation effect.
Taken together, the visual evidence complements
Table 3 by showing three distinct features of the results: a broad upward shift in the four routinely observed outcomes, especially the process-oriented usual grade; clear redistribution toward higher grade bands in knowledge, practice, and routine/AI-related performance; and an innovation indicator that remains too sparse for reliable group-level inference. The figures are descriptive summaries of the supplied end-of-course records and do not remove the design limitations noted above, including cohort nonrandomization, absent baseline measures, and uncertain cross-semester measurement equivalence.
6. Discussion
6.1. Principal Findings and Alignment with the Research Questions
The findings address the three research questions at different levels of evidence. For RQ1, the framework integrates constructive alignment, TPACK, AI literacy research, the research–teaching nexus, DBR-informed design principles, and validity and governance considerations; these are design resources rather than experimentally demonstrated necessary conditions. For RQ2, the framework operationalizes these principles through a responsible problem-solving core, two governed research–teaching channels, and coordinated changes to content, pedagogy, assessment, and the educational environment. For RQ3, the archival cohort comparison showed higher intervention-cohort scores on four routinely observed outcomes, but the evidential strength of those differences varied substantially by measure.
New Agricultural Science contributes an authentic application context by foregrounding rural infrastructure, agricultural-residue variability, green construction, life-cycle reasoning, and local resource constraints. The present study demonstrates how these concerns can be incorporated into course design, but it does not independently measure New Agricultural Science competence or transfer to rural engineering decision making. Claims about that dimension should therefore remain curricular rather than outcome-based.
6.2. Interpretation of the Cohort Outcomes
The common final examination provides the strongest basis for cross-cohort comparison because assessment content and scoring scale were held constant. The large difference in experimental grades is also consistent with stronger recorded practical performance, although the measure remains dependent on local rubric implementation. The continuous-assessment score showed by far the largest separation, but this result is partly structurally linked to the intervention because AI-related tasks constituted a substantial component of the score and were not equivalently embedded in the comparison condition. It therefore provides evidence of a difference in the recorded learning process under the two instructional conditions, not a clean estimate of an AI literacy treatment effect.
6.3. Implications for AI-Enabled Engineering Education
RQ3 reveals a coherent multidimensional pattern suggesting that the reform was associated with development beyond simple score improvement. The intervention cohort performed more strongly in knowledge mastery, practical performance, continuous-assessment performance, and course satisfaction, with the largest between-cohort difference observed in the continuous-assessment composite. Improved examination performance is consistent with stronger command of material properties, mechanisms, and engineering principles, while higher practical performance suggests a greater ability to connect experimental measurement, data interpretation, material-performance evaluation, and engineering decision making. The continuous-assessment pattern further indicates stronger recorded engagement in course tasks, evidence retrieval, verification, critical evaluation, and responsible AI-supported problem solving. Taken together, these outcomes suggest a shift from predominantly content-oriented learning toward a more integrated form of disciplinary knowledge, experimental competence, engineering reasoning, and evidence-based professional judgment.
Three broader implications follow from this pattern. First, AI integration is most educationally defensible when each use serves a clearly defined disciplinary purpose, requires verification, and is made visible in assessment rather than being treated as unrestricted tool access. Second, research-informed teaching requires transformation and scaffolding: research materials become educationally useful when students are supported in connecting evidence, uncertainty, standards, and engineering decisions at an appropriate level. Third, authentic rural and green-construction contexts can provide meaningful constraints for Building Materials learning without displacing the underlying material-science sequence. These implications are consistent with the intended mechanism of the framework: authentic engineering problems may create a demand for evidence; laboratory and analytical tasks may strengthen the transition from material testing to application decisions; structured AI use may support information retrieval, comparison, and critique; and verification requirements may encourage students to identify uncertainty rather than accept generated information uncritically. However, this mechanism remains plausible rather than demonstrated because the study includes no mediation analysis, mechanism-specific pre–post measures, or independent implementation-fidelity dataset. The framework should therefore be understood as a testable intervention logic rather than a confirmed causal mechanism. The documented publication provides an illustrative example of how course knowledge may be further transformed into a scholarly output, but the sparse project-participation data do not support a stable cohort-level innovation effect.
6.4. Limitations and Future Research
The main limitations include nonrandomized cohort allocation, absence of participant-level baseline measures, reliance on administrative records, and unequal measurement quality across outcomes. Consistency in non-examination scoring remains a particular concern for the usual-grade composite, and the intervention and comparison conditions did not provide identical structured AI-learning opportunities. The project innovation indicator was highly sparse, with zero values potentially representing non-participation rather than low performance. In addition, the current study does not independently measure New Agricultural Science competence or the fidelity of the two-channel and ecosystem mechanisms. These limitations weaken causal inference and constrain mechanism claims.
Vickers and Altman (
2001) and the Standards for Educational and Psychological Testing (
American Educational Research Association et al., 2014) support stronger baseline handling and measurement comparability; future studies should therefore collect common baseline and final measures, retain participant-level distributions, document rubric implementation and rater consistency, and use common cross-condition performance tasks. For AI literacy specifically,
Koch et al. (
2024) provide a recent review and factor-analytic evaluation of available AI literacy scales, supporting the use of validated standalone instruments rather than course-specific composites. They should also include direct scenario-based measures of New Agricultural Science-related engineering judgment, such as rural infrastructure, agricultural-residue utilization, green-material selection, and life-cycle decision making. To evaluate the proposed mechanism rather than only outcomes, future research should preserve design logs, version histories, fidelity measures, structured observations, and artifacts, consistent with the iterative documentation emphasis of DBR (
Design-Based Research Collective, 2003;
Wang & Hannafin, 2005). If interviews or open-ended artifacts are analyzed qualitatively,
Braun and Clarke (
2006) provide a well-established thematic-analysis approach. For project innovation, datasets should distinguish “not participated” from scored performance, provide comparable project opportunities to both cohorts, and report participation and performance separately.
AI tools and policies change rapidly. The framework therefore specifies functions and principles rather than dependence on one proprietary platform.
Des Jarlais et al. (
2004) recommend transparent reporting for nonrandomized evaluations; future reporting should therefore document tool versions, permitted uses, privacy procedures, deviations from the instructional protocol, and any changes in assessment implementation across cohorts. Confirmatory studies should also use multiple independent classes or cohorts per condition where feasible and apply baseline-adjusted, cluster-aware analyses appropriate to the sampling structure.
7. Conclusions
This study developed and evaluated the “One Core, Two Channels, and Four-Dimensional Drivers” model to align AI, research-informed teaching, and New Agricultural Science within an undergraduate Building Materials course. Three findings answer the research questions.
First, coherent reform was found to require the integrated use of constructive alignment, TPACK, AI literacy, the research–teaching nexus, DBR, and explicit validity and ethical principles. AI was therefore positioned not as an independent tool, but as a means of supporting responsible, evidence-based engineering problem solving.
Second, these principles were operationalized through a responsible problem-solving core, two governed research–teaching channels, and coordinated reform of content, pedagogy, assessment, and the educational ecosystem. Research outputs were transformed into verified and scaffolded teaching resources, while aggregated learning evidence was used to refine course design.
Third, the descriptive evaluation based on the supplied records was organized into five complementary outcome domains. The intervention cohort had higher end-of-course means in final-examination knowledge performance (93.60 vs. 85.60), experimental skills (82.93 vs. 78.25), the usual-grade composite reflecting AI literacy and routine participation (85.38 vs. 49.82), and course satisfaction (69.49 vs. 67.85 on a 70-point scale). Official student teaching-evaluation summaries were consistent with the satisfaction data (69.5/70 vs. 68.0/70). For project innovation, one intervention-group student had a documented score of 100, whereas no comparison-group student had a nonzero score.
This study is limited by nonrandomized grouping, absent baseline measures, uncertain cross-semester measurement equivalence, and unequal evidential strength across outcomes. Findings should therefore be viewed as preliminary. Future research should standardize assessments and rubrics, collect baseline and final measures, monitor implementation fidelity, account for cohort effects, and ensure comparable project innovation participation and opportunities.