1. Introduction
The rapid advancement of large language models (LLMs) and generative artificial intelligence (Gen-AI) has profoundly reshaped the landscape of education [
1]. From personalized tutoring and automated feedback to intelligent resource generation and collaborative problem-solving, it can be found that AI-augmented systems are increasingly integrated into teaching and learning environments [
2]. The transformation, however, is not merely technological; alternatively, it alters the nature of human–machine interaction, re-distributes cognitive labor between learners and AI agents, and raises critical questions about trust, autonomy, ethics, and system design [
3,
4].
Recognizing the urgency of these issues, the journal Systems launched a Special Issue entitled “AI-Augmented Human–Machine Systems: Engineering and Design in Education”. The Special Issue aimed to bring together empirical, theoretical and systems-oriented research methods and applications that address how AI tools, especially LLMs and generative models, can be engineered and designed to enhance educational outcomes while mitigating risks such as over-reliance, trust erosion and ethical misalignment [
5,
6]. To date, the Special Issue has published nine full-length articles, covering diverse educational levels including K-12, vocational, and higher education; methodological approaches including quasi-experiments, system development, surveys, grounded theory, and systematic reviews; and thematic foci like trust formation, feedback systems, resource generation, ethical governance, and pedagogical design [
7,
8,
9,
10].
The article provides a concise thematic review of these nine contributions. It synthesizes key findings, identifies cross-cutting themes and emerging contradictions, and outlines directions for future research without the use of tables, following the narrative style of a thematic analysis.
2. An Overview of Published Articles
2.1. Trust, Performance Expectancy and the Dual-Dimension Model
Trust is widely recognized as a cornerstone of AI adoption, yet most prior research treated trust as a single dimension. Zhang et al. (contribution 1) challenge this view by proposing a dual-dimension framework rooted in interpersonal trust theory: system-like trust (AST), based on rational evaluation of functionality and reliability, and human-like trust (AHT), based on emotional bonds and ethical confidence. Using survey data from 466 Chinese university students and a hybrid PLS-SEM + ANN approach, they found that performance expectancy (PE) is the strongest driver of both trust dimensions, especially for subjective tasks, e.g., creative ideation rather than objective tasks e.g., data analysis. Notably, objective tasks did not significantly affect trust, suggesting that advanced LLMs may have overcome earlier algorithm aversion in subjective domains. Furthermore, AST strongly predicted AHT (β = 0.416), confirming that rational trust is the cognitive foundation for affective trust. Both trust dimensions positively influenced continuance usage intention. The ANN analysis revealed that PE is the most important predictor of AST, AST is the most important predictor of AHT, and facilitating conditions (FC) is the most important predictor of continued use. The study extends interpersonal trust theory to human–AI interaction and provides actionable design guidance to foster trust, prioritize performance expectancy and system reliability, sustain long-term engagement, and ensure robust facilitating conditions such as AI literacy and technical support.
2.2. Designing LLM-Based Learning Environments: Autonomy and Moderate Constraints
How should LLMs be integrated into project-based learning (PBL) to optimize learning outcomes? Yi et al. (contribution 2) addressed this question through a one-semester quasi-experiment (N = 120) with a 2 × 2 factorial design such as individual vs. shared LLM use and restricted vs. unrestricted interaction. Grounded in self-determination theory, the results were striking. Individual use significantly outperformed shared use on engagement, psychological needs satisfaction, higher-order cognitive interactions and project scores. Moreover, restricted interaction like moderate constraints served as a meta-cognitive scaffold, promoting deliberate planning and deeper processing, whereas unrestricted use encouraged superficial answer gathering. Importantly, individual autonomy did not undermine collaboration; it enhanced by improving the quality of individual contributions to group work. Students also developed robust critical verification habits in response to LLM hallucinations. This study identifies individual autonomy as the core mechanism and moderate constraint as a crucial design principle, offering an empirically supported framework for harnessing Gen-AI in engineering PBL.
2.3. Neuro-Symbolic Systems for Actionable Feedback
One persistent challenge in automated writing feedback is that LLM responses are often generic, while traditional essay scoring provides only holistic scores. Yang and Zhao (contribution 3) present ARGUS, a novel architecture that combines LLM semantic understanding with graph neural network (GNN) structural reasoning. The system comprises three modules, including (1) an LLM-based parser that converts essays into structured argument graphs (F1 = 90.4% for components, 86.1% for relations); (2) an R-GCN that identifies seven types of logical/structural flaws (macro F1 = 0.83); and (3) a conditional LLM that generates feedback explicitly targeting those flaws. Human evaluation showed that ARGUS feedback was rated significantly more specific, accurate, actionable and helpful than fine-tuned LLMs or zero-shot GPT-4. This work demonstrates that hybrid architectures combining neural and symbolic components can overcome the limitations of monolithic LLMs, providing a template for intelligent tutoring systems that deliver diagnostic, prescriptive feedback at scale.
2.4. Generative AI for Multimodal Resource Construction
Yu, Song, and Lu (contribution 4) apply Gen-AI to a specific pedagogical bottleneck distinguishing similar Chinese characters. They propose a semi-automated framework that prioritizes image illustrations, with future expansion to micro-videos, self-test questions and basic information. An experimental comparison with traditional text-based resources showed that multimodal resources significantly improved performance on simple character discrimination and increased motivation especially for micro-videos. However, they were not effective for non-homophones such as visually similar characters with different pronunciations, indicating boundary conditions. This study illustrates how Gen-AI can be harnessed for targeted resource creation while revealing that not all learning objectives benefit equally from multimodal augmentation.
2.5. Domain-Specific Fine-Tuning and Vocational Education
Two articles focus on vocational education. Huang et al. (contribution 5) fine-tuned Llama 3 on a specialized knowledge base of circuit principles and fault diagnosis. The system achieved 89.7% accuracy on validation tasks, a 12.4% improvement over baseline, and an F1 of 0.87 for domain adaptability, but showed clear limitations, a 78.4% mathematical accuracy (Cohen’s d = 1.15 vs. human experts) and latency of 6.3 s under peak loads, deteriorating beyond 50 concurrent users. These realistic boundary conditions are essential for deployment decisions.
Scott and Dwight (contribution 6) surveyed 60 vocational teachers in England. Most used AI tools infrequently (0–10 times/month) yet rated them highly useful (4/5). Resource and assessment creation dominated use, where as administrative applications were less common. The novel “choice architecture” framing highlights that GPTs implicitly guide teachers toward certain resources, shaping pedagogy in non-transparent ways. Qualitative concerns included quality, over-reliance and diminished professional agency. Together, these studies underscore the need for vocational AI systems that are reliable, transparent, and respectful of teacher expertise.
2.6. AI in K-12 Programming Education
Tang et al. (contribution 7) conducted a six-week quasi-experiment (N = 103 Grade 7 students) comparing LLM-assisted instruction with traditional instruction. The experimental group outperformed the control on programming performance, cognitive interest and programming self-efficacy. Qualitative interviews revealed mechanisms that enhanced self-directed learning, a sense of real-time human–machine interaction and exploratory learning behaviors. The study frames LLMs as co-participants in a human–AI learning system, arguing that the question is not AI or human, but how to orchestrate their complementary capabilities.
2.7. Ethical Governance and Systematic Review
Barthwal, Campbell, and Shrestha (contribution 8) employ grounded theory to examine privacy concerns across three stakeholder groups including young digital citizens emphasizing autonomy and agency, parents and educators prioritizing oversight and AI literacy, and AI professionals balancing ethics with performance. The resulting Privacy-Ethics Alignment in AI (PEA-AI) model conceptualizes privacy as a dynamic negotiation rather than a static rule set. Significant gaps in transparency and digital literacy underscore the need for inclusive, stakeholder-driven frameworks.
Finally, Martínez-Peláez et al. (contribution 9) provide a compliant systematic review of 22 studies on LLMs in higher education. They conclude that LLMs can transform teaching through active learning, personalized feedback in large classes, and applied problem-solving assessments. Their effects are transversal with potential to improve educational equity, workforce readiness and cross-disciplinary innovation.
3. Synthesis and Cross-Cutting Themes
Across the nine articles, several themes recur:
3.1. Multidimensional and Developmental Trust
System-like trust (cognitive) precedes and enables human-like trust (affective) (contribution 1). Performance expectancy and facilitating conditions are primary drivers. Task type moderates trust formation with subjective tasks now eliciting positive trust responses, a reversal of earlier algorithm aversion findings, likely due to improved LLM capabilities.
3.2. Design Choices of Profound Pedagogical Consequences
Individual autonomy and moderate constraints optimize PBL outcomes (contribution 2), but unbounded access may be counterproductive. Hybrid neuro-symbolic architectures outperform monolithic LLMs for feedback generation (contribution 3).
3.3. Domain-Specificity and Technical Limitations
Realistic benchmarking (78% math accuracy, latency issues) provides a model for responsible system evaluation (contribution 5). Not all tasks or character types benefit equally from Gen-AI (contribution 4).
3.4. Ethical Considerations Integral to System Design
The PEA-AI model emphasizes stakeholder negotiation over top-down rules (contribution 8). Concerns about professional agency (contribution 6) and appropriate trust calibration (contribution 1) require governance mechanisms, not just technical solutions.
4. Innovations
The nine articles collectively introduce several theoretical, methodological, and design-oriented innovations that advance the field beyond prior work.
Theoretically, the most significant innovation is the dual-dimension trust model extended from interpersonal trust to human–AI interaction (contribution 1).While previous technology acceptance models treated trust as unidimensional, the Special Issue empirically demonstrates that system-like trust (rational or functional) and human-like trust (affective or ethical) are distinct yet causally linked, and rational trust precedes and enables emotional trust. Moreover, the finding that subjective tasks now elicit positive trust responses challenges the long-standing algorithm aversion hypothesis, suggesting that modern LLMs have qualitatively changed human perceptions of AI capability in creative and opinion-based domains. Another theoretical contribution is identification of individual autonomy as the core mechanism and moderate constraint as a critical design principle in LLM-augmented PBL (contribution 2), a nuanced departure from either laissez-faire or highly restrictive approaches.
Methodologically, the Special Issue showcases hybrid and neuro-symbolic approaches that break from conventional single-method designs. The combination of PLS-SEM with artificial neural networks (contribution 1) overcomes SEM’s linearity assumption and quantifies predictor importance, offering a template for future educational technology research. More radically, ARGUS (contribution 3) integrates GNNs with LLMs in a neuro-symbolic architecture that explicitly models argument structure and logical flaws, producing feedback that is demonstrably more actionable than LLM-alone systems. This represents a shift from monolithic end-to-end LLM applications to modular, explainable hybrid intelligence systems.
Design-wise, several innovations directly inform system engineering. The moderate constraints principle (contribution 2) shows that deliberately limiting AI queries can promote meta-cognitive reflection, transforming LLMs from answer-delivery tools into cognitive scaffolds. The PEA-AI model (contribution 8) introduces a stakeholder-centric, dynamic negotiation framework for privacy governance—moving beyond static checklists to adaptive, participatory design. Transparent reporting of quantitative boundary conditions (e.g., 78.4% mathematical accuracy, latency thresholds) (contribution 5) sets a new standard for realistic evaluation, countering the prevailing tendency to report only positive results.
Finally, a semi-automated multimodal resource framework (contribution 4) demonstrates how Gen-AI can be harnessed for specific pedagogical bottlenecks while also identifying clear boundary conditions such as ineffectiveness for non-homophones. This granular, task-specific design logic stands in contrast to generic AI for education proposals. Collectively, these innovations provide both theoretical depth and actionable engineering principles for building effective, trustworthy, and ethically aligned AI-augmented educational systems.
5. Limitations and Future Directions
The Special Issue, while valuable, has limitations. Most studies are cross-sectional or short-term while longitudinal designs are needed to capture trust evolution and sustained effects. Geographic concentration predominantly comes from Chinese samples, which limits cross-cultural application, though UK, Canadian, and Latin American contributions partially mitigate this. Vocational education remains under-represented. Future research should prioritize: (1) longitudinal and dynamic modeling of trust; (2) cross-cultural comparative studies; (3) design-based research systematically varying constraint types and hybrid architectures; (4) integration of physiological and behavioral measures; (5) studies of under-resourced and non-dominant populations; and (6) systems-level institutional and policy analysis.
6. Conclusions
The Special Issue “AI-Augmented Human–Machine Systems: Engineering and Design in Education” offers a timely and rigorous collection of research. It demonstrates that trust is dual-dimensional, that design principles such as individual autonomy and moderate constraints matter, that neuro-symbolic hybrids can overcome LLM limitations, and that ethical governance requires stakeholder-driven frameworks. Collectively, these nine articles advance both theoretical understanding and practical guidance for integrating AI into education. They also make clear that the challenge is not whether to adopt AI, but how and under what design principles, with what constraints, and governed by which ethical structures. Systems thinking, integrating cognitive, pedagogical, technical, and ethical perspectives, remains the essential lens for navigating this transformation.
Author Contributions
Writing—original draft preparation, S.Z.; writing—review and editing, S.Z. and F.Z.; supervision, S.Z. and F.Z. All authors have read and agreed to the published version of the manuscript.
Funding
This research was funded by the Planning Fund Project for Humanities and Social Sciences of the Ministry of Education of China, grant no. 23YJA880084 (Development of ICT Digital Education in European Universities), and by the Fundamental Research Funds for the Central Universities in China, grant no. CUC25CGJ04 (Comparative Study of Digital Education).
Acknowledgments
We would like to take this opportunity to express our sincere gratitude to dedicated efforts of the outstanding authors, professional and rigorous reviewers, and the editorial team of Systems. Meanwhile, congratulations are extended to all authors.
Conflicts of Interest
The authors declare no conflicts of interest.
List of Contributions
Zhang, Y.; Guo, J.; Wang, Y.; Li, S.; Yang, Q.; Zhang, J.; Lu, Z. Understanding Trust and Willingness to Use GenAI Tools in Higher Education: A SEM-ANN Approach Based on the S-O-R Framework.
Systems 2025,
13, 855.
https://doi.org/10.3390/systems13100855.
Yi, X.; Feng, W.; He, Y.; Wang, F. How to Harness LLMs in Project-Based Learning: Empirical Evidence for Individual Autonomy and Moderate Constraints in Engineering Education.
Systems 2025,
13, 1112.
https://doi.org/10.3390/systems13121112.
Yang, L.; Zhao, S. ARGUS: A Neuro-Symbolic System Integrating GNNs and LLMs for Actionable Feedback on English Argumentative Writing.
Systems 2025,
13, 1079.
https://doi.org/10.3390/systems13121079.
Yu, J.; Song, J.; Lu, Y. Harnessing Generative Artificial Intelligence to Construct Multimodal Resources for Chinese Character Learning.
Systems 2025,
13, 692.
https://doi.org/10.3390/systems13080692.
Huang, Y.-C.; Tsai, H.-J.; Liang, H.-T.; Chen, B.-S.; Chu, T.-H.; Ho, W.-S.; Huang, W.-L.; Tseng, Y.-J. Development of an Automotive Electronics Internship Assistance System Using a Fine-Tuned Llama 3 Large Language Model.
Systems 2025,
13, 668.
https://doi.org/10.3390/systems13080668.
Tang, B.; Liang, J.; Hu, W.; Luo, H. Enhancing Programming Performance, Learning Interest, and Self-Efficacy: The Role of Large Language Models in Middle School Education.
Systems 2025,
13, 555.
https://doi.org/10.3390/systems13070555.
Martínez-Peláez, R.; Mena, L.J.; Toral-Cruz, H.; Ochoa-Brust, A.; Potes, A.G.; Flores, V.; Ostos, R.; Pacheco, J.C.R.; Félix, R.A.; Félix, V.G. Can Large Language Models Foster Critical Thinking, Teamwork, and Problem-Solving Skills in Higher Education? A Literature Review.
Systems 2025,
13, 1013.
https://doi.org/10.3390/systems13111013.
References
- Yang, L.; Zhao, S. ARGUS: A Neuro-Symbolic System Integrating GNNs and LLMs for Actionable Feedback on English Argumentative Writing. Systems 2025, 13, 1079. [Google Scholar] [CrossRef]
- Scott, H.; Dwight, A. GPTs and the Choice Architecture of Pedagogies in Vocational Education. Systems 2025, 13, 872. [Google Scholar] [CrossRef]
- Alfarwan, A. Generative AI Use in K-12 Education: A Systematic Review. Front. Educ. 2025, 10, 1647573. [Google Scholar] [CrossRef]
- Flores Romero, P.; Fung, K.N.N.; Rong, G.; Cowley, B.U. Structured human-LLM interaction design reveals exploration and exploitation dynamics in higher education content generation. npj Sci. Learn. 2025, 10, 40. [Google Scholar] [CrossRef] [PubMed]
- Guizani, S.; Mazhar, T.; Shahzad, T.; Ahmad, W.; Bibi, A.; Hamam, H. A Systematic Literature Review to Implement Large Language Model in Higher Education: Issues and Solutions. Discov. Educ. 2025, 4, 35. [Google Scholar] [CrossRef]
- Shi, Y.; Yu, K.; Dong, Y.; Chen, F. Large Language Models in Education: A Systematic Review of Empirical Applications, Benefits, and Challenges. Comput. Educ. Artif. Intell. 2025, 10, 100529. [Google Scholar] [CrossRef]
- Yi, X.; Feng, W.; He, Y.; Wang, F. How to Harness LLMs in Project-Based Learning: Empirical Evidence for Individual Autonomy and Moderate Constraints in Engineering Education. Systems 2025, 13, 1112. [Google Scholar] [CrossRef]
- Zhang, Y.; Guo, J.; Wang, Y.; Li, S.; Yang, Q.; Zhang, J.; Lu, Z. Understanding Trustand Willingness to Use GenAI Tools in Higher Education: ASEM-ANN Approach Based on the S-O-R Framework. Systems 2025, 13, 855. [Google Scholar] [CrossRef]
- Yu, J.; Song, J.; Lu, Y. Harnessing Generative Artificial Intelligence to Construct Multimodal Resources for Chinese Character Learning. Systems 2025, 13, 692. [Google Scholar] [CrossRef]
- Al-Ali, S.; Miles, R. Upskilling Teachers to Use Generative Artificial Intelligence: The TPTP Approach for Sustainable Teacher Support and Development. Australas. J. Educ. Technol. 2025, 41, 88–106. [Google Scholar] [CrossRef]
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |