Next Article in Journal
GNSS Spoofing Attacks and Countermeasures in Military Missile Systems: A Scoping Review
Previous Article in Journal
MindVoice: A Web-Based Multi-Dataset Speech Depression Screening System with Pipeline-Aware Grad-CAM Explainability
Previous Article in Special Issue
A Provider-Independent LLM Architecture for Adaptive Oracle SQL Tutoring: A Design Study
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Tool-Augmented Agentic AI Pipeline for Reliable Circuit-Analysis Tutoring with Local Language Models

1
Porto School of Engineering (ISEP), Polytechnic Institute of Porto, Rua Dr. António Bernardino de Almeida 431, 4249-015 Porto, Portugal
2
Faculty of Engineering, University of Porto (FEUP), Rua Dr. Roberto Frias, 4200-465 Porto, Portugal
3
INESC TEC—Institute for Systems and Computer Engineering, Technology and Science, 4200-465 Porto, Portugal
*
Author to whom correspondence should be addressed.
Computers 2026, 15(10), 647; https://doi.org/10.3390/computers15100647
Submission received: 5 August 2026 / Revised: 3 September 2026 / Accepted: 16 September 2026 / Published: 24 September 2026

Abstract

LLM-based systems used for tutoring can produce fluent and apparently plausible explanations, but they do not guarantee Technical Correctness, which is paramount in STEM fields such as Electrical Circuit Analysis. This study investigates how three LLM-based response systems—parametric model knowledge (Pure LLM), Retrieval-Augmented Generation (RAG) and Agentic AI—affect response quality in the field of Electrical Circuit Analysis. It also introduces the HALO Agentic AI Pipeline, a local neuro-symbolic architecture that integrates LLMs, retrieval from curated pedagogical resources and deterministic algorithms for circuit solving, and the production of verifiable artefacts (such as theoretical context, step-by-step analysis, visual representations, numerical values and pedagogical hints). We conducted a blind evaluation involving 135 students comparing Pure LLM, RAG and Agentic AI across five dimensions: Correctness, Clarity, Learning Value, Relevance and Style. The three conditions were implemented using two distinct free and open-weight LLMs: GPT-OSS:20B and Llama-3.3:70B. These were compared with ChatGPT, using GPT-5.5 Instant as a commercial benchmark. The study covered five prompts related to the Loop Current Method, seven configurations, and thirty-five anonymised responses. In addition, we conducted a technical evaluation of 72 responses generated for the 12 questions used in this study. Overall, the evaluation demonstrates that both enriched architectures (RAG and Agentic AI) were rated significantly higher than Pure LLM, with Agentic AI achieving the highest overall mean rating. Agentic AI was rated significantly higher than both alternatives in Perceived Correctness, while architecture-related differences were limited to circuit-specific prompts, where both enriched architectures outperformed Pure LLM overall, but only Agentic AI showed a significant advantage in Correctness. In the complementary technical evaluation, the Agentic AI system obtained the highest mean Technical Correctness score, compared to RAG (2.63) and Pure LLM (1.79), and only 1 of the 24 responses received a rating of 2 or lower. Additionally, for the prompts under evaluation, no statistically significant differences were detected between the highest-rated local configuration and the commercial reference, globally or across the five dimensions. Our findings shed light on the suitability of the different architectures for different tutoring tasks, from generic (merely conceptual or procedural) to circuit-specific questions. They also support the feasibility of local, tool-augmented architectures for developing tutoring applications for electrical engineering education.

1. Introduction

The use of Natural Language Processing (NLP) systems in education settings has gone from a futuristic mirage to a concrete reality and an increasingly complex challenge. Given the exponential growth of Large Language Model (LLM)-based applications, accessible using only a mobile phone, students are integrating these tools into their studying habits [1], although the extent of this use remains difficult to ascertain [2,3]. In Science, Technology, Engineering, and Mathematics (STEM) contexts and, more specifically, in Electrical Engineering (EE), this reality brings opportunities and serious risks simultaneously. On the one hand, LLMs are very capable of understanding students in their natural language, reformulating concepts, and responding in natural language in adaptive tones, which enables them to adapt explanations to the students’ level.
However, on the other hand, these models’ ability to present impressive textual fluency does not guarantee grounded, factual responses, or the Technical Correctness required in STEM fields. In problems involving deep reasoning, physical models, methodological choices, and numerical fidelity, a written answer might sound perfect, i.e., clear, coherent, convincing, and pedagogically useful, but in fact be deceptive and ultimately false or misleading.
Electrical Circuit Analysis constitutes a particularly relevant domain with which to delve into this problem. Circuit analysis involves much more than applying mathematical formulas; before performing any calculations, students must correctly formulate the problem by identifying relevant circuit patterns, interpreting schematic diagrams, recognizing relationships between the circuit elements, selecting the most appropriate analytical strategy, and interpreting numerical values. Additionally, they must reflect on the physical plausibility of the results and verify their consistency with expected circuit behaviour.
This means that acquiring competence in circuit analysis requires not only mathematical proficiency but also the conceptual understanding, modelling and reasoning needed to interpret the physical behaviour and develop critical thinking. These are foundational and transversal skills in engineering. The Loop Current Method (LCM) is used in this work as an empirical case study because it builds on a the research team’s established body of work, including pedagogical human-curated datasets, studies of student errors, and deterministic mechanisms previously developed and tested, such as U=RIsolve solvers. Simultaneously, it concentrates skills (and difficulties) in a way that is transversal to engineering problem solving, including the interpretation of technical representations, modelling, variable selection, equation formulation, and the physical interpretation of the results.
Previous research from our team [4] showed that, contrary to what one might intuitively expect, the most relevant student difficulties in Electrical Circuit Analysis are not related to mathematics. In our analysis of 260 final exams from the first semester of the first year of an Electrical and Computer Engineering (ECE) course, 71.5% of the students were unable to solve the LCM problem under analysis without making any mistakes. The most frequent errors were related to conceptual misunderstandings, such as circuit interpretation, topological circuit pattern identification, and, especially, miscounting independent Kirchhoff’s Voltage Law (KVL) equations and mishandling current sources in auxiliary meshes. This evidence suggests that the design of electrical-engineering-oriented pedagogical applications should address structural reasoning steps, focusing on the intermediate steps and their explanations.
Pure LLMs have natural limitations in this regard, which, at first glance, seem to make them an impractical solution. Even when capable of answering conceptual (generic) questions, their performance vanishes as problem complexity increases. Answering questions that involve never-before-seen data, the correlation of multiple concepts or specific circuits requires higher reasoning capabilities and multimodal analysis (for example, to solve a circuit based on its image). Retrieval-Augmented Generation (RAG) mitigates issues related to a model’s knowledge cut-off and/or grounding, but the model’s intrinsic characteristics, such as mathematical skills, multimodal capabilities or reasoning level, remain unchanged. Consequently, the gap addressed in this work lies in access to pedagogical documentation and the production of problem-specific evidence. Retrieval content can provide conceptual data and explain a circuit-analysis method but does not guarantee correct method application when instantiated to a specific circuit, nor does it guarantee that the topological interpretation, equations, values or representations are consistent.
In this work, we propose a tool-augmented Agentic Artificial Intelligence (AI) architecture for circuit-analysis tutoring. The approach combines Local LLMs, RAG, Simulation Program with Integrated Circuit Emphasis (SPICE) netlist normalisation, deterministic solvers, circuit rendering for visual aid, and complete pedagogical circuit solutions. We build over lightweight, free and open-source LLMs that can run locally in a commodity hardware as the fundamental premise of our research framework.
This work builds on U=RIsolve as a teaching and self-learning tool for Electrical Circuit Analysis [5,6] and the HALO research on the use of pedagogical human-curated data and local models to support engineering education learning activities [4,7]. The study compares three AI systems that employ different regimes of response generation: (i) Pure LLM, (ii) RAG and (iii) Agentic AI. All the systems are implemented with the use of local free and open-weight models. The objective is to investigate whether a Local LLM, when supported by verifiable artefacts and a pedagogical pipeline build on empirical evidence, can produce more reliable, transparent and useful responses, compared with isolated uses of foundational LLMs. An additional comparison was made against a top-tier commercial model to better frame the overall quality of our architectures.

Research Questions and Contributions

This study investigates the perceived quality of the responses generated by three LLM-based systems that progressively incorporate more contextual information/data: Pure LLM, RAG and Agentic AI. The analysis examines the global differences between the systems, the prompt type influence and the behaviour in confrontation with a State of the Art (SOTA) commercial model. The main student-based analysis is complemented by the technical evaluation of locally generated responses, while keeping the Perceived Correctness separate from the Correctness evaluation conducted by the authors.
The study was guided by the following Research Questions (RQs)s:
  • RQ1—Architecture-Level Response Quality: How do students evaluate the responses produced by the Pure LLM, RAG and Agentic AI systems, both overall and across the following dimensions: Perceived Correctness, Clarity, Learning Value, Relevance and Style?
  • RQ2—Prompt-Type Dependence: How does response quality vary when comparing conceptual prompts to circuit-based questions that require structured technical information?
  • RQ3—Local Versus Commercial Reference: What is the response quality difference between the proposed local systems and a more powerful top-tier commercial model?
Accordingly, the main contributions of this work are as follows:
  • We propose a local Agentic AI architecture conceived for tutoring tasks in engineering education, instantiated and assessed in the electrical-circuit-analysis domain. The architecture is leveraged by a Local LLMs and integrates the automatic retrieval of pedagogical documentation curated by experts, deterministic modules for solving electrical circuits using different methods, enhanced visual layers for facilitating the analysis comprehension, and the production of pedagogically structured artefacts.
  • We propose a tutoring workflow centred in verifiable artefacts. Unlike a standard chatbot, the system goes far beyond text generation based on parametric model knowledge, general-purpose theory document retrieval, and/or reasoning supported by on-the-fly Python code. Instead, the system’s final response is conditioned by verifiable intermediate results built by deterministic electrical circuit solvers, according to an elaborated graph.
  • We define a controlled comparison protocol between three modes of LLM-based system operation: System A—Pure LLM; System B—RAG; System C—Agentic AI. These conditions are tested for two reference Local LLMs. We also conducted a comparison with a top-tier commercial model to better frame the systems’ qualities.
  • We conducted a blind evaluation involving 135 students, covering 5 prompts and 7 generation configurations, establishing a total of 35 distinct responses. The responses assessment considers the Perceived Correctness, Clarity, Learning Value, Relevance and Style, comparing the different architectures globally, per dimension and between (merely textual) prompts, covering theoretical circuit-analysis topics and circuit-dependent prompts.
  • We present a complementary technical evaluation based on the 72 local responses produced for the original 12-question battery. This analysis allows us to compare the Perceived Correctness assessments conducted by the students with the Technical Correctness assessments, and examine the task types where the circuit-determinist-processed evidence adds value when compared to RAG.
  • We present an in-depth statistical analysis per prompt type, distinguishing conceptual from procedural tasks. This helped to examine the task’s performance dependency, and to assess the adequacy of the different systems.
In addition to these contributions, the article discusses the feasibility of employing lightweight, free and open-weight Local LLMs in the construction of pedagogical applications. This options contributes for a scenario less (or even not) dependent on commercial Application Programming Interface (API) and decentralised processing. This also facilitates institutional data integration and control over the data while opening space for building applications for autonomous learning, capable of evolving based on human-curated data gathered in educational contexts.

2. Background and Related Work

2.1. LLMs and RAG in Engineering Education

The usage of LLMs in educational settings rapidly evolved from a curiosity-alike phase to a real integration into the most relevant triad of stages in education: teach, learn and assess [8]. In engineering education, this (r)evolution is particularly relevant since, in our labs, we assist students consulting conversational LLM-based applications to obtain quick explanations, alternative examples, programming support, summarization of reports and pedagogical content, (surprisingly) performing computations (over recent years, we have noticed an increase in the use of web LLM-based applications as calculators; we discourage our students from using these applications, since the models that power them are not deterministic and therefore do not provide assuredly correct computation results; nevertheless, students tend to use them, since it is easier to copy–paste an image with a mathematical formula than type it into a calculator), and for assistance in troubleshooting-related tasks. Recent literature reviews show that LLMs have been explored in diverse engineering areas, providing valuable support on project and programming tasks, writing technical documents, and gaining formative feedback. However, doubts have been raised regarding the ways in which such uses of LLMs pose persistent risks to reliability, academic integrity, and privacy, while perpetuating inequality through imbalanced access to these resources among students, as well as potentially producing over-reliance on responses among users [9,10]. Empirical studies with undergraduate engineering students demonstrate that the adoption of these tools is now significant, but the lack of trust that people have in the precision/reliability of the generated content remains a critical issue. As such, students must engage in extra effort in developing strategies for content verification, which generally involves back-and-forth interactions [11] to reach satisfactory results.
The above-mentioned problems are particularly critical in the STEM fields, since a response is often written in a coherent manner but can be completely inaccurate, e.g., on numerical values, real world phenomena descriptions, or settled facts. Unlike general writing tasks, engineering problems require the articulation of multi-layer domains such as symbolic representation, physical models, algorithmic procedures, numerical computations and conceptual interpretation of the computed values. Thus, the pedagogical usefulness of an LLM depends not only on its linguistic skills, but also on its ability to support verifiable reasoning, respect technical conventions, identify problem constraints, and produce feedback aligned with learning objectives. This point of view is consistent with previous work on programming domains, where tools such as Codex, ChatGPT or GitHub Copilot demonstrated impressive skills in code generation, explaining and commenting upon code blocks, and in performing debugging tasks. Nevertheless, some doubts were raised regarding difficulties in assessment and authorship attribution and, more importantly for educational settings, the risk of over-reliance on these tools and a lack of comprehension of coding processes [12,13,14]. In microcontroller programming education, it was also demonstrated that the potential of NLP does not rely merely on coding ability but rather on the comprehension of programming architectures, comprehension of electronic and physical aspects such as peripherals control (I/O pots, timers, AC/DC converters, etc.), debouncing strategies, configuration sequences and debugging processes [15].
Programming is an instructive analogy for this work, since the generated code can be executed, tested, and debugged, transforming a textual response from the model into a verifiable artefact. Circuit analysis presents an analogous, although not identical, situation: the answers may include technical artefacts, such as netlists, systems of equations, or calculated quantities, although these cannot always be independently verified without the input data or the intermediate steps of the analysis. This distinction is pedagogically relevant, since the most common difficulties of students in circuit analysis are not limited to algebraic manipulation, but also involve circuit interpretation, identification of interconnection patterns, nodes, branches, and meshes, selection of variables, formulation of equations, and physical interpretation of results [4]. This allows them to better understand the dynamics of the electrical circuit and to create critical thinking.
A frequent response to mitigate the occurrence of factually incorrect outputs (LLM hallucinations) consists in approximating the foundation model’s knowledge to a specific area of domain knowledge, usually validated by experts. Given the prohibitive cost of retraining a whole model, the most common option is to perform a fine-tuning which intervenes in a fraction of the original model’s parameter count. This technique allows the model to adapt to new domains, to adopt specific writing styles (for example, explain as a specific professor would do), or to instruct the model to achieve specific goals. In this regard, the most significant technique is Low-Rank Adaptation (LORA) [16] (usually its more efficient version, Quantized Low-Rank Adaptation (QLORA) [17]), which reduces the Video Random Access Memory (VRAM) footprint by storing the parameters of the pre-trained model in a quantized, low-precision format, such as 4-bit NormalFloat, while maintaining a minimal accuracy loss. In the context of the present research [4], it was previously demonstrated that, when fine-tuned on carefully curated (by human experts), domain-specific documents, lightweight LLMs (Meta Llama 3.2 1B and 3.1 8B) perform within one rating point (on a five-point Likert scale) of the high-end commercial model (GPT 4.5), despite a substantial difference in parameter count (a single and affordable Graphics Processing Unit (GPU) and a sub-400-second training pipeline are sufficient to achieve the reported quality levels), reinforcing the viability of local models for specialized educational applications.
Other relevant alternative is RAG, which consists in combining the parametric knowledge of a foundational model with external knowledge sources retrieved dynamically [18]. The objective is to improve the factuality and traceability of the generated content on knowledge-dependent tasks [19]. A more recent work (Self-RAG [20]) introduces other mechanisms such as the sequential conjunction of retrieval, generate, critique and self-reflect, seeking to employ deeper control in the model’s generation by introducing an iterative process of reasoning. Nevertheless, RAG mainly offers means to improve documentary grounding, as observed in our previous research experiments on this regard [7]. Although important in the STEM fields, it may be insufficient when the answer depends on on-demand data production based on a technical query. For example, the system may be able to accurately explain all the steps of a circuit-analysis method but be incapable of solving it and might use the intermediate results to iteratively adapt the explanation to the user’s question. For these cases, more sophisticated systems are needed, since the answer requires the instantiation of the procedures listed in the steps, which means the model must interpret complex circuit topologies, execute various computations and interpret the results based on arbitrary decisions (such as the branch or loop currents). The most recent research regarding tool-augmented LLMs and Agentic AI addresses this limitation, deviating the LLM from the unique source of truth for a subsidiary role where it orchestrates the execution of various tools to achieve a desired goal. Although in a distinct field, works such as Toolformer [21], ReAct [22] and Reflexion [23] set the roots for the transition to systems that combine reasoning, action, and external feedback.
Recent works expand the RAG utilization to visual information. PixelRAG represents and retrieves web pages directly from its visual space, preserving elements of spacial disposition and formatting the elements that can be converted to text [24]. Other studies evaluate the controlled combination of textual, visual and graph-based evidence, showing that the usefulness of each modality is dependent on the question type, the retrieval, and the generator’s ability to interpret figures and tables [25]. Although these works are being used in different scopes, the same techniques are promising candidates for circuit data extraction and interpretation. In this domain, specific circuit schematic recognition and data extraction seek to preserve not only the components data but also their topological disposition, terminals and interconnection relationships, allowing us to build structured circuit representations [26,27,28]. These contributions are particularly relevant for a future layer in the input block of the HALO system, allowing us to submit circuit images and transform them in structured, text-based representations, which is necessary to trigger the deterministic processes. Nevertheless, the current study focus is different: it starts from an already-structured formal representation and studies the transformation of that data into technical and pedagogical evidence for response composition.

2.2. Local Models and Sustainability

In the NLP field, improving the systems’ performances usually involves the employment of a disproportionate increase in computational requirements, which brings pressure on the economical, institutional and environmental dimensions of the field. These systems imply enormous energy, (fresh) water, hardware, and other environmental and infrastructural costs [29,30,31]. This phenomenon reinforces the discussion regarding the usage of local models versus the commercial models that run in large-scale, high-end data centres. Top-tier commercial models offer high-quality and easy-to-integrate modules, typically as Model-as-a-Service (MaaS) systems, where users access pre-trained machine learning models without directly managing the underlying infrastructure. There are various pricing options, but usually access and use are token-based, in a pay-as-you-go pricing model. This approach has some liabilities, such as recurrent costs per token, high exposure to external APIs, vendor lock-in, exposure to pricing changes, low control over the underlying infrastructure and service availability, and dependence on the provider’s future product roadmap (including API changes, model updates, and service discontinuation). As an illustrative example, considering public API prices for the OpenAI flagship model GPT-5.5, a software application with 10,000 monthly interactions, containing 8000 input tokens and 1000 output tokens would have a cost of approximately USD 700 per month (80 M input tokens × 5 USD/M + 10 M output tokens × 30 USD/M) [32]. This is without counting the costs for tools, calls, storage, retries, utilization growth, or lengthier contexts. In this example, we see that each interaction is cheap (around 0.07 USD), but for educational applications with frequent utilization, lengthier answers or RAG-enriched answers are difficult to scale up if they are dependent on remote commercial models.
Alongside the commercial models’ growth, there is a whole ecosystem of free and open-weight resources being offered, such as Llama, Mistral, Qwen, Gemma, DeepSeek, and Falcon [33]. We can find these alternatives in hubs like Hugging Face [34], and their performance can be compared using dynamic and public platforms such as Chatbot Arena [35]. This ecosystem creates opportunities for reproducible research, local deployment, institution model adaptations and inference marginal cost reductions. Logically, we must acknowledge that, in this scenario, the infrastructure cost, including hardware, energy, maintenance, and technical support are transferred to the user/institution. Furthermore, the performance of these models can vary across tasks, especially when considering using it to build educational applications. As such, we cannot rely on general model rankings, which means special care must be taken for choosing the model that best fits each application task’s requirements (e.g., we can use a smaller model for language detection and a larger one for reasoning).
The cost discussion should not be limited to institutional/individual financial budgets. The literature regarding AI shows that large language models and their supporting infrastructures have significative environmental impacts associated with electrical energy consumption, carbon emissions, water consumption for cooling and energy production, hardware manufacturing pollution and end-of-life disposal, impacts on local flora and fauna in areas where critical materials are extracted, as well as the geographical locations of data centres. As such, the natural commitments of educational and open and public research with environmental issues must be aligned with the technological choices made in these contexts. Therefore, the scientific question is not whether a local model is better than a commercial version, but whether adequate pedagogical architectures can approximate the educational utility of smaller and open models on specialized tasks, simultaneously reducing economical, institutional and infrastructural dependencies.

2.3. Previous Work on Circuit Analysis

In the specific circuit-analysis domain, the literature presents different sets of tool families. On one hand, we have simulators that are effective in outputting the values of electrical quantities such as voltage, current and power. This approach has a tendency for privileging the final output over pedagogy. In this sense, general-purpose simulators function as computational support tools under a black-box logic: they receive a circuit representation and return the value of electrical quantities, but do not expose the subjacent analysis, including the decisions sequence, variable assignment, equations creation and results interpretations that a student must learn. On the other hand, we have intelligent tutoring systems that seek to overcome this limitation by precisely addressing the process of solving electrical circuits. They operate under a transparent-box logic: instead of simply presenting the results, they follow and evaluate the intermediary steps of the solution process [36,37]. The Circuit Tutor constitutes an interesting example of this approach, centring its action on step-based tutoring for linear circuit analysis, automatic problem generation, immediate feedback and progressive mastery of skills [38,39,40], and empirical evaluation of its impact on learning [41]. More recently, our investigation introduced the U=RIsolve ecosystem, that is driven by the same goals, proposing web-based tools for teaching and self-learning circuit analysis, with focus on efficient algorithmic implementation of circuit-analysis methods while presenting its solution in a rich, pedagogical manner [5,6].
More recently, studies on the use of LLMs in circuit analysis show that these models can support evaluation and feedback, but have limitations in diagram interpretation, topological analysis and coherent calculations, particularly without formal circuit representations or reference solutions [42,43]. AITEE, an Agentic tutor closely related to this work, combines hand-drawn circuit schematics, structure-based retrieval, SPICE simulation and Socratic dialogue [44]. The system shows that textual RAG may be insufficient in symbolic domains, using topological representations and SPICE simulation to improve its computational reliability, while pedagogical mediation is conducted through Socratic dialogue.
This pedagogical approach should be considered in light of Bloom’s Taxonomy and Cognitive Load Theory. In introductory circuit-analysis courses, students are still developing technical vocabulary, mental representations and fundamental procedures, so prolonged sequences of open-ended questions can impose a high cognitive load [44,45,46,47]. Inquiry-based strategies therefore tend to benefit from explicit scaffolding adapted to the student’s level of competence [8,48]. In this respect, HALO adopts a complementary approach to AITEE: it retains the support of deterministic tools, but uses the evidence produced to support direct and pedagogically structured explanations.

2.4. Tool-Augmented and Agentic AI Systems

The pedagogical and technical limitations identified above motivate the adoption of tool-augmented and/or Agentic systems [49]. This paradigm shift also brings the discussion to a classic AI debate: the opposition, and complementarity, between connectionist and symbolic approaches [50,51]. The first is associated with artificial neural networks, where the system learns by analysing the patterns in data and are particularly effective in tasks regarding patterns recognition, statistical generalization, distributed representation and natural language generation. The LLMs are framed under this approach: they are highly competent systems in the modelisation of linguistic regularities, capable of producing explanations adapted to a given context, reformulation of content, adjusting the detail level and well-suited for interacting with students in a natural manner. In this sense, they function as “Linguistic Maestros”, coordinating the syntax, semantic, context, style and communication intention.
Nevertheless, referred linguistic competence does not guarantee mathematical accuracy, physical consistence or the rigorous implementation of formal procedures (as required in this regard). In STEM domain tasks, such as those in Electrical Circuit Analysis, a pedagogically relevant answer must respect physical laws, topological relationships, and signal conventions, must produce correct numerical computations, and must provide means to understand the reasoning behind every decision. It is precisely in this point that the symbolic approach remains relevant, since it relies on strict rules, deterministic algorithms, formal knowledge representation and verifiable inference mechanisms. Historically, this approach is related to expert systems, in which domain knowledge is encoded in a structured form to support decisions or diagnoses. Even the initial systems of interaction between humans and machines in natural language, such as ELIZA, illustrate a practical implementation that shows that apparent intelligent behaviour can emerge from simple, rule-based mechanisms [52].
In the educational context, this distinction is particularly important, since each paradigm offers different advantages and limitations. On one hand, the connectionist component allows for a flexible dialogue, discursive adaptation, explanation generation, content summarization, and interaction customization. On the other hand, the symbolic component contributes with the rigid control of data, traceability, predictability, and validation and a guarantee of formal methodological alignment with step-based solutions. As such, instead of establishing that an LLM should substitute deterministic tools or expert systems, the present work starts from the hypothesis that pedagogical utility emerges from the combination of these two realms. The generative model is not treated as a unique source of truth; rather, it operates as a mediating operative and linguistic layer supported by a pedagogically aligned, structured source of knowledge, powered by deterministic engines, domain rules, verifiable artefacts and human-curated data.
The here-presented work stands out by exploring a complementary logic of Agentic transparent-box tutoring, in which the deterministic tools do not produce final outputs or verifiable numerical solutions that are equivalent to those produced by a Circuit Analysis Simulator such as SPICE. Although the simulation remains essential for the validation of the electrical quantities, it represents just a fraction of the pedagogical potential of a hybrid architecture. The deterministic part acts as producer of pedagogically enriched data, such as explanatory, structured, and reusable artefacts, used to both guide the LLM response and provide means for a student to know more about how the answer has been produced. These artefacts may include theoretical general information regarding the method under study, explanations customized to a concrete circuit, visual identification of components, nodes, branches and loops (including a catalogue of all available loops in the circuit), symbolic equations (according to the notation used in the schematic), a step-by-step equation solver, Python code to verify the solution, and explanation of sign (+ or −) interpretation on circuit currents. The deterministic engine does not act as a numerical oracle to confirm whether a solution is correct or not; rather, it acts as a fully autonomous mechanism to produce pedagogically injected evidence.
From this perspective, the proposed architecture assumes a neural–symbolic nature: it combines the linguistic flexibility of an LLMs with formal representations, pedagogical rules, topological artefacts and validation through deterministic tools. The gap explored in this work is more than a mere integration of an LLM, RAG and external tools. Rather, it introduces an Agentic AI architecture centred in the pedagogical value of transforming deterministic tools into producers of explanatory artefacts, capable of making the circuit-solving processes more transparent, verifiable and adjusted to students’ expertise/knowledge level. Its Agentic AI neural–symbolic approach transforms the textual generation of an explanation into evidence-supported pedagogical answers, simultaneously enhancing the base model’s output quality and providing a means for the student to better understand the content under study.

3. Materials and Methods

This section describes the system architecture and experimental design used to generate the evaluated responses. The objective is not only to present a novel software paradigm, but also to introduce a methodology that can be adapted to different engineering education contexts (e.g., electronics, electromechanical systems and power electronics) involving similar characteristics: problems that can be represented using formal structures, step-by-step solution strategies, intermediate verification procedures, and the need to explain the underlying methodology and provide meaning to numerical results. This study validates the proposed methodology through a representative task in Electrical Circuit Analysis, building on previous work involving human-curated data, a taxonomy of student errors, deterministic solvers and pedagogically consolidated artefacts [4].

3.1. LLM-Based System Architectures

The AI responses were generated in three distinct experimental conditions: System A—Pure LLM; System B—RAG; System C—Agentic AI. System C implements the complete runtime of the HALO Agentic AI Pipeline described in Section 3.3. Responses for System A and System B were generated in the same run using a comparative mode implemented within the same service. However, explicit restrictions were applied to the information available to the model. In System A, the model received only the question and, when applicable, a textual representation of the circuit, such as a netlist. In System B, the model received the same input, together with a human-curated pedagogical dataset covering Electrical Circuit Analysis theory (available in Markdown format at https://urisolve.pt/app/nlp/, accessed on 18 September 2026). System C received the same initial input as Systems A and B but processed it through the complete Agentic AI pipeline.
This methodology was designed to compare the performance of different base models under three response-generation regimes, supported by progressively increasing levels of external knowledge, evidence and methodological support: (i) parametric model knowledge; (ii) document-retrieval-based knowledge; and (iii) in-depth pedagogical support combined with deterministic computational components. This design enables a broader interpretation of the results by highlighting the contribution of each additional level of support to the observed response quality. The comparison should, however, be interpreted at the level of the complete architectures. The progression from Pure LLM, RAG and Agentic AI does not constitute an ablation of the individual pipeline modules.

3.2. Experimental Conditions and Response Generation

To assess the incremental value of the proposed approach, responses were generated under the three experimental conditions within the same run. For each prompt, all three conditions used the same question and, when applicable, the same formal textual representation of the circuit. The information and evidence made available to the LLM differed according to the system, as detailed in Table 1.
All systems responses have been generated by the AI-agent service in comparison mode to guarantee similar experimental conditions. They used exactly the same prompt (and circuit data, when applicable) and LLM. Each system was evaluated using both GPT-OSS:20B and Llama-3.3:70B, resulting in six independent response artefacts per prompt ( 3   systems × 2   models = 6   responses ).
The two models were chosen to align with the study objectives, considering three main criteria: the availability of the model weights for local deployment; the compatibility with the available computational infrastructure; complementary computational profiles. GPT-OSS:20B is a Mixture of Experts (MoE) reasoning model with 21 billion parameters, where only 3.6 billion are active per token, conceived for instruction following, tool use, structured output generation and Agentic workflows (https://openai.com/index/gpt-oss-model-card, accessed on 18 September 2026). Llama-3.3:70B is a significantly higher-level model that contains 70 billion parameters, is optimized for multi-lingual dialogue, and is aligned for following instructions (https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md, accessed on 18 September 2026).
Since the models also belong to different model families, their use across Systems A, B, and C enables an assessment of whether the observed architectural effects remain consistent across distinct base-model characteristics rather than being specific to a single model family. Therefore, the purpose of this selection was not to establish a general ranking of local language models, but rather to evaluate the proposed approaches under two distinct and practically relevant local-deployment profiles. Although GPT-OSS:20B contains substantially fewer total parameters, the number of parameters does not determine, per se, its performance in specific tasks. The eventual differences between the models should be interpreted in the framework of their deployed configurations, how they used the retrieved content, how they structured evidence, and their respective generation styles, rather than in terms of the general superiority of one model family over another.
Additionally, GPT-5.5 Instant was used as an external commercial reference. For this condition, we used the ChatGPT web application rather than the local AI-agent service. For circuit-dependent prompts, the prompt included both the circuit image and its formal textual representation to maximize comparability with System C.
The software implementation used to generate the local responses is publicly available in the project’s GitHub repository, https://github.com/urisolve/halo, (accessed on 18 September 2026).

3.3. HALO Agentic AI Pipeline

The HALO Agentic Pipeline is a modular architecture designed to transform a student’s question, optionally accompanied by formally structured circuit data (e.g., a netlist), into a tutorial response supported by pedagogical knowledge, specialized tools, deterministic artefacts and natural language composition. The pipeline does not rely on the LLM as the sole mechanism for generating the response. Instead, it combines conditional decisions, tool execution, validation, knowledge retrieval, artefact production and configurable calls to language models. Its Agentic nature arises not from the exclusive use of AI modules, but from the automatic orchestration of the student question (prompt), circuit data, rule-based routing with optional LLM-assisted preprocessing, deterministic tools and AI-based evidence verification and textual generation.
From an architectural point of view, this approach is consistent with the neuro-symbolic logic discussed earlier:
  • The deterministic (symbolic) component is responsible for the formalization, normalisation, circuit solving and evidence production;
  • The LLM-based (connectionist) component is used as a linguistic mediator, routing mechanism, response composer and, when applicable, assisting with the preprocessing of ambiguous or incomplete inputs.
Depending on the question, the LLMs may intervene in three levels of the proposed architecture:
  • Routing: evaluating the user question and selecting the most appropriate path through the pipeline, i.e., determining which modules should be activated and in what sequence.
  • Experimental flows: supporting the comparison of different response-generation systems, namely:
    • Pure LLM: using solely the LLM’s parametric knowledge;
    • RAG: augmenting the LLM’s parametric knowledge with information retrieved from a human-curated pedagogical dataset.
    • Agentic AI: combining deterministic and AI-based modules through an orchestrated process.
  • LLM composer: generating the final response.
From a functional point of view, the input is decomposed into two distinct processing lines:
  • Linguistic and tutorial line: corresponds to the student question and it is used to identify the scope of the tutorial response and guide the activation of evidence-generation nodes.
  • Circuit representation line: This corresponds to the circuit data associated with the query, when applicable. The pipeline currently supports the circuit representations considered appropriate for the scope of this study, which are converted, when necessary, into a common structured representations compatible with the deterministic processing components:
    • Netlist provided as inline text or as a file;
    • JavaScript Object Notation (JSON) representation of the circuit, compatible with the U=RIsolve Editor;
    • Quite Universal Circuit Simulator (QUCS) schematic file, which is automatically converted into a valid JSON representation.
    Furthermore, previously used circuit data can be reused as input and can also be automatically produced by the pipeline itself using the iGen module. Regardless of the initial input modality, the pipeline ultimately operates on a common structured circuit representation. Thus, future modalities, such as handwritten circuit diagrams, can be supported by introducing an appropriate translation module before the processing chain.
In the evaluated configuration, the input data on circuit-specific questions were supplied through structured representations. The netlist is a general format used in a variety of circuit simulators. While capturing the most relevant circuit information for software solvers, it does not preserve the topological circuit information. The QUCS schematic file and JSON format from the U=RIsolve Circuit Editor allow us to preserve this information. Nevertheless, both pertain to specific representations. A general input based on the visual recognition of circuit diagrams was not considered in the present study.
This separation prevents the question and circuit data from being treated as an undifferentiated block of text passed directly to the LLM. The two lines follow different paths to process and enrich the input independently before rejoining in the final phase, specifically through the evidence and context packs sent to the LLM.
Before further detailing the “nodes” present in the pipeline and the execution flow, Table 2 synthesizes the principal artefacts manipulated by the pipeline.
Figure 1 illustrates a macro-scale visualization of the pipeline, presenting the most important aspects, including input, routing, evidence sources, context construction, response composition and traceability.
The first block of the pipeline is responsible for input normalisation and for classifying the types of information submitted by the user. It enables the detection of the following types of input:
  • Question with no circuit data: a question with no circuit data available (not even by inference), usually related to theoretical/conceptual aspects.
  • Question with circuit data: a question accompanied with the circuit data model, in one of the following forms:
    –
    Generic netlist description: for example, a netlist written by the student (it can be incomplete—in this case, the system will try to repair it).
    –
    Netlist file: a formal netlist file generated by a software (e.g., as used in QUCS).
    –
    Schematic file: a file with a visual representation of the circuit. The pipeline accepts a QUCS schematic file or a U=RIsolve Editor JSON file.
    –
    Pedagogical document: a Markdown file containing pedagogical data from the circuit. Some systems can use it for RAG to complement LLM context.
The above-mentioned classification is used by the router to determine the modules that should be activated and which type of evidence should be produced to support the final response generation. Hence, a solely conceptual question may follow the path for pedagogical documents retrieval and textual composition, while a question containing circuit data activates the formalization chain, validation, solver and artefact generation.
The architecture includes exploratory routing components driven by the LLMs. Organized as a graph-based logic, these components can be used to detect the presence of a circuit in free text, propose a netlist from a transcription, and attempt to repair an invalid netlist. Because each candidate is subsequently revalidated deterministically, the outcome of this process may update the circuit classification and consequently alter the downstream modules and their execution sequence. Even in this scenario, all subsequent nodes are activated to the resulting execution path. If an error occurs, a fallback message is passed to the LLM’s response composer, for example, to request more specific circuit data from the student.
The architecture also includes an optional branch for LLM-assisted preprocessing. This branch is relevant when the input contains netlist data or a circuit description in free text. In these cases, the model supports circuit detection, candidate netlist construction and, when the circuit representation fails the initial validation, an attempt to repair the representation. This mechanism does not substitute the deterministic validation. The candidate netlist, including any repaired version, is always checked by the deterministic validation layer. If validation fails, the solver is not executed and the system proceeds directly to the evidence-pack stage, attaching a note indicating that the circuit validation has failed. This mechanism allows the LLM composer to identify and communicate the need for more complete or compatible circuit data rather than fabricating results.
When the input contains a valid circuit schematic, the pipeline does not pass the raw data directly to the subsequent nodes. The system first performs a deterministic extraction of the electrical connectivity. This process analyses the electrical components, terminals, wires and connection points, organising them according to the electrical nodes; an electrical node is defined as any point or continuous wire at which three or more electrical components (or branches) are electrically connected and therefore share the same voltage. This representation ensures compatibility with the U=RIsolve solver. This process is intentionally conservative because the resulting data are passed to deterministic modules that impose specific structural requirements. When successful, this phase produces a valid netlist, a mapping between the visual representation of interconnection points and electrical nodes, components metadata, labels for wires and nodes, and warnings about any limitations encountered on extraction.
The construction, validation and normalisation layer acts as a bridge between the structural representation of the circuit and the deterministic modules. The netlist either supplied directly in the prompt or constructed a posteriori is validated and converted to a normalised form. This stage includes the normalisation of node names, mass reference allocation (if needed), and component value normalisation (e.g., U = 10 V in some netlists is converted to E = 10 V), among other modifications needed to turn it compatible with solver modules. When this stage fails to comply with the needed standard, the system avoids its use in the evidence pack.
The deterministic analysis node integrates the U=RIsolve ecosystem modules including solvers for different analysis methods, such as Node Voltage Method (NVM), Branch Current Method (BCM) and LCM). The normalised netlist produced by the previous stage is then passed to the solver corresponding to the chosen analysis method.
Given that the objective of this study is to evaluate response-generation architectures rather than to compare different circuit-analysis methods, the experimental pipeline, evaluation tasks and all supporting resources were deliberately designed around a single analysis method. The Loop Current Method (LCM) was selected as the representative method because it is among those reported to present greater difficulty for students. This choice provides a well-defined scope for the present evaluation while keeping the architecture extensible to other circuit-analysis methods.
For the purposes of this study, the solver receives the normalised netlist and returns a complete pedagogically structured solution in Markdown format, containing the following:
  • Fundamental variables: counts of branches, nodes, and current sources.
  • Loop analysis: the counting of loops used in the analysis and its identification.
  • Equation system: the equation system used to solve the circuit and all steps of the symbolic solution (using the variable names present in the circuit).
  • Loop currents: the numerical values for each loop current.
  • Branch currents arbitration: the information regarding the branch currents direction.
  • Branch currents computation: the establishment of the equations for compute the branch currents based on the loop currents values and directions.
  • Solution tips: various interpretation tips on different stages of the solution.
This output, in Markdown format, is a valuable source of information; it both provides reliable numerical values and, more importantly in the pedagogical regard, it returns an in-depth representation of the solution process, similar to the steps a student would perform in a manuscript solution. This enables a detailed understanding of how the circuit-analysis method can be applied to a specific circuit. Subsequently, the system extracts relevant metadata for the tutorial explanation.
Additionally, when a schematic file is submitted, the pipeline activates the Renderer. This module uses the circuit’s structure and the solver data to produce informative visual artefacts, such as the circuit base with overlays for displaying nodes, branches, currents and loops. It produces individual representations—for example, one-node-at-a-time and complete representations—containing information regarding various aspects in the same illustration. Figure 2 shows examples of a single branch overlay, a currents overlay, a loop overlay, and an all-nodes overlay.
The Renderer aims to convert formal circuit structure (textual) data into visual representations that help students to correlate the output textual explanations with the visual marks in a circuit, in a manner consistent with how circuit analysis is taught in class.
Another important pedagogical component is the Pedagogical Markdown Builder (PMB) module. It aggregates the information produced by the deterministic processing chain into a single, pedagogically structured document in Markdown format. It includes relevant theory for the method, a visual representation of the circuit, the circuit netlist, table of components, nodes, branches and loops, as well as a step-by-step solution, including equations, numerical values, final results, interpretation notes, pedagogical tips, references and visual artefacts. The objective of this module is to transform the technical output of the deterministic modules into pedagogical artefacts that can be used by both humans and machines. The Markdown structure enables immediate rendering for careful human verification (by the end user or during debugging tasks) as well as indexing data for retrieval processes, before the content is passed to the LLM composer.
The RAG component acts as an auxiliary source of knowledge. While the deterministic modules produce specific evidence of the circuit under analysis, RAG allows for retrieving definitions, circuit-analysis method explanations, terminology, conventions, interpretation rules and fragments of the curated, human-verified dataset. This component is useful when the question requires theoretical contextualization or concepts comparison. Within the pipeline, the RAG does not replace the deterministic analysis. It can be used when no circuit data are provided or in combination with circuit artefacts to establish a more complete evidence base.
After all artefacts have been produced, the pipeline constructs a context plan and an evidence pack. At this stage, the student question and the processed data rejoin. The system builds a compact selection that best matches the user’s question by correlating the initial query with relevant portions of the produced artefacts. This is crucial to control the number of input tokens and to avoid LLM hallucinations or deviation from the user query, since it mitigates the probability of the model entering in erroneous chains of thought.
The final generative node is the LLM composer, which is responsible for converting technical facts, deterministic results and pedagogical materials into a less extensive yet comprehensive explanation for the user’s question. The model is instructed to rely solely on the available evidence and not to invent components, branches, loops, values, equations, Uniform Resource Locators (URLs) or file links. More information about this mechanism can be found in Appendix A. This constraints limit the model’s natural creative behaviour, positioning it primarily as a pedagogical and linguistic mediation layer, while structural and numerical correctness depends on the artefacts produced by deterministic tools.
The pipeline also includes mechanisms of validation and traceability. The principal verification is based on contract rules and the available context, designed to detect relevant omissions or errors, including missing references to selected artefacts (e.g., the path to an overlay file used to identify circuit currents), inadequate use of symbols or references to non-existent elements. In parallel, the system stores artefacts associated with the execution, including the original input data, the circuit-type classification, the extracted netlist and other relevant node outputs. This information supports the identification and analysis of the potential causes of erroneous responses.
The proposed Agentic AI architecture enables the same pipeline to be executed with distinct LLMs, which is essential for performance benchmarking. In the questions used in the study that incorporate circuits concrete data, the pipeline followed the logic detailed in the flowchart of Figure 1. As such, the final responses result from a technical sequence that transforms the initial circuit representation into validated formal, visual and pedagogical evidence before it is used for natural-language response generation.

3.4. Prompt Set and Tutoring Task Categories

The prompts covered both general conceptual reasoning and more structured circuit-analysis tasks. They were divided into the following two categories:
  • Generic prompts—questions addressing abstract circuit aspects through textual descriptions, without supplying an explicit circuit netlist.
  • Circuit-specific prompts—questions concerning a specific circuit represented by an explicit netlist.
The evaluation included two generic prompts and three circuit-specific prompts. The first set focused on the identification of loops and equations in circuits described textually, whereas the circuit-specific prompts required the interpretation of an explicit circuit netlist. This distinction made it possible to examine whether the additional technical evidence and processing provided by more structured architectures, such as RAG and Agentic AI, offered greater benefits when the task involved explicit technical artefacts.
Table 3 presents a concise description of the prompts used in the evaluation. The complete prompt statements, response-generation configurations and a summary of the questionnaire organisation are provided in Appendix B.

3.5. Student-Based Evaluation (Blind Study)

The student evaluation was designed to compare the three AI response systems—Pure LLM, RAG and Agentic AI—in the context of circuit-analysis learning tasks. These architectures were evaluated through responses generated for five prompts related to the LCM. The prompts were intentionally selected to represent two levels of task complexity: two generic conceptual or procedural prompts, without an explicit circuit netlist, and three circuit-specific prompts, which required the system to interpret structured technical information before producing the answer.
In the local experimental setup, the three architectures were implemented using two base models, resulting in six local response-generation configurations: GPT-OSS-20b–Pure LLM, GPT-OSS-20b–RAG, GPT-OSS-20b–Agentic AI, Llama 3.3-70B–Pure LLM, Llama 3.3-70B–RAG and Llama 3.3-70B–Agentic AI. In addition, ChatGPT, powered by GPT-5.5 Instant at the time of evaluation (to facilitate communication throughout the article, we may use ChatGPT and GPT-5.5 Instant interchangeably), was included as the commercial reference model. Thus, the evaluation comprised seven response-generation configurations in total, although the main statistical analysis focused primarily on the comparison between the three architectural approaches (Systems A, B and C).
The 5 prompts and the 7 response-generation configurations produced a total of 35 ( 7 × 5 ) different responses. To keep the evaluation task manageable for students, we did not require that every participant evaluate all 35 responses. Instead, the responses were distributed across five questionnaire forms. Each student answered only one form and evaluated seven responses each. The origins of each response were hidden from the students, so that the evaluation was blind with respect to the model and architecture that generated the answer.
A total of 135 students participated in the evaluation. The five questionnaire forms were distributed as follows: 26 students answered Form 1, 25 answered Form 2, 25 answered Form 3, 29 answered Form 4 and 30 answered Form 5. This arrangement ensured that all response-generation configurations were represented across the dataset while avoiding excessive workload for each participant.

3.6. Evaluation Metrics

Each response was evaluated using a five-point Likert scale, ranging from 1 (strongly disagree) to 5 (strongly agree). The questionnaire included five complementary evaluation dimensions (D1–D5). Table 4 summarises the five evaluation dimensions used in the questionnaire.

3.7. Statistical Analysis

The statistical analysis was conducted using IBM Statistical Package for the Social Sciences (SPSS) Statistics (version 32.0 (IBM Corp., Armonk, NY, USA)). The analysis combined descriptive statistics with non-parametric inferential tests in order to compare the students’ perceptions of the three AI response architectures: Pure LLM, RAG and Agentic AI.
The first stage of the analysis was descriptive. Mean ratings and standard deviations were computed for each architecture, both globally and by evaluation dimension.
Since each student evaluated multiple responses, the data were aggregated at the student level before inferential testing. Depending on the analysis, aggregation was performed by architecture, by architecture and evaluation dimension, or by architecture, prompt type and evaluation dimension. This procedure ensured that the statistical comparisons were based on student-level summaries rather than on multiple individual ratings from the same participant.
Architecture-related differences were assessed using Friedman tests. When significant differences were found, pairwise Wilcoxon signed-rank tests were used as post hoc comparisons, with Bonferroni correction used for multiple testing.
The analysis was organised around three complementary objectives. First, the study examined whether students perceived a global hierarchy among the three response architectures. Second, it investigated whether this pattern depended on the type of prompt, distinguishing between generic prompts and circuit-specific prompts. Third, the highest-rated local configuration according to student ratings was compared with the external commercial reference model, ChatGPT.
Although the evaluation included seven response-generation configurations in total, the main statistical analysis focused on the three architectural approaches. The six local configurations were used to study the effect of architecture and to identify the best-performing local configuration. The comparison with ChatGPT was then conducted separately, using the local configuration with the highest mean rating.

4. Performance Evaluation

4.1. Overall Comparative Performance

The first analysis examined whether students’ evaluations showed a global preference pattern across the three systems: Pure LLM, RAG, and Agentic AI. For this purpose, ratings were aggregated by student and architecture, producing one global mean score per student for each architectural condition.
The global descriptive results showed progressively higher mean ratings in students’ evaluations as the response-generation architecture incorporated additional sources of knowledge and processing support. Pure LLM obtained the lowest global mean score ( M = 3.982 , S D = 0.660 ), followed by RAG ( M = 4.117 , S D = 0.656 ), while Agentic AI obtained the highest mean rating ( M = 4.181 , S D = 0.611 ). This pattern suggests an overall preference for architectures that incorporate additional retrieval or reasoning mechanisms, as expected.
The Friedman test indicated statistically significant differences between the three architectures, χ 2 ( 2 ) = 24.054 , p < 0.001 . The mean ranks also followed the same ordering: Pure LLM obtained the lowest mean rank (1.70), followed by RAG (2.03), while Agentic AI obtained the highest mean rank (2.26).
The post hoc Wilcoxon signed-rank tests with Bonferroni correction showed that both enriched architectures were significantly better evaluated than Pure LLM. The difference between Pure LLM and RAG was statistically significant ( Z = − 3.388 , p = 0.001 ), as was the difference between Pure LLM and Agentic AI ( Z = − 4.121 , p < 0.001 ).
The pairwise difference between RAG and Agentic AI did not reach statistical significance after correction ( Z = − 1.792 , p = 0.073 ). This global result should be interpreted together with the dimension- and prompt-level analyses below, which identify more specific architecture-related effects. The complete pairwise Wilcoxon comparisons for the global architecture analysis are reported in Appendix D.
These results support a descriptive hierarchy in students’ evaluations, with Pure LLM receiving the lowest scores, RAG occupying an intermediate position, and Agentic AI obtaining the highest mean rating and mean rank. The most distinctive result associated with Agentic AI emerged in the dimension-level analysis, particularly in perceived Correctness.

4.2. Dimension-Level Comparison Between Architectures

The second analysis examined whether the differences between architectures were consistent across the five evaluation dimensions. For this purpose, ratings were aggregated by student, architecture and dimension.
The results show that the effect of architecture was not uniform across all dimensions. Statistically significant differences were found for D1 (Correctness), D2 (Clarity) and D5 (Style), but not for D3 (Learning Value) or D4 (Relevance). Table 5 summarises the mean ratings and Friedman test results for each dimension.
For D1, related to Perceived Correctness, Agentic AI obtained the highest mean rating ( M = 4.452 ), whereas Pure LLM and RAG obtained identical mean scores ( M = 4.267 ). The Friedman test indicated a statistically significant difference between architectures, χ 2 ( 2 ) = 13.050 , p = 0.001 . The post hoc Wilcoxon tests showed that Agentic AI was significantly better evaluated than Pure LLM and RAG. No difference was observed between Pure LLM and RAG.
For D2, related to Clarity, both enriched architectures obtained higher scores than Pure LLM. RAG achieved the highest mean rating ( M = 4.033 ), followed closely by Agentic AI ( M = 4.000 ), while Pure LLM obtained the lowest score ( M = 3.678 ). The Friedman test was statistically significant, χ 2 ( 2 ) = 25.010 , p = 0.000 . The post hoc comparisons indicated that both RAG and Agentic AI were significantly better evaluated than Pure LLM.
For D3, related to Learning Value, the mean scores were very close across architectures: Pure LLM ( M = 3.911 ), RAG ( M = 3.967 ) and Agentic AI ( M = 3.974 ). The Friedman test was not statistically significant, χ 2 ( 2 ) = 2.116 , p = 0.347 . Therefore, the data do not support architecture-related differences in students’ perception of how much they learned from the responses.
For D4, related to whether the response addressed what was asked, all three architectures obtained high and similar mean ratings. Pure LLM obtained M = 4.448 , RAG obtained M = 4.374 and Agentic AI obtained M = 4.444 . The Friedman test was not statistically significant, χ 2 ( 2 ) = 1.600 , p = 0.449 . This suggests that the students generally considered the responses from all architectures to be aligned with the prompt.
For D5, related to response Style, the pattern again favoured the enriched architectures. Pure LLM obtained the lowest mean rating ( M = 3.604 ), while RAG ( M = 3.944 ) and Agentic AI ( M = 4.033 ) obtained higher scores. The Friedman test indicated statistically significant differences, χ 2 ( 2 ) = 23.839 , p = 0.000 . The post hoc comparisons showed that both RAG and Agentic AI were significantly better evaluated than Pure LLM.
Overall, the dimension-level analysis indicates that both enriched architectures improved perceived Clarity and response Style relative to Pure LLM, while Agentic AI obtained a distinct advantage in Perceived Correctness.
By contrast, the three architectures were similarly evaluated in terms of perceived Learning Value and Relevance to the question asked. The absence of significant differences in D4 suggests that students generally considered the responses from all architectures to be aligned with the prompts. Likewise, the absence of differences in D3 should be interpreted in light of the students’ background: participants had already attended circuit-analysis course (1st semester) and were familiar with the LCM. Therefore, D3 may have reflected the perceived novelty or reinforcement value of the response rather than actual learning gains. This may explain why architectural differences were more visible in Correctness, Clarity and Style than in Relevance or perceived Learning Value. Detailed post hoc Wilcoxon comparisons by evaluation dimension are provided in Appendix D.

4.3. Effect of Prompt Type

The third analysis examined whether the effect of architecture depended on the type of prompt. For this purpose, the prompts were divided into two groups: generic prompts, which did not include an explicit circuit netlist, and circuit-specific prompts, which required the interpretation of structured circuit information.
Descriptive statistics were first computed separately for each prompt-type subgroup. As shown in Table 6, mean ratings were high across the three architectures in generic prompts, with only small descriptive differences between Pure LLM, RAG and Agentic AI. In contrast, in circuit-specific prompts, Pure LLM obtained a lower mean rating than both enriched architectures, suggesting that architecture-related differences became more visible when the prompt required the interpretation of structured circuit information.
For the subsequent Friedman tests, complete paired observations across the three architectures were required. Therefore, the valid inferential sample size differed from the descriptive N reported in Table 6: N = 26 for generic prompts and N = 110 for circuit-specific prompts.
In the generic prompts, no statistically significant differences were found between the three architectures in the global comparison, although the result was close to the conventional significance threshold, χ 2 ( 2 ) = 5.952 , p = 0.051 .
The mean ranks suggested a slight advantage for RAG and Agentic AI over Pure LLM, but this pattern was not statistically robust. Therefore, for generic conceptual or procedural prompts, the data do not provide sufficient evidence of a clear architecture-related difference.
By contrast, for circuit-specific prompts, the Friedman test indicated statistically significant differences between architectures, χ 2 ( 2 ) = 14.747 , p = 0.001 . In this condition, Pure LLM obtained the lowest mean rank, whereas Agentic AI obtained the highest. The post hoc Wilcoxon tests with Bonferroni correction showed that both RAG and Agentic AI were significantly better evaluated than Pure LLM. The difference between RAG and Agentic AI was not statistically significant.
These results suggest that the type of prompt influenced the extent to which architectural differences became visible in student evaluations. When the task consisted mainly of conceptual or procedural explanation, the three architectures were evaluated similarly. However, when the prompt included a circuit netlist and therefore required the interpretation of structured technical information, the enriched architectures were more clearly favoured.
A dimension-level analysis confirmed this pattern. In generic prompts, no significant differences between architectures were observed in any of the five evaluation dimensions. In circuit-specific prompts, significant differences were found in D1 (Correctness), D2 (Clarity) and D5 (Style), whereas D3 (Learning Value) and D4 (Relevance) remained non-significant.
Table 7 summarises the dimension-level results by prompt type.
In the circuit-specific prompts, the strongest effects were observed in Clarity and response Style. Both RAG and Agentic AI were significantly better evaluated than Pure LLM in these dimensions, suggesting that architectures incorporating retrieval or Agentic processing may produce responses that students perceive as clearer and better structured when the task involves explicit circuit representations.
For Correctness, Agentic AI obtained the highest mean rating and was significantly better evaluated than Pure LLM. Agentic AI was therefore the only enriched architecture to show a statistically supported advantage over Pure LLM in Perceived Correctness for circuit-specific prompts. This result is particularly relevant because circuit-specific prompts require the system to interpret circuit structure before generating an explanation. The advantage of Agentic AI over Pure LLM in this dimension suggests that additional reasoning and processing steps may contribute to greater perceived technical reliability in more structured tasks.
As in the previous analysis, no significant differences were observed in Learning Value or Relevance. The lack of differences in Relevance indicates that students generally considered all architectures capable of addressing the prompt. The lack of differences in perceived Learning Value may again be related to the students’ prior exposure to circuit-analysis content, which may have limited the perceived novelty of the responses.
Overall, the prompt-type analysis indicates that architectural differences were more evident when the task involved netlist interpretation. This supports the idea that enriched architectures add value particularly in tasks requiring the processing of structured technical artefacts, rather than in simpler conceptual prompts, where a Pure LLM may already provide an acceptable explanation. The post hoc comparisons for the prompt-type analyses are reported in Appendix E.

4.4. Selection of the Students’ Highest-Rated Local Configuration

Although the main objective of the statistical analysis was to compare the three response architectures, the evaluation also included two base models for each architecture. Therefore, before comparing the local system with the external commercial reference model, a descriptive analysis was conducted to identify the students’ highest-rated local configuration.
For this analysis, the data were aggregated by student and local configurations. Each aggregated score corresponded to the mean rating assigned by a student to a given model–architecture configuration across the five evaluation dimensions. This resulted in 135 student-level scores for each of the six local configurations.
Table 8 presents the global mean ratings and standard deviations for the six local configurations.
The highest mean rating was obtained by Llama 3.3-70B–Agentic AI ( M = 4.262 , S D = 0.716 ), followed by GPT-OSS-20b–RAG ( M = 4.164 , S D = 0.712 ) and GPT-OSS-20b–Agentic AI ( M = 4.099 , S D = 0.752 ). The lowest mean rating was obtained by GPT-OSS-20b–Pure LLM ( M = 3.893 , S D = 0.942 ).
These results indicate that the Llama 3.3-70B model combined with the Agentic AI architecture was the students’ highest-rated local configuration in terms of global student evaluation. For this reason, Llama 3.3-70B–Agentic AI was selected as the local reference configuration for the subsequent comparison with ChatGPT.
It should be noted that this analysis was used for selection purposes and should be interpreted descriptively. The purpose was not to establish a full statistical ranking among the six local configurations, but to identify the strongest local configuration to be compared with the external commercial model. The selection and subsequent comparison use the same student evaluation; as such, this analysis should be regarded as exploratory, not an independent validation. Differences in output Style, Clarity and response organization could equally contribute to the overall evaluation of each model.

4.5. Comparison with GPT-5.5 Instant

The final analysis compared the highest-rated local configuration, Llama 3.3-70B–Agentic AI, with the external commercial reference model, ChatGPT. This comparison was conducted separately from the architecture-level analysis, since ChatGPT was not part of the controlled local architecture set, and its internal architecture was not experimentally manipulated.
For this analysis, the data were aggregated by student and model, producing paired student-level scores for Llama 3.3-70B–Agentic AI and ChatGPT. Each global score corresponded to the mean rating assigned by each student across the five evaluation dimensions. Wilcoxon signed-rank tests were then used to compare the two models globally and across the five dimensions.
In the global comparison, Llama 3.3-70B–Agentic AI obtained a mean rating of 4.262 ( S D = 0.716 ), while ChatGPT obtained a mean rating of 4.222 ( S D = 0.798 ). The Wilcoxon signed-rank test indicated no statistically significant difference between the two models, Z = − 0.740 , p = 0.459 . Thus, although the local Agentic AI configuration showed a slightly higher descriptive mean, this difference was not statistically significant.
The same pattern was observed at the dimension level. No statistically significant differences were found in Perceived Correctness, Clarity, Learning Value, Relevance to the question, or response Style. Descriptively, Llama 3.3-70B–Agentic AI obtained slightly higher mean scores in Correctness, Learning Value and response Style, whereas ChatGPT obtained slightly higher mean scores in Clarity and Relevance. However, none of these differences reached statistical significance.
Table 9 summarises the global and dimension-level comparisons between Llama 3.3-70B–Agentic AI and ChatGPT.
Overall, the results indicate that the local Llama 3.3-70B–Agentic AI configuration was perceived by students as statistically comparable to ChatGPT in the evaluated circuit-analysis tasks. This finding is relevant because it suggests that a locally deployed Agentic architecture, supported by structured processing mechanisms, can achieve student-perceived response quality close to that of an external commercial reference model.
This result should not be interpreted as evidence of formal equivalence between the two systems, since no equivalence test was conducted. Rather, within the present evaluation design and sample, no statistically significant differences were detected between Llama 3.3-70B–Agentic AI and ChatGPT, either globally or in any of the five evaluation dimensions.

4.6. Complementary Criterion-Based Technical Evaluation

The evaluation conducted by the students characterizes the perceived quality of the responses, and, therefore, is particularly relevant to study its adequacy in an educational context. However, a clear, well-structured and linguistically convincing response does not necessarily correspond to a technically correct response. To complement this perspective, an additional technical evaluation was conducted, focusing on the conformity with the LCM method and adequacy as learning support. The two evaluation were kept as independent evaluations since they correspond to distinct constructs: the first characterizes student perception; the second verifies the responses content according to the technical criteria defined for the studied domain.
A total of 72 responses generated for the 12 questions prepared for this study were evaluated, considering the conditions Pure LLM, RAG, and Agentic AI, using the models GPT-OSS:20b and Llama-3.3:70b. The questions Q02, Q03, Q05, Q06, and Q09 correspond to the five prompts included in the student evaluation. The remaining seven questions broaden the technical coverage of the analysis, including topological aspects, equation formulation, sign conventions, current interpretation and generic conceptual doubts. The complete formulation of the 12 questions is presented in Appendix C.
Table 10 summarizes the main questions’ focuses and the type of competences addressed per question group.
The evaluation was conducted by the author team, which has teaching experience on this subjects, and knowledge of both the method and the architecture. A common five-level ordinal quality scale was used: 1—very poor; 2—poor; 3—acceptable; 4—good; 5—excellent. A total of eight complementary dimensions were considered (Table 11). Technical Correctness (TC) was used as main indicator of the technical profile of the responses, while the remaining dimensions allow us to characterize relevant response aspects in an educational setting.
The results are analysed descriptively to identify recurrent patterns between the architectures, models, and question types, as well as cases in which this patterns are weaker, diffuse or even contradictory.

4.6.1. Overall Architecture Profiles

Table 12 presents the aggregated profile for the three architectures. The most evident contrast occurs in the TC. The Pure LLM condition achieved a mean of 1.79, the RAG condition achieved a mean of 2.63, and the Agentic AI condition achieved a mean of 3.71. This progression is more expressive when comparing the ratings distribution. Only 1 of the 24 responses from Pure LLM achieved TC ≥ 4 , compared to 8 responses in RAG and 13 responses in Agentic AI. Conversely, 21 of the 24 responses from Pure LLM received TC ≤ 2 , a value that decreased to 14 in RAG and only 1 in Agentic AI. Thus, 23 of the 24 responses from the Agentic AI condition were rated as acceptable (TC = 3) at least, and none was rated as very poor (TC = 1).
The same pattern, while less expressive, occurred in the QA, SC and PR criteria. This suggests that the difference is not only limited to the Correctness of the final value: the Agentic AI condition tended to also produce responses which were more aligned with the request (prompt), with more consistent notation and a solution closer to what a student may do in a paper-and-pencil solution. On the other hand, CT present a different behaviour: RAG obtained the highest mean, 3.83, above Agentic AI, with 3.46. Providing deterministic evidence, intermediate results and a panoply of additional artefacts may increase the input context, which may lead the composer to present extra information and even mislead it to select unrelated electrical quantities. This suggests that the final composition stages may benefit from more restrictive guidelines. Thus, the more favourable technical profile of Agentic AI does not imply an automatic or guaranteed improvement over the other architectures in every dimension of tutoring quality.
Table 13 presents the TC ratings per question, model and architecture. This visual representation allows us to simultaneously observe the global trend, while showing cases where the advantage of the Agentic AI architecture is weakened, disappears, or even reverses (e.g., as visualized in questions Q08, Q11, and Q12).

4.6.2. Consistency of the Pattern Across Base Models

The aggregate analysis was complemented by a separate Technical Correctness analysis for both models’ responses, as presented in Table 14. The purpose of this separation is not to establish a general ranking between the models; rather, it is to verify whether the observed progression pattern depends exclusively on one of the employed models.
The same descriptive progression order emerges in both models: Pure LLM < RAG < Agentic AI. In GPT-OSS:20b, the Agentic AI condition surpassed RAG in 8 of the 12 questions and tied in the remaining 4. In Llama-3.3:70b, it outperformed RAG in 8 questions, tied in 2 and achieved an inferior rating in the remaining (Q08 and Q12). Regarding Pure LLM, the Agentic AI condition was higher in 11 of the 12 questions in each model and equal in the remaining question, with no observation of inferior performance.
The GPT-OSS:20b configuration achieve higher TC means than Llama-3.3:8b in the RAG and Agentic AI conditions, while the Pure LLM condition presented similar values in both models. This performance difference should be observed with precaution since we cannot generalize the superiority of one model over another. The models differ in their architecture, post-training (e.g., the tuning of the model for Agentic scenarios or for tool use), the knowledge cut-off, the instruction-following procedures, and the generation style. The here-observed difference is consistent with more efficient resource utilization, from content retrieval to the determinist evidence collected prior to the final generation of an output from the model.

4.6.3. Effect of Question Type

The analysis by question group helps us to understand where the Agentic architecture adds more value. Table 15 presets the mean TC values per category. Agentic AI presents the highest mean in the first three groups and ties with RAG in the Q11–Q12 group. The magnitudes of the differences were not uniform.
The largest advantage emerged in the Q05-06 group, which requires us to distinguish and select loops based on the circuit structure. In these questions, the response benefits directly from the topological formalization and from the artefacts produced by the deterministic pipeline. A similar pattern occurs in Q09 and Q10. The former requires the articulation of auxiliary and principal loops, equation count, and the reconstruction of branch currents; the TC mean was 1.00 for Pure LLM, 2.50 for RAG and 4.00 for Agentic AI. In Q10, which requires a branch current and interpretation of its (mathematical) sign, the means were 1.00, 1.50 and 4.00, respectively. This is the type of situation where the functional difference between document retrieval and the production of circuit-specific evidence becomes clearer. Retrieved theory can correctly explain a methodological procedure, yet it does not substitute the need to instantiate this knowledge in a given circuit; this might be performed by interpreting the circuit patterns and topological structure, performing computations to obtain electrical quantities, and ensuring coherence is maintained in the solution (e.g., taking the current flow direction into consideration).
The advantage is not, however, uniform. Q07 and Q07 are particularly informative since these questions involve current flow directions, polarities, and dealing with the contribution of loop currents in shared branches. In Q07, Agentic AI obtained the highest mean, with TC = 3.00, compared with 2.50 for Pure LLM and 1.50 for RAG. Although this represents a relative improvement, this remains in the acceptance level, which suggest this kind of question is challenging for the models in use. In Q08, the RAG obtained the highest mean, 4.00, compared with 3.50 for Agentic AI. These cases reveal that merely providing technically correct representations and results does not eliminate the possibility of precision loss during content selection, combination or even evidence understanding and verbalization.
In the predominantly conceptual Q11–Q12 group, RAG and Agentic AI obtained the same aggregated mean, 3.50. The analysis separated by model show, however, different behaviours. In GPT-OSS:20b, the Agentic AI condition obtained 4.00, compared to 3.50 from RAG; in Llama-3.3:70b, RAG obtained 3.50, while 3.00 was achieved for Agentic AI. This is compatible with the architectures’ details, since both benefit from similar information and use the same models—thus, they use the same parametric knowledge.
Across the 24 comparisons, paired by question and model, Agentic AI presented a higher TC than RAG in 16 cases, while being equal in 6 and inferior in 2. Both lower ratings occurred in when using the Llama-3.3:70b, in Q08 and Q12. Relative to Pure LLM, the Agentic AI condition outperformed it in 22 cases, with equal ratings in the remaining 2. It is worth noting that these values are not treated as inferential tests, but show that the aggregated pattern does not result from a single favourable question. At the same time, the cases where Agentic AI does not prevail are methodologically important since they help to find limitations in the synthesis mechanism and avoid an excessive interpretation on the use of deterministic tools as a means to, for example, fully eliminate error in LLM-based systems.

4.6.4. Relationship with the Student Evaluation

The joint reading of the two evaluations was limited to five common prompts, corresponding to questions Q02, Q03, Q05, Q06, and Q09. Table 16 present the means per architecture. The values are shown side-by-side for pattern characterization; however, they are not aggregated, nor are they compared using inferential tests, since they result from scales with different descriptors, evaluators and units of analysis.
The Agentic AI architecture obtained the highest mean in the two perspectives, which demonstrates a convergence regarding the highest-rated architecture. Nevertheless, the high Perceived Correctness ratings for all three conditions coexist with significantly lower technical ratings for the Pure LLM and RAG architectures. This difference does not represent an inconsistency between the studies. Rather, it suggests that the fluency, organization and formal text quality may lead students to perceive a response that contains technical imprecisions as being correct. The evaluation with the students characterizes the perception and acceptability of the responses, while the complementary technical evaluation adds a disciplinary verification of their content.

4.6.5. Architectural Interpretation and Study Limitations

The progression of Pure LLM < RAG < Agentic AI can be understood as a contrast at an architectural level. The first condition depends essentially on the parametric model’s knowledge; the second adds selected pedagogical documentation; and the third adds a processing chain that includes circuit structure formalization, results validation, deterministic tools, pedagogical artefact construction and evidence selection. The results suggest that RAG is useful for providing terminology, response structural patterns and pedagogical context but remains dependent on the model’s capabilities to apply this information to concrete circuits. In Agentic AI, the architecture changes the relationship, since part of the necessary information is produced by deterministic tools—including topology analysis, equations, numerical results and visual representations. These artefacts are then supplied to the model for final response generation. The TC pattern is consistent with this interpretation, while the LLM continues to be responsible for selection, combination and evidence verbalization, which, in some cases, could lead to the production of partially incorrect answers. Therefore, this architecture should be understood as a mechanism for reducing the model’s error surface, rather than an absolute mechanism to eliminate the possibility of errors being made.
The study was deliberately designed around the LCM, a method that involves various electrical engineering skills. Given its limited scope, the generalization of its conclusions to the full scope of circuit analysis should be observed with care. The circuit-specific cases also require input data in the structured form with which the model was tested. While a netlist is a general format used by various simulators, it does not conserve its topological or geometric notions. The QUCS and the U=RIsolve schematic files preserve that information but correspond to specific representations. A general input based in visual recognition was not contemplated in this work.
The comparison between RAG and Agentic AI identifies a pipeline aggregated effect, not the contemplation of a component-by-component ablation; therefore, it does not isolate the contributions of the solver, Renderer, validation mechanisms, or evidence selection achieved in our other method. Each combination of question, model and architecture was evaluated based on responses from the same experimental run. The students evaluated the questions in a pre-prepared, blind-evaluation setting, accounting for both the rendered question and the response. The evaluation was conducted by the author team, who have teaching experience on Electrical Circuit Analysis-related courses and knowledge of the methods and architectures used.
Taken together, the results support a substantial improvement in Technical Correctness. Agentic AI reduces the model’s exposition to errors in circuit-specific tasks, while not completely eliminating the probability of error. The joint reading of the two evaluations also reinforces the notion that the high-quality, coherent texts generated by an LLMs can make a technically imperfect response appear plausible, reinforcing the relevance of architectures that mediate their technical content and user interactions.

4.7. Final Considerations

The statistical analysis of the student questionnaires provides a consistent view of how the three AI response architectures were perceived in the context of Loop Current Method learning tasks. Overall, students tended to assign higher ratings to the enriched architectures, RAG and Agentic AI, than to the Pure LLM approach. The descriptive ordering was Pure LLM, followed by RAG and Agentic AI, with Agentic AI obtaining the highest global mean rating.
At the global level, both RAG and Agentic AI were significantly better-evaluated than Pure LLM. This suggests that adding retrieval mechanisms or Agentic processing steps can improve the perceived quality of AI-generated responses in circuit-analysis tasks. The dimension-level results provide a more specific distinction, with Agentic AI receiving higher ratings for Perceived Correctness than both Pure LLM and RAG.
The dimension-level analysis showed that the effect of architecture was not uniform across all evaluation criteria. Significant differences were found in Correctness, Clarity and response Style. In these dimensions, the enriched architectures were generally favoured over Pure LLM. By contrast, no significant differences were found in perceived Learning Value or Relevance to the question asked. The absence of differences in Relevance suggests that students generally considered the responses from all architectures to be aligned with the prompts. The absence of differences in perceived Learning Value may be related to the students’ previous exposure to circuit-analysis content, since participants had already attended courses where the Loop Current Method had been addressed. In this context, D3 may have reflected the perceived novelty or reinforcement value of the response rather than actual learning gains.
The analysis by prompt type further refined these conclusions. For generic conceptual or procedural prompts, no significant differences were found between the architectures. Although the generic subgroup contains fewer items, the evaluation remained high and close in the three architectures. This suggests that, when the task mainly requires a general explanation of familiar concepts, a Pure LLM may already provide responses that students consider acceptable. In contrast, for circuit-specific prompts, significant differences emerged, particularly in Perceived Correctness, Clarity and response Style. For circuit-specific prompts, Agentic AI was the only enriched architecture to show a statistically significant advantage over Pure LLM in Perceived Correctness. Overall, these results indicate that enriched architectures add more value when the task requires the interpretation of structured technical artefacts, such as circuit netlist files.
The exploratory comparison with ChatGPT showed that the students’ highest-rated local configuration, Llama 3.3-70B–Agentic AI, achieved student evaluations statistically comparable to those obtained by the external commercial reference model. No statistically significant differences were found between the two systems, either globally or in any of the five evaluation dimensions. This result suggests that a locally deployed Agentic architecture, supported by structured processing mechanisms, can reach a level of student-perceived response quality close to that of a commercial model in the evaluated circuit-analysis tasks.
The complementary technical evaluation introduces a distinct perspective. In the five common prompts, Agentic AI obtained the highest mean in both Perceived Correctness (D1) and Technical Correctness (TC). The difference between the two readings suggests that the fluency, organization and formal text quality of the responses may lead to students over-relying on the content. This observation reinforces the importance of supporting linguistic composition with reliable artefacts, while not diminishing the value of students’ evaluation in the other dimensions, including response Clarity, usefulness and their willingness to use these kinds of AI-based systems. In the complete response dataset, Agentic AI presented the most favourable TC profile.
Taken together, the student questionnaires support three main conclusions. First, enriched architectures are preferable to a Pure LLM approach when students evaluate the quality of responses to circuit-analysis prompts. Second, the distinctive contribution of Agentic AI was most evident in Perceived Correctness and in tasks involving structured circuit information, such as a circuit netlist. Third, the students’ highest-rated local Agentic AI configuration was not statistically distinguishable from ChatGPT in the students’ evaluations. These findings support the use of local, architecture-enhanced AI systems as credible candidates for educational support tools in introductory circuit analysis.

5. Conclusions

The present work proposed and evaluated Local LLM architectures for the production of tutoring responses for electrical circuits. The quality of the generated responses was assessed through a blind study involving 135 students, covering five prompts, seven response-generation configurations, and 35 distinct anonymized responses.
The HALO Agentic AI Pipeline combines language models executed locally, pedagogical documentation retrieval, and deterministic tools for formalization, validation and circuit solving. These tools are not used just to present the circuit solution, but also to produce verifiable intermediate artefacts, including normalized SPICE-like netlists, topological information, equations, calculate electrical quantities, visual representations and pedagogically structured (Markdown) documents. These data are then used as a knowledge base for the LLM responsible for composing the final response.
In regard to RQ1, results show that the LLM architecture influenced the quality as perceived by the students. The Pure LLM condition obtained the lowest global mean rating, followed by RAG, while Agentic AI obtained the highest descriptive mean score. Both RAG and Agentic AI were rated significantly higher than Pure LLM. A fine-grained analysis across the evaluation dimensions revealed a more specific distinction: Agentic AI was rated significantly higher than the other two architectures in Perceived Correctness. By contrast, RAG and Agentic AI were rated similarly in Clarity and Style, with both outperforming Pure LLM. On the remaining dimensions, namely Relevance and Learning Value, no statistically significant differences were observed.
Regarding the RQ2, the differences between the architectures were related to the prompt type. In conceptual and/or procedural prompts (with no circuit data), it was not possible to infer statistically relevant differences between the three architectures. On the contrary, in prompts containing circuit data the Agentic AI architecture presented the most favourable profile. Both RAG and Agentic AI were rated significantly higher than Pure LLM when assessing the response on the Clarity and Style dimensions. Nevertheless, in the Correctness dimension, just Agentic AI demonstrated a significant superiority in relation to Pure LLM system. The descriptive complementary technical analysis supports this result: in generic prompts, the means were high and similar in the three architectures; in circuit-specific prompts, RAG and Agentic AI obtained higher means than the Pure LLM condition.
This pattern suggests a functional distinction between the two enriched architectures. The RAG condition enhances pedagogical knowledge and is able to impose a more organized response structure. Nevertheless, it continues to be dependent on the LLM’s skills to address the students’ questions, which, for modest local models, is challenging. The Agentic AI system, on the other hand, adds reliable circuit evidence while requiring just a fraction of the resources an LLM needs, since it is produced by deterministic algorithms. However, the correctness of this evidence does not guarantee, per se, a final correct answer as the final response is composed by the LLM, which is responsible for selecting and combining appropriate input content to deliver the response. It offers a model a panoply of data regarding the submitted circuit, organized in a pedagogical structure. As such, the (weak) reasoning capabilities of the model have a lower impact on response quality. This difference is consistent with the more favourable rating of Agentic AI on questions accompanied by circuit data. While not superior in all dimensions, this system gains ground in ratings when technical reliability is required. Without being superior in every dimension or question, this architecture presented a technical profile that was more favourable in the evaluated set, particularly for tasks dependent on specific circuit structure and numerical quantities.
This result can also be interpreted in light of the complementarity of Symbolic AI and Connectionist AI. In prompts containing circuit data, the symbolic component formalizes and processes the circuit in a determinist manner, applying domain rules to produce reliable technical evidence. As a complement, the connectionist components, operationalized by the LLM, play a different role; they mediate human interactions and transform the data into pedagogically enriched explanations, presented in natural language. The proposed architecture materializes a neuro-symbolic approach that brings together the best of both worlds: (i) the linguistic flexibility of the LLM assisted (and constrained) by formal circuit representations, and (ii) pedagogically structured artefacts created by deterministic tools.
With respect to RQ3, the configuration of Agentic AI supported by the Llama-3.3:70B model obtained the highest mean rating among the six configurations under evaluation. In comparison with the top-tier commercial system, no statistically significant differences were found, in terms of both global evaluation and fine-grained evaluation (by dimension). This result does not show formal equivalence between the two, since it was not specified, and no formal equivalency test was executed. Nevertheless, it indicates that, under the considered tasks and conditions, a local architecture supported by deterministic tools and curated pedagogical knowledge is capable of achieving a perceived quality similar to that of a top-tier commercial system.
Taken together, the results demonstrate that response quality depends not only on the LLM’s intrinsic characteristics, but also on the supporting tools and the quality of the data it uses as context. The model response quality can be highly enhanced in STEM contexts when supported by a system capable of retrieving pedagogical knowledge and producing formal problem representations and technical evidence. As such, the central contribution of this work resides in the transition from a general, ad hoc generation setup to a pedagogical, field-specific enriched environment, where the model’s composition can benefit from deterministic tools to produce reliable and pedagogically aligned data.
The previous conclusions demonstrate that a local implementation of the HALO tutoring system may exist and will function correctly without external, per-query costs. As an additional feature, the local RAG system allows for the traceability of the materials given to students, so students can pinpoint the consultation on the official documents of the courses at stake. In the opinion of the authors of this article, this is a valuable side feature because it reinforces the use of the “official” materials of the course at stake, hopefully reinforces possible previous knowledge, hopefully reduces the likelihood of confusing explanations and, again, hopefully further promotes the importance of local professors and locally produced materials.
The study focused on a proof-of-concept application to empirically evaluate the students’ perceptions of the response quality. Although the study focuses on a deliberately bounded set of tasks related to LCM, the combination of the student evaluation with the complementary technical analysis allows us to observe the architectures’ behaviours through two distinct perspectives that converge on the relevant aspects of architecture progression. In future iterations, we intend to specify an assessment environment to measure learning gains. We also intend to extend the scope to other circuit-analysis methods and extend the evaluation to an expert panel with greater institutional diversity, study the individual contribute of the different components through ablation, and explore new input forms, such as circuit recognition using computer vision. These strategies will pave the way for effective longitudinal evaluations, where we can measure the impact on learning, particularly in terms of knowledge acquisition and study habits.

Author Contributions

Conceptualization, A.R.; methodology, A.R. and J.F.; software, A.R. and J.F.; validation, P.C.O.; formal analysis, M.A. and A.S.; investigation, A.R.; resources, A.R.; data curation, P.C.O. and J.F.; writing—original draft preparation, A.R., J.F. and P.C.O.; writing—review and editing, A.R., J.F., M.A. and A.S.; supervision, M.A. and A.S. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by FCT—Fundação para a Ciência e a Tecnologia, I.P., by project reference 2025.21758.PRT and DOI identifier https://doi.org/10.54499/2025.21758.PRT. This work is partially funded by national funds through FCT—Fundação para a Ciência e a Tecnologia, I.P., under the support UID/50014/2025 (https://doi.org/10.54499/UID/50014/2025).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Acknowledgments

The authors would like to gratefully acknowledge the students who participated in the evaluation.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ISEPInstituto Superior de Engenharia do Porto
FEUPFaculdade de Engenharia da Universidade do Porto
LLMLarge Language Model
AIArtificial Intelligence
NLPNatural Language Processing
RAGRetrieval-Augmented Generation
GPTGenerative Pre-trained Transformer
STEMScience, Technology, Engineering, and Mathematics
VRAMVideo Random Access Memory
SOTAState-of-the-Art
APIApplication Programming Interface
DCDirect Current
ACAlternating Current
ECEElectrical and Computer Engineering
EEElectrical Engineering
KVLKirchhoff’s Voltage Law
LCMLoop Current Method
NVMNode Voltage Method
BCMBranch Current Method
SIInternational System of Units
MoEMixture of Experts
PMBPedagogical Markdown Builder
KGKnowledge Graph
SPSSStatistical Package for the Social Sciences
LoRALow-Rank Adaptation
QLoRAQuantized Low-Rank Adaptation
GPUGraphics Processing Unit
MaaSModel-as-a-Service
URLUniform Resource Locator
RQResearch Question
QUCSQuite Universal Circuit Simulator
SPICESimulation Program with Integrated Circuit Emphasis
JSONJavaScript Object Notation
TCTechnical Correctness

Appendix A. HALO Agentic AI Composer Prompt

The final LLM composer node does not rely on a single static prompt that fits all scenarios. The prompt is generated dynamically from the context pack produced by the pipeline’s previous stages. This context pack contains the student question, deterministic facts, normalized netlist, solver evidence, selected theory sections, selected PMB sections, visual references, known limitations and the answer contract. (During execution, the rendered composer prompt is stored for later inspection).
The LLM call uses a system message and a user message. The system message is:
{You are} HALO Agentic AI. Answer as a rigorous but student-friendly circuit-analysis tutor. Use only the supplied compact evidence packand visual references.
The user message is rendered from the context pack. An example of a generated composer prompt is shown below. Fields between angle brackets are populated at runtime from the artefacts generated by the pipeline.
# HALO Agentic LLM Composer Prompt
## System role
You are HALO Agentic AI, a pedagogical circuit-analysis assistant.
Produce a clear, student-friendly answer grounded only in the evidence pack below.
## Non-negotiable constraints
- 
Answer only from deterministic facts, selected sections and visual references in this context pack.
- 
Do not invent circuit values, equations, currents, branches, loops or assets.
- 
Explain negative currents as direction/reference-direction information, not as automatic errors.
- 
Distinguish fictitious loop currents from physical branch currents.
- 
Reference only available artefacts, theory datasets, PMB sections or visual assets by real label/path; do not invent URLs or document names.
- 
Define LCM symbols such as B, N, C, Ma and Mp before using their formulas.
- 
Keep the answer concise enough for a tutoring turn; cite real theory/artefact labels or paths only when they are present.
- 
Do not merely list evidence. Synthesize it into a clear explanation with the essential numeric facts.
- 
When available, include B, N, C, Ma, Mp, the equation count, loop-current results and branch-current results.
- 
When using B, N, C, Ma or Mp, define the symbol before using the formula.
- 
Prefer a student-facing shape: short answer, setup, key results, sign interpretation, and what to inspect next.
- 
If evidence is insufficient, say what is missing and point to the relevant artefact instead of guessing.
## Student question
<student_question>
## Knowledge guidance
<knowledge_guidance_items>
Each knowledge-guidance item may include:
- 
intent;
- 
description;
- 
prerequisites;
- 
likely misconceptions;
- 
recovery hints.
## Deterministic facts
```json
<compact_deterministic_facts>
```
## LCM symbol glossary
<lcm_symbol_glossary>
Examples:
- 
B: number of branches.
- 
N: number of principal nodes.
- 
C: number of ideal current sources.
- 
Ma: number of auxiliary loops.
- 
Mp: number of principal loops.
## Solver evidence
<selected_solver_sections>
This section may include:
- 
validation status;
- 
normalized netlist;
- 
LCM counts;
- 
KVL equation count;
- 
loop-current results;
- 
branch-current results;
- 
solver warnings.
## Official course theory sections
<selected_theory_sections>
## Selected PMB sections
<selected_pedagogical_markdown_sections>
## Visual references
<available_visual_references>
Each visual reference must correspond to an existing generated artefact.
The model must not invent image names, file paths, URLs or links.
## Required output
Return Markdown only. Use a student-facing structure appropriate to the question type.
Do not change or invent facts: use only claims supported by the supplied context.
Use this structure when useful:
1. 
Short answer to the student question.
2. 
Reasoning grounded in the supplied context.
3. 
Symbols, results or references only when they are actually present.
The final Required Output block may be adapted according to the classification given to the question type. This adaptation is only responsible for the presentation of the answer; it does not authorize the model to introduce new circuit facts such as values, equations or artefact references. For example, if the question is related to auxiliary or principal loops, the generated prompt asks the model to include the relevant circuit facts and to incorporate the information regarding the application of the course method. If the question concerns branch currents or current signs (direction of flow), the generated prompt asks for the requested value, evidence provided by the solver and additional information in relation to the interpretation of the sign (with respect to the assumed reference direction). If there is no solver evidence, visual reference or PMB section available in the artefacts, the prompt explicitly instructs the model not to make false claims, i.e., to not infer that such evidence exists.

Appendix B. Prompts and Questionnaire Organisation

This appendix presents the complete set of prompts and the organisation of the student questionnaires. The evaluation included five prompts related to the Loop Current Method: two generic prompts and three circuit-specific prompts. Each prompt was answered by 7 response-generation configurations, producing a total of 35 anonymized responses. These responses were distributed across five questionnaire forms, with each student evaluating seven responses.
Table A1. Prompts used in the student evaluation.
Table A1. Prompts used in the student evaluation.
Prompt CodePrompt TypePrompt Statement
Q1Generic promptIn a circuit with 3 nodes, 5 branches and 1 ideal current source, using the Loop Current Method, determine: (1) the total number of independent loops; (2) the number of auxiliary loops; (3) the number of main loops; and (4) the minimum number of KVL equations to be written. Explain your reasoning.
Q2Generic promptConsider a circuit with two nodes, consisting of one ideal current source in parallel with four resistors R1, R2, R3 and R4. According to the Loop Current Method described in the course, determine: (1) how many independent loops exist; (2) how many auxiliary loops should be selected; (3) how many main loops remain; and (4) how many KVL equations should be written.
Q3Circuit-specific promptThere are several possible loops in this circuit. Explain the difference between: (1) possible loops; (2) independent loops; (3) auxiliary loops; and (4) main loops. Why is a KVL equation not written for all possible loops? The prompt included a circuit netlist.
Q4Circuit-specific promptIn this circuit, identify a valid choice of auxiliary loop. Explain why this loop is auxiliary and why its current is no longer an unknown. The prompt included a circuit netlist.
Q5Circuit-specific promptExplain the solution of this circuit using the Loop Current Method described in the course, identifying: (1) the auxiliary loops; (2) the main loops; (3) which loop currents are already known due to current sources; (4) how many KVL equations must be written; and (5) how to obtain the real branch currents from the loop currents. The prompt included a circuit netlist.
Table A2. Response-generation configurations included in the evaluation.
Table A2. Response-generation configurations included in the evaluation.
ConfigurationBase ModelArchitecture/SystemConfiguration Label
C1GPT-OSS-20bPure LLMGPT-OSS-20b–Pure LLM
C2GPT-OSS-20bRAGGPT-OSS-20b–RAG
C3GPT-OSS-20bAgentic AIGPT-OSS-20b–Agentic AI
C4Llama 3.3-70BPure LLMLlama 3.3-70B–Pure LLM
C5Llama 3.3-70BRAGLlama 3.3-70B–RAG
C6Llama 3.3-70BAgentic AILlama 3.3-70B–Agentic AI
C7ChatGPTExternal commercial referenceChatGPT external
Table A3. Distribution questionnaire forms to students.
Table A3. Distribution questionnaire forms to students.
FormNumber of StudentsNumber of Responses Evaluated Per Student
Form 1267
Form 2257
Form 3257
Form 4297
Form 5307
Total135—

Appendix C. Questions Used in the Complementary Technical Evaluation

This section presents the 12 questions used in the complementary technical evaluation. For circuit-dependent questions, the prompt text was accompanied by the structured circuit representation. The original Portuguese version is provided together with an English translation.
OriginalEnglish translation
Q01:Q01:
Num circuito conexo com B ramos, N nós e C fontes ideais de corrente, usando o Método das Correntes nas Malhas descrito na disciplina, indique: (1) o número total de malhas independentes; (2) o número de malhas auxiliares; (3) o número de malhas principais; (4) o número mínimo de equações KVL a escrever. Explique também porque a corrente de uma malha auxiliar deixa de ser uma incógnita.In a connected circuit with B branches, N nodes and C ideal current sources, using the Loop Current Method described in the course, state: (1) the total number of independent loops; (2) the number of auxiliary loops; (3) the number of main loops; (4) the minimum number of KVL equations to be written. Also explain why the current of an auxiliary loop is no longer an unknown.
Q02:Q02:
Num circuito com 3 nós, 5 ramos e 1 fonte ideal de corrente, utilizando o Método das Correntes nas Malhas, determine: (1) o número total de malhas independentes; (2) o número de malhas auxiliares; (3) o número de malhas principais; e (4) o número mínimo de equações KVL a escrever. Explique o seu raciocínio.In a circuit with 3 nodes, 5 branches and 1 ideal current source, using the Loop Current Method, determine: (1) the total number of independent loops; (2) the number of auxiliary loops; (3) the number of main loops; and (4) the minimum number of KVL equations to be written. Explain your reasoning.
Q03:Q03:
Considere um circuito com dois nós, constituído por uma fonte ideal de corrente em paralelo com quatro resistências R1, R2, R3 e R4. De acordo com o Método das Correntes nas Malhas descrito na disciplina, determine: (1) quantas malhas independentes existem; (2) quantas malhas auxiliares devem ser selecionadas; (3) quantas malhas principais restam; e (4) quantas equações KVL devem ser escritas.Consider a circuit with two nodes, consisting of one ideal current source in parallel with four resistors R1, R2, R3 and R4. According to the Loop Current Method described in the course, determine: (1) how many independent loops exist; (2) how many auxiliary loops should be selected; (3) how many main loops remain; and (4) how many KVL equations should be written.
Q04:Q04:
Neste circuito, indique: 1. o número de nós elétricos; 2. o número de ramos; 3. o número total de malhas independentes; 4. o número de fontes ideais de corrente; 5. o número de malhas auxiliares; 6. o número de malhas principais; 7. o número mínimo de equações KVL a escrever. Explique o raciocínio.In this circuit, state: 1. the number of electrical nodes; 2. the number of branches; 3. the total number of independent loops; 4. the number of ideal current sources; 5. the number of auxiliary loops; 6. the number of main loops; 7. the minimum number of KVL equations to be written. Explain your reasoning.
Q05:Q05:
Existem várias malhas possíveis neste circuito. Explique a diferença entre: (1) malhas possíveis; (2) malhas independentes; (3) malhas auxiliares; (4) malhas principais. Porque não se escreve uma equação KVL para todas as malhas possíveis? O prompt incluía uma netlist do circuito.There are several possible loops in this circuit. Explain the difference between: (1) possible loops; (2) independent loops; (3) auxiliary loops; and (4) main loops. Why is a KVL equation not written for all possible loops? The prompt included a circuit netlist.
Q06:Q06:
Neste circuito, identifique uma escolha válida de malha auxiliar. Explique porque esta malha é auxiliar e porque a sua corrente deixa de ser uma incógnita. O prompt incluía uma netlist do circuito.In this circuit, identify a valid choice of auxiliary loop. Explain why this loop is auxiliary and why its current is no longer an unknown. The prompt included a circuit netlist.
Q07:Q07:
Considere um circuito constituído por 1 fonte de corrente If1, em paralelo com 3 resistências, R1, R2, R3 e um último ramo constituído pela série da fonte de tensão V1 e uma resistência R4. A fonte de corrente If1 encontra-se no extremo esquerdo, com a corrente a circular do nó A, em baixo, para o nó B, em cima. Considere que a polaridade da fonte V1 tem o polo negativo voltado para R4 e o polo positivo voltado para o nó B. Se formarmos uma malha nos dois ramos mais à direita, R3 em paralelo com a série de V1 e R4, com sentido de circulação de A para B quando observado no ramo de R3, como fica a equação KVL dessa malha (Im1)?Consider a circuit consisting of 1 current source If1 in parallel with 3 resistors, R1, R2, R3, and a final branch consisting of the series combination of voltage source V1 and resistor R4. The current source If1 is located at the far left, with the current flowing from node A, at the bottom, to node B, at the top. Consider that voltage source V1 has its negative terminal facing R4 and its positive terminal facing node B. If we form a loop using the two rightmost branches, with R3 in parallel with the series combination of V1 and R4, and with a direction of circulation from A to B when viewed along the R3 branch, what is the KVL equation for this loop (Im1)?
Nota: considere que em R3 passa também a corrente da malha Im2 com o mesmo sentido da malha que estamos a analisar.Note: Consider that the current of loop Im2 also flows through R3 in the same direction as the loop being analysed.
Q08:Q08:
Considere um circuito constituído por 1 fonte de corrente If1, em paralelo com 3 resistências, R1, R2, R3 e um último ramo constituído pela série da fonte de tensão V1 e uma resistência R4. A fonte de corrente If1 encontra-se no extremo esquerdo, com a corrente a circular do nó A, em baixo, para o nó B, em cima. Considere que a polaridade da fonte V1 tem o polo positivo voltado para R4 e o polo negativo voltado para o nó B. Se formarmos uma malha nos dois ramos mais à direita, R3 em paralelo com a série de V1 e R4, com sentido de circulação de A para B quando observado no ramo de R3, como fica a equação KVL dessa malha (Im1)?Consider a circuit consisting of 1 current source If1 in parallel with 3 resistors, R1, R2, R3, and a final branch consisting of the series combination of voltage source V1 and resistor R4. The current source If1 is located at the far left, with the current flowing from node A, at the bottom, to node B, at the top. Consider that voltage source V1 has its positive terminal facing R4 and its negative terminal facing node B. If we form a loop using the two rightmost branches, with R3 in parallel with the series combination of V1 and R4, and with a direction of circulation from A to B when viewed along the R3 branch, what is the KVL equation for this loop (Im1)?
Nota: considere que em R3 passa também a corrente da malha Im2 com o mesmo sentido da malha que estamos a analisar.Note: Consider that the current of loop Im2 also flows through R3 in the same direction as the loop being analysed.
Q09:Q09:
Explique a resolução deste circuito utilizando o Método das Correntes nas Malhas descrito na disciplina, identificando: (1) as malhas auxiliares; (2) as malhas principais; (3) quais as correntes de malha que já são conhecidas devido às fontes de corrente; (4) quantas equações KVL devem ser escritas; e (5) como obter as correntes reais dos ramos a partir das correntes de malha. O prompt incluía uma netlist do circuito.Explain the solution of this circuit using the Loop Current Method described in the course, identifying: (1) the auxiliary loops; (2) the main loops; (3) which loop currents are already known due to current sources; (4) how many KVL equations must be written; and (5) how to obtain the real branch currents from the loop currents. The prompt included a circuit netlist.
Q10:Q10:
Neste circuito, indique o valor da corrente no ramo da resistência R8 e explique o significado do seu sinal. Use o Método das Correntes nas Malhas para a resolução.In this circuit, determine the value of the current in the branch containing resistor R8 and explain the meaning of its sign. Use the Loop Current Method to solve the circuit.
Q11:Q11:
Porque é que, se o circuito tem M malhas independentes, eu não escrevo M equações KVL quando existe uma fonte ideal de corrente?Why, if the circuit has M independent loops, do I not write M KVL equations when an ideal current source is present?
Q12:Q12:
Vi uma explicação a dizer que se deve usar supermalha quando há uma fonte de corrente. Isso é a mesma coisa que o Método das Correntes nas Malhas usado nesta disciplina, com malhas auxiliares e malhas principais?I saw an explanation stating that a supermesh should be used when there is a current source. Is that the same as the Loop Current Method used in this course, with auxiliary loops and main loops?

Appendix D. Pairwise Comparisons Between System Architectures

This appendix reports the pairwise Wilcoxon signed-rank comparisons used as post hoc tests after the Friedman analyses. Bonferroni correction was applied to control for multiple comparisons. For three pairwise comparisons, the corrected significance threshold was α = 0.017 .
Table A4. Pairwise Wilcoxon comparisons for the global architecture comparisons.
Table A4. Pairwise Wilcoxon comparisons for the global architecture comparisons.
ComparisonZp-ValueInterpretation
RAG vs. Pure LLM−3.3880.001Significant after Bonferroni correction
Agentic AI vs. Pure LLM−4.121<0.001Significant after Bonferroni correction
Agentic AI vs. RAG−1.7920.073Not significant after Bonferroni correction
Table A5. Pairwise Wilcoxon comparisons by evaluation dimension.
Table A5. Pairwise Wilcoxon comparisons by evaluation dimension.
DimensionComparisonZp-ValueInterpretation
D1—CorrectnessRAG vs. Pure LLM−0.0200.984Not significant
D1—CorrectnessAgentic AI vs. Pure LLM−2.8810.004Significant after Bonferroni correction
D1—CorrectnessAgentic AI vs. RAG−2.9580.003Significant after Bonferroni correction
D2—ClarityRAG vs. Pure LLM−4.279<0.001Significant after Bonferroni correction
D2—ClarityAgentic AI vs. Pure LLM−3.499<0.001Significant after Bonferroni correction
D2—ClarityAgentic AI vs. RAG−0.2440.807Not significant
D3—Learning ValueRAG vs. Pure LLM−0.7360.462Not significant
D3—Learning ValueAgentic AI vs. Pure LLM−1.3640.172Not significant
D3—Learning ValueAgentic AI vs. RAG−0.5750.565Not significant
D4—RelevanceRAG vs. Pure LLM−1.3320.183Not significant
D4—RelevanceAgentic AI vs. Pure LLM−0.1890.850Not significant
D4—RelevanceAgentic AI vs. RAG−1.3820.167Not significant
D5—StyleRAG vs. Pure LLM−3.498<0.001Significant after Bonferroni correction
D5—StyleAgentic AI vs. Pure LLM−4.528<0.001Significant after Bonferroni correction
D5—StyleAgentic AI vs. RAG−1.3360.182Not significant

Appendix E. Pairwise Comparisons by Prompt Type

This appendix reports the post hoc Wilcoxon signed-rank comparisons for the prompt-type analyses. Pairwise comparisons are reported mainly for analyses in which the Friedman test indicated statistically significant differences between architectures.
Table A6. Pairwise Wilcoxon comparisons for global circuit-specific prompts.
Table A6. Pairwise Wilcoxon comparisons for global circuit-specific prompts.
Prompt TypeComparisonZp-ValueInterpretation
Circuit-specificRAG vs. Pure LLM−2.6220.009Significant after Bonferroni correction
Circuit-specificAgentic AI vs. Pure LLM−3.554<0.001Significant after Bonferroni correction
Circuit-specificAgentic AI vs. RAG−1.4840.138Not significant
Table A7. Pairwise Wilcoxon comparisons for circuit-specific prompts by dimension.
Table A7. Pairwise Wilcoxon comparisons for circuit-specific prompts by dimension.
DimensionComparisonZp-valueInterpretation
D1—CorrectnessRAG vs. Pure LLM−0.1870.851Not significant
D1—CorrectnessAgentic AI vs. Pure LLM−2.6140.009Significant after Bonferroni correction
D1—CorrectnessAgentic AI vs. RAG−2.3140.021Not significant after Bonferroni correction
D2—ClarityRAG vs. Pure LLM−4.067<0.001Significant after Bonferroni correction
D2—ClarityAgentic AI vs. Pure LLM−3.999<0.001Significant after Bonferroni correction
D2—ClarityAgentic AI vs. RAG−0.5470.584Not significant
D5—StyleRAG vs. Pure LLM−3.0470.002Significant after Bonferroni correction
D5—StyleAgentic AI vs. Pure LLM−3.655<0.001Significant after Bonferroni correction
D5—StyleAgentic AI vs. RAG−1.7370.082Not significant

Appendix F. System Configuration

This appendix provides the complete system configuration used for all training and inference tasks, including hardware details, software versions, and package information.

Appendix F.1. Hardware Platform

  • CPU: Intel(R) Core(TM) i7-9700K @ 3.60 GHz (8 cores, up to 4.9 GHz boost clock).
  • Memory: 64 GB DDR4–2 × 32 GB (KF3200C16D4/32GX) running at 2400 MT/s.
  • Storage: Kingston NV2 NVMe SSD (1 TB)—for OS and the following applications:
    –
    Interface: PCIe 4.0 × 4 NVMe.
    –
    Sequential Read/Write Speeds: Up to 3500 MB/s read, 2800 MB/s write.
  • Graphics Cards: 2 × Asus GeForce RTX 3090 GAMING (24 GB GDDR6X):
    –
    Memory: 24 GB GDDR6X (384-bit interface).
    –
    Memory Speed: 19.5 Gbps.
    –
    Boost Clock: 1800 MHz (factory overclocked).
    –
    Core Count: 10,496 CUDA cores.
    –
    Cooling Solution: Triple-fan.
    –
    Interface: PCIe 4.0 ×16.
    –
    DLSS Support: Yes (2nd-gen RT cores, 3rd-gen Tensor cores).
  • Power Supply: Cooler Master V1200 Platinum:
    –
    Output Capacity: 1200 W.
    –
    Certification: 80 Plus Platinum (93% efficiency).

Appendix F.2. Software Platform

  • Operating System: Ubuntu 24.04.1 LTS (Codename: noble).
  • Containerization Platform: Docker 26.1.3 (build 26.1.3-0ubuntu1 24.04.1).
  • Python Environment: Python 3.11.4, managed with Conda (base environment).
  • Python Packages:
    –
    TensorFlow: 2.13.1.
    –
    PyTorch: 2.3.1.
    –
    Transformers: 4.45.2.
  • Jupyter Environment:
    –
    IPython: 8.15.0.
    –
    JupyterLab: 4.0.6.
    –
    Jupyter Notebook: 7.0.4.
    –
    Jupyter Core: 5.3.2.
    –
    Jupyter Server: 2.7.3.
    –
    Nbconvert: 7.8.0.
    –
    Nbclient: 0.8.0.
  • Machine Learning Acceleration:
    –
    CUDA Version: 12.1.
    –
    NVIDIA Driver: 530.41.03.
    –
    CUDNN Version: 8.9.2.
  • Version Control: Git 2.43.0.
  • Compiler: GCC 13.3.0.

References

  1. Citil, C. Student perspectives on integrating generative artificial intelligence into dual higher education curriculum. Discov. Educ. 2026, 5, 518. [Google Scholar] [CrossRef] [Scilit]
  2. Walton, J.; Bearman, M.; Crawford, N.; Tai, J.; Boud, D. How university students work on assessment tasks with generative artificial intelligence: Matters of judgement. In Assessment & Evaluation in Higher Education; Taylor & Francis: Abingdon, UK, 2025; pp. 1–17. [Google Scholar]
  3. Qu, Y.; Loo, H.E.; Wang, J. Generative artificial intelligence in higher education: Emotional tensions and ethical declaration. Br. J. Educ. Technol. 2025, 1–20. [Google Scholar] [CrossRef] [Scilit]
  4. Rocha, A.; Ferreira, J.; Oliveira, P.; Alves, M.; Sousa, A. Fine-Tuning Lightweight LLMs With Human-Curated Data on Electrical Circuit Fundamentals for E-Learning. Comput. Appl. Eng. Educ. 2026, 34, e70176. [Google Scholar] [CrossRef] [Scilit]
  5. Sousa, L.; Rocha, A.; Alves, M.; Pereira, F. U= RIsolve: A web-based application for learning electrical circuit analysis [Education]. IEEE Circuits Syst. Mag. 2021, 21, 66–95. [Google Scholar] [CrossRef] [Scilit]
  6. Sousa, L.; Rocha, A.; Alves, M.; Pereira, F. Revisiting the nodal voltage method for both human comprehension and software implementation: Towards a teaching/self-learning simulation tool. Comput. Appl. Eng. Educ. 2021, 29, 1642–1664. [Google Scholar] [CrossRef] [Scilit]
  7. Rocha, A.; Oliveira, P.; Ferreira, J.; Alves, M.; Sousa, A. Lightweight Llama Models with Experts’ Curated RAG for Electrical Engineering Education: An Exploratory Comparison. Appl. Sci. 2026, 16, 7745. [Google Scholar] [CrossRef] [Scilit]
  8. Albadarin, Y.; Saqr, M.; Pope, N.; Tukiainen, M. A systematic literature review of empirical research on ChatGPT in education. Discov. Educ. 2024, 3, 60. [Google Scholar] [CrossRef] [Scilit]
  9. Filippi, S.; Motyl, B. Large language models (LLMs) in engineering education: A systematic review and suggestions for practical adoption. Information 2024, 15, 345. [Google Scholar] [CrossRef] [Scilit]
  10. Kasneci, E.; Seßler, K.; Küchemann, S.; Bannert, M.; Dementieva, D.; Fischer, F.; Gasser, U.; Groh, G.; Günnemann, S.; Hüllermeier, E.; et al. ChatGPT for good? On opportunities and challenges of large language models for education. Learn. Individ. Differ. 2023, 103, 102274. [Google Scholar] [CrossRef] [Scilit]
  11. Li, R.; Li, M.; Qiao, W. Engineering students’ use of large language model tools: An empirical study based on a survey of students from 12 universities. Educ. Sci. 2025, 15, 280. [Google Scholar] [CrossRef] [Scilit]
  12. Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Pinto, H.P.D.O.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al. Evaluating large language models trained on code. arXiv 2021, arXiv:2107.03374. [Google Scholar]
  13. Finnie-Ansley, J.; Denny, P.; Becker, B.A.; Luxton-Reilly, A.; Prather, J. The robots are coming: Exploring the implications of openai codex on introductory programming. In Proceedings of the 24th Australasian Computing Education Conference, Virtual, 14–18 February 2022; pp. 10–19. [Google Scholar]
  14. Denny, P.; Kumar, V.; Giacaman, N. Conversing with copilot: Exploring prompt engineering for solving cs1 problems using natural language. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1, Toronto, ON, Canada, 15–18 March 2023; pp. 1136–1142. [Google Scholar]
  15. Rocha, A.; Sousa, L.; Alves, M.; Sousa, A. The underlying potential of NLP for microcontroller programming education. Comput. Appl. Eng. Educ. 2024, 32, e22778. [Google Scholar] [CrossRef] [Scilit]
  16. Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. Lora: Low-rank adaptation of large language models. arXiv 2024, arXiv:2106.09685. [Google Scholar]
  17. Dettmers, T.; Pagnoni, A.; Holtzman, A.; Zettlemoyer, L. Qlora: Efficient finetuning of quantized llms. Adv. Neural Inf. Process. Syst. 2023, 36, 10088–10115. [Google Scholar] [CrossRef] [Scilit]
  18. Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.t.; Rocktäschel, T.; et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Adv. Neural Inf. Process. Syst. 2020, 33, 9459–9474. [Google Scholar]
  19. Karpukhin, V.; Oguz, B.; Min, S.; Lewis, P.; Wu, L.; Edunov, S.; Chen, D.; Yih, W.t. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, 16–20 November 2020; pp. 6769–6781. [Google Scholar] [CrossRef] [Scilit]
  20. Asai, A.; Wu, Z.; Wang, Y.; Sil, A.; Hajishirzi, H. Self-rag: Learning to retrieve, generate, and critique through self-reflection. In Proceedings of the International Conference on Learning Representations, Vienna, Austria, 7–11 May 2024; Volume 2024, pp. 9112–9141. [Google Scholar]
  21. Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Hambro, E.; Zettlemoyer, L.; Cancedda, N.; Scialom, T. Toolformer: Language models can teach themselves to use tools. Adv. Neural Inf. Process. Syst. 2023, 36, 68539–68551. [Google Scholar] [CrossRef] [Scilit]
  22. Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. REACT: SYNERGIZING REASONING AND ACTING IN LANGUAGE MODELS. In Proceedings of the 11th International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, 1–5 May 2023; ICLR: Dublin, Ireland, 2023. [Google Scholar]
  23. Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: Language agents with verbal reinforcement learning. Adv. Neural Inf. Process. Syst. 2023, 36, 8634–8652. [Google Scholar] [CrossRef] [Scilit]
  24. Wang, Y.; Li, Z.; Wang, Z.; Teiletche, P.; Jin, L.; Zaharia, M.; Gonzalez, J.E.; Min, S. PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation. arXiv 2026, arXiv:2606.28344. [Google Scholar]
  25. Jonah, S. When Do Multimodal and Graph-Augmented RAG Help? A Controlled Evaluation for Document Question Answering. arXiv 2026, arXiv:2607.16604. [Google Scholar]
  26. Huang, C.Y.; Chen, H.I.; Ho, H.W.; Kang, P.H.; Lin, M.P.H.; Liu, W.H.; Ren, H. Netlistify: Transforming circuit schematics into netlists with deep learning. In Proceedings of the 2025 ACM/IEEE 7th Symposium on Machine Learning for CAD (MLCAD), Santa Cruz, CA, USA, 8–10 September 2025; IEEE: New York, NY, USA, 2025; pp. 1–8. [Google Scholar]
  27. Hu, W.; Zhan, X.; Tong, M. Parsing netlists of integrated circuits from images via graph attention network. Sensors 2023, 24, 227. [Google Scholar] [CrossRef] [Scilit]
  28. Peker, Ö.B.; Toker, E.; Öcal, D.; Dalyan, T.; Afacan, E.; Gökdel, Y.D. A Fully Automated SPICE-Compatible Netlist Extraction From Image Using Deep Learning and Image Preprocessing Techniques. IEEE Access 2026, 14, 19750–19765. [Google Scholar] [CrossRef] [Scilit]
  29. Luccioni, A.S.; Viguier, S.; Ligozat, A.L. Estimating the carbon footprint of bloom, a 176b parameter language model. J. Mach. Learn. Res. 2023, 24, 1–15. [Google Scholar]
  30. Oviedo, F.; Kazhamiaka, F.; Choukse, E.; Kim, A.; Luers, A.; Nakagawa, M.; Bianchini, R.; Ferres, J.M.L. Energy use of AI inference, efficiency pathways, and test-time scaling. Joule, 2026; in press. [CrossRef] [Scilit]
  31. Xiao, T.; Nerini, F.F.; Matthews, H.D.; Tavoni, M.; You, F. Environmental impact and net-zero pathways for sustainable artificial intelligence servers in the USA. Nat. Sustain. 2025, 8, 1541–1553. [Google Scholar] [CrossRef] [Scilit]
  32. OpenAI. OpenAI API Pricing. 2026. Available online: https://developers.openai.com/api/docs/pricing (accessed on 28 June 2026).
  33. Minaee, S.; Mikolov, T.; Nikzad, N.; Chenaghlu, M.; Socher, R.; Amatriain, X.; Gao, J. Large language models: A survey. arXiv 2024, arXiv:2402.06196. [Google Scholar]
  34. Osborne, C.; Ding, J.; Kirk, H.R. The AI community building the future? A quantitative analysis of development activity on Hugging Face Hub. J. Comput. Soc. Sci. 2024, 7, 2067–2105. [Google Scholar] [CrossRef] [Scilit]
  35. Chiang, W.L.; Zheng, L.; Sheng, Y.; Angelopoulos, A.N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M.I.; Gonzalez, J.E.; et al. Chatbot arena: An open platform for evaluating LLMs by human preference. In Proceedings of the 41st International Conference on Machine Learning, JMLR.org, ICML’24, Vienna, Austria, 21–27 July 2024. [Google Scholar]
  36. VanLehn, K. The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educ. Psychol. 2011, 46, 197–221. [Google Scholar] [CrossRef] [Scilit]
  37. Butz, B.P.; Duarte, M.; Miller, S.M. An intelligent tutoring system for circuit analysis. IEEE Trans. Educ. 2006, 49, 216–223. [Google Scholar] [CrossRef] [Scilit]
  38. Skromme, B.J.; Rayes, P.J.; McNamara, B.E.; Seetharam, V.; Gao, X.; Thompson, T.; Wang, X.; Cheng, B.; Huang, Y.F.; Robinson, D.H. Step-based tutoring system for introductory linear circuit analysis. In Proceedings of the 2015 IEEE Frontiers in Education Conference (FIE), Erie, PA, USA, 21–24 October 2015; IEEE: New York, NY, USA, 2015; pp. 1–9. [Google Scholar]
  39. Skromme, B.J.; Bansal, S.K.; Barnard, W.M.; O’Donnell, M.A. Step-based tutoring software for complex procedures in circuit analysis. In Proceedings of the 2019 IEEE Frontiers in Education Conference (FIE), Covington, KY, USA, 16–19 October 2019; IEEE: New York, NY, USA, 2019; pp. 1–5. [Google Scholar]
  40. Skromme, B.J.; Wong, M.; Redshaw, C.; O’donnell, M. Teaching series and parallel connections. IEEE Trans. Educ. 2021, 65, 461–470. [Google Scholar] [CrossRef] [Scilit]
  41. Skromme, B.J.; Seetharam, V.; Gao, X.; Korrapati, B.; McNamara, B.E.; Huang, Y.F.; Robinson, D.H. Impact of step-based tutoring on student learning in linear circuit courses. In Proceedings of the 2016 IEEE Frontiers in Education Conference (FIE), Erie, PA, USA, 12–15 October 2016; IEEE: New York, NY, USA, 2016; pp. 1–9. [Google Scholar]
  42. Chen, L.; Qin, Z.; Guo, Y.; Rohde, J.; Zhang, Y. Benchmarking large language models on homework assessment in circuit analysis. Int. J. Artif. Intell. Educ. 2025, 35, 3294–3355. [Google Scholar] [CrossRef] [Scilit]
  43. Chen, L.; Xie, H.; Rohde, J.; Zhang, Y. WIP: Large Language Model-Enhanced Smart Tutor for Undergraduate Circuit Analysis. In Proceedings of the 2025 IEEE Frontiers in Education Conference (FIE), Nashville, TN, USA, 2–5 November 2025; IEEE: New York, NY, USA, 2025; pp. 1–5. [Google Scholar]
  44. Knievel, C.; Bernhardt, A.; Bernhardt, C. Recognition, Retrieval, and Response: The AITEE Framework for Socratic Tutoring in Electrical Engineering. IEEE Access 2026, 14, 50109–50126. [Google Scholar] [CrossRef] [Scilit]
  45. Sweller, J. Cognitive load during problem solving: Effects on learning. Cogn. Sci. 1988, 12, 257–285. [Google Scholar] [CrossRef] [Scilit]
  46. Kircher, P.; Sweller, J.; Clark, R.E. Why minimal guidance during instruction does not work: An analysis of the failure of constructivist, discovery, problem-based, experiential, and inquiry-based teaching. Educ. Psychol. 2006, 4, 75–86. [Google Scholar] [CrossRef] [Scilit]
  47. Hmelo-Silver, C.E.; Duncan, R.G.; Chinn, C.A. Scaffolding and achievement in problem-based and inquiry learning: A response to Kirschner, Sweller, and. Educ. Psychol. 2007, 42, 99–107. [Google Scholar] [CrossRef] [Scilit]
  48. Kalyuga, S. The expertise reversal effect. In Managing Cognitive Load in Adaptive Multimedia Learning; IGI Global Scientific Publishing: Hershey, PA, USA, 2009; pp. 58–80. [Google Scholar]
  49. Ding, K.; Yu, J.; Huang, J.; Yang, Y.; Zhang, Q.; Chen, H. SciToolAgent: A knowledge-graph-driven scientific agent for multitool integration. Nat. Comput. Sci. 2025, 5, 962–972. [Google Scholar] [CrossRef] [Scilit]
  50. Haykin, S. Neural Networks and Learning Machines, 3/E; Pearson Education: Noida, India, 2009. [Google Scholar]
  51. Russell, S.J. Artificial Intelligence a Modern Approach; Pearson Education, Inc.: London, UK, 2010. [Google Scholar]
  52. Weizenbaum, J. ELIZA—A computer program for the study of natural language communication between man and machine. Commun. ACM 1966, 9, 36–45. [Google Scholar] [CrossRef] [Scilit]
Figure 1. General architecture of the HALO Agentic AI Pipeline.
Figure 1. General architecture of the HALO Agentic AI Pipeline.
Computers 15 00647 g001
Figure 2. Examples of Renderer module overlays.
Figure 2. Examples of Renderer module overlays.
Computers 15 00647 g002
Table 1. Experimental conditions for response generation.
Table 1. Experimental conditions for response generation.
SystemAdditional Processing Resources and Evidence Made Available to the Composer
A—Pure LLMNo document retrieval or tool-generated evidence.
B—RAGText excerpts retrieved from an expert-curated pedagogical corpus.
C—Agentic AIPedagogical excerpts, normalized representation, topological information, equations, results, validations, and visual artefacts produced and selected by the pipeline according to the question.
Table 2. Main artefacts produced or manipulated by the HALO Agentic AI Pipeline.
Table 2. Main artefacts produced or manipulated by the HALO Agentic AI Pipeline.
ArtefactOriginMain ContentUse
User questionTextual inputDoubt, tutorial intention, and response focusGuides the system routing, evidence selection, and the final composition
Input classificationMain runtimeInput type, circuit mode, and available artefactsDefines the conditional branches of the pipeline
Circuit JSONInput or frontendComponents, values, wires, (inter)connection points, and geometric informationBasis for connectivity extraction, netlist construction, and rendering
JSON–netlist mapJSON circuit extractorRelationship between components, electrical nodes, and netlist linesExplains how the drawing was converted into a formal representation
Extracted netlistJSON circuit extractorInitial textual representation obtained from the circuit connectivityInput for validation and normalisation
Normalised netlistNetlist builder/validatorFormal and computable circuit representation, with normalised nodes and componentsInput for the deterministic solver and the evidence package
Solver outputU=RIsolveCircuit patterns, equations, currents, intermediate values, and final resultsTechnical basis for circuit-dependent explanations
Topological metadataExtractor, parser, and solverNodes, branches, loops, sources, relationships between variables, and reference directionsSupports structural explanation and method interpretation
Layered renderingsCircuit RendererBase circuit and overlays of nodes, branches, loops, and currentsVisual support for the response and for human review
PMBPedagogical Markdown BuilderTheory, circuit, netlist, solution steps, equations, results, tips and notesStructured pedagogical artefact for consultation, and LLM retrieval
Evidence packageEvidence Pack BuilderCompact selection of facts, sections, results, visual references, limitations, and constraintsBasis for constructing the context sent to the LLM
LLM contextContext Pack BuilderQuestion, deterministic facts, normalised netlist, solver evidence, pedagogical sections, and response contractStructured prompt for composing the final response
Final responseLLM composerTutorial explanation in natural language, generated by the modelOutput presented to the student
Response verificationContract checksWarnings about omissions, non-existent references, lack of sign explanation, or incomplete use of evidenceQuality control and response traceability
Trace packagePipeline orchestratorInput, artefacts used, selected evidence, context, response, and verificationReview, auditing, and subsequent analysis
Table 3. Summaryof prompts used in the evaluation. More details in Table A1.
Table 3. Summaryof prompts used in the evaluation. More details in Table A1.
Prompt CodePrompt TypeMain FocusBrief Description
Q1GenericStructural identification of the circuitDetermination of the number of independent loops, main loops, auxiliary loops and minimum number of KVL equations in a circuit described textually.
Q2GenericLoops in a circuit with a current sourceAnalysis of a circuit with an ideal current source in parallel with resistors, requiring the identification of relevant loops and equations.
Q3Circuit-specificDistinction between loops (auxiliary and independent)Explanation of the difference between possible loops, independent loops, auxiliary loops and main loops, based on a circuit netlist.
Q4Circuit-specificSelection of an auxiliary loopIdentification of a valid auxiliary loop and explanation of why the associated loop current is no longer unknown.
Q5Circuit-specificStructured solution using the Loop Current MethodExplanation of the solution procedure, including auxiliary loops, main loops, known loop currents, KVL equations and real (branch) currents.
Table 4. Student questionnaire evaluation dimensions.
Table 4. Student questionnaire evaluation dimensions.
CodeDimensionStatement
D1Perceived CorrectnessI believe the answer is correct.
D2ClarityThe answer is clear and easy to understand.
D3Learning ValueI learned something from this answer.
D4RelevanceThe answer addresses what was asked.
D5StyleI like the style of the answer.
Table 5. Dimension-level comparison between response architectures.
Table 5. Dimension-level comparison between response architectures.
DimensionPure LLM MRAG MAgentic AI MFriedman χ 2 ( 2 ) p-ValuePost hoc Interpretation
D1—Perceived Correctness4.2674.2674.45213.0500.001Agentic AI rated higher than Pure LLM and RAG
D2—Clarity3.6784.0334.00025.0100.000RAG and Agentic AI rated higher than Pure LLM
D3—Learning Value3.9113.9673.9742.1160.347No significant differences
D4—Relevance4.4484.3744.4441.6000.449No significant differences
D5—Style3.6043.9444.03323.8390.000RAG and Agentic AI rated higher than Pure LLM
Table 6. Descriptive statistics by prompt type and response architecture.
Table 6. Descriptive statistics by prompt type and response architecture.
Prompt TypeArchitectureNMeanSD
GenericPure LLM1104.2110.656
GenericRAG824.2680.662
GenericAgentic AI1044.2580.678
Circuit-specificPure LLM1353.8480.868
Circuit-specificRAG1104.1100.717
Circuit-specificAgentic AI1354.1160.749
Note: N refers to aggregated student-level observations available for each prompt type and architecture. The Friedman tests used only complete paired observations across the three architectures.
Table 7. Dimension-level comparison by prompt type.
Table 7. Dimension-level comparison by prompt type.
Prompt TypeDimensionFriedman χ 2 ( 2 ) p-ValuePost hoc Interpretation
GenericD1—Perceived Correctness0.0570.972No significant differences
GenericD2—Clarity3.1380.208No significant differences
GenericD3—Learning Value4.0370.133No significant differences
GenericD4—Relevance1.0000.607No significant differences
GenericD5—Style5.0690.079No significant differences
Circuit-specificD1—Perceived Correctness10.6070.005Agentic AI rated higher than Pure LLM
Circuit-specificD2—Clarity22.6050.000RAG and Agentic AI rated higher than Pure LLM
Circuit-specificD3—Learning Value1.8910.388No significant differences
Circuit-specificD4—Relevance1.2640.532No significant differences
Circuit-specificD5—Style13.9100.001RAG and Agentic AI rated higher than Pure LLM
Table 8. Global student-level evaluation by local configuration.
Table 8. Global student-level evaluation by local configuration.
Local ConfigurationNMeanStandard Deviation
GPT-OSS-20b–Pure LLM1353.8930.942
GPT-OSS-20b–RAG1354.1640.712
GPT-OSS-20b–Agentic AI1354.0990.752
Llama 3.3-70B–Pure LLM1354.0700.740
Llama 3.3-70B–RAG1354.0700.854
Llama 3.3-70B–Agentic AI1354.2620.716
Table 9. Comparison between Llama 3.3-70B–Agentic AI and ChatGPT.
Table 9. Comparison between Llama 3.3-70B–Agentic AI and ChatGPT.
ComparisonLlama 3.3-70B Agentic AI MChatGPT MZp-ValueInterpretation
Global score4.2624.222−0.7400.459No significant difference
D1—Perceived Correctness4.5104.340−1.7390.082No significant difference
D2—Clarity4.1204.160−0.3270.743No significant difference
D3—Learning Value4.0704.050−0.3460.729No significant difference
D4—Relevance4.5004.520−0.4030.687No significant difference
D5—Style4.1204.040−0.9450.345No significant difference
Table 10. Organization of the questions used in the complementary technical evaluation.
Table 10. Organization of the questions used in the complementary technical evaluation.
QuestionsMain FocusType of Competence Evaluated
Q01–Q04Topological counting and method structureLCM structure and concepts
Q05–Q06Mesh identification and selectionApplication of topological structure
Q07–Q10Circuit equations and quantitiesApplication and interpretation
Q11–Q12Current sources and loop typesConceptual understanding
Table 11. Criteria used in the complementary technical evaluation.
Table 11. Criteria used in the complementary technical evaluation.
CriterionDimensionEvaluation Focus
TCTechnical CorrectnessTechnical Correctness of the content and results
QAQuestion AlignmentAdequacy of the response to the prompt request
SCSymbol ClarityClarity and consistency of symbols, quantities and references
PRPaper-and-Pencil ReasoningAdequacy of the reasoning for a manual solution of the problem
MHMisconception HandlingAbility to avoid or clarify incorrect interpretations
VRVisual Reference UsefulnessUsefulness and adequacy of the visual references presented
CTConciseness for TutoringConciseness and focus of the response in a tutoring-like context
NLNext Learning StepUsefulness of the response in guiding further learning
Table 12. Descriptive profile of the complementary technical evaluation.
Table 12. Descriptive profile of the complementary technical evaluation.
MetricPure LLMRAGAgentic AI
TC1.792.633.71
QA2.673.213.83
SC3.053.543.79
PR2.503.103.71
MH1.772.303.13
VR1.872.003.65
CT3.083.833.46
NL2.001.803.50
TC ≥ 4 1/24 (4.2%)5/24 (20.8%)13/24 (54.2%)
TC ≤ 2 21/24 (87.5%)14/24 (58.3%)1/24 (4.2%)
Table 13. Technical Correctness (TC) ratings by question, base model, and response architecture. Ratings range from 1 (very poor) to 5 (excellent).
Table 13. Technical Correctness (TC) ratings by question, base model, and response architecture. Ratings range from 1 (very poor) to 5 (excellent).
QuestionGPT-OSS:20bLlama-3.3:70b
Pure
LLM
RAGAgentic
AI
Pure
LLM
RAGAgentic
AI
Q01255323
Q02235223
Q03124122
Q04223224
Q05124224
Q06455114
Q07323213
Q08233254
Q09135123
Q10113125
Q11144233
Q12234243
1—Very poor 2—Poor 3—Acceptable
4—Good 5—Excellent
Table 14. Mean Technical Correctness by base model and architecture.
Table 14. Mean Technical Correctness by base model and architecture.
Base ModelPure LLMRAGAgentic AI
GPT-OSS:20b1.832.924.00
Llama-3.3:70b1.752.333.42
Table 15. Mean Technical Correctness by question group.
Table 15. Mean Technical Correctness by question group.
GroupQuestionsPure LLMRAGAgentic AI
Counting and method structureQ01–Q041.882.503.63
Mesh identification and selectionQ05–Q062.002.504.25
Equations and quantitiesQ07–Q101.632.383.63
Conceptual doubtsQ11–Q121.753.503.50
Table 16. Student-reported Perceived Correctness and Technical Correctness in the five common prompts.
Table 16. Student-reported Perceived Correctness and Technical Correctness in the five common prompts.
ArchitectureStudent-Reported Perceived Correctness (D1)Technical Correctness (TC)
Pure LLM4.2671.60
RAG4.2672.40
Agentic AI4.4523.90
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Rocha, A.; Ferreira, J.; Oliveira, P.C.; Alves, M.; Sousa, A. A Tool-Augmented Agentic AI Pipeline for Reliable Circuit-Analysis Tutoring with Local Language Models. Computers 2026, 15, 647. https://doi.org/10.3390/computers15100647

AMA Style

Rocha A, Ferreira J, Oliveira PC, Alves M, Sousa A. A Tool-Augmented Agentic AI Pipeline for Reliable Circuit-Analysis Tutoring with Local Language Models. Computers. 2026; 15(10):647. https://doi.org/10.3390/computers15100647

Chicago/Turabian Style

Rocha, André, João Ferreira, Paulo C. Oliveira, Mário Alves, and Armando Sousa. 2026. "A Tool-Augmented Agentic AI Pipeline for Reliable Circuit-Analysis Tutoring with Local Language Models" Computers 15, no. 10: 647. https://doi.org/10.3390/computers15100647

APA Style

Rocha, A., Ferreira, J., Oliveira, P. C., Alves, M., & Sousa, A. (2026). A Tool-Augmented Agentic AI Pipeline for Reliable Circuit-Analysis Tutoring with Local Language Models. Computers, 15(10), 647. https://doi.org/10.3390/computers15100647

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop