Next Article in Journal
Numerical Investigation for 3D Branches of the Lyapunov Families in the Hill’s Problem with Radiation Pressure
Next Article in Special Issue
A Comparative Performance Evaluation of Classical and Quantum-Based Deep Learning Models in the Classification of Breast Cancer Histopathological Images
Previous Article in Journal
AGNAE: An Augmented-Driven Graph Network with Adaptive Exploration for Real-Time Fraud Detection in Dynamic Financial Networks
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Pharmacokinetics-Informed Agentic Architecture for Drug Dynamics with LLM-Driven In Silico Patients

School of Engineering, CUNEF Universidad, 28040 Madrid, Spain
Mathematics 2026, 14(10), 1627; https://doi.org/10.3390/math14101627
Submission received: 16 April 2026 / Revised: 1 May 2026 / Accepted: 8 May 2026 / Published: 11 May 2026

Abstract

This research presents the architecture, implementation, and validation of a novel Agentic Large Model (ALM) framework for in silico patient simulation, synergistically combining pharmacokinetic modeling with autonomous clinical reasoning. The multi-agent architecture enables emerging collaboration between specialized components for physiological simulation, medication management, and safety monitoring, coordinated through a central orchestrator to ensure behavioral alignment. Leveraging LLaMA3 via Ollama, the system demonstrates clinically plausible strategic planning across 12 diverse patient scenarios while maintaining trustworthiness through structured JSON-constrained outputs. Key innovations include a modular agent design with clear separation of concerns, integrated drug interaction checking with hierarchical severity assessment, and multi-layer safety assurance. Rigorous evaluation demonstrates effective pain management (39–58% reduction) with appropriate safety flagging in complex cases, particularly for geriatric and polypharmacy patients. The system achieves an average decision latency of 2.1 s while maintaining 98.7% structured output compliance, advancing the foundations of reliable agentic systems in critical healthcare domains.

1. Introduction

In-silico patient systems represent an innovative class of clinical simulation tools that combine physiological modeling with interactive decision-making. These systems create dynamic digital representations of human patients that respond realistically to interventions while maintaining clinically accurate physiological behaviors.
Agentic in silico patient systems extend this foundation by incorporating pharmacokinetic modeling, multi-agent decision-making, and constrained LLM reasoning. These advanced systems enable pharmaceutical companies to simulate diverse patient populations with unprecedented accuracy, revolutionizing three key areas: drug development, clinical training, and post-market surveillance.
For Big Pharma, this technology offers a powerful tool to optimize clinical trial design, predict real-world drug performance, and accelerate time-to-market. in silico patients can model complex scenarios—such as polypharmacy in elderly populations or rare disease presentations—that are difficult to capture in traditional trials, enabling more robust safety profiling and personalized dosing strategies before human testing begins.
One of the most compelling applications lies in adaptive clinical trial simulation. Pharmaceutical companies can use agentic systems to virtually test different dosing regimens, inclusion criteria, and combination therapies in thousands of simulated patients, identifying optimal trial parameters before committing to costly real-world studies. This reduces the risk of late-stage trial failures and allows for more efficient resource allocation. In addition, these systems can model drug-drug interactions in patient subgroups (e.g., renal impairment, genetic metabolizers) that may not be fully represented in Phase III trials, helping mitigate post-approval safety concerns.
Beyond R&D, agentic in silico patients can enhance medical affairs and commercial strategies. They enable dynamic training platforms for healthcare providers, simulating patient responses to new therapies under real-world conditions. Pharma sales teams can use these simulations to demonstrate drug efficacy and safety in diverse clinical scenarios, improving stakeholder engagement. Furthermore, post-market surveillance can be augmented with AI-driven in silico cohorts that continuously monitor for adverse events, providing early signals that may require further investigation.
Looking ahead, the integration of digital twins, high-fidelity virtual representations of individual patients, could enable truly personalized medicine. Pharmaceutical companies could partner with healthcare providers to simulate treatment outcomes for specific patients, optimizing therapy selection before administration. This not only improves patient care but also strengthens the value proposition of precision therapeutics.
Ultimately, agentic in silico patient systems offer Big Pharma a competitive edge by reducing development costs, improving trial success rates, and enabling data-driven commercialization. As regulatory bodies begin to accept in silico trials as supplementary evidence, early adopters of this technology will lead the next wave of innovation in drug development and patient care.
Modern clinical education and decision support systems face three fundamental challenges: providing realistic patient interactions, adapting to individual learner decisions, and maintaining strict clinical accuracy. This LLM-Driven in silico patients and Doctors system introduces a novel approach to medical simulation by combining deterministic pharmacokinetics with constrained Large Language Model reasoning and modular agent-based design. Its main contribution lies in bridging the gap between precise drug modeling and flexible clinical decision support. The proposal ensures clinical realism while maintaining deterministic, reproducible behavior, a combination that is rarely achieved on existing in silico patient platforms. Specifically, it addresses critical challenges in medical simulation:
  • Clinical accuracy vs LLM flexibility: Clinical reasoning from LLaMA3 (via Ollama) is constrained using JSON schemas and augmented with retrieval from a knowledge graph built from authoritative sources such as Goodman & Gilman’s Pharmacology, ensuring accurate yet contextually rich outputs.
  • Deterministic-Probabilistic balance: Physiologically-based pharmacokinetic models provide exact predictions of drug concentrations, while the LLM handles qualitative clinical judgments within strictly defined boundaries.
  • Complex regimen safety: The system performs hierarchical drug interaction checks with severity stratification (contraindication > major > moderate > minor), supporting safe and realistic simulation of multi-drug regimens.
The remainder of this paper is organized as follows. Section 2 reviews the relevant background and related work, with particular attention to existing approaches in in silico patients and agentic AI. Section 3 presents the proposed methodology, detailing the model formulation and underlying computational framework. Section 4 describes the experimental setup and discusses the results obtained, highlighting the performance and practical implications of the approach. Finally, Section 5 concludes the paper by summarizing the main findings and outlining directions for future research.

2. Related Works

2.1. In Silico Patients in Medical Training

In silico patients have emerged as valuable tools in health professions education, addressing challenges such as reduced patient contact hours and the need for safe practice environments. The MedBiquitous Consortium established an XML-based standard (MedBiquitous in silico patient Standard) to enable interoperability and sharing of in silico patient cases across institutions [1]. This standard defines components for patient data, media resources, and activity models that structure clinical encounters [1]. The standardization effort has been crucial for creating reusable educational content and facilitating comparative research across institutions.
Systematic reviews and meta-analyses demonstrate that in silico patient simulations can effectively improve clinical skills compared to traditional education methods [2]. Studies have shown particular effectiveness in developing clinical reasoning abilities, procedural skills, and team-based competencies [2]. A comprehensive meta-analysis of 37 studies revealed that in silico patient interventions produced moderate to large effect sizes (Hedges’ g = 0.71) in clinical reasoning outcomes compared to control conditions [2]. The technology has proven valuable across both high-income and low-to-middle-income countries, showing global applicability in medical education [2]. Research in resource-limited settings has demonstrated that in silico patients can partially compensate for limited access to clinical training opportunities while maintaining educational quality.
The effectiveness of in silico patients for the training of medical doctors appears to be mediated by several key factors. Scaffolded feedback mechanisms significantly improve learning outcomes, particularly when providing specific and timely guidance on clinical decision-making [3]. Furthermore, the authenticity of clinical scenarios and the fidelity of patient interactions are strongly correlated with skill transfer to real-world settings [4]. In-silico patients incorporating physiological modeling—such as cardiovascular or respiratory simulations—demonstrate particularly strong effects on diagnostic accuracy and treatment planning skills. Furthermore, high-fidelity simulation features have been shown to improve outcomes when learners engage in realistic, immersive scenarios [4].
Recent advancements have incorporated Large Language Models to create more sophisticated in silico patient interactions. Systems like ChatGPT-powered chatbots serve as simulated patients for history-taking practice [5], while frameworks such as EvoPatient utilize multi-agent coevolution to create more realistic standardized patient interactions [6]. These developments represent significant progress beyond earlier screen-based in silico patient systems. The integration of natural language processing enables dynamic, unscripted conversations that more closely mimic clinical encounters, addressing a major limitation of previous generation systems [6].
Emerging research indicates that AI-enhanced in silico patients can adapt to learner performance, providing increasingly complex challenges as competence develops [7,8]. This adaptive capability, combined with automated assessment of clinical reasoning patterns, offers unprecedented opportunities for personalized medical education. Furthermore, the integration of emotional and behavioral modeling in in silico patients shows promise for teaching communication skills and empathy, aspects previously difficult to simulate effectively [9].
Previous work has also shown that interpretable neuro-symbolic models such as Fuzzy Cognitive Maps can support medical decision-making under scarce data conditions, for example in rheumatoid arthritis diagnosis [10].
Despite these advances, significant challenges remain in validating the clinical competence developed through in silico patient interactions and ensuring equitable access to these technologies across diverse educational settings. The field continues to grapple with questions of assessment validity, cost-effectiveness, and integration with existing curricular structures [3]. Nevertheless, the trajectory of development suggests increasingly sophisticated in silico patients will play an essential role in the future of medical education.

2.2. Agent-Based Healthcare Models

Multi-agent systems have shown considerable promise in healthcare applications due to their ability to model complex, distributed clinical environments [11]. These systems leverage autonomous agents with capabilities for collaboration, negotiation, and distributed decision-making [11]. The inherent parallelism and specialization of multi-agent architectures make them particularly suitable for healthcare domains requiring coordinated expertise across multiple medical specialties [12,13].
In clinical contexts, agent-based approaches have been applied to various domains including emergency medical services, patient monitoring, and resource management [11]. Recent advances have expanded into chronic disease management, where agent-based models facilitate personalized treatment planning through continuous monitoring and adaptive intervention strategies [14]. These systems demonstrate particular strength in managing complex, comorbid conditions that require coordinated care across multiple healthcare providers and settings [14].
The ClinicalAgent system represents a notable advancement, employing a multi-agent framework for clinical trial tasks with specialized agents for drug information retrieval, disease analysis, and explanatory reasoning [15]. This system achieved impressive performance in clinical trial outcome prediction (0.7908 PR-AUC), demonstrating the potential of multi-agent approaches in complex clinical domains [15]. The architecture employs a hierarchical coordination mechanism where a master agent orchestrates specialized sub-agents, each trained on domain-specific medical literature and clinical guidelines [15].
The Multi-Agent Conversation (MAC) framework further illustrates the power of agent-based approaches in healthcare. Inspired by clinical Multi-Disciplinary Team discussions, MAC significantly outperformed single models in diagnostic accuracy for rare diseases, achieving optimal performance with four doctor agents and a supervisor agent [16]. This approach demonstrates how multi-agent systems can mimic real-world clinical collaboration patterns. The framework incorporates a dynamic weighting system that adjusts agent contributions based on their confidence levels and domain expertise [16].
Recent research has explored the integration of reinforcement learning with multi-agent systems for adaptive treatment planning. Wang [17] demonstrated that multi-agent reinforcement learning systems can optimize chemotherapy regimens in oncology by simultaneously considering tumor response, toxicity management, and patient quality of life. These systems employ deep Q-networks with reward functions that balance clinical efficacy against treatment-related adverse effects [17].
The emergence of federated learning architectures has further enhanced the capabilities of healthcare multi-agent systems. Kaissis et al. [18] developed a privacy-preserving multi-agent framework that enables collaborative learning across healthcare institutions without sharing sensitive patient data. This approach maintains data sovereignty while allowing agents to benefit from diverse clinical experiences across multiple healthcare settings [18].
Despite these advances, several challenges remain in deploying multi-agent systems in clinical practice. Issues of system verification, safety assurance, and regulatory compliance require careful consideration [19]. Additionally, the explainability of multi-agent decisions remains a critical concern, particularly for high-stakes medical decisions where understanding the reasoning process is as important as the outcome itself [20]. Future research directions include the development of more sophisticated coordination mechanisms, improved uncertainty quantification, and better integration with existing clinical workflows [17].

2.3. LLMs in Clinical Reasoning

Large language Models have demonstrated remarkable capabilities in medical question answering and clinical reasoning tasks. The Med-PaLM 2 system represents a significant milestone, achieving up to 86.5% accuracy on the MedQA dataset and showing substantial improvements across multiple medical question-answering benchmarks [21]. This performance notably exceeds the previous state-of-the-art and approaches human expert-level performance on standardized medical examinations [21]. Through techniques including ensemble refinement and chain of retrieval, Med-PaLM 2 answers were preferred to those from physicians on eight of nine clinical axes in evaluations, including factuality, precision, medical consensus, and reduced likelihood of harm [21]. The system’s architecture combines an improved base LLM (PaLM 2) with medical domain-specific fine-tuning and novel prompting strategies that enhance reasoning and grounding capabilities [21].
Studies have shown that both closed-source and open-source LLMs can pass medical licensing examinations, with GPT-3.5 achieving 60.2% accuracy on MedQA-USMLE and Llama 2 70B reaching 62.5% accuracy These results demonstrate that LLMs can mobilize expert medical knowledge and reasoning capabilities to address challenging medical questions Beyond multiple-choice formats, LLMs have shown proficiency in long-form medical question answering, with human evaluations indicating that their responses often exceed physician answers in quality and empathy for consumer health queries [21]. The development of comprehensive evaluation frameworks like MultiMedQA, which spans professional medical exams, medical research, and consumer queries, has been instrumental in benchmarking these advancements [21].
The integration of retrieval augmentation techniques has further enhanced LLM performance in medical domains. Chain of retrieval approaches enable models to ground their responses in relevant medical sources, improving factuality and reducing hallucinations [22]. This is particularly important in clinical contexts where accuracy and reliability are paramount. Retrieval-Augmented Generation (RAG) systems address critical limitations of LLMs, including domain knowledge gaps, factuality issues, and outdated information, by augmenting them with external knowledge from databases and medical literature [22]. These systems have evolved from Naive RAG to Advanced RAG and Modular RAG architectures, incorporating improvements in retrieval quality, embedding optimization, and post-retrieval processing [22,23].
Recent research has explored specialized frameworks that enhance clinical reasoning through structured rationalization processes. Kwon et al. [24] proposed a reasoning-aware diagnosis framework that generates diagnostic rationales through prompt-based learning, creating Clinical Chain-of-Thought (Clinical CoT) pathways that mirror clinical reasoning processes. This approach addresses the challenge of expensive rationale annotation by leveraging prompt-generated rationales that provide insights into patient data interpretation and diagnostic reasoning paths. Empirical demonstrations show that LLMs can effectively generate clinically plausible rationales and improve diagnostic accuracy through this reasoning-aware approach.
The application of LLMs in clinical reasoning extends beyond question answering to practical clinical workflows [24]. In pilot studies using real-world medical questions from consultation services, specialists preferred Med-PaLM 2 answers over generalist physician answers 65% of the time, while both specialists and generalists rated the model’s answers as equally safe as physician responses [21]. This suggests emerging utility for LLMs in supporting information needs where specialist access is limited, though specialist answers remain superior for complex cases [21]. The development of multimodal medical LLMs like Med-PaLM M further expands these capabilities by incorporating imaging data from chest X-rays, mammograms, and other diagnostic modalities alongside textual information [21].
Despite these advances, significant challenges remain in ensuring the safe and effective deployment of LLMs in clinical environments. Issues of hallucination, bias amplification, and security vulnerabilities require careful mitigation through rigorous validation, ethical considerations, and appropriate guardrails [21]. Furthermore, current benchmarks may not fully capture performance in real-world clinical workflows, which requires the continued development of comprehensive evaluation frameworks that address the nuances and complexities of actual clinical practice [21]. Future research directions include improving specialization for clinical subdomains, improving multimodal reasoning capabilities, and developing more sophisticated safety assurance protocols for clinical implementation [21].

2.4. Gap Analysis and Novel Contribution

While existing research has made significant advances in in silico patients, agent-based systems, and clinical LLMs separately, our review identifies a critical gap in integrating pharmacokinetic modeling with LLM-powered agent systems. Current in silico patient systems primarily focus on clinical reasoning and diagnostic training but lack sophisticated pharmacological simulation capabilities [2,5]. Similarly, agent-based healthcare models typically address coordination and decision-making without incorporating detailed pharmacokinetic/pharmacodynamic (PK/PD) modeling [5,15].
Although LLMs have demonstrated impressive capabilities in medical knowledge and reasoning, they remain limited in their ability to simulate physiological drug responses and predict pharmacological outcomes [6,21]. For instance, the Med-PaLM system, while exceptional in medical question-answering, does not incorporate dynamic pharmacological modeling [21].
This research proposes an innovative integration of sophisticated pharmacokinetic modeling with LLM-powered clinical agents, creating an in silico patient simulation that combines physiological drug response prediction with natural language clinical interactions. This integration enables more comprehensive clinical training scenarios that address both diagnostic reasoning and pharmacological management, filling a critical gap in medical simulation technology. It is important to emphasize that the proposed use of ALMs in pharmacokinetic research is not intended to replace classical pharmacokinetic analysis, clinical trials, regulatory evaluation, or expert pharmacological judgment. In this framework, the ALM does not determine drug efficacy or safety at the population level. Instead, it operates as a constrained reasoning and orchestration component within an in silico simulation environment. The pharmacokinetic behavior is governed by explicit compartmental equations, while the ALM supports contextual interpretation, clinical workflow coordination, and structured decision generation. This distinction is particularly important because pharmacological decisions can affect large patient populations; therefore, the ALM component is restricted by deterministic PK/PD models, safety validation, drug-interaction checks, and schema-constrained outputs.
In addition to existing approaches, recent systems such as AutoPK have explored the use of Large Language Models for pharmacokinetic data extraction from complex tables and documents. While these approaches demonstrate strong capabilities in information retrieval and structured data extraction, they do not address dynamic pharmacokinetic simulation or clinical decision-making. In contrast, the proposed system integrates pharmacokinetic modeling with agent-based reasoning and safety validation, enabling end-to-end simulation of treatment decisions rather than static data extraction. This distinction highlights the complementary nature of retrieval-focused and simulation-focused approaches in the broader landscape of AI-assisted pharmacology.
Recent approaches such as AutoPK [25] leverage Large Language Models and hybrid similarity metrics to extract pharmacokinetic parameters from complex scientific tables. While these methods achieve high accuracy in structured data extraction tasks, they focus on information retrieval rather than dynamic simulation or clinical decision-making. In contrast, the proposed system integrates pharmacokinetic modeling with agent-based reasoning and safety validation, enabling end-to-end simulation of treatment decisions. This highlights a fundamental distinction between retrieval-oriented and simulation-oriented approaches in AI-assisted pharmacology (see Table 1).

3. Methodological Proposal

3.1. Agentic Architecture

The in silico patient system employs a multi-agent architecture that cleanly separates physiological modeling from clinical decision-making while maintaining tight integration through well-defined interfaces. At its core, the design follows a hub-and-spoke pattern with the simulation orchestrator acting as the central coordinator that maintains system state and mediates all inter-agent communication [26]. This architectural choice ensures deterministic behavior in the pharmacokinetic core while allowing flexible reasoning in the clinical decision agents.
Three primary agent classes constitute the system’s intelligence layer. The physiological agents, comprising the pharmacokinetic model and patient state tracker, form the simulation backbone with strict mathematical foundations. These components maintain a continuous, time-stepped representation of drug concentrations and vital signs, updated through differential equations that account for patient-specific parameters. The clinical agents, including the doctor and patient actors, handle qualitative aspects of care through constrained natural language interactions. Between these layers sits the safety monitor, which acts as a guardrail system that validates all proposed clinical actions against pharmacological knowledge bases.
The architecture makes several deliberate design tradeoffs. First, it prioritizes modularity over monolithic integration, allowing components like the LLM interface or knowledge retrieval system to be upgraded independently. Second, it enforces unidirectional data flow from physiological to clinical agents to prevent feedback loops that could destabilize the simulation. Third, the system implements a hybrid synchronization model where the physiological simulation runs on fixed time steps while clinical agents operate in event-driven fashion, balancing computational efficiency with responsiveness.
Communication between agents follows a strict schema-based protocol where all messages must conform to predefined JSON structures containing both clinical content and control metadata. This design ensures interoperability while enabling comprehensive logging and audit capabilities. The system maintains a clear separation between its deterministic core (handling drug concentrations and vital signs) and probabilistic periphery (managing clinical decisions), with validation gates at each transition point between these domains.
Knowledge integration occurs at multiple levels. The RAG system provides evidence-based context to clinical decisions, while the drug interaction database informs safety checks. This layered approach to knowledge representation allows the system to combine static pharmacological data with dynamically retrieved clinical guidelines. The architecture positions the knowledge base as a first-class component rather than an implementation detail, reflecting its critical role in maintaining clinical accuracy.
The agent design reflects careful consideration of clinical workflow patterns. The doctor agent implements a three-phase decision cycle (assessment, planning, validation) mirroring real-world clinical reasoning, while the patient agent models symptom progression and reporting behaviors. This workflow-aware design helps maintain plausibility in simulated clinical encounters and ensures decisions follow medically coherent sequences rather than isolated actions [27].
The proposed architecture (see Figure 1) employs a hybrid design pattern combining reactive agents with a centralized orchestrator:

3.1.1. Physiological Simulation Layer

The physiological simulation layer forms the foundation of the in silico patient system, comprising two core components that work in tandem. The PKModel implements a sophisticated two-compartment pharmacokinetic model that accounts for both patient-specific and drug-specific parameters. For individualized dosing, it dynamically adjusts calculations based on clinical parameters like creatinine clearance when determining renal drug elimination. The model incorporates fundamental pharmacokinetic properties including the absorption rate ( k a ), intercompartmental transfer rates ( k 12 , k 21 ), apparent central volume of distribution ( V c ), and drug clearance ( C l ), which are specific to each medication. These parameters drive continuous concentration updates through first-order kinetic calculations, enabling precise simulation of drug metabolism over time.
Complementing the pharmacokinetic model, the PatientState component maintains a comprehensive temporal record of the in silico patient’s status. This includes continuous monitoring of vital signs such as heart rate (HR), blood pressure (BP), respiratory rate (RR), and oxygen saturation (SpO2) as time-series data. The component also tracks evolving symptom profiles with severity gradations and maintains a complete chronological record of medication administrations, including precise timestamps for accurate pharmacokinetic modeling.

3.1.2. Clinical Decision Layer

The clinical decision layer orchestrates the system’s intelligent behaviors through three specialized components. The DoctorAgent serves as the primary clinical decision-maker, employing a structured three-phase reasoning process that mirrors professional medical workflow: comprehensive patient assessment, therapeutic planning, and rigorous validation. It enhances decision quality by incorporating relevant patient demographics and comorbidities into its contextual prompting strategy, while maintaining robust fallback procedures to handle potential LLM non-compliance scenarios.
Working in concert with the physician agent, the PatientAgent manages all aspects of symptom reporting and patient interaction. It employs dynamic models of symptom progression that evolve based on both disease processes and therapeutic interventions. The agent generates natural language descriptions of symptoms that are strictly constrained to the patient’s actual physiological state, with additional modulation based on simulated affective states to increase behavioral realism.
Completing the clinical triad, the SafetyMonitor provides continuous medication surveillance through multiple protective mechanisms. It performs real-time screening for potential drug-drug interactions, generates dosing alerts adjusted for specific organ function parameters, and monitors therapeutic ranges while analyzing temporal trends to identify developing concerns before they reach critical levels.

3.1.3. Infrastructure Layer

The infrastructure layer provides essential supporting services for the entire system. The OllamaClient manages all LLM communications with enterprise-grade reliability features, including intelligent connection pooling and automated retry logic for fault tolerance. It enforces strict schema compliance through structured prompt templating and validates all responses against established clinical ontologies to ensure semantic correctness.
For knowledge-intensive operations, the Medical RAG system delivers evidence-based clinical information through advanced retrieval techniques. The infrastructure layer provides the foundational services enabling evidence-based clinical operations. At its core, the Medical RAG system leverages a comprehensive knowledge graph constructed from authoritative pharmacological sources, most notably [28]. This knowledge graph encodes multiple layers of pharmacological information. It incorporates drug mechanism-of-action relationships derived from primary pharmacology literature, therapeutic indications and contraindications as documented in FDA labeling, and metabolic pathways curated from PharmGKB as well as clinical pharmacokinetic studies.
The Medical RAG system employs a multi-stage retrieval process that begins with graph traversal through ontological relationships. This foundational stage navigates conceptual hierarchies, such as the pharmacological pathway connecting specific medications like morphine to their drug class (opioids) and ultimately to their physiological effects (CNS depression).
Following conceptual identification, the system executes precise evidence retrieval from source materials, extracting verbatim excerpts with exact page references to maintain academic rigor. This ensures all generated recommendations remain tethered to primary sources while supporting clinical verification.
The final stage performs intelligent relevance scoring through a dual-mechanism approach. First, it calculates semantic similarity using neural embeddings to evaluate conceptual alignment. Second, it applies clinical context matching that weights evidence based on current patient parameters, giving priority to information most pertinent to the active case scenario. This combined scoring method enables dynamic adaptation to diverse clinical contexts while maintaining evidence-based fidelity.
For the management of warfarin, the system is capable of simultaneously retrieving multiple sources of relevant information. This includes dosing protocols as recommended by the CHEST guidelines [29], pharmacodynamic details documented in [28], and genetic considerations outlined in the CPIC guidelines [30] for CYP2C9 and VKORC1. This architecture ensures all generated recommendations maintain direct lineage to citable sources while enabling real-time synthesis across multiple evidence streams. The knowledge graph currently contains 17,432 conceptual nodes and 48,917 relational edges, covering all major drug classes in the referenced compendia.
It processes clinical guidelines into high-dimensional vector embeddings, employs hierarchical document chunking to maintain contextual relationships, and synthesizes evidence with relevance weighting to surface the most pertinent information for each clinical scenario. This sophisticated retrieval architecture enables the system to combine the breadth of medical literature with the precision needed for clinical decision support.

3.2. Agent Communication Protocol

The system employs a structured JSON message format with schema validation at each hop as follows. The system architecture incorporates several critical design elements to ensure robust clinical functionality. First, a comprehensive temporal context is maintained through the precise timestamping of all clinical data entries, enabling sophisticated temporal reasoning about disease progression and treatment effects. This temporal awareness supports both retrospective analysis and real-time decision making.
For measurement integrity, the system enforces strict International System of Units compliance across all data inputs and outputs, supplemented by automated conversion utilities. This normalization eliminates potential errors from unit mismatches while maintaining flexibility for clinical preferences in data entry.
A foundational feature is the detailed provenance tracking system, which meticulously records the complete lineage of clinical actions from initial symptom identification through final treatment decisions. This audit capability provides full transparency into the decision-making process, supporting both quality assurance and regulatory compliance requirements. The provenance framework captures not just final actions but all considered alternatives and the evidence behind each choice.

3.3. Doctor Agent

3.3.1. Design

The Doctor Agent embodies a hybrid clinical reasoning architecture that combines protocol-driven decision making with context-aware judgment. Designed to simulate an experienced physician, the agent operates through a multi-stage cognitive process incorporating long-horizon reasoning, strategic planning, and safety validation. This structured approach ensures clinically coherent decisions while maintaining the flexibility needed for complex patient presentations [31].
At its core, the agent implements a dual-knowledge system integrating both explicit clinical protocols and implicit pattern recognition. The protocol system contains codified treatment guidelines for common scenarios like pain management and vital sign abnormalities, providing deterministic decision pathways. Complementing this, the contextual reasoning system employs retrieval-augmented generation to incorporate relevant medical knowledge based on the patient’s current state, allowing for nuanced adaptation to complex cases.
The decision-making process follows a safety-constrained workflow where each proposed intervention undergoes multiple validation checks. These include allergy screening, organ function adjustments, and drug interaction analysis. The system employs defensive clinical reasoning by default, prioritizing medication safety over aggressive treatment even when presented with severe symptoms. This conservative bias reflects real-world medical practice where “first, do no harm” remains paramount.
A key innovation in the agent’s design is its hierarchical error recovery system. When primary decision pathways fail, the agent gradually returns to simpler heuristics while maintaining clinical validity. This graceful degradation ensures that the system remains operational even with incomplete information or external service interruptions, a critical feature for clinical environments. The agent maintains comprehensive decision provenance through structured logging that captures not just final actions but the complete reasoning chain. Each decision includes the clinical context considered, alternative options evaluated, and safety checks performed. This audit trail enables both retrospective analysis and real-time decision transparency.
The agent adheres to established clinician–patient communication protocols, dynamically adjusting its interaction patterns based on symptom severity. For acute conditions, it employs closed-ended, focused questioning to rapidly identify critical issues. In chronic or complex cases, it shifts to an exploratory approach to uncover subtle contributing factors.
The design emphasizes cognitive load management through information prioritization. Vital signs, active medications, and known allergies receive automatic prominence in decision contexts, while less critical details remain available but don’t dominate the reasoning process. This reflects how human physicians triage clinical data during patient assessments.

3.3.2. Architecture

The Doctor agent implements a three-layer reasoning architecture that combines symbolic rule-based systems with statistical learning approaches. At the infrastructure level, the system utilizes a microservices design pattern in which discrete clinical functions operate as independent but coordinated services. This modular approach enables parallel processing of vital sign analysis, drug interaction checking, and treatment planning through dedicated subsystems. The Doctor Agent implements a three-layer reasoning structure combining probabilistic, protocol-driven, and safety-constrained decision-making:
  • LLM-based reasoning layer (high-level clinical reasoning): This layer uses LLaMA3 to generate clinical decisions (assessment, treatment plan, monitoring) based on the assembled patient context. It handles complex, context-dependent reasoning and produces structured outputs.
  • Protocol-driven layer (deterministic clinical logic): This layer encodes clinical guidelines and decision trees. It is used to validate or replace LLM outputs when necessary, ensuring that decisions remain consistent with established medical practice.
  • Rule-based safety layer (hard constraints): This layer enforces strict safety constraints, including drug–drug interactions, dose limits, allergies, and organ function adjustments. It acts as a final guardrail to block or modify unsafe actions.
These layers are integrated within the overall architecture as follows. The Simulation Orchestrator provides the global patient state and coordinates the decision cycle. The Doctor Agent processes this input through the three reasoning layers. All proposed actions are then validated by the SafetyMonitor, which interacts primarily with the rule-based layer to enforce safety constraints before execution.
The system utilizes a sophisticated hybrid ontology-based model that synthesizes multiple standardized frameworks to enable comprehensive clinical reasoning. At its core, it integrates SNOMED CT’s formal clinical terminology system [32] to ensure precise symptom classification and phenotypic characterization. This provides a robust foundation for capturing nuanced patient presentations using internationally recognized codes and relationships.
For pharmacological accuracy, the model incorporates RxNorm’s standardized medication concepts from the U.S. National Library of Medicine [33]. This inclusion enables unambiguous drug identification across different naming conventions (brand names, generic names, and various dosage forms), while maintaining connections to molecular mechanisms and therapeutic categories.
Complementing these established terminologies, the system embeds customizable rule sets that encode institutional protocols and local practice guidelines. These rules operate within the ontological framework to adapt general medical knowledge to specific organizational workflows and evidence-based treatment pathways. The integration of these three knowledge sources—SNOMED CT’s clinical semantics, RxNorm’s pharmacological precision, and localized protocol rules—creates a dynamic representation that supports both standardized decision-making and context-aware clinical judgment.
Decision pathways follow a probabilistic graphical model structure as follows
P ( T | S ) = i = 1 n P ( t i | s i ) · P ( s i | p a ( s i ) )
where T represents treatment options, S denotes variables of patient state, and  p a ( s i ) indicates parent nodes on the clinical decision graph. The safety validation subsystem implements real-time constraint checking using a temporal logic framework:
( allergy _ check ddi _ check organ _ function _ check )
where □ denotes the temporal logic always operator, indicating that the specified conditions must hold at every decision step. The predicates allergy _ check , ddi_check, and organ_function_check represent validation functions ensuring, respectively, the absence of patient-specific allergies, drug–drug interactions, and contraindications related to organ function (e.g., renal or hepatic impairment). This formulation defines a global safety invariant enforced by the SafetyMonitor, ensuring that all medication orders satisfy these constraints before execution.
For clinical context retrieval, the system utilizes vector embeddings defined within a 768-dimensional clinical concept space, where 768 corresponds to the embedding size of the underlying transformer-based model used to encode clinical text and concepts (e.g., BERT-based architectures). In this space, patient states and clinical knowledge are represented as dense vectors in R 768 , enabling similarity-based retrieval. This representation is complemented by hierarchical attention mechanisms to score the relevance of retrieved information. Additionally, dynamic query expansion is applied based on UMLS semantic types, enhancing the precision and comprehensiveness of information retrieval in complex clinical scenarios.
The agent’s reasoning process achieves logarithmic time complexity, O ( log n ) , for most clinical decisions. This efficiency is attained through the use of pre-computed therapeutic ranges stored in hash-mapped structures, bitmask indexing for allergy-related contraindications, and memoization of frequently encountered decision patterns.
Error handling is implemented through a tiered fallback protocol. The primary layer consists of LLM-based reasoning validated against predefined schemas. If this layer fails or produces uncertain outputs, the system reverts to protocol-driven decision trees. Finally, a symptom-triggered rule-based fallback ensures continued operation and clinical plausibility in cases where both higher-order reasoning strategies are insufficient.
The system incorporates a multi-stage error handling mechanism across layers. At the LLM level, outputs are validated through JSON schema checks, semantic verification of clinical entities, and safety evaluation via the SafetyMonitor. Typical errors include malformed JSON, missing required fields, or unsafe drug recommendations. When such errors are detected, the system transitions to a hierarchical fallback strategy rather than iterative re-prompting. Specifically, it moves from LLM-based reasoning to protocol-driven decision trees, and finally to rule-based safety heuristics that enforce conservative and clinically safe actions, such as blocking contraindicated drugs or adjusting doses. This design ensures robust and safe operation through deterministic fallback. The hierarchical error recovery system consists of three layers, organized as a tiered fallback mechanism with progressive simplification of reasoning:
  • Primary layer (LLM-based reasoning): Clinical decisions are first generated using the LLM (LLaMA3), constrained by structured prompting and validated through JSON schemas.
  • Secondary layer (protocol-driven decision trees): If the LLM output fails validation or is uncertain, the system falls back to deterministic clinical protocols encoded as decision trees.
  • Tertiary layer (rule-based fallback): If both previous layers fail, a final rule-based mechanism is activated, using symptom-triggered heuristics to ensure safe and clinically plausible actions.
This design ensures graceful degradation, allowing the system to maintain safe operation even under incomplete information or LLM failure scenarios.
The architecture supports hot-swappable knowledge components through versioned API endpoints, allowing for real-time updates to clinical guidelines without service interruption. All decision pathways maintain full reproducibility through deterministic random number seeding and comprehensive state snapshotting.

3.4. Patient Agent

3.4.1. Patient Agent Design

The Patient Agent implements a physiologically-grounded simulation architecture that models drug pharmacokinetics, symptom progression, and subjective experience with clinical fidelity [34]. Designed as a closed-loop biological system, the agent maintains dynamic internal state variables including drug concentrations, pain tolerance thresholds, and cumulative side effect burdens that evolve according to first-order pharmacokinetic principles.
The core simulation operates through a multi-layer symptom generation system where observed clinical manifestations emerge from the interaction between baseline pathology and drug effects. Physiological modeling follows concentration-response curves with drug-specific thresholds for adverse effects, creating realistic nonlinear dose-response relationships [35]. The agent implements adaptive homeostasis through mechanisms like opioid tolerance development, where repeated morphine exposure progressively reduces analgesic efficacy.
A key innovation in the agent’s design is its dual-representation symptom model that distinguishes between objective physiological changes and subjective experience. While drug concentrations deterministically generate physiological effects through the threshold-based system, the translation to patient-reported symptoms incorporates stochastic variability and individual difference factors. This separation allows the agent to realistically simulate the gap between measurable drug levels and patient experience.
The natural language generation subsystem employs a constrained creativity approach, where LLM outputs are tightly anchored to the physiological simulation state through prompt engineering and response validation. Descriptions follow a deterministic symptom-to-language mapping that ensures clinical coherence while maintaining natural expression variability. The system implements multiple validation layers including symptom presence verification, severity consistency checks, and grammatical structure analysis.

3.4.2. Patient Agent Architecture

The Patient Agent’s architecture implements a three-phase state transition model comprising pharmacokinetic absorption, physiological effect calculation, and symptom expression phases. The pharmacokinetic layer models drug metabolism using first-order elimination kinetics:
C t = C 0 · e k t
where C t represents current drug concentration, C 0 the initial dose, and k the elimination rate constant specific to each compound. The physiological effect calculator employs a sigmoidal response model for drug actions:
E = E m a x · C n E C 50 n + C n
where E represents the physiological effect magnitude, E m a x the maximum possible effect, E C 50 the half-maximal concentration, and n the Hill coefficient governing response steepness. The output of Equation (4) can be zero when the current drug concentration is zero, i.e.,  C = 0 . Under the usual assumptions that E max > 0 , E C 50 > 0 , and  n > 0 , the numerator becomes zero and therefore E = 0 . This corresponds to the absence of a drug-induced physiological effect. As C increases, the effect rises nonlinearly toward E max , with  E C 50 determining the concentration at which half of the maximum effect is reached. Side effect generation follows a threshold-triggered system where each adverse effect activates when its specific drug concentration exceeds a predetermined level:
S E a c t i v e = 1 if C d T s e 0 otherwise
with S E a c t i v e indicating the side effect presence, C d the current drug concentration, and  T s e the substance-specific threshold.
The natural language interface is built as a hybrid template-retrieval system. Base symptom expressions are first generated using deterministic templates, then enriched with qualitative descriptors retrieved from a Large Language Model. Finally, all outputs are validated for consistency with the simulated physiological state.
The clinical decision support integration uses a bidirectional coupling mechanism where treatment changes from external systems immediately update the agent’s internal drug concentrations, while symptom reports from the agent influence subsequent therapeutic decisions. This creates a closed-loop simulation environment that realistically models clinician–patient interactions.
The architecture maintains complete reproducibility through deterministic random number generation seeded by the physiological state, ensuring identical inputs always produce the same symptom reports and behavioral responses.

3.5. Pharmaceutical Data Model

The system incorporates a comprehensive pharmacological database structured around a Drug class that captures essential pharmacokinetic and pharmacodynamic properties. Each medication is characterized by a standardized set of attributes that enable precise simulation of drug behavior in in silico patients. The model encompasses twelve major drug classes including analgesics, anticoagulants, antibiotics, and psychotropic medications.
Pharmacokinetic parameters are specified with clinical precision, including molecular weight (ranging from 129.2 g/mol for metformin to 435.9 g/mol for rivaroxaban), bioavailability (from 30% for morphine to 95% for amoxicillin), and protein binding (varying from 0% for metformin to 99% for warfarin). Volume of distribution values reflect clinical observations, from 0.14 L/kg for warfarin indicating limited tissue distribution to 20 L/kg for sertraline demonstrating extensive tissue penetration.
Therapeutic ranges are carefully calibrated to clinical standards, such as 0.01–0.1 mg/L for morphine and 4–12 mg/L for carbamazepine. Toxic thresholds are similarly evidence-based, with acetaminophen hepatotoxicity set at 25 mg/L and warfarin bleeding risk escalating above 1.0 mg/L. Clearance rates span three orders of magnitude from warfarin’s 0.0026 L/h/kg to metoprolol’s 1.1 L/h/kg, accurately reflecting metabolic pathways.
Drug-drug interactions are cataloged with attention to clinical significance, with warfarin showing the most extensive interaction profile (seven documented combinations) while albuterol maintains a more selective pattern. The probabilities of side effects are stratified by severity, ranging from a common but mild (10–15% incidence of diarrhea with amoxicillin or metformin) to a rare but serious (0.01% risk of necrosis with warfarin) [36].
Specialized parameters address class-specific concerns: opioid respiratory depression (15% baseline risk for morphine), SSRI-induced serotonin syndrome (5% for sertraline), and antipsychotic extrapyramidal symptoms (40% for haloperidol). The model differentiates between concentration-dependent effects (such as warfarin anticoagulation) and idiosyncratic reactions (such as carbamazepine-induced rash).
The available drugs include agents like ondansetron (antiemetic) and omeprazole (PPI), demonstrating the system’s extensibility. Each new agent follows the same rigorous parameterization, with omeprazole’s interaction with clopidogrel and warfarin specifically flagged due to clinically significant CYP450 effects [37].

3.6. Pharmacokinetic Modeling

The pharmacokinetic core is based on a standard two-compartment model with first-order gastrointestinal absorption, central–peripheral distribution, and systemic clearance, following classical compartmental pharmacokinetic formulations [38,39]. Similar mass-balance formulations are widely used to describe drug exchange between absorption, central, and peripheral compartments [35]. The renal-function adjustment in Equation (9) is used as a simplified simulation rule to account for reduced clearance in patients with impaired renal function, consistent with pharmacokinetic dose-adjustment principles based on renal function [40,41]. The model is defined as:
d C g d t = k a C g
d C p d t = k a C g k 12 C p + k 21 C t C l a d j V c C p
d C t d t = k 12 C p k 21 C t
where C g = Drug amount (or concentration) in the gastrointestinal absorption compartment, C p = Plasma concentration in the central compartment (mg/L), C t = Tissue concentration in the peripheral compartment (mg/L), k a = Absorption rate constant (h−1), k 12 and k 21 = Intercompartmental rate constants (h−1), V c = Apparent volume of the central compartment (L), and  C l a d j = Clearance adjusted for organ function (L/h):
C l a d j = C l n o r m a l × 0.6 + 0.4 × eGFR 90
The model handles non-linear kinetics, such as Michaelis-Menten metabolism [42] for relevant drugs, protein binding for unbound fraction calculations, and accumulation for steady-state prediction for chronic dosing.

3.7. LLM Integration Architecture

The Doctor Agent employs a multi-stage prompting strategy:
Context Assembly
Clinical Context:
- Patient: 78 yo F, 62 kg, CrCl 42 mL/min
- Active Issues: Post-op hip fracture, nausea
- Current Meds: Morphine 2 mg IV q4 h,
warfarin 5 mg daily
- Latest Vitals: HR 88, BP 132/84, RR 18
Decision Prompt
Generate ONLY JSON output per schema:
{
  "actions": [{
    "type": "medication|monitoring|procedure",
    "details": {…},
    "priority": "emergency|urgent|routine"
  }],
  "rationale": {
    "pathophys": "text",
    "evidence": ["citation1", "citation2"]
  },
  "monitoring_plan": [{
    "parameter": "string",
    "frequency": "string"
  }]
}
The validation pipeline performs a series of structured checks to ensure correctness and safety. Initially, a syntax check is conducted using JSON schema validation to confirm that the data structure adheres to the expected format. This is followed by a semantic check, which verifies that all referenced medications exist within the formulary. Finally, a safety check is executed by cross-referencing the data with the SafetyMonitor module to detect potential risks or contraindications.
For example, a valid response generated by the Doctor Agent for a simulated elderly patient with postoperative pain and renal impairment may take the following form:
{
  "actions": [
    {
      "type": "medication",
      "details": {
        "drug": "morphine",
        "dose": "reduced dose",
        "route": "IV",
        "reason": "renal impairment and advanced age"
      },
      "priority": "routine"
    },
    {
      "type": "monitoring",
      "details": {
        "parameter": "respiratory rate and sedation score",
        "frequency": "every 4 h"
      },
      "priority": "urgent"
    }
  ],
  "rationale": {
    "pathophys": "Pain control is required,
    but opioid exposure should be limited
    because of age-related and renal clearance concerns.",
    "evidence": ["renal dosing adjustment", "opioid safety monitoring"]
  },
  "monitoring_plan": [
    {
      "parameter": "respiratory rate",
      "frequency": "every 4 h"
    },
    {
      "parameter": "pain score",
      "frequency": "every 4 h"
    }
  ]
}
Typical incorrect outputs include malformed JSON, missing required fields, unsupported medication names, or unsafe medication recommendations. For instance, an output that omits the monitoring_plan field is rejected by schema validation. A recommendation containing a medication not present in the formulary is rejected by semantic validation. A recommendation involving a contraindicated drug or excessive dose is blocked by the SafetyMonitor. In these cases, the system does not execute the LLM output directly; instead, it falls back to protocol-driven decision rules or conservative rule-based safety actions.
The selection of LLaMA3 via Ollama was driven primarily by architectural considerations rather than model benchmarking. The main contribution of this work lies in the integration of pharmacokinetic modeling with agentic LLM-based reasoning. Using a general-purpose LLM enables clearer attribution of system performance to the proposed architecture, rather than to domain-specific pretraining.
In addition, many specialized medical LLMs are only accessible through proprietary APIs, which limits their applicability in privacy-sensitive environments. In contrast, LLaMA3 deployed via Ollama supports local execution, which is particularly relevant in healthcare contexts where data governance and regulatory compliance are critical. Furthermore, the system is intentionally designed to be model-agnostic, allowing future integration of specialized medical LLMs without architectural modifications.
The use of LLaMA3 via Ollama provides several practical advantages:
  • local deployment, ensuring data privacy and regulatory compliance;
  • reduced latency, enabling real-time interaction within the multi-agent architecture;
  • improved reproducibility through the use of open-weight models; and
  • reliable integration via schema-constrained outputs, achieving 98.7% structured compliance in our experiments.
At the same time, this choice introduces several limitations. As a general-purpose model, LLaMA3 lacks domain-specific medical fine-tuning, requiring compensation through retrieval-augmented generation and rule-based safety mechanisms. It is also susceptible to hallucinations and remains sensitive to prompt design and context construction. To mitigate these issues, clinical correctness in the proposed system does not rely solely on the LLM, but is enforced through deterministic pharmacokinetic models, safety validation, and knowledge-based constraints.

4. Experimental Approach

4.1. Patient Synthetic Scenarios

The Doctor Agent was evaluated across 11 fully synthetic patient scenarios representing common clinical challenges. All patient scenarios used in this evaluation are fully synthetic and were manually designed for simulation purposes. They do not correspond to real patient records, do not include identifiable clinical information, and therefore did not require ethics approval or informed consent. Table 2 shows a summary of the different patient scenarios. The goal of each patient scenario is detailed as follows.
  • Elderly patient with multiple comorbidities. A 78-year-old female (65 kg) presents with moderate chronic pain, hypertension, diabetes, and chronic kidney disease (GFR 60 mL/min). This scenario evaluates the system’s ability to manage polypharmacy risks in geriatric patients while balancing pain control with renal dosing adjustments. The patient’s penicillin allergy further tests the agent’s allergy screening capabilities when suggesting antibiotic alternatives for potential infections.
  • Young healthy adult. A 25-year-old male (75 kg) with severe acute pain and no comorbidities serves as a baseline control scenario. This tests the system’s fundamental analgesic decision-making in an uncomplicated case, where optimal dosing can be achieved without special considerations for organ function or drug interactions.
  • Middle-aged patient with liver disease. A 55-year-old male (80 kg) with cirrhosis (liver function 45%) and alcohol use disorder presents with moderate abdominal pain. This scenario challenges the system’s hepatic dosing adjustments and its ability to avoid hepatotoxic medications while managing pain in a patient with a substance use history.
  • Pediatric patient with asthma. A 12-year-old female (42 kg) with moderate asthma and allergic rhinitis presents with wheezing and shortness of breath. This evaluates pediatric-specific dosing calculations and the system’s awareness of contraindications in patients with respiratory conditions and environmental allergies (dust mites, pollen).
  • Post-cardiac surgery patient. A 68-year-old male (82 kg) status-post cardiac surgery with severe coronary artery disease presents with mild chest pain and severe incision pain. This complex case tests the system’s ability to balance analgesic needs against cardiovascular risks while respecting the patient’s sulfa drug allergy.
  • Cancer patient on chemotherapy. A 45-year-old female (58 kg) with breast cancer and chemotherapy-induced neutropenia presents with severe nausea and fatigue. This scenario evaluates the system’s management of chemotherapy side effects while avoiding medications that might exacerbate neutropenia or interact with cancer therapies.
  • Geriatric patient with polypharmacy. An 82-year-old female (61 kg) with osteoporosis, hypertension, diabetes, and renal impairment (GFR 50 mL/min) presents with confusion and dizziness. This tests the system’s ability to identify medication-induced cognitive effects in elderly patients taking multiple medications while adjusting for renal function.
  • Trauma patient with multiple injuries. A 32-year-old male (88 kg) with fractures, lacerations, and headache presents with morphine allergy. This acute trauma scenario evaluates the system’s ability to prioritize pain management while working around opioid allergies and coordinating care for multiple concurrent injuries.
  • HIV patient with opportunistic infection. A 38-year-old male (63 kg) with advanced HIV and oral candidiasis presents with fever and severe fatigue. This tests the system’s knowledge of antiretroviral interactions and its ability to select appropriate treatments for opportunistic infections while avoiding the patient’s sulfamethoxazole allergy.
  • Pregnancy with hypertensive complications. A 29-year-old pregnant female (72 kg) with preeclampsia and gestational diabetes presents with headache and mild abdominal pain. This scenario evaluates the system’s understanding of medication risks in pregnancy, particularly for antihypertensive selection in preeclampsia while managing gestational diabetes.
  • Obese Patient with Sleep Apnea. A 41-year-old male (120 kg) with severe obesity, sleep apnea, and hypertension presents with daytime sleepiness. This tests the system’s dosing adjustments for obesity and its avoidance of respiratory depressants in patients with sleep-disordered breathing.

4.2. Clinical Evaluation

The pain reduction metric is computed from simulated patient states and does not use real patient outcome data. Pain is represented as a numerical symptom severity score within the in silico patient model. For each scenario, pain reduction is calculated as the relative decrease between the initial pain score before intervention and the final pain score after the simulated treatment period:
Pain Reduction ( % ) = P initial P final P initial × 100 ,
where P initial denotes the simulated baseline pain score at the beginning of the scenario, and P final denotes the simulated pain score after the agent-selected intervention. Thus, the values reported in Table 3 reflect simulated treatment response within the in silico patient environment rather than real clinical outcomes.
Descriptive statistics were computed over the six representative scenarios reported in Table 3. For each metric, we report the mean, standard deviation (SD), observed range, and 95% confidence interval (CI) using the Student t distribution because of the small number of representative scenarios. These statistics are descriptive and are intended to summarize the behavior of the simulated scenarios rather than to support clinical generalization. Across the six representative scenarios, pain reduction showed a mean of 46.33% (SD = 8.33; range = 37–58%; 95% CI: 37.59–55.08). The number of safety flags showed a mean of 1.67 alerts per scenario (SD = 1.21; range = 0–3; 95% CI: 0.40–2.94). Decision time showed a mean of 2.17 s (SD = 0.28; range = 1.8–2.5 s; 95% CI: 1.87–2.46).
No inferential statistical test was applied in the present study because the evaluation was based on synthetic scenarios and was not designed as a comparative trial against a baseline method or expert-annotated real-world cohort. Therefore, the reported results should be interpreted as descriptive evidence of internal system behavior and operational feasibility. Formal statistical testing will require larger repeated simulation batches or real-world annotated cases, which are planned as future work.
The system was evaluated across 11 clinically diverse scenarios (Table 3) with rigorous outcome tracking.
To provide an expert-reference comparison, representative system outputs were contrasted with expected clinical recommendations derived from standard pharmacological reasoning, drug safety principles, and the cited guideline sources used in the knowledge base. This comparison does not constitute an external independent clinical validation, but it provides a structured assessment of whether the system behavior is aligned with expected expert decision patterns in the synthetic scenarios (see Table 4).
These six scenarios were selected to provide coverage of geriatric, pediatric, trauma, pregnancy, hepatic impairment, and obesity-related decision contexts.
Key performance observations reveal several important trends. Age correlation was evident, with elderly patients exhibiting 22% more safety flags compared to younger adults. The impact of patient complexity was also pronounced, as those with three or more comorbidities required 2.6 times more dose adjustments. Finally, the system demonstrated high LLM performance, achieving 98.7% compliance with structured output requirements across 1243 clinical decisions. The system shows the ability to generate clinically appropriate management strategies across diverse scenarios.

4.3. Geriatric Case Management (78F)

The system exhibited appropriate pharmacological judgment in managing a 78-year-old female patient with multiple comorbidities. For opioid-induced nausea, it correctly prescribed metoclopramide 10 mg intravenously, following evidence-based guidelines for antiemetic therapy in elderly patients. The system also appropriately adjusted warfarin dosing in response to elevated INR values, demonstrating effective anticoagulation management. Notably, it implemented a 25% reduction in morphine dosing to account for age-related decline in renal function, showcasing its ability to integrate pharmacokinetic principles with clinical decision-making.

4.4. Therapeutic Monitoring Performance

The system’s therapeutic monitoring capabilities proved particularly robust, identifying 11 clinically significant drug interactions through its hierarchical screening process. These included 4 major contraindicated combinations that would have required immediate medication changes, 5 moderate interactions necessitating enhanced monitoring, and 2 minor interactions of limited clinical significance.
Furthermore, the system flagged 7 instances requiring renal dose adjustments, with an average dose reduction of 32%. This demonstrates its ability to appropriately modify medication regimens based on patient-specific factors, particularly in cases of renal impairment. The precision of these adjustments suggests effective implementation of established dosing guidelines for patients with compromised kidney function.

4.5. Safety System Performance

The integrated safety architecture demonstrated robust performance across three critical dimensions of medication safety monitoring.
First, the system achieved perfect detection sensitivity, correctly identifying 100% of known contraindications in the test scenarios. This flawless performance indicates comprehensive coverage of drug-drug and drug-condition interactions in the knowledge base, ensuring no hazardous combinations went undetected. The system’s hierarchical severity classification (contraindicated, major, moderate, minor) proved particularly valuable in clinical prioritization.
Second, analysis revealed a false positive rate of 8%, with most unnecessary alerts originating from the conservative dosing rules applied to geriatric patients. While this conservative approach generates some non-critical alerts, it reflects the system’s precautionary design principle, favoring patient safety over alert minimization. The false positives primarily involved excessive caution with CNS-active medications in elderly patients, mirroring real-world clinical practice patterns.
Third, the system maintained excellent responsiveness with 92% of safety checks completing in under 500 milliseconds. This fast response time enables real-time decision support during clinical workflow without disruptive delays. The remaining 8% of checks requiring longer processing involved complex polypharmacy scenarios with 5 or more concurrent medications, where comprehensive interaction screening demands greater computational resources.
The safety system’s architecture, combining rule-based checks, pharmacokinetic modeling, and LLM-powered context assessment, proved particularly effective at balancing sensitivity and specificity while maintaining clinically acceptable response times.

4.6. Limitations

Despite the promising results, several limitations of the proposed approach must be acknowledged. First, the system inherits known limitations of Large Language Models, particularly the risk of hallucinations, where the model may generate plausible but factually incorrect clinical statements. Although this risk is mitigated through JSON schema constraints, safety validation, and retrieval-augmented generation, it cannot be fully eliminated.
Second, the use of a general-purpose LLM (LLaMA3) introduces domain knowledge limitations compared to specialized medical models. While the integration of a knowledge graph and evidence-based retrieval partially compensates for this, the system’s clinical reasoning remains dependent on the quality, coverage, and updating of external knowledge sources.
Third, the approach exhibits dependency on prompt design and context construction. Variations in prompt structure, incomplete patient information, or suboptimal retrieval may affect the consistency and robustness of generated decisions. This sensitivity is particularly relevant in multi-agent settings, where intermediate outputs influence downstream reasoning.
Fourth, although the architecture incorporates multiple safety mechanisms, including pharmacokinetic constraints and interaction checks, it has not been validated in real clinical environments. Therefore, the system should be interpreted as a research-oriented simulation and decision-support framework rather than a substitute for clinical judgment.
A further limitation concerns the translation of ALM-supported reasoning into pharmacokinetic research. While ALMs can assist in scenario interpretation, decision workflow coordination, and simulation of clinical reasoning, they should not be interpreted as substitutes for validated pharmacometric models, regulatory pharmacokinetic studies, or human expert review. This is especially relevant because drug-related decisions may affect populations rather than isolated individuals. Consequently, the present system should be regarded as an exploratory in silico simulation framework whose outputs require further validation before any clinical, regulatory, or population-level pharmacological application.
Finally, the evaluation is limited to synthetic patient scenarios. While these scenarios were designed to cover diverse clinical conditions, they do not fully capture the variability, uncertainty, and complexity of real-world patient data. The number of scenarios used in the present evaluation is also limited. The 11 synthetic cases were selected to provide initial coverage of clinically relevant high-risk contexts, including geriatric care, renal impairment, hepatic impairment, pregnancy, pediatric disease, trauma, polypharmacy, and respiratory vulnerability. However, this sample size is not sufficient to establish full mathematical or clinical reliability. The results should therefore be interpreted as preliminary evidence of internal plausibility and operational feasibility rather than as definitive validation of clinical performance.
Future work will expand the evaluation to a substantially larger scenario library with repeated simulation runs, broader drug classes, wider demographic variation, and systematically varied comorbidity profiles. This larger evaluation will enable more robust estimation of variability, sensitivity, specificity, confidence intervals, and failure modes. In addition, expert-annotated cases and, where ethically approved, anonymized real-world clinical data will be required to assess external validity. Moreover, future work should include validation using real clinical datasets and prospective evaluation in realistic clinical workflows. Future research will extend the evaluation to real-world clinical data through a staged validation process. First, the system will be tested retrospectively using anonymized patient records and medication histories to evaluate pharmacokinetic plausibility, safety-flag accuracy, and agreement with expert clinical judgment. Second, prospective studies may be conducted in simulated or supervised clinical environments, where clinicians review system-generated recommendations without direct autonomous execution. Any study involving real patient data would require institutional ethics approval, informed consent where applicable, appropriate data governance procedures, and strict clinician oversight.

5. Conclusions

This system represents a significant contribution to clinical simulation technology, successfully integrating deterministic pharmacokinetics with LLM-based clinical reasoning while maintaining rigorous safety standards. Specifically, this research advances the field through three key technical contributions:
First, it validates a hybrid neuro-symbolic architecture where the Safety Monitor acts as a deterministic guardrail, achieving 100% sensitivity in detecting contraindications and effectively mitigating LLM hallucinations through rigorous JSON-schema validation, which maintained a 98.7% compliance rate. Second, it introduces a novel dual-layer in silico patient model that distinctively separates physiological state from subjective symptom expression, allowing for realistic, non-linear dose-response simulations that mirror biological complexity. Finally, the empirical evaluation confirms the system’s operational viability for real-time interaction, demonstrating an average decision latency of just 2.1 s while achieving significant clinical efficacy, evidenced by a 39–58% reduction in pain scores across scenarios.
The modular architecture has demonstrated scalability across diverse patient scenarios, with particular strength in complex geriatric and polypharmacy cases. Future development will focus on expanding therapeutic domains, incorporating genetic factors, and enhancing real-time performance, moving toward a new standard for accessible, high-fidelity medical decision support.
These results position the proposed approach as a complementary advancement to existing LLM-based pharmacological tools, extending beyond data retrieval systems toward integrated simulation and decision-support frameworks. The proposed framework should therefore be interpreted as a controlled simulation architecture for exploring ALM-assisted pharmacokinetic reasoning, rather than as a direct clinical or regulatory decision-making system.

Funding

This research received no external funding.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The author declares no conflicts of interest.

References

  1. Triola, M.M.; Campion, N.; McGee, J.B.; Albright, S.; Greene, P.; Smothers, V.; Ellaway, R. An XML Standard for Virtual Patients: Exchanging Case-Based Simulations in Medical Education. AMIA Annu. Symp. Proc. 2007, 2007, 741–745. [Google Scholar]
  2. Kononowicz, A.A.; Woodham, L.A.; Edelbring, S.; Stathakarou, N.; Davies, D.; Saxena, N.; Car, L.T.; Carlstedt-Duke, J.; Car, J.; Zary, N. Virtual Patient Simulations in Health Professions Education: Systematic Review and Meta-Analysis by the Digital Health Education Collaboration. J. Med. Internet Res. 2019, 21, e14676. [Google Scholar] [CrossRef]
  3. Cook, D.A.; Erwin, P.J.; Triola, M.M. Computerized virtual patients in health professions education: A systematic review and meta-analysis. Acad. Med. 2010, 85, 1589–1602. [Google Scholar] [CrossRef] [PubMed]
  4. Issenberg, S.B.; McGaghie, W.C.; Petrusa, E.R.; Lee Gordon, D.; Scalese, R.J. Features and uses of high-fidelity medical simulations that lead to effective learning: A BEME systematic review. Med. Teach. 2005, 27, 10–28. [Google Scholar] [CrossRef] [PubMed]
  5. Holderried, F.; Stegemann-Philipps, C.; Herschbach, L.; Moldt, J.A.; Nevins, A.; Griewatz, J.; Holderried, M.; Herrmann-Werner, A.; Festl-Wietek, T.; Mahling, M. A Generative Pretrained Transformer (GPT)-Powered Chatbot as a Simulated Patient to Practice History Taking: Prospective, Mixed Methods Study. JMIR Med. Educ. 2024, 10, e53961. [Google Scholar] [CrossRef]
  6. Liévin, V.; Hother, C.E.; Motzfeldt, A.G.; Winther, O. Can large language models reason about medical questions? Patterns 2024, 5, 100943. [Google Scholar] [CrossRef]
  7. Lee, K.; Lee, S.; Kim, E.H.; Ko, Y.; Eun, J.; Kim, D.; Cho, H.; Zhu, H.; Kraut, R.E.; Suh, E.; et al. Adaptive-VP: A Framework for LLM-Based Virtual Patients that Adapts to Trainees’ Dialogue to Facilitate Nurse Communication Training. arXiv 2025, arXiv:2506.00386. [Google Scholar]
  8. Hicke, Y.; Geathers, J.; Rajashekar, N.; Chan, C.; Jack, A.G.; Sewell, J.; Preston, M.; Cornes, S.; Shung, D.; Kizilcec, R. MedSimAI: Simulation and Formative Feedback Generation to Enhance Deliberate Practice in Medical Education. arXiv 2025, arXiv:2503.05793. [Google Scholar]
  9. Roux, P.; Okuya, Y.; Morel, C.; Soulès, M.; Bottemanne, H.; Brunet-Gouet, E.; Frileux, S.; Passerieux, C.; Younes, N.; Martin, J.C. Effectiveness of a Web-Based Virtual Simulation to Train Nursing Students in Suicide Risk Assessment: Randomized Controlled Investigation. JMIR Serious Games 2025, 13, e69347. [Google Scholar] [CrossRef]
  10. Salmeron, J.L.; Rahimi, S.A.; Navali, A.M.; Sadeghpour, A. Medical diagnosis of Rheumatoid Arthritis using data driven PSO–FCM with scarce datasets. Neurocomputing 2017, 232, 104–112. [Google Scholar] [CrossRef]
  11. Safdari, R.; Shoshtarian Malak, J.; Mohammadzadeh, N.; Danesh Shahraki, A. A Multi Agent Based Approach for Prehospital Emergency Management. Bull. Emerg. Trauma 2017, 5, 171–178. [Google Scholar]
  12. Chávez-Juárez, F.; Hackett, L.; Trujillo, G.; Blasco, A. A Multi-purpose Agent-Based Model of the Healthcare System. In Advances in Social Simulation; Ahrweiler, P., Neumann, M., Eds.; Springer: Cham, Switzerland, 2021; pp. 409–413. [Google Scholar] [CrossRef]
  13. Vemuri, A.; Decker, K.; Saponaro, M.; Dominick, G. Multi Agent Architecture for Automated Health Coaching. J. Med. Syst. 2021, 45, 95. [Google Scholar] [CrossRef]
  14. Li, Y.; Lawley, M.A.; Siscovick, D.S.; Zhang, D.; Pagán, J.A. Agent-Based Modeling of Chronic Diseases: A Narrative Review and Future Research Directions. Prev. Chronic Dis. 2016, 13, E69. [Google Scholar] [CrossRef]
  15. Yue, L.; Xing, S.; Chen, J.; Fu, T. ClinicalAgent: Clinical Trial Multi-Agent System with Large Language Model-based Reasoning. arXiv 2024, arXiv:2404.14777. [Google Scholar]
  16. Chen, X.; Yi, H.; You, M.; Liu, W.; Wang, L.; Li, H.; Zhang, X.; Guo, Y.; Fan, L.; Chen, G.; et al. Enhancing Diagnostic Capability with Multi-Agents Conversational Large Language Models. Npj Digit. Med. 2025, 8, 159. [Google Scholar] [CrossRef] [PubMed]
  17. Wang, J. Shapley Value Based Multi-Agent Reinforcement Learning: Theory, Method and Its Application to Energy Network. arXiv 2024, arXiv:2402.15324. [Google Scholar] [CrossRef]
  18. Kaissis, G.A.; Makowski, M.R.; Rückert, D.; Braren, R.F. Secure, privacy-preserving and federated machine learning in medical imaging. Nat. Mach. Intell. 2020, 2, 305–311. [Google Scholar] [CrossRef]
  19. Dong, Z.; Omidshafiei, S.; Everett, M. Collision Avoidance Verification of Multiagent Systems with Learned Policies. arXiv 2024, arXiv:2403.03314. [Google Scholar] [CrossRef]
  20. Hildt, E. What Is the Role of Explainability in Medical Artificial Intelligence? A Case-Based Approach. Bioengineering 2025, 12, 375. [Google Scholar] [CrossRef]
  21. Singhal, K.; Tu, T.; Gottweis, J.; Sayres, R.; Wulczyn, E.; Hou, L.; Clark, K.; Pfohl, S.; Cole-Lewis, H.; Neal, D.; et al. Toward expert-level medical question answering with large language models. Nat. Med. 2025, 31, 943–950. [Google Scholar] [CrossRef]
  22. Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; Wang, M.; Wang, H. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv 2023, arXiv:2312.10997. [Google Scholar]
  23. Ke, Y.H.; Jin, L.; Elangovan, K.; Abdullah, H.R.; Liu, N.; Sia, A.T.H.; Soh, C.R.; Tung, J.Y.M.; Ong, J.C.L.; Kuo, C.F.; et al. Retrieval augmented generation for 10 large language models and its generalizability in assessing medical fitness. npj Digit. Med. 2025, 8, 187. [Google Scholar] [CrossRef]
  24. Kwon, T.; iunn Ong, K.T.; Kang, D.; Moon, S.; Lee, J.R.; Hwang, D.; Sim, Y.; Sohn, B.; Lee, D.; Yeo, J. Large Language Models Are Clinical Reasoners: Reasoning-Aware Diagnosis Framework with Prompt-Generated Rationales. Proc. AAAI Conf. Artif. Intell. 2024, 38, 18417–18425. [Google Scholar] [CrossRef]
  25. Sholehrasa, H.; Ghanaatian, A.; Caragea, D.; Tell, L.A.; Riviere, J.E.; Jaberi-Douraki, M. AutoPK: Leveraging LLMs and a Hybrid Similarity Metric for Advanced Retrieval of Pharmacokinetic Data from Complex Tables and Documents. In Proceedings of the IEEE International Conference on Tools with Artificial Intelligence (ICTAI), Athens, Greece, 3–5 November 2025; pp. 338–346. [Google Scholar] [CrossRef]
  26. Beckenbauer, L.; Loewe, J.L.; Zheng, G.; Brintrup, A. Orchestrator: Active Inference for Multi-Agent Systems in Long-Horizon Tasks. arXiv 2025, arXiv:2509.05651. [Google Scholar]
  27. Almansoori, M.; Kumar, K.; Cholakkal, H. Self-Evolving Multi-Agent Simulations for Realistic Clinical Interactions. arXiv 2025, arXiv:2503.22678. [Google Scholar]
  28. Goodman, L.S.; Gilman, A.; Brunton, L.L.; Chabner, B.; Knollmann, B.C. The Pharmacological Basis of Therapeutics, 13th ed.; McGraw-Hill Education: New York, NY, USA, 2018. [Google Scholar]
  29. Stevens, S.M.; Woller, S.C.; Kreuziger, L.B.; Bounameaux, H.; Doerschug, K.; Geersing, G.J.; Huisman, M.V.; Kearon, C.; King, C.S.; Knighton, A.J.; et al. Antithrombotic Therapy for VTE Disease: Second Update of the CHEST Guideline and Expert Panel Report. Chest 2021, 160, e545–e608. [Google Scholar] [CrossRef]
  30. Johnson, J.A.; Caudle, K.E.; Gong, L.; Whirl-Carrillo, M.; Stein, C.M.; Scott, S.A.; Lee, M.T.M.; Gage, B.F.; Kimmel, S.E.; Perera, M.A.; et al. Clinical Pharmacogenetics Implementation Consortium (CPIC) Guideline for Pharmacogenetics-Guided Warfarin Dosing: 2017 Update. Clin. Pharmacol. Ther. 2017, 102, 397–404. [Google Scholar] [CrossRef] [PubMed]
  31. Hong, S.; Xiao, L.; Zhang, X.; Chen, J. ArgMed-Agents: Explainable Clinical Decision Reasoning with LLM Discussion via Argumentation Schemes. arXiv 2024, arXiv:2403.06294. [Google Scholar]
  32. El-Sappagh, S.; Franda, F.; Ali, F.; Kwak, K.S. SNOMED CT standard ontology based on the ontology for general medical science. BMC Med. Inform. Decis. Mak. 2018, 18, 76. [Google Scholar] [CrossRef]
  33. Nelson, S.J.; Zeng, K.; Kilbourne, J.; Powell, T.; Moore, R. Normalized names for clinical drugs: RxNorm at 6 years. J. Am. Med. Inform. Assoc. 2011, 18, 441–448. [Google Scholar] [CrossRef]
  34. Chang, R.; Jiao, H.; Nie, W.; Guo, H.; Xie, K.; Wu, Z.; Zhao, L.; Bai, Y.; Ma, Y.; Wang, L.; et al. Organ-Agents: Virtual Human Physiology Simulator via LLMs. arXiv 2025, arXiv:2508.14357. [Google Scholar] [CrossRef]
  35. Zou, H.; Banerjee, P.; Leung, S.S.Y.; Yan, X. Application of Pharmacokinetic-Pharmacodynamic Modeling in Drug Delivery: Development and Challenges. Front. Pharmacol. 2020, 11, 997. [Google Scholar] [CrossRef]
  36. Holbrook, A.M.; Pereira, J.A.; Labiris, R.; McDonald, H.; Douketis, J.D.; Crowther, M.; Wells, P.S. Systematic overview of warfarin and its drug and food interactions. Arch. Intern. Med. 2005, 165, 1095–1106. [Google Scholar] [CrossRef] [PubMed]
  37. U.S. Food and Drug Administration. Prilosec (Omeprazole) Delayed-Release Capsules: Prescribing Information. 2012. Available online: https://www.accessdata.fda.gov/drugsatfda_docs/label/2012/019810s096lbl.pdf (accessed on 7 May 2026).
  38. Gibaldi, M.; Perrier, D. Pharmacokinetics, 2nd ed.; Marcel Dekker: New York, NY, USA, 1982. [Google Scholar]
  39. Derendorf, H.; Schmidt, S. Rowland and Tozer’s Clinical Pharmacokinetics and Pharmacodynamics: Concepts and Applications, 5th ed.; Wolters Kluwer: Philadelphia, PA, USA, 2020. [Google Scholar]
  40. Therapeutic Goods Administration. Guideline on the Evaluation of the Pharmacokinetics of Medicinal Products in Patients with Decreased Renal Function. EMA/CHMP/83874/2014. 2024. Available online: https://www.tga.gov.au/resources/resources/international-scientific-guidelines-adopted-australia/guideline-evaluation-pharmacokinetics-medicinal-products-patients-decreased-renal-function (accessed on 6 March 2026).
  41. Food and Drug Administration. Pharmacokinetics in Patients with Impaired Renal Function: Study Design, Data Analysis, and Impact on Dosing. Guidance for Industry. 2024. Available online: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/pharmacokinetics-patients-impaired-renal-function-study-design-data-analysis-and-impact-dosing (accessed on 6 March 2026).
  42. Saganuwan, S. Application of modified Michaelis–Menten equations for determination of enzyme inducing and inhibiting drugs. BMC Pharmacol. Toxicol. 2021, 22, 57. [Google Scholar] [CrossRef] [PubMed]
Figure 1. High-level architecture showing UI/orchestrator, specialized agents, LLM backend, and knowledge base with data/control (solid) and reasoning/RAG (dashed) links.
Figure 1. High-level architecture showing UI/orchestrator, specialized agents, LLM backend, and knowledge base with data/control (solid) and reasoning/RAG (dashed) links.
Mathematics 14 01627 g001
Table 1. Comparison of the proposed system with related approaches.
Table 1. Comparison of the proposed system with related approaches.
FeatureAutoPK [25]Medical LLMs [21]Agent-Based Systems [15]Proposal
PK data extraction×
PK simulation×××
LLM-based reasoningLimitedLimited
Multi-agent architecture××
Safety validation (DDI, dosing)×LimitedLimited
Real-time decision support×LimitedLimited
Table 2. Summary of patient scenarios for system evaluation.
Table 2. Summary of patient scenarios for system evaluation.
ScenarioAge/SexWeight (kg)Comorbidities/
Key Conditions
Clinical PresentationMain Testing Focus
Elderly patient with multiple comorbidities78 F65Hypertension, diabetes, chronic kidney disease (GFR 60), penicillin allergyModerate chronic painPolypharmacy management, renal dosing adjustments, allergy screening
Young healthy adult25 M75NoneSevere acute painBaseline analgesic decision-making
Middle-aged patient with liver disease55 M80Cirrhosis (45% liver function), alcohol use disorderModerate abdominal painHepatic dosing adjustments, hepatotoxicity avoidance, substance use considerations
Pediatric patient with asthma12 F42Asthma, allergic rhinitisWheezing, shortness of breathPediatric dosing, contraindication awareness for respiratory conditions and allergies
Post-cardiac surgery patient68 M82Coronary artery disease, sulfa allergyMild chest pain, severe incision painBalancing analgesia with cardiovascular risk, allergy considerations
Cancer patient on chemotherapy45 F58Breast cancer, chemotherapy-induced neutropeniaSevere nausea, fatigueManagement of chemotherapy side effects, drug interactions, neutropenia avoidance
Geriatric patient with polypharmacy82 F61Osteoporosis, hypertension, diabetes, renal impairment (GFR 50)Confusion, dizzinessIdentifying medication-induced cognitive effects, renal dosing adjustments
Trauma patient with multiple injuries32 M88Morphine allergyFractures, lacerations, headacheAcute pain management with opioid allergy, coordination for multiple injuries
HIV patient with opportunistic infection38 M63Advanced HIV, oral candidiasis, sulfamethoxazole allergyFever, severe fatigueAntiretroviral interaction knowledge, appropriate treatment selection for opportunistic infections
Pregnancy with hypertensive complications29 F72Preeclampsia, gestational diabetesHeadache, mild abdominal painMedication safety in pregnancy, antihypertensive selection, gestational diabetes management
Obese patient with sleep apnea41 M120Severe obesity, sleep apnea, hypertensionDaytime sleepinessDosing adjustments for obesity, avoidance of respiratory depressants
Table 3. Representative simulation outcomes for 6 of the 11 evaluated clinical scenarios.
Table 3. Representative simulation outcomes for 6 of the 11 evaluated clinical scenarios.
ScenarioAge/SexPain ReductionSafety FlagsDecision Time
Elderly Complex78 F42%32.4 s
Trauma Multiple32 M49%11.8 s
Pregnancy HTN29 F58%02.1 s
Obesity OSA41 M39%22.3 s
Liver Cirrhosis55 M37%32.5 s
Pediatric Asthma12 F53%11.9 s
Table 4. Comparison between system behavior and expert-reference recommendations in representative synthetic scenarios.
Table 4. Comparison between system behavior and expert-reference recommendations in representative synthetic scenarios.
ScenarioExpert-Reference ExpectationSystem BehaviorAgreement
Elderly complex patientAvoid aggressive opioid escalation; adjust dosing according to renal function; monitor polypharmacy risks.Reduced morphine dosing, flagged safety risks, and generated monitoring recommendations.Consistent
Trauma patient with morphine allergyAvoid morphine and select non-contraindicated analgesic alternatives; prioritize acute pain control.Detected opioid allergy and avoided unsafe morphine recommendation.Consistent
Pregnancy with hypertensionAvoid medications with pregnancy-related safety concerns; prioritize safer monitoring and conservative treatment.Generated conservative recommendations without high-risk medication escalation.Consistent
Liver cirrhosisAvoid hepatotoxic drugs and consider reduced hepatic clearance.Flagged hepatic impairment and avoided hepatotoxic treatment paths.Consistent
Obesity with sleep apneaAvoid respiratory depressants or recommend cautious dosing and monitoring.Flagged respiratory-risk context and applied conservative safety constraints.Consistent
Pediatric asthmaConsider age-appropriate dosing and avoid treatments that may worsen respiratory symptoms.Applied pediatric-specific safety checks and respiratory-condition awareness.Consistent
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Salmeron, J.L. Pharmacokinetics-Informed Agentic Architecture for Drug Dynamics with LLM-Driven In Silico Patients. Mathematics 2026, 14, 1627. https://doi.org/10.3390/math14101627

AMA Style

Salmeron JL. Pharmacokinetics-Informed Agentic Architecture for Drug Dynamics with LLM-Driven In Silico Patients. Mathematics. 2026; 14(10):1627. https://doi.org/10.3390/math14101627

Chicago/Turabian Style

Salmeron, Jose L. 2026. "Pharmacokinetics-Informed Agentic Architecture for Drug Dynamics with LLM-Driven In Silico Patients" Mathematics 14, no. 10: 1627. https://doi.org/10.3390/math14101627

APA Style

Salmeron, J. L. (2026). Pharmacokinetics-Informed Agentic Architecture for Drug Dynamics with LLM-Driven In Silico Patients. Mathematics, 14(10), 1627. https://doi.org/10.3390/math14101627

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop