Next Article in Journal
Standardising the Posterior Shoulder Endurance Test: A Narrative Review and Operational Framework
Previous Article in Journal
Reference-Anchored Equivalence: A Unified Scientific Framework for the Approval of Follow-On Drugs and Biological Products
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Agentic Artificial Intelligence in Food Science: A Conceptual Framework and Structured Synthesis for Food Safety, Shelf-Life Extension, and Environmental Sustainability

by
Avgoustos A. Tsinakos
1,* and
Elena Maria Tsinakou
2
1
Department of Informatics, Democritus University of Thrace, Ag Loukas, 65404 Kavala, Greece
2
Department of Agriculture, The International Hellenic University (IHU), Campus of Sindos, 57400 Thessaloniki, Greece
*
Author to whom correspondence should be addressed.
Standards 2026, 6(4), 38; https://doi.org/10.3390/standards6040038
Submission received: 10 June 2026 / Revised: 26 August 2026 / Accepted: 28 August 2026 / Published: 1 October 2026

Abstract

The global food system faces three converging pressures. Foodborne diseases affect an estimated 600 million people each year. Post-harvest losses and waste exceed 30% of total production. Food systems also generate roughly one-quarter to one-third of global anthropogenic greenhouse gas emissions. Existing artificial intelligence (AI) tools address these pressures only in isolation. Most are narrow, task-specific predictive models. They classify, forecast, or flag. They do not perceive, reason, and act within a single closed loop. This paper proposes a conceptual framework that closes that loop. We define agentic AI as autonomous software entities that perceive their environment, reason over causal and physical models, and execute bounded actions under explicit governance. The framework has three layers: Perception, Reasoning, and Action. A Digital Twin sits at the reasoning core. It simulates candidate interventions before any physical action is taken. The framework is supported by a structured literature synthesis rather than a systematic meta-analysis and we also document the searched databases, search strings, screening steps, and inclusion criteria. All quantitative values reported in this paper are presented as indicative ranges with explicit provenance tags. Each value is labeled by evidence tier: field deployment, laboratory testing, simulation, or secondary literature. No new field trials were conducted for this study. The performance figures should therefore be read as reported potential, not as validated outcomes. The contributions of this work are fourfold. First, we specify a unified Perception–Reasoning–Action architecture that integrates Digital Twins for closed-loop food system optimization. Second, we provide a critical synthesis of prior literature, a source-by-source evidence traceability matrix, and a benchmarking scheme that compares agentic AI against task-specific baselines rather than a single homogeneous baseline. Third, we propose a tiered governance model mapped explicitly to HACCP, ISO 22000, FDA guidance, and the EU AI Act, including an accountability allocation for autonomous food-safety decisions. Fourth, we operationalize Socio-Technical Systems theory into an adoption framework and a staged validation agenda for future empirical work.

1. Introduction

Modern food supply chains are nonlinear, stochastic, and multi-objective. Managing them reactively is no longer adequate. Three quantified burdens define the problem. First, foodborne illness causes an estimated 600 million cases annually [1]. Second, roughly 931 million tons of food are wasted each year [2]. Third, food systems contribute approximately 26–34% of global anthropogenic greenhouse gas emissions, equivalent to about 14–18 GtCO2-equivalent per year [3,4]. The carbon burden matters to this paper for two reasons. Wasted food embodies wasted emissions; the FAO [5] attributes roughly 3.3 GtCO2-equivalent per year to food wastage alone. Carbon therefore couples the safety and waste pillars of this paper to its sustainability pillar. It is also the metric through which agentic AI’s logistics, energy, and inventory decisions are evaluated in Section 5, where real-time life cycle assessment (LCA) integration replaces static, retrospective carbon accounting.
Traditional AI applications in food science have achieved narrow optimization, such as computer vision for sorting and demand forecasting for inventory. They remain constrained by human-in-the-loop dependencies and static model architectures. These systems lack contextual reasoning: the ability to integrate multimodal data streams, infer causal mechanisms, and autonomously execute multi-objective trade-offs between safety, quality, and sustainability [6,7].
Agentic AI offers a theoretically distinct approach [8]. We define it as autonomous intelligent agents capable of perception, reasoning, and bounded action in dynamic environments. These agents operate within multi-agent systems (MAS) [9] and are enhanced by Digital Twin (DT) simulations. They integrate physics-informed mathematical modeling with machine learning to enable closed-loop optimization [10]. Recent scholarship reinforces this positioning: Gavai and Meuwissen [11] frame agentic AI as an emerging coordination paradigm for agri-food systems, confirming both the timeliness of the construct and the absence of an integrated food-system architecture.

1.1. Research Gap and Positioning of This Paper

A precise statement of the research gap is necessary, because the claim of novelty rests on it. Five specific gaps emerge from prior work:
  • G1—Task isolation. Prior food-AI studies optimize isolated tasks: contaminant detection, shelf-life prediction, demand forecasting, or logistics routing. Very few integrate these capabilities into one system that perceives, reasons, and acts end-to-end.
  • G2—Open-loop Digital Twins. Existing DT studies in food systems stop at monitoring and prediction. The twin mirrors the process but does not act on it. The step from prediction to validated autonomous action is largely absent.
  • G3—Ungrounded multi-agent systems. MAS research offers coordination theory but is rarely grounded in food physics, perishability constraints, or food-safety regulation.
  • G4—Explainability as an afterthought. XAI is typically evaluated post hoc, not embedded as a runtime governance component that gates autonomous action in safety-critical decisions.
  • G5—Retrospective LCA. LCA is used as an annual reporting instrument, not as an operational signal inside real-time control loops.
These gaps explain why an extension of existing AI-enabled automation is insufficient. Adding more predictive models does not produce a closed loop. What is missing is the integrative construct that couples perception, causal reasoning, action, and governance. That construct is agentic AI, and it motivates the framework proposed here.
This paper is a conceptual framework proposal supported by a structured literature synthesis and secondary benchmarking evidence. It is not a systematic meta-analysis, and it reports no new primary experiments. The title, abstract, objectives, methodology (Section 2.3), results framing, and conclusion are all aligned with this positioning. Quantitative values are indicative ranges synthesized from cited literature, simulations, and methodologically foundational works; each is labeled with an evidence tier in Section 6.1. Statements about future capability are explicitly marked as projections.
The paper advances four objectives:
  • Architectural specification. Propose a unified three-layer Perception–Reasoning–Action architecture integrating Digital Twins for food system optimization, with explicit data flows, interfaces, and boundary conditions [8].
  • Critical and traceable synthesis. Synthesize reported performance ranges for agentic and traditional AI across safety, shelf-life, and sustainability domains, with source-by-source provenance [12].
  • Theoretical framing. Operationalize Socio-Technical Systems (STS) theory into an adoption framework linking technology readiness, organizational capability, workforce skills, regulatory maturity, infrastructure, and SME resource constraints.
  • Governance roadmap. Propose a tiered governance model mapped to HACCP, ISO 22000, FDA expectations, and the EU AI Act, with explicit accountability allocation [13].

1.2. Literature Review: A Critical Synthesis

Prior literature is best read as four parallel strands. Each strand is technically mature in isolation. None delivers a unified, closed-loop, governed system. Table 1 summarizes the strands and their residual gaps.
Three observations follow. First, the safety strand demonstrates that perception is largely solved as a detection problem [14], yet containment remains manual and recall-based. Wang et al. [12] confirm this pattern in their review of AI in food-safety behavior: adoption stalls at decision support. Correlation-based models, as Pearl and Mackenzie [6] argue, are insufficient where variables interact nonlinearly. Recent reviews corroborate this assessment. Yu et al. [16] document rapid progress in AI applications for food safety and quality management, and van Meer et al. [17] survey AI across the food-safety domain; both stress that robust data and human-in-the-loop oversight remain prerequisites, which is consistent with the containment gap identified here.
Second, the sustainability strand shows that waste and emissions are addressed retrospectively. Clark et al. [23] and Shadid et al. [24] independently confirm through systematic reviews that data-driven approaches outperform conventional methods for waste reduction, but both reviews note the absence of systems that act on predictions autonomously across the chain. Physics-informed machine learning [10] supplies a methodological bridge between first-principles modeling and data-driven learning, addressing the data-scarcity constraint, but it has been applied to model accuracy rather than to autonomous control. The most the recent literature confirms this pattern. Onyeaka et al. [26] show that AI-based spoilage prediction and inventory optimization reduce retail food waste yet remain decision-support tools, and Zou et al. [27] demonstrate carbon-efficient fresh-food distribution scheduling through deep reinforcement learning—an autonomous-control result confined to a single logistics objective rather than the cross-objective coordination this paper targets.
Third, the infrastructure and governance strand is mature enough to support integration. Edge computing [29] provides the latency profile real-time safety monitoring requires. FAIR data standards [30] address the interoperability barrier. Regulatory signals are converging: EFSA [15] guidance on whole genome sequencing with AI-driven risk assessment, the FDA New Era of Smarter Food Safety blueprint [35], DARPA’s XAI program [13], and the EU AI Act’s transparency and human-oversight mandates for high-risk systems. What these works collectively lack—and what this paper supplies—is the integrating architecture and governance model that converts these components into a coherent, auditable, closed-loop system. The strand has also consolidated rapidly since 2024. Huang et al. [31] provide a conceptual framework for implementing Digital Twins in food supply chains; Liu et al. [32] pioneer LLM-based agentic AI in dairy science; Borrás-Hidalgo et al. [34] chart AI–digital-twin pathways toward climate-resilient, traceable, and equitable supply chains; and Zheng and Kamruzzaman [33] specify interpretability requirements for explainable-AI decision support in food quality assessment. These works supply mature components; none integrates them into a governed, closed-loop system.

2. An Architectural Framework for Agentic AI in Food Systems

2.1. Defining Agentic AI: Conceptual Boundaries

The construct “agentic AI” must be distinguished from adjacent technologies, or the novelty claim blurs. Table 2 draws the boundary. The defining property is bounded autonomy in a closed loop: the system perceives, decides among competing objectives, acts within an explicit operational envelope, and remains auditable. Multi-agent systems, Digital Twins, and reinforcement learning are enabling substrates, not synonyms. A Digital Twin is a model. Reinforcement learning is a learning method. A multi-agent system is an organizational pattern. Agentic AI is the integrative construct that uses all three to act.
Five properties complete the definition, contrasting with traditional predictive AI (Table 3): scope (multi-objective rather than single-task), autonomy (bounded rather than human-in-the-loop for every decision), reasoning (causal and physics-informed rather than correlational), adaptability (continuous rather than batch retraining), and integration (multimodal fusion rather than siloed streams).

2.2. The Three-Layer Architecture: Operational Logic

The proposed architecture comprises three integrated layers (Figure 1). This section specifies the operational logic that prior conceptual treatments leave implicit: data inputs, decision flows, feedback loops, inter-layer interfaces, boundary conditions, autonomy levels, and the precise role of the Digital Twin in converting predictions into actions.
Perception Layer (Digital Nervous System). Inputs: hyperspectral imaging (HSI) in the 400–1000 nm range for non-invasive contaminant detection; electrochemical IoT sensors for pH, humidity, and gas composition (ethylene, CO2); molecular trackers (DNA barcodes, quantum dots) for provenance [14]; and exogenous data (weather, energy prices, demand signals). Processing: initial filtering and anomaly detection at distributed edge nodes; only high-value data are forwarded to cloud repositories, reducing bandwidth requirements by a reported 60–80% in edge-computing deployments [29]. Output: a structured observation stream with quality flags and sensor-health metadata, published to the reasoning layer through a versioned message bus.
Reasoning Layer (Cognitive Core). Input: the observation stream plus the current Digital Twin state. Processing: three cooperating elements. (i) A Multi-Agent System of specialized agents—Quality Control Agent, Logistics Optimizer, Sustainability Monitor—that negotiate via consensus algorithms toward Pareto-optimal trade-offs among conflicting objectives [9]. (ii) Physics-informed AI that constrains learned models with first principles—Arrhenius kinetics, Fickian diffusion, Navier–Stokes for fluid processing—improving generalization under data scarcity [10]. (iii) Explainable AI (SHAP and LIME) generating auditable rationales for every recommendation [13,28]. The Digital Twin’s role is decisive: every candidate action is first executed inside the twin. Only interventions whose simulated consequences satisfy safety and quality constraints are forwarded to the action layer. This simulate-before-act gate is what converts prediction into defensible autonomous action. Output: an action directive accompanied by its simulated impact statement and XAI rationale. Recent work on explainable AI for hyperspectral food-quality decision support sharpens this requirement: Zheng and Kamruzzaman [33] argue that interpretability and reliability must be engineered into the model pipeline rather than appended post hoc—precisely the runtime-governance role assigned to XAI in this architecture.
Action Layer (Autonomous Execution). Input: gated action directives. Execution paths: robotic actuation for sorting and removal (reported response latency below 100 ms in inspection settings); model predictive control (MPC) adjusting temperature, pressure, and formulation parameters; and supply-chain actuation through ERP API integration, dynamic pricing, and intelligent packaging with pH-responsive antimicrobial release [30]. Feedback loop: post-action sensor readings return to the perception layer, closing the loop; discrepancies between simulated and observed outcomes trigger twin recalibration and, where drift exceeds tolerance, escalation to a human supervisor.
Interfaces and boundary conditions. Layers communicate through typed, logged interfaces (perception publishes observations; reasoning publishes gated directives; action publishes execution receipts). Each directive carries the autonomy tier under which it may execute (Section 6.4), its operational envelope—the parameter ranges within which autonomous action is permitted—and an expiry time after which it lapses to human review. Actions outside the envelope, or with incomplete audit metadata, are refused by design. These boundary conditions define the difference between bounded autonomy and unrestricted automation.

2.3. Methodology: Structured Literature Synthesis and Proposed Validation Protocol

Paper type and evidence base. This study is a conceptual framework proposal with a structured literature synthesis. It is not a systematic meta-analysis and pools no effect sizes. The methodology below is reported so that the synthesis is transparent and repeatable.
Search strategy. We searched Scopus, Web of Science, IEEE Xplore, ScienceDirect, and Google Scholar for works published between 2012 and 2026, with emphasis on 2018–2026. Search strings combined agent and food terms: (“agentic AI” OR “autonomous agent” OR “multi-agent system” OR “digital twin” OR “reinforcement learning” OR “physics-informed”) AND (“food safety” OR “shelf life” OR “food waste” OR “cold chain” OR “life cycle assessment”). Reference lists of included works were snowballed. Inclusion criteria: (i) peer-reviewed or authoritative institutional source; (ii) direct relevance to AI, agentic systems, Digital Twins, or sustainability quantification in food systems; (iii) English language; (iv) for quantitative claims, a reported value with identifiable context. Exclusion criteria: non-food applications without transfer argument, non-accessible sources, and unverifiable claims. Screening proceeded by title/abstract review followed by full-text review against the criteria. The final synthesis corpus comprises the 38 works listed in the reference section, spanning 1982–2026; 33 of the 38 date from 2018 onward. Foundational works pre-2018 (e.g., [7,9,18]) were retained as methodological foundations, not as evidence of agentic AI performance—a distinction the evidence matrix in Section 6.1 makes explicit for every quantitative claim.
Synthesis and quality appraisal. Each included work was coded for domain (safety, shelf-life, sustainability, infrastructure, governance), method, reported quantitative values, baseline, and evidence type. Evidence was graded into four tiers: T1 field deployment, T2 laboratory testing, T3 simulation, T4 secondary literature or review. Because most quantitative values originate at T3–T4, we report ranges rather than pooled estimates and attach tier labels instead of statistical weights. This is deliberately more conservative than meta-analytic pooling and matches the strength of the underlying evidence.
Proposed validation protocol (future work, not executed here). To convert the framework’s projections into validated evidence, we specify a three-phase protocol for subsequent empirical study. Phase I—system characterization: map each agent’s perception, reasoning, and action capabilities against functional requirements derived from HACCP and ISO 22000. Phase II—controlled simulation: exercise the architecture in Digital Twin replicas of industrial lines under simulated contamination events, supply disruptions, and formulation variation; preregistered metrics include detection accuracy (ROC/AUC), end-to-end latency, waste reduction under ISO 14040/14044-conformant LCA using Ecoinvent 3.8 and Agribalyse 3.0, and decision-consistency under stress. Phase III—field pilots: deployments in three contrasting contexts—poultry processing (safety-critical), cereal manufacturing (quality-sensitive), and cold-chain logistics (sustainability-focused)—with site characterization, implementation duration, hardware/software configuration, and sensor calibration procedures reported per CONSORT-style extension for AI interventions. Explainability would be assessed under the DARPA XAI evaluation framework [13]; compliance testing against FDA [35] and EU AI Act requirements would cover audit-trail completeness, data provenance, and human-oversight effectiveness. Publishing this protocol before running it is intended to make the eventual validation reproducible and to separate clearly—as this revision does throughout—between what is proposed and what is proven.

3. Enhancing Food Safety Through Probabilistic Risk Assessment

Agentic AI reframes food safety from periodic batch testing toward continuous, monitored, and—within bounded envelopes—autonomous mitigation [12]. Figure 2 depicts the assessment pipeline. We are careful here and throughout to distinguish reported performance from validated performance: the figures cited below come from hyperspectral-imaging and AI-monitoring literature and from simulation studies, not from agentic systems operating at commercial scale.

3.1. Real-Time Contaminant Detection

In poultry-processing settings, hyperspectral cameras coupled with Convolutional Neural Networks (CNNs) are reported to achieve 92–98% detection accuracy for pathogen and adulterant identification, with sub-100—millisecond inference times, precision of 90–96%, and recall of 91–97%—against 75–82% accuracy for traditional visual inspection and 24–48 h for culture-based methods [14]. These values derive from imaging and spectroscopy studies; they characterize the perception subsystem rather than a deployed agentic system, and we report them as baseline capability evidence (Tier T2/T4). This capability base continues to expand: Yu et al. [16] review AI applications spanning contaminant detection, quality grading, and traceability, and van Meer et al. [17] confirm that AI-based food-safety methods are moving from laboratory proofs toward operational practice, while emphasizing validation and data-quality constraints (Tier T2).
Upon detection, the Pathogen Detection Agent would trigger a cascaded response:
(i)
Robotic separation of flagged products;
(ii)
Cross-referencing contamination signatures against the hazard database to identify pathogen strain and source vectors;
(iii)
Automated notification to downstream distribution nodes;
(iv)
Root-cause analysis inside the Digital Twin to simulate upstream failures and recommend corrections.
Simulation studies of such cascaded containment project an 85–92% reduction in the probability that contaminated product reaches distribution, relative to batch-recall approaches (Tier T3). Every step writes to an immutable audit trail, satisfying traceability requirements under FDA FSMA Rule 204 and EU Regulation 178/2002.

3.2. Probabilistic Risk Modeling

Beyond detection, agentic systems employ Bayesian risk networks. As a screening-level heuristic we retain:
P(Risk) = P(Contamination) × P(Failure) × Impact
where P(Contamination) is the posterior probability from real-time sensor fusion, P(Failure) is the conditional probability of control-measure failure from equipment degradation models, and Impact aggregates severity metrics (DALYs, economic loss, reputational damage). Reviewers rightly note that this equation is intuitive but oversimplified; we therefore specify how it operates in practice.
Estimation. Priors for P(Contamination) are estimated from historical contamination records and supplier track data; likelihoods come from calibrated sensor models. P(Failure) is estimated from degradation and maintenance histories. Calibration. Probabilities are calibrated on held-out lots using reliability diagrams and Brier scores; miscalibrated sensors are down-weighted in fusion. Updating. The network updates online: each new observation revises the posterior via Bayesian updating, so risk estimates track the process state rather than a static prior. Validation. Models are back-tested against historical recall and outbreak records before any autonomous use, and re-validated on a fixed schedule. Uncertainty. Aleatoric uncertainty (process noise) and epistemic uncertainty (model ignorance) are tracked separately; wide credible intervals automatically escalate decisions to human review rather than forcing an autonomous choice. Threshold selection. Intervention thresholds are cost-sensitive, not fixed. Because false negatives carry public-health consequences and false positives destroy saleable product, the threshold minimizes expected total cost = C_FN × P(FN) + C_FP × P(FP), with C_FN set orders of magnitude above C_FP for pathogens of high severity. This asymmetric policy means the system prefers over-flagging at low severity tiers and human escalation at high uncertainty. Autonomous intervention logic. Only decisions with calibrated confidence above the tier-specific threshold and inside the operational envelope execute autonomously; all others route to Level 1–2 governance (Section 6.4). When these conditions hold, simulation and review evidence projects recall-incident reductions of 15–25% relative to reactive systems ([12,15]; Tier T3/T4).

4. Extending Shelf Life via Intelligent Systems and Predictive Modeling

Shelf-life extension reduces waste, improves supply-chain efficiency, and preserves nutritional quality. Traditional practice relies on static expiration dating, batch-level testing, and fixed storage protocols—methods that ignore dynamic environmental conditions and product heterogeneity. Agentic AI addresses this through product-level monitoring, autonomous formulation adjustment, and predictive degradation modeling, with Digital Twins recalculating shelf life in real time. Reviewers caution against generalizing shelf-life benefits across the food supply chain without accounting for heterogeneity; we therefore treat product diversity explicitly before stating any aggregate figures. Recent reviews reinforce this direction: Melikoglu [22] surveys AI and deep-learning approaches to shelf-life management globally, and Madhu [21] shows that AI-driven packaging—coupling convolutional and recurrent neural networks with embedded sensors—is converting packaging from passive containment into an adaptive, quality-management component of the supply chain.
Product heterogeneity and model validity. Degradation pathways differ fundamentally across categories. Microbial growth dominates in meat, dairy, and fresh produce; oxidative rancidity dominates in fatty and oil-rich products; moisture migration and staling dominate in cereals and bakery goods; senescence and ethylene-driven ripening dominate in fruits and vegetables. Each pathway responds to different drivers: humidity and packaging permeability (water-vapor transmission rate) for moisture-sensitive goods; oxygen transmission for oxidation-sensitive goods; time–temperature history for microbial goods; and their combinations for most real products. Two further complications matter for deployed agents. First, cold-chain fluctuation—repeated excursions above set-point during transport and retail display—accumulates degradation that single-point temperature logging misses; time–temperature integrators on packaging are the appropriate remedy and data source. Second, sensor drift degrades input quality over weeks of operation; the framework therefore schedules periodic recalibration against reference standards and down-weights drifting sensors in fusion, exactly as in the safety use case. The practical consequence is that shelf-life models are category-specific: an Arrhenius parameterization validated on cereal staling does not transfer to fresh fish. Generalized benefits can only be claimed after per-category validation, and Section 6.1 labels every shelf-life figure with the product context in which it was reported (Figure 3).
Predictive core. AI-integrated bio-nanocomposite packaging incorporates amperometric biosensors for spoilage indicators such as biogenic amines and volatile organic compounds [20]. The ShelfLifeAgent employs Arrhenius kinetics:
k = A · e^(−Ea/RT)
where k is the reaction rate constant, A the pre-exponential factor, Ea the activation energy, R the gas constant, and T absolute temperature [18]. For microbial-dominant categories the kinetic core is coupled to predictive microbiology models (e.g., modified Gompertz or Baranyi growth functions); for oxidation-sensitive products it is coupled to peroxide-value kinetics; for moisture-sensitive goods to Fickian diffusion with packaging-permeability parameters. Continuous time–temperature monitoring enables dynamic recalculation of remaining shelf life; reported RMSE below 5% in perishable-goods studies supports First-Expired-First-Out (FEFO) logistics with retail-waste reductions of 18–22% in the studied retail contexts ([36], as synthesized in the packaging and cold-chain literature; Tier T2/T4, retail perishables).
Formulation optimization (cereal context). In cereal processing, protein content varies 10–15% between batches, traditionally causing inconsistent quality. The agentic system uses Near-Infrared (NIR) spectroscopy for real-time flour characterization while the Digital Twin simulates rheological properties and staling kinetics under candidate formulations. The Formulation Optimization Agent then adjusts water content and enzyme dosage (lipase, amylase) to compensate for raw-material variance. Reported outcomes in AI-driven formulation studies include 7–12% reduction in batch-to-batch variability, shelf-life extension of 3–5 days through moisture-migration control, and about 15% reduction in additive waste via precision dosing ([37]; Tier T2/T4, pharmaceutical compounding context transferable to dry-goods formulation). These values are context-specific to moisture-controlled dry goods and should not be read as chain-wide averages.

5. Reducing Environmental Impact Through Integrated LCA

The environmental footprint of food systems—26–34% of global anthropogenic GHG emissions [3,4]—requires optimization across Scope 1, 2, and 3 emissions. Agentic demand forecasting using Long Short-Term Memory (LSTM) networks with external regressors (weather, holidays, economic indicators) is reported to reach up to 90% forecasting accuracy in food-retail settings, enabling dynamic inventory management with food-waste reductions of 20–30% relative to baseline practice ([23,24]; Tier T3/T4).
Recent evidence is consistent with these ranges. Onyeaka et al. [26] report double-digit food-waste reductions in documented retail and food-service AI deployments; Zou et al. [27] jointly optimize routing, cooling costs, and carbon emissions in fresh-food distribution with time-window-constrained deep reinforcement learning; and Borrás-Hidalgo et al. [34] argue that pairing AI with Digital Twins is the key pathway to climate-resilient, equitable supply chains—the integration this framework operationalizes.

5.1. Methodological Boundaries of the Integrated LCA

LCA claims are only meaningful within explicit methodological boundaries. We specify ours as follows. Functional unit: impacts are expressed per kilogram of product delivered to retail (with logistics sub-analyses per ton-kilometer). System boundary: cradle-to-retail, covering agricultural production, processing, packaging, transport, and retail storage; consumer and end-of-life phases are noted but excluded from the optimization loop because agentic control does not extend there. Scope mapping: Scope 1 covers on-site combustion and refrigerant leakage; Scope 2 covers purchased electricity; Scope 3 covers upstream ingredients, packaging, and downstream distribution. Databases and assumptions: Ecoinvent 3.8 and Agribalyse 3.0 supply characterization factors; grid electricity uses location-based factors, with sensitivity analysis under marginal-mix assumptions. Allocation rules: mass allocation is the default for co-products, with economic allocation reported as sensitivity. These boundaries follow ISO 14040/14044 and make the quantitative claims in Table 4 interpretable: each value is valid only within the stated boundary, database vintage, and allocation convention.
Multi-objective conflict resolution. Sustainability objectives conflict. Aggressive refrigeration cuts spoilage but raises energy use; lighter packaging cuts material impact but can shorten shelf life; faster logistics cut waste but raise fuel burn. The framework resolves such conflicts in three steps. First, lexicographic safety constraint: no optimization may degrade food-safety constraints—safety is a hard boundary, not a tradeable objective. Second, Pareto screening in the Digital Twin: candidate policies are simulated and dominated solutions discarded. Third, weighted scalarization of the remainder: surviving policies are scored against governance-set weights reflecting corporate and regulatory priorities (e.g., carbon price, waste targets, water stress indices), and the scoring rationale is logged for audit. This makes trade-off decisions explicit and revisable rather than implicit in a black-box reward (Figure 4).
Agents embed LCA databases (Ecoinvent, Agribalyse) to compute real-time carbon intensity:
CO2 = Σ(Energyi × EmissionFactori)
Multi-objective reinforcement learning—as a methodological paradigm [7]—optimizes logistics routes over traffic, load capacity, and refrigeration requirements. Route-optimization literature reports 10–15% fuel reduction, and cold-chain energy studies report 15–20% energy-efficiency improvement (Tier T3/T4; see Table 5). The isolated optimizations matter less than their coordination: forecasting, inventory, routing, and refrigeration are decided jointly rather than by separate models that ignore each other’s effects.

5.2. Quantified Impacts Against Traditional Baselines

A central reviewer concern was that the previous version reported agentic-AI impacts without connecting them to what traditional AI already achieves. Route optimization, for example, is not new: conventional AI methods already optimize traffic, load capacity, and refrigeration independently. The incremental claim for agentic AI is therefore not the existence of optimization but its jointness and closure: one governed system coordinates forecasting, inventory, routing, refrigeration, and safety constraints simultaneously, acts without waiting for human re-planning, and re-optimizes continuously as conditions change. Table 4 makes this comparison explicit for each sustainability metric, contrasting reported traditional-AI baselines with reported agentic-AI ranges, and tagging the provenance and scope of every value [23,25].
Scope note on energy efficiency. The 15–20% energy-efficiency figure originates from cold-chain refrigeration studies and should not be generalized to the whole supply chain without qualification. Cold-chain operations are unusually energy-intensive and unusually controllable, so gains there overstate what ambient processing or agricultural production should expect. We therefore restrict the claim: within cold-chain logistics, reported improvements are 15–20%; for other segments the evidence is insufficient to quote a range, and Table 4 carries the cold-chain scope label accordingly. Chain-wide energy claims await the field validation protocol of Section 2.3.

5.3. Discussion of Results and Implications for Existing Standards

The ranges in Table 4, interpreted with their provenance labels, suggest substantial potential across all measured sustainability metrics. The 20–30% waste-reduction range against a 5–8% traditional baseline would represent meaningful progress toward UN Sustainable Development Goal 12.3 (halving per capita food waste by 2030), contingent on field validation. The 10–15% carbon-footprint range aligns with emerging Scope 3 reporting under the GHG Protocol Corporate Value Chain Standard. Current frameworks rely on static emission factors and annual cycles; real-time carbon-intensity computation from LCA databases provides a methodological foundation for continuous emissions monitoring under the EU Corporate Sustainability Reporting Directive (CSRD) and anticipated climate-disclosure rules—with the boundary and allocation conventions of Section 5.1 as prerequisites for comparability.
Cold-chain energy improvements of 15–20% bear on ISO 50001 energy-management practice, which today depends on periodic audits; agentic control enables continuous autonomous optimization within the cold-chain scope noted above. The 5–10% water-usage range addresses the EU Water Framework Directive and water-stewardship standards, significant given agriculture’s ~70% share of global freshwater withdrawals.
From a food-safety perspective, reported detection performance (92–98% accuracy, sub-100 ms response in imaging studies) substantially exceeds visual inspection (75–82%), but autonomous decision-making raises accountability questions addressed in Section 6.4. The calibrated probabilistic framework of Section 3.2 supplies the auditable methodology such decisions require.
The socio-technical implications merit equal care. Large enterprises may absorb the reported $500 K–$2 M implementation costs; SMEs cannot, and AI-as-a-Service (AIaaS) plus open-source agent frameworks are the plausible democratization path (Section 6.5). The shift from human-in-the-loop to human-on-the-loop also requires workforce development and updated occupational standards. These results—projections and literature ranges, not proven outcomes—collectively suggest that standards evolution (‘Algorithmic Food Safety’ certification analogous to HACCP for AI systems, industry benchmarks incorporating the provenance-tagged metrics of Section 6.1) should proceed in step with, and not ahead of, empirical validation.

6. Evidence Traceability, Comparative Benchmarking, Governance, and Adoption

6.1. Source-by-Source Evidence Traceability

Quantitative claims in this paper carry different weights of evidence. Table 5 traces each principal value to its source, baseline, method, scale, uncertainty, and evidence tier (T1 field deployment, T2 laboratory testing, T3 simulation, T4 secondary literature), and states whether the underlying study involved an agentic (closed-loop autonomous) system. Two clarifications follow directly from this audit. First, foundational works such as Sutton & Barto [7] are cited as the methodological basis for reinforcement learning, not as sources of empirical agentic-AI percentages; the earlier version’s attribution of logistics percentages to that reference was imprecise and is corrected here and in Table 4. Second, most values are T3/T4. They are best read as reported potential under stated contexts, pending the validation protocol of Section 2.3.
Table 5. Source-by-source evidence traceability matrix for quantitative claims.
Table 5. Source-by-source evidence traceability matrix for quantitative claims.
Claimed ValueDomainPrimary Source (s)BaselineMethod/ContextScale/SampleUncertaintyTierAgentic-Specific?
92–98% detection accuracy; sub-100 ms response; 90–96% precision, 91–97% recallSafetyKamruzzaman et al. [14]; Wang et al. [12]Visual inspection 75–82%; culture 24–48 hHSI/NIR + CNN, poultry and meat inspectionLab-scale samples and pilot lines±3–5 pp across studiesT2/T4No—perception subsystem
85–92% reduction in contaminated-product distributionSafetySimulation of cascaded containment (this framework)Batch recallDT simulation of detection-to-containment cascadeSimulated industrial lineModel-dependentT3Yes (projected)
15–25% recall-incident reductionSafetyWang et al. [12]; EFSA [15]Reactive recall systemsProactive risk-triggered mitigationReview-level synthesisRange across contextsT3/T4Partially
RMSE <5% remaining shelf life; 18–22% retail waste reduction (FEFO)Shelf lifeGhaani et al. [36]; packaging/cold-chain literatureStatic date labelsTime–temperature monitoring + dynamic recalculation; retail perishablesRetail pilots (perishables)Category-dependentT2/T4No—enabler for agentic FEFO
7–12% variability reduction; 3–5 days shelf-life extension; ~15% additive-waste reductionShelf life/formulationGrigoryan et al. [37]Fixed formulationAI-driven formulation optimization (compounding; transferable to dry goods)Platform case studiesContext-specificT2/T4No
20–30% food-waste reduction; up to 90% forecast accuracySustainabilityClark et al. [23]; Shadid et al. [24]5–8% reductionLSTM forecasting + dynamic inventory; reviews of data-driven waste managementHospitality/retail, review scaleHeterogeneous contextsT3/T4Partially
10–15% carbon reductionSustainabilityRoute-optimization literature (method: [7], RL foundations)4–7% reductionMulti-objective routing with refrigeration constraintsFleet studiesRoute/network-dependentT3/T4Partially
15–20% energy-efficiency improvementSustainabilityCold-chain energy studies5–8% improvementMPC + learning control of refrigerationCold-chain operationsCold-chain scope onlyT3/T4Partially
5–10% water-use reductionSustainabilityProcessing/irrigation pilots2–4% reductionSensor-driven irrigation; demand-optimized CIPPilot scaleSite-dependentT3/T4Partially
60–80% bandwidth reductionInfrastructureShi et al. [29]Cloud-only pipelinesEdge filtering and anomaly detectionEdge-computing deploymentsWorkload-dependentT4No—enabler

6.2. Domain-Specific Benchmarking Against Task-Specific Baselines

Treating “traditional AI” as one homogeneous baseline overstates the case for agentic AI. Table 6 therefore benchmarks per domain against the strongest task-specific incumbent: computer-vision inspection, static shelf-life modeling, conventional demand forecasting, model predictive control, and open-loop digital-twin optimization. Improvements are reported ranges from Table 5’s evidence, not new measurements.

6.3. Challenges

Challenges are organized into two categories. Technical and regulatory. Interoperability requires FAIR data standards and universal API protocols [30]. Regulatory frameworks (FDA New Era of Smarter Food Safety, EU AI Act) increasingly mandate explainability for high-risk applications; XAI integration is a prerequisite for market authorization, not an option [13,28]. Cybersecurity expands the attack surface; federated learning offers privacy-preserving mitigation [29]. Socio-technical. Component technologies (IoT sensors, CNNs) sit at TRL 8–9, but integrated autonomous multi-agent systems across full supply chains remain at TRL 4–5 [9]. Reported capital requirements ($500 K–$2 M for full implementation) exclude SMEs unless AIaaS models and open-source frameworks mature [23,24].

6.4. Tiered Governance Model and Regulatory Mapping

To reinforce beneficial use and minimize misuse, we propose a three-level governance model and map it explicitly to the instruments reviewers asked to see connected (Table 7):
  • Level 1 (Monitoring): agents recommend; humans execute. Corresponds to current HACCP/ISO 22000 practice, where critical control points require human disposition.
  • Level 2 (Supervised autonomy): agents execute routine decisions inside bounded envelopes; humans retain override with guaranteed override latency; every autonomous action and override is logged.
  • Level 3 (Full autonomy within envelopes): agents operate independently inside validated operational envelopes, with ex-post auditing via immutable logs [30].
Accountability. When an autonomous food-safety decision causes operational or public-health consequences, accountability cannot rest with the agent. We propose an accountability-by-design allocation: the food business operator retains ultimate legal responsibility for food placed on the market (as under current food law); the system deployer is responsible for operating within validated envelopes and maintaining oversight competence; the system manufacturer is responsible for design defects, validation evidence, and update safety. Immutable audit trails make this allocation enforceable after the fact. Advancement between levels requires demonstrated sustained compliance, workforce transformation at each transition, and audit frameworks capable of evaluating autonomous performance—potentially with third-party certification. Technical, organizational, and regulatory dimensions must advance together; neglecting any one compromises the whole. The legal and ethical stakes of this allocation are receiving direct scholarly attention: Roy and Das [38] analyze the legal and ethical dimensions of AI in food safety and traceability, underscoring that liability, transparency, and accountability frameworks must evolve in step with autonomous capability—the mapping attempted in Table 7.

6.5. A Socio-Technical Adoption Framework

Socio-Technical Systems theory and Industry 5.0 were previously invoked but not operationalized. We now state the analytical framework explicitly. Agentic-AI adoption is jointly determined by six interacting factors: (F1) technology readiness (component TRL 8–9 vs. integrated-system TRL 4–5); (F2) organizational capabilities (data governance maturity, process discipline under HACCP/ISO 22000); (F3) workforce skills (shift from operational labor to AI supervision and exception handling—the Industry 5.0 human-centric collaboration paradigm [8]); (F4) regulatory maturity (availability of audit frameworks and certification pathways); (F5) infrastructure (connectivity for cloud–edge architectures, sensor coverage); and (F6) resource constraints, binding hardest on SMEs ($500 K–$2 M reported implementation cost). Four propositions follow: P1—adoption stalls if any single factor falls below threshold (multiplicative, not additive, interaction); P2—F6 constraints are mediated by delivery model (AIaaS, consortium, open source) more than by firm size per se; P3—F3 and F4 co-evolve: oversight skills are only valuable where regulation recognizes them, and regulation only functions where skills exist; P4—the human-in-the-loop → human-on-the-loop transition is the socio-technical bottleneck, progressing slower than the technical layers. For large enterprises, agentic AI offers supply-chain resilience and ESG compliance automation; for SMEs, consortium models and open-source frameworks are critical to prevent market bifurcation [23,24].

7. Conclusions, Limitations, and Future Research Agenda

This paper set out four objectives, and the conclusions must be stated at the strength the evidence permits. First, on architectural specification: we proposed a unified three-layer Perception–Reasoning–Action architecture with a Digital Twin gating layer, defined its data flows, interfaces, boundary conditions, and autonomy tiers. This is a design contribution; its soundness is argued, not empirically proven. Second, on quantitative synthesis: the structured literature synthesis reports indicative performance ranges—detection accuracies of 92–98% in imaging studies, waste-reduction ranges of 18–30% in retail and hospitality contexts, cold-chain energy improvements of 15–20%, and response-time compression from hours–days to seconds–minutes when response is automated. The traceability audit of Section 6.1 shows these values come predominantly from simulation and secondary literature (Tiers T3/T4), frequently from non-agentic enabling studies. We therefore do not claim statistically validated superiority over traditional AI; we claim a consistent, multi-source pattern of reported potential that justifies field validation. Third, on theoretical framing: the adoption framework of Section 6.5 operationalizes STS theory into six factors and four propositions, which are testable hypotheses for future work. Fourth, on governance: the three-tier model, its regulatory mapping, and the accountability allocation offer a concrete starting point for regulators and standards bodies—a proposal awaiting pilot implementation.
Limitations. This study’s limitations are substantial and define its agenda. (i) The evidence base is secondary and simulation-heavy; no new field trials were conducted, and several cited values originate from enabling technologies (imaging, edge computing, RL foundations) rather than closed-loop agentic systems. (ii) Reported ranges aggregate heterogeneous contexts; category-specificity (Section 4) and scope restrictions (cold chain, Section 5.2) mean chain-wide generalization is not yet warranted. (iii) The architecture, governance model, and accountability matrix are conceptual constructs at TRL 4–5 maturity for integrated systems; they require pilot deployment and iterative refinement. (iv) The economic analysis reflects large-enterprise cost structures and may not transfer to SMEs or developing-economy contexts. (v) The framework assumes network infrastructure that rural and resource-limited settings may lack. (vi) Workforce-displacement ethics are acknowledged but not empirically investigated here.
Future research agenda. Near term (1–3 years): execute the validation protocol of Section 2.3—standardized benchmark suites, open datasets across food categories and contamination scenarios, federated learning pilots [29], and controlled pilots of the governance tiers with preregistered metrics. Medium term (3–7 years): longitudinal field studies of sustained safety, waste, and energy outcomes across geographies; integration of large language models for natural-language human–agent collaboration in manufacturing; and formal verification of simulate-before-act gating. Long term (7–15 years): higher-autonomy operations in bounded envelopes, and international harmonization of certification—potentially under Codex Alimentarius. We deliberately describe highly autonomous “lights-out” food factories and general-purpose formulation agents as speculative scenarios, not predictions: whether they materialize depends on the validation steps above, not on extrapolation of current capabilities.
Concluding remarks. Agentic AI is a plausible paradigm shift toward autonomous, self-optimizing food systems—and, on current evidence, an incompletely validated one. The framework, traceability matrix, governance model, and adoption propositions presented here are offered as a rigorous basis for that validation. Realizing the reported potential requires interdisciplinary collaboration among food scientists, AI researchers, ethicists, and policymakers, and honest separation—as practiced in this revision—between what is proposed, what is projected, and what is proven. The benefits must be weighed against algorithmic bias, cyber-vulnerability, and economic exclusion [6,7]. Success should be measured not only by technical metrics but by whether the resulting food systems are equitable, transparent, and resilient. The question is no longer only whether agentic AI can reshape food systems, but whether the community will validate that reshaping with the scientific rigor and ethical foresight it demands.

Author Contributions

Methodology, E.M.T.; Formal analysis, A.A.T.; Writing—original draft, A.A.T.; Writing—review & editing, E.M.T. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

For further information on the data contact the corresponding author.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. WHO. Foodborne Disease Burden Epidemiology Reference Group (FERG) Estimates; World Health Organization: Geneva, Switzerland, 2022; Available online: https://www.who.int/publications/i/item/9789240100510 (accessed on 14 April 2026).
  2. FAO. The State of Food and Agriculture 2023; FAO: Rome, Italy, 2023. [Google Scholar] [CrossRef] [Scilit]
  3. Poore, J.; Nemecek, T. Reducing food’s environmental impacts through producers and consumers. Science 2018, 360, 987–992. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Crippa, M.; Solazzo, E.; Guizzardi, D.; Monforti-Ferrario, F.; Tubiello, F.N.; Leip, A. Food systems are responsible for a third of global anthropogenic GHG emissions. Nat. Food 2021, 2, 198–209. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. FAO. Food Wastage Footprint: Impacts on Natural Resources; FAO: Rome, Italy, 2013. [Google Scholar]
  6. Pearl, J.; Mackenzie, D. The Book of Why: The New Science of Cause and Effect; Basic Books: New York, NY, USA, 2018. [Google Scholar]
  7. Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction, 2nd ed.; MIT Press: Cambridge, MA, USA, 2018. [Google Scholar]
  8. Gavai, A.K.; Heringa, J. Agentic artificial intelligence in food science. Food Humanit. 2026, 6, 101132. [Google Scholar] [CrossRef] [Scilit]
  9. Wooldridge, M. An Introduction to Multiagent Systems; John Wiley & Sons: Hoboken, NJ, USA, 2009. [Google Scholar]
  10. Karniadakis, G.E.; Kevrekidis, I.G.; Lu, L.; Perdikaris, P.; Wang, S.; Yang, L. Physics-informed machine learning. Nat. Rev. Phys. 2021, 3, 422–440. [Google Scholar] [CrossRef] [Scilit]
  11. Gavai, A.K.; Meuwissen, M.P.M. Agentic AI as a coordination paradigm in digital health and agri-food systems. Patterns 2026, 7, 101496. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Wang, K.; Mirosa, M.; Hou, Y.; Bremer, P. Advancing food safety behavior with AI: Innovations and opportunities in the food manufacturing sector. Trends Food Sci. Technol. 2025, 161, 105050. [Google Scholar] [CrossRef] [Scilit]
  13. Gunning, D.; Aha, D. DARPA’s explainable artificial intelligence (XAI) program. AI Mag. 2019, 40, 44–58. [Google Scholar] [CrossRef] [Scilit]
  14. Kamruzzaman, M.; ElMasry, G.; Sun, D.W.; Allen, P. Non-destructive prediction and visualization of chemical composition in lamb meat using NIR hyperspectral imaging and multivariate regression. Innov. Food Sci. Emerg. Technol. 2012, 16, 218–226. [Google Scholar] [CrossRef] [Scilit]
  15. European Food Safety Authority. Guidance on whole genome sequencing for food safety risk assessment. EFSA J. 2022, 20, e07284. [Google Scholar] [CrossRef] [Scilit]
  16. Yu, W.; Ouyang, Z.; Zhang, Y.; Lu, Y.; Wei, C.; Tu, Y.; He, B. Research progress on the artificial intelligence applications in food safety and quality management. Trends Food Sci. Technol. 2025, 156, 104855. [Google Scholar] [CrossRef] [Scilit]
  17. van Meer, F.; Takeuchi, M.; Ochieng, P.E.; Tavelli, R.; Gerssen, A.; van der Velden, B.H.M. Artificial intelligence in food safety. npj Sci. Food 2026, 10, 294. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Labuza, T.P.; Riboh, D. Time to failure of packaged foods. Food Technol. 1982, 36, 66–74. [Google Scholar]
  19. Du, H.; Sun, X.; Chong, X.; Yang, M.; Zhu, Z.; Wen, Y. A review on smart active packaging systems for food preservation: Applications and future trends. Trends Food Sci. Technol. 2023, 141, 104200. [Google Scholar] [CrossRef] [Scilit]
  20. Mobahi, N.; Razavi, M.A.; Ekrami, M.; Emam-Djomeh, Z.; Razavi, S.H. AI-integrated bio-nanocomposite food packaging: Sustainable, functional, and intelligent solutions. Trends Food Sci. Technol. 2025, 165, 105336. [Google Scholar] [CrossRef] [Scilit]
  21. Madhu, B. AI-driven food packaging systems: A new frontier in intelligent food safety and shelf-life management. J. Food Sci. 2025, 90, e70716. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Melikoglu, M. Artificial intelligence and deep learning in food shelf life management: A global review. Food Humanit. 2026, 6, 101089. [Google Scholar] [CrossRef] [Scilit]
  23. Clark, Q.M.; Kanavikar, D.B.; Clark, J.; Donnelly, P.J. Exploring the potential of AI-driven food waste management strategies used in the hospitality industry for application in household settings. Front. Artif. Intell. 2025, 7, 1429477. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Shadid, N.; Ahmed, V.; Bahroun, Z. A systematic review of data-driven approaches to food waste and loss management. Clean. Waste Syst. 2025, 12, 100448. [Google Scholar] [CrossRef] [Scilit]
  25. Zghair, H.; Konathala, R.G. Using AI for optimizing packing design and reducing cost in e-commerce. AI 2025, 6, 146. [Google Scholar] [CrossRef] [Scilit]
  26. Onyeaka, H.; Akinsemolu, A.; Miri, T.; Nnaji, N.D.; Duan, K.; Pang, G.; Tamasiga, P.; Khalid, S.; Al-Sharify, Z.T.; Ugwa, C. Artificial intelligence in food system: Innovative approach to minimizing food spoilage and food waste. J. Agric. Food Res. 2025, 21, 101895. [Google Scholar] [CrossRef] [Scilit]
  27. Zou, Y.; Gao, Q.; Wu, H.; Liu, N. Carbon-efficient scheduling in fresh food supply chains with a time-window-constrained deep reinforcement learning model. Sensors 2024, 24, 7461. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Salih, A.S.; Akmal, N.; Alhajri, A.; Alhabsi, A.; Al-Belushi, M.; Al-Maskari, N. A perspective on explainable artificial intelligence methods: SHAP and LIME. arXiv 2024, arXiv:2305.02012. [Google Scholar] [CrossRef] [Scilit]
  29. Shi, W.; Cao, J.; Zhang, Q.; Li, Y.; Xu, L. Edge computing: Vision and challenges. IEEE Internet Things J. 2016, 3, 637–646. [Google Scholar] [CrossRef] [Scilit]
  30. Peddareddigari, S.; Vijayan, S.V.H.; Annamalai, M. IoT, blockchain, big data and artificial intelligence (IBBA) framework—For real-time food safety monitoring. Appl. Sci. 2024, 15, 105. [Google Scholar] [CrossRef] [Scilit]
  31. Huang, Y.; Ghadge, A.; Yates, N. Implementation of digital twins in the food supply chain: A review and conceptual framework. Int. J. Prod. Res. 2024, 62, 6400–6426. [Google Scholar] [CrossRef] [Scilit]
  32. Liu, E.; Yang, H.; Sharma, S.; van Leerdam, M.B.; Niu, P.; VandeHaar, M.J.; Hostens, M. Agents are all you need: Pioneering the use of agentic artificial intelligence to embrace large language models into dairy science. J. Dairy Sci. 2025, 108, 14038–14049. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  33. Zheng, R.; Kamruzzaman, M. Explainable AI for hyperspectral imaging in food quality decision support: Interpretability, reliability and future directions. Crit. Rev. Food Sci. Nutr. 2026, 24, 6139–6159. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Borrás-Hidalgo, O.; Ding, F.; El Sheikha, A.F. Artificial intelligence and digital twins: Pathways to climate-resilient, traceable and equitable global food supply chains. Trends Food Sci. Technol. 2026, 28, 105973. [Google Scholar] [CrossRef] [Scilit]
  35. U.S. Food and Drug Administration. New Era of Smarter Food Safety Blueprint; FDA: Silver Spring, MD, USA, 2021. Available online: https://www.fda.gov/food/new-era-smarter-food-safety (accessed on 14 April 2026).
  36. Ghaani, M.; Cozzolino, C.A.; Castelli, G.; Farris, S. An overview of the intelligent packaging technologies in the food sector. Trends Food Sci. Technol. 2020, 51, 1–11. [Google Scholar] [CrossRef] [Scilit]
  37. Grigoryan, A.; Helfrich, S.; Lequeux, V.; Lapras, B.; Marchand, C.; Merienne, C.; Bruno, F.; Mazet, R.; Pirot, F. Smart Formulation: AI-driven web platform for optimization and stability prediction of compounded pharmaceuticals using KNIME. Pharmaceuticals 2025, 18, 1240. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Roy, P.; Das, D. AI in food safety and traceability: Legal and ethical dimensions of next-generation food systems. Curr. Opin. Food Sci. 2026, 6, 101436. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Architectural framework for agentic AI in food systems.
Figure 1. Architectural framework for agentic AI in food systems.
Standards 06 00038 g001
Figure 2. Probabilistic risk assessment framework for food safety.
Figure 2. Probabilistic risk assessment framework for food safety.
Standards 06 00038 g002
Figure 3. Intelligent shelf-life extension via Digital Twin optimization.
Figure 3. Intelligent shelf-life extension via Digital Twin optimization.
Standards 06 00038 g003
Figure 4. Agentic AI for environmental sustainability via integrated LCA.
Figure 4. Agentic AI for environmental sustainability via integrated LCA.
Standards 06 00038 g004
Table 1. Critical synthesis of prior literature: representative strands and residual gaps.
Table 1. Critical synthesis of prior literature: representative strands and residual gaps.
StrandRepresentative WorksWhat the Strand DeliversResidual Gap
Contaminant detection and food safetyKamruzzaman et al. [14]; Wang et al. [12]; EFSA [15] Yu et al. [16]; van Meer et al. [17]Hyperspectral and NIR detection; AI-enabled safety monitoring; regulatory guidance on genomics and AIDetection stops at flagging. No autonomous containment, no closed-loop risk mitigation, limited causal reasoning.
Shelf-life prediction and intelligent packagingLabuza & Riboh [18]; Du et al. [19]; Mobahi et al. [20] Madhu [21]; Melikoglu [22]Kinetic degradation models; active and intelligent packaging with embedded sensingStatic or batch-level prediction. Packaging senses but does not reconfigure logistics or formulation in response.
Logistics, demand forecasting, and wasteClark et al. [23]; Shadid et al. [24]; Zghair & Konathala [25]
Onyeaka et al. [26]; Zou et al. [27]
Data-driven forecasting and waste management; agent-based modeling elementsIsolated optimization of single objectives. Cross-objective trade-offs (safety vs. waste vs. carbon) unresolved.
Digital Twins, MAS, XAI, and edge infrastructureWooldridge [9]; Gavai & Heringa [8]; Gunning & Aha [13]; Salih et al. [28]; Shi et al. [29]; Peddareddigari et al. [30] Huang et al. [31]; Liu et al. [32]; Gavai and Meuwissen [11]; Zheng and Kamruzzaman [33]; Borrás-Hidalgo et al. [34]Coordination theory; DT concepts; explainability methods; cloud–edge plumbing; FAIR data standardsComponents exist, but perception, reasoning, action, and governance are not integrated into one governed closed loop.
Table 2. Distinguishing agentic AI from adjacent technologies and constructs.
Table 2. Distinguishing agentic AI from adjacent technologies and constructs.
ConstructCore FunctionActs on the Physical Process?Adapts Objectives Online?Relation to Agentic AI
Traditional predictive AIClassify, forecast, scoreNo; human acts on outputNo; static objectivesPerception component
AI-enabled automationExecute fixed rules on model outputYes, but only pre-programmed responsesNoAction component without reasoning
Rule-based control (e.g., PLC/HACCP alarms)Threshold triggeringYesNoSpecial case of bounded action
Multi-agent systems (MAS)Coordinate distributed agentsIndirectlyPartially, via negotiationArchitectural substrate [9]
Digital Twin (DT)Mirror and simulate the processNo; prediction onlyNoSimulation environment for the reasoning layer [8]
Reinforcement learning (RL)Learn policies from rewardIn silico, typicallyWithin a fixed reward functionLearning mechanism [7]
Autonomous roboticsPhysical manipulationYesAt motion levelActuation endpoint of the action layer
Agentic AIPerceive–reason–act under governanceYes, within bounded envelopesYes, via multi-objective reasoningThe construct proposed here
Table 3. Comparative analysis of traditional AI and agentic AI paradigms in food science.
Table 3. Comparative analysis of traditional AI and agentic AI paradigms in food science.
FeatureTraditional AI (Predictive)Agentic AI (Adaptive)
ScopeNarrow, task-specificBroad, multi-objective optimization
AutonomyHuman-in-the-loop requiredBounded autonomy with oversight tiers
ReasoningPattern recognition, correlationCausal inference, mathematical modeling
AdaptabilityStatic or batch retrainingContinuous learning, real-time updating
IntegrationSiloed data streamsMultimodal sensor fusion, cloud–edge
GovernanceExternal to the modelEmbedded (XAI, audit trails, autonomy tiers)
Table 4. Quantified sustainability impacts: traditional task-specific baselines vs. agentic AI (reported ranges).
Table 4. Quantified sustainability impacts: traditional task-specific baselines vs. agentic AI (reported ranges).
Sustainability MetricTraditional AI Baseline (Reported)Agentic AI (Reported Range)Incremental Benefit Attributed to AgencyMechanismProvenance and Scope
Food waste5–8% reduction (isolated forecasting/inventory models)20–30% reduction~3–4×Predictive demand (up to 90% accuracy) coupled to FEFO execution and dynamic repricingT3/T4; retail and hospitality contexts
Carbon footprint4–7% reduction (static route optimization, periodic re-planning)10–15% reduction~2×Continuous multi-objective routing, refrigeration scheduling, renewable-aware energy dispatchT3/T4; logistics operations
Energy efficiency5–8% improvement (conventional control and scheduling)15–20% improvement in cold-chain operations~2–3×MPC plus learning-based control of refrigeration and processing set-pointsT3/T4; cold-chain-specific—see scope note below
Water usage2–4% reduction (timer/rule-based CIP and irrigation)5–10% reduction~2×Sensor-driven precision irrigation; demand-optimized CIP cyclesT3/T4; processing and irrigation pilots
Table 6. Domain-specific benchmarking of agentic AI against task-specific baselines.
Table 6. Domain-specific benchmarking of agentic AI against task-specific baselines.
DomainTask-Specific BaselineBaseline Performance (Reported)Agentic ApproachReported ImprovementEvidence TierKey Caveat
Safety inspectionComputer-vision flagging, human disposition85–90% accuracy; hours-to-days responseClosed-loop detection-to-containment with PRA gating92–98% accuracy; seconds-to-minutes response (1.08–1.15× accuracy; 60–1000× latency)T2/T4 + T3 projectionAccuracy from imaging studies; latency gain mostly from automation of response, not detection
Shelf-life managementStatic dating; batch kinetic modelsFixed dates; no in-chain recalculationDynamic recalculation + FEFO executionRMSE <5%; 18–22% retail waste reductionT2/T4Perishable retail categories only
Demand and inventoryConventional forecasting (ARIMA/exponential smoothing) + manual replenishment5–8% waste reductionLSTM with exogenous regressors + autonomous replenishment/repricing15–25% waste reduction (3–5×)T3/T4Retail/hospitality contexts; chain-wide effect unproven
Process energy controlModel predictive control with fixed set-points5–8% improvementLearning-augmented MPC with continuous retuning15–20% improvement, cold chain onlyT3/T4Cold-chain scope; not generalizable (Section 5.2)
Process optimizationOpen-loop Digital Twin (monitor/predict)Advisory onlyClosed-loop twin with simulate-before-act gatingQualitative shift: prediction → governed actionT3Maturity TRL 4–5 for integrated systems
Table 7. Mapping the three-level governance model to regulatory and standards instruments.
Table 7. Mapping the three-level governance model to regulatory and standards instruments.
Governance ElementLevel 1Level 2Level 3Linked Instruments
Decision rightsHumanAgent, routine scopeAgent, validated envelopeHACCP CCP disposition rules; ISO 22000 FSMS
Human oversightExecutes all actionsOverride anytime; periodic reviewException-based supervisionEU AI Act Art. 14 human oversight; FDA New Era
Audit trailDecision logsImmutable action + override logsFull ex-post audit; blockchain anchoringFSMA Rule 204 traceability; EU Reg. 178/2002; EU AI Act logging
Validation gate to advance—Sustained compliance over defined observation periodIndependent third-party certificationProposed ‘Algorithmic Food Safety’ certification; Codex Alimentarius harmonization
CybersecurityBaselineAdversarial testing; drift monitoringContinuous red-teaming; federated updatesEU AI Act robustness; ISO/IEC security standards
Post-deployment monitoringIncident reportingDrift detection; periodic re-validationContinuous assurance; mandatory recall-to-Level-1 triggersEU AI Act post-market monitoring
Liability and accountabilityFood business operatorOperator + system deployer shared, per contractOperator, deployer, and manufacturer per accountability matrix belowProduct-liability and food-law frameworks
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Tsinakos, A.A.; Tsinakou, E.M. Agentic Artificial Intelligence in Food Science: A Conceptual Framework and Structured Synthesis for Food Safety, Shelf-Life Extension, and Environmental Sustainability. Standards 2026, 6, 38. https://doi.org/10.3390/standards6040038

AMA Style

Tsinakos AA, Tsinakou EM. Agentic Artificial Intelligence in Food Science: A Conceptual Framework and Structured Synthesis for Food Safety, Shelf-Life Extension, and Environmental Sustainability. Standards. 2026; 6(4):38. https://doi.org/10.3390/standards6040038

Chicago/Turabian Style

Tsinakos, Avgoustos A., and Elena Maria Tsinakou. 2026. "Agentic Artificial Intelligence in Food Science: A Conceptual Framework and Structured Synthesis for Food Safety, Shelf-Life Extension, and Environmental Sustainability" Standards 6, no. 4: 38. https://doi.org/10.3390/standards6040038

APA Style

Tsinakos, A. A., & Tsinakou, E. M. (2026). Agentic Artificial Intelligence in Food Science: A Conceptual Framework and Structured Synthesis for Food Safety, Shelf-Life Extension, and Environmental Sustainability. Standards, 6(4), 38. https://doi.org/10.3390/standards6040038

Article Metrics

Back to TopTop