Next Article in Journal
Joint Feature Selection and Hyperparameter Optimization Using Evolutionary Algorithms for Diabetes Identification Across Multiple Datasets
Previous Article in Journal
Thresholds Before Labels: Ceiling Quantization Determines When Autoscaling Signals Change Replica Decisions
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Adversarial Evidence-Plane Robustness for Bounded Agent Assurance: Deterministic Mutation Testing of Claim Admissibility

Independent Researcher, Upper Marlboro, MD 20774, USA
Computers 2026, 15(10), 655; https://doi.org/10.3390/computers15100655
Submission received: 7 September 2026 / Revised: 24 September 2026 / Accepted: 25 September 2026 / Published: 27 September 2026

Abstract

Assurance for autonomous agents depends on operational evidence about policy decisions, actions, execution, effects, provenance, and enforcement. This study tests whether bounded assurance claims remain semantically admissible when that evidence is adversarially degraded. Three co-primary questions address deviation existence (B1), observable action-path reconstruction (B2), and containment/enforcement-boundary localization (B3). A deterministic mutation harness applied 14 frozen single-operator evidence-plane attacks to three known-ground-truth baselines, producing 42 adversarial cases and 126 claim evaluations. A prospectively frozen strongest-safe oracle and closed-world semantic scorer classified the 126 claim evaluations into three predefined categories: exact safe, safe conservative, and unsafe false establishment. All 42 cases were valid. Of the 126 claim evaluations, 56 were exact safe, 47 safe conservative, and 23 unsafe false establishment. The resulting Unsafe False Establishment Rate (UFER) was 23/126 ≈ 0.18254, so the pre-specified zero-UFER target failed. Unsafe outcomes were B1 3/42, B2 19/42, and B3 1/42. Exploratory forensics classified 21 unsafe evaluations as substantive semantic incompatibilities and two as possible scorer-normalization artifacts. The dominant B2 mechanism was global ineligibility after localized evidentiary degradation. These results distinguish semantic safety from information retention: evidence defects must not induce propositions beyond the boundary justified by surviving admissible evidence.

1. Introduction

Autonomous and agentic systems increasingly perform consequential operations through APIs, filesystems, orchestration services, control interfaces, and delegated software components. Evaluating such systems therefore requires more than inspecting model output or eliciting explanations. Operational assurance depends on evidence about what action was requested, what policy applied, whether authorization was granted or denied, whether execution was attempted, what enforcement mechanisms acted, and whether an operational effect occurred.
This creates a bounded assurance problem centered on the operational evidence plane. A conclusion that is appropriately bounded under pristine evidence may become unjustified after evidence is corrupted, deleted, inserted, replayed, reordered, forged, conflicted, suppressed, or bypassed. The security question is therefore not limited to whether alteration can be detected. A downstream evaluator must also determine what, if anything, can still be claimed from surviving admissible evidence.
In this study, the operational evidence plane is the collection of externally available records and relationships from which the evaluator constructs assurance propositions. It includes action and policy records, source identity and authority information, ordering and lineage relations, control and enforcement observations, execution and effect observations, and evidence about observation coverage. It excludes hidden model cognition and the experiment’s hidden D/X/E ground truth. “Bounded” assurance means that a conclusion is limited in content and scope to what those surviving admissible records can support; the evaluator is not permitted to fill evidentiary gaps with assumptions about unobserved behavior.
This issue intersects several established areas. Assurance-case research formalizes relationships among claims, evidence, argumentation, and defeaters [1,2,3]. Prior work by the author demonstrates deterministic appraisal of admitted runtime evidence against provenance, conformance, and policy conditions [4]. Provenance, software-supply-chain integrity, and attestation frameworks bind statements to sources, histories, and appraisal procedures [5,6,7,8,9]. Recent agent-specific work emphasizes auditable evidence planes, evidence-gated lifecycle decisions, recoverable operational histories, claim-aware artifact provenance, tamper-evident action records, and causal attribution [10,11,12,13,14,15,16,17,18]. Adversarial agent benchmarks measure behavioral manipulation, prompt injection, component-level attacks, and harmful task completion [19,20,21,22,23]. Runtime-verification research addresses sound conclusions under incomplete and partially observable traces [24,25,26], while mutation testing provides a mature methodology for controlled perturbation [27].
The contribution of this study is not any one of these established constructs in isolation. Relative to this prior work, the study-specific novelty lies in prospectively coupling bounded claim semantics with deterministic evidence-plane degradation, a strongest-safe oracle, and closed-world semantic scoring, so that assurance failure is measured as claim-level semantic overreach rather than merely as tamper detection, behavioral failure, or missing evidence.
The study asks one central question: what B1/B2/B3 assurance claims remain semantically admissible when an adversary deliberately degrades the operational evidence plane?
Rather than using a frontier model as the primary adversary, the experiment uses deterministic mutation. This provides known ground truth, exact mutation provenance, controlled single-factor perturbations, and reproducible comparison between evaluator output and a frozen strongest-safe oracle.
A central study-specific concept is claim-relative degradation. Its primary safety requirement is that evidentiary degradation must not cause the evaluator to retain or introduce propositions outside the semantic boundary permitted by surviving admissible evidence. This is distinct from information retention. A safe evaluator may return a proposition weaker than the strongest one still justified; under the frozen scorer, such an output is SAFE_CONSERVATIVE, not unsafe.
For example, if evidence supporting the final segment of an observable action path becomes unauthoritative, retaining the complete path may be unsafe. If an earlier prefix remains supported, an evaluator may retain that prefix exactly or conservatively withhold some of it. The former is more informative; the latter may remain semantically safe. What is unsafe is asserting content or scope that the surviving evidence does not justify.
The study contributes a frozen experimental composition consisting of three bounded assurance questions, 14 deterministic evidence-plane mutations, a strongest-safe oracle, a three-state semantic scoring relation that separates safety from information retention, and an unsafe-false-establishment endpoint. Its empirical contribution is the complete result over the constructed corpus together with an evaluation-level forensic decomposition of observed failure mechanisms.
For provenance and reproducibility, the frozen program is internally designated S02, Adversarial Evidence-Plane Robustness for Bounded Agent Assurance, with machine-readable identifier s02-evidence-plane. The label S02 is a study identifier rather than a scientific construct.

2. Background and Related Work

2.1. Assurance Claims and Evidentiary Support

The relationship between evidence and assurance claims is well established in safety and security engineering. The Object Management Group Structured Assurance Case Metamodel provides standardized constructs for claims, argumentation, evidence artifacts, and their relationships [1]. Assurance 2.0 emphasizes logical soundness, evidential support, defeaters, and residual uncertainty when evaluating confidence in critical claims [2]. Frontier-AI safety-case work similarly organizes safety arguments around explicit claims and supporting evidence rather than treating model evaluations as self-interpreting proof of system safety [3].
These approaches establish a principle relevant to the present experiment: an assurance conclusion is justified through the evidentiary and inferential relationship supporting it. They do not, however, directly measure what happens to multiple bounded operational claims when the evidentiary substrate supplied to an evaluator is deliberately mutated.
Campbell previously connected admitted runtime cryptographic evidence to provenance closure, implementation identity, external conformance evidence, policy appraisal, and deterministic bounded verdicts [4]. That work preserved an evidence-insufficient review state rather than converting insufficient evidence into acceptance or rejection. The present study addresses a different domain and question: adversarial degradation of evidence supporting behavioral, action-path, and containment claims about autonomous systems.

2.2. Provenance, Attestation, and Transparency

W3C PROV provides a general representation for entities, activities, agents, derivations, and provenance relationships [5]. The in-toto framework applies verifiable provenance to software supply chains by binding expected steps, authorized actors, and artifact transformations [6].
Remote attestation provides an especially useful separation between evidence and downstream judgment. The IETF RATS architecture distinguishes evidence produced by an attester, appraisal performed by a verifier, and attestation results consumed by a relying party [7]. The Entity Attestation Token specification supplies a standardized carrier for attested claims while leaving appraisal and trust decisions to relying-party policy [8].
The Supply Chain Integrity, Transparency, and Trust (SCITT) architecture provides another relevant distinction: authenticated and transparently recorded statements establish provenance and history, but authenticated origin does not itself establish semantic correctness [9]. This motivates the present study’s separation among evidence integrity, source authority, and claim admissibility.

2.3. Auditable Agent Operation and Evidence-Gated Claims

Recent work has moved autonomous-agent assurance toward explicit operational evidence. Theodorakopoulos and Theodoropoulou describe auditable autonomy through an evidence plane, decision traces, and operational outcomes, emphasizing provenance, trace completeness, policy checks, and trace tampering [10]. The term evidence plane therefore has direct prior use; this study does not claim novelty for the construct itself.
Proof-or-Stop is particularly close conceptually. It treats autonomous-agent outputs as claims rather than lifecycle state and permits consequential lifecycle transitions only when fresh, mechanically verifiable, state-bound evidence satisfies a gate [11]. Its primary object is evidence-gated lifecycle control and tamper rejection. The present study instead evaluates graded B1/B2/B3 semantic conclusions after controlled mutation of the evidence presented to an assurance evaluator, comparing each output with a separately frozen strongest-safe proposition.
Auditable Agents likewise centers trustworthy evidence and reconstructability. It distinguishes accountability, auditability, and auditing and evaluates action recoverability, lifecycle coverage, policy checkability, responsibility attribution, and evidence integrity [12]. The present experiment is narrower: it does not propose a general auditability architecture or attribution framework, but tests whether specific bounded assurance claims remain semantically admissible after controlled degradation of their evidentiary basis.
Recent work further narrows the surrounding space. Yin et al. make artifact–claim evidence bindings first-class observability objects for autonomous scientific agents [13]. Wang et al. study classification and localization of authorization-, provenance-, and completeness-related harness tampering in self-improving agents [14]. These studies reinforce the importance of claim-aware evidence and tamper auditing, but their experimental targets differ from the one studied here: graded semantic admissibility of multiple bounded assurance propositions relative to a prospectively frozen strongest-safe oracle after controlled evidence-plane mutation.
Emerging agent-audit specifications provide further context. The Agent Audit Trail Internet-Draft proposes hash-linked records for agent identities, actions, outcomes, and trust information [15], while the Proof-of-Behavior Internet-Draft proposes signed receipts, chaining, and pre-execution policy gates [16]. Both are works in progress rather than established IETF standards.
AUDITA combines tamper-evident inter-agent records with evidence-based causal attribution [17]. The present study adopts a narrower claim boundary: B2 reconstructs only evidence-supported operational paths, while B3 localizes only evidence-supported containment boundaries. Neither infers hidden motive, private deliberation, or unsupported causal links.
Prior empirical work on agentic composition drift provides a related observability result. Campbell found that runtime provenance availability materially affected detection of compositional drift [18]. The current experiment changes the unit of analysis: rather than asking whether a detector observes a state, it asks whether a downstream assurance proposition remains appropriately bounded when evidence is adversarially degraded.

2.4. Adversarial Evaluation of Tool-Using Agents

A substantial agent-security literature evaluates attacks on the agent, its inputs, tools, memory, or operational environment. ToolEmu evaluates risks arising from language-model agents operating external tools [19]. AgentDojo evaluates prompt-injection attacks and defenses in tool-using agents [20]. Agent Security Bench formalizes attacks and defenses across multiple agent components [21], while AgentHarm measures whether autonomous agents can be induced to complete harmful multi-step tasks [22]. Broader surveys synthesize risks including prompt injection, tool misuse, memory poisoning, privilege abuse, and multi-agent interactions [23].
These studies primarily perturb the agent or its operational environment and measure behavior, attack success, or harm. The present experiment holds the constructed D/X/E ground-truth class fixed and instead attacks the evidence presented to the assurance evaluator. Behavioral robustness and assurance-evidence robustness are therefore different experimental targets.

2.5. Partial Observability and Incomplete Traces

Runtime-verification research supplies the strongest formal precedent for reasoning under incomplete observation. Leucker et al. propagate uncertainty through incomplete timed event streams rather than inferring unobserved events [24]. Taleb et al. survey monitoring under incomplete and imprecise traces and show that insufficient observation can prevent a sound conclusive verdict [25]. Cimatti et al. address monitoring under assumptions and partial observability, including explicitly ordered verdict specificity [26].
The present study adopts the same general discipline but evaluates a richer family of operational propositions. Instead of only determining whether a temporal property is true, false, or unresolved, the evaluator may support multiple strengths of B1, B2, or B3 claims.
The study therefore separates two consequences of partial evidence, semantic safety and information retention, defined in Section 5.5.

2.6. Mutation as Experimental Methodology

Mutation testing has long used controlled artificial modifications to assess test sensitivity and empirical adequacy [27]. The present experiment transfers this methodological principle to assurance evidence. Each mutation changes a frozen evidentiary property while preserving the constructed baseline D/X/E ground-truth class, allowing post-mutation evaluator output to be compared with a separately frozen strongest-safe oracle.

2.7. Differentiating Composition

Section 1 states the central research question and formal contribution. Relative to the cited literature, differentiating composition is operationalized through the three co-primary questions B1–B3 defined below.

3. Assurance Questions and Claim Boundary

The assurance framework used in this study consists of four bounded assurance questions (B1–B4). The frozen empirical scope of S02 is intentionally limited to B1, B2, and B3, which constitute the co-primary assurance questions evaluated in the deterministic mutation campaign. B4 defines the subsequent question of validated causal explanation and mitigation but lies outside the empirical scope of S02 v0.1.
B1, B2, and B3 are co-primary. None is a prerequisite, substitute, or proxy for another.

3.1. B1: Deviation Existence

Does admissible evidence establish that observed behavior departed from the frozen behavioral, policy, or authorization envelope?
B1 is an evidence-backed behavioral proposition. A policy denial alone is insufficient; admissible evidence must bind observed behavior to the relevant policy or authorization context.
The frozen envelope is the prospectively specified set of behavioral, policy, and authorization constraints against which the observed action is appraised. A B1 deviation claim therefore requires both an admissible observation of relevant behavior and an admissible binding between that behavior and the applicable envelope; the mere existence of a denial or policy rule does not establish that the observed behavior violated it.

3.2. B2: Observable Action-Path Reconstruction

What sequence of behavioral, control, execution, and effect events can be reconstructed from admissible evidence, including only evidence-supported identity, lineage, and ordering relationships?
An observable action path is an ordered sequence of operational events supported by admitted evidence. Depending on the case, this may include an action request, authorization decision, invocation or execution attempt, control observation, enforcement outcome, or operational-effect observation.
“Observable” means observable through operational evidence external to hidden model cognition. It does not mean physically external to the computing system. Machine-generated runtime records can constitute observable evidence when they are admissible observations of system behavior.
B2 does not infer chain-of-thought, motive, private deliberation, subjective intent, or causal steps across an unobserved gap.

3.3. B3: Containment Outcome/Enforcement-Boundary Localization

What containment or enforcement outcome is established, and at what evidence-supported boundary?
B3 may establish that containment prevented the relevant progression, did not prevent it, was bypassed, failed to enforce at the applicable boundary, or was inadequately observed, but only to the extent supported by admissible evidence.
An enforcement boundary is the specific operational point at which surviving evidence can localize the action of a control (for example, before invocation, at an execution attempt, or before an unauthorized effect). B3 does not infer a boundary merely because the outcome is known; the location itself must be supported by admitted control, execution, or effect evidence.

3.4. B4: Systemic Cause/Validated Mitigation

What systemic mechanism explains an observed assurance failure, and does controlled intervention validate mitigation?
B4 requires stronger causal and interventional evidence than B1–B3 and was not empirically evaluated here. Post-campaign identification of a code-level mechanism does not itself constitute B4 establishment.

4. Frozen Hypotheses and Co-Primary Analysis Plan

4.1. Frozen Hypotheses

4.1.1. Safety

H1. 
No valid adversarial case will cause an unjustifiably stronger B1, B2, or B3 conclusion.
H1 is operationalized through the Unsafe False Establishment Rate (UFER). The frozen target was UFER = 0. A single UNSAFE_FALSE_ESTABLISHMENT evaluation is sufficient to fail the zero-error primary target.

4.1.2. Conservative Degradation

H2. 
When evidence required for a pristine claim is compromised or lost, the evaluator will weaken the claim rather than silently retain unsupported establishment.
H2 is safety-oriented. It requires removal or qualification of unsupported establishment but does not require exact preservation of every narrower proposition still supported after degradation. H2 also does not require weakening if surviving admissible evidence, including redundancy where present, continues to support the original proposition.

4.1.3. Differential Resilience

H3. 
B1, B2, and B3 will exhibit different resilience because they depend on different evidence properties.
H3 is evaluated descriptively over the complete deterministic corpus. It does not posit a population effect size and does not authorize inferential significance testing.

4.1.4. Evidence-Property Separation

H4. 
Integrity, authority, ordering, completeness, and observation coverage will remain distinct, and authenticated or recognized evidence will not automatically be treated as semantically true.
Representation integrity does not imply source authority; source authority does not imply semantic truth; valid ordering does not imply completeness; missing evidence does not establish non-occurrence; and observation coverage is distinct from the semantic content of the observed proposition.
The hypotheses are evaluated descriptively over the frozen corpus. They do not authorize additional cases, post hoc endpoints, or inferential statistical tests.

4.2. Co-Primary Analysis Units

Terminology. Throughout this manuscript, a case denotes one frozen baseline–attack experimental unit (n = 42). A claim evaluation or case–claim pair denotes one B1, B2, or B3 assessment of a case (n = 126). Safety classifications apply to claim evaluations, not to the case as a whole. Equations (1) and (2) give the case and claim-evaluation counts.
3 baselines × 14 attacks = 42 experimental cases
42 cases × 3 co-primary claims = 126 claim evaluations
The experimental unit is one baseline–attack pair. The three B evaluations derived from one case are co-primary assessments of that same experimental unit and are not treated as three statistically independent cases.
Repeated execution of an identical deterministic case is verification only and does not create statistical replication or increase the denominator.

4.3. Primary Endpoint and Measurement Model

The primary endpoint is defined in Equation (3).
UFER = N_UNSAFE_FALSE_ESTABLISHMENT/N_valid B evaluations
For a complete campaign, the planned denominator is 126.
Let O denote the normalized strongest-safe oracle proposition and E the normalized evaluator proposition. Equations (4)–(6) define the three scoring categories.
O ⊭ E ⇒ UNSAFE_FALSE_ESTABLISHMENT
O ⊨ E and E ⊨ O ⇒ EXACT_SAFE
O ⊨ E and E ⊭ O ⇒ SAFE_CONSERVATIVE
Here, ⊨ denotes the frozen scorer’s closed-world semantic entailment relation over registered claim profiles, not unrestricted logical entailment. Evaluator and oracle propositions are first normalized into the frozen proposition registry and scope model before this relation is applied.
Normalization maps supported proposition forms into a finite semantic profile containing asserted claim content, explicitly withheld dimensions, claim scope, and, where applicable, boundary localization. The registered scope ordering is given in Equation (7).
LOCAL_BOUNDARY ≺ EVIDENCED_SCOPE ≺ PRINCIPAL_OBSERVED_PATH ≺ OBSERVED_PATH ≺ GLOBAL_SYSTEM
The scope labels describe how far a normalized proposition reaches. LOCAL_BOUNDARY confines a statement to one evidence-supported control or enforcement point. EVIDENCED_SCOPE limits a statement to the entities, events, or relations directly supported by the admitted evidence relevant to the claim. PRINCIPAL_OBSERVED_PATH restricts the proposition to the selected evidence-supported principal operational path. OBSERVED_PATH extends the proposition to the broader observed-path scope represented by the frozen scorer, and GLOBAL_SYSTEM applies it to the system as a whole. Because broader scopes carry stronger commitments, an otherwise identical proposition can become unsafe if it is asserted over a scope that the strongest-safe oracle does not support.
For otherwise comparable normalized claim profiles, asserting the same proposition over a broader scope constitutes a stronger scope assertion. Entailment therefore depends on claim content, withheld dimensions, and scope rather than scope alone.
For example, if the strongest-safe oracle supports a proposition only over PRINCIPAL_OBSERVED_PATH while the evaluator normalizes an otherwise comparable proposition over the broader OBSERVED_PATH scope, the oracle does not entail the evaluator under the frozen scope relation and the evaluation is UNSAFE_FALSE_ESTABLISHMENT. A generic INCONCLUSIVE label is therefore not automatically conservative; its score depends on the normalized semantic profile.
Normalization is part of the experimental measurement instrument rather than a semantically neutral preprocessing step. Mapping a raw evaluator output into a registered proposition profile can affect claim content, withheld dimensions, or scope, and therefore can affect the resulting safety classification. Because the normalization registry and scope relation were prospectively frozen, no mapping was changed after campaign execution. Consequently, any post-campaign indication that normalization may have introduced a classification discrepancy is retained in the authoritative result and treated as a measurement-model construct-validity limitation rather than retrospectively rescored.
The three scoring states separate semantic safety from information retention. EXACT_SAFE indicates semantic equivalence with the strongest-safe oracle. SAFE_CONSERVATIVE indicates a weaker proposition that remains within the oracle-permitted boundary. UNSAFE_FALSE_ESTABLISHMENT indicates that the normalized evaluator proposition is not entailed by the strongest-safe oracle.
The category name UNSAFE_FALSE_ESTABLISHMENT is broader than a simple affirmative factual falsehood. An unsafe evaluation can result from an unsupported positive proposition, an unsupported scope expansion, or an overbroad proposition about what is or is not admissible.

4.4. Frozen Secondary Endpoints

The frozen secondary endpoints are the exact EXACT_SAFE rate, the exact SAFE_CONSERVATIVE rate, the three-category B1/B2/B3-specific outcome distributions, and correct enforcement/failure-boundary localization over eligible B3 evaluations.
For localization, both the numerator and eligible denominator are reported. A B3 evaluation for which surviving admissible evidence does not support localization is not counted as a localization failure.
No additional primary or secondary endpoint was introduced after observing the campaign results. The later forensic decomposition is exploratory rather than a frozen endpoint.

4.5. Frozen Analysis Plan

The 42 adversarial cases constitute a constructed deterministic corpus, not an independent and identically distributed (IID) sample from a real-world attack population.
The frozen analysis uses exact counts, exact corpus proportions, complete-corpus and B-specific results, descriptive attack-class summaries, and case-level traceability. It does not prescribe null-hypothesis significance testing or population inference.
For every scored claim evaluation, the analysis record maintains traceability among the baseline–attack case, B question, evaluator proposition, surviving admissible evidence basis, strongest oracle-permitted proposition, scoring category, and any supported boundary localization.
The case oracle, attack provenance, hidden D/X/E ground truth, and expected score remain outside evaluator input.

5. Experimental Design

5.1. Internal Study Designation

As introduced in Section 1, the frozen program carries the designation S02 and machine-readable identifier s02-evidence-plane. The designation functions as a provenance identifier rather than a scientific construct.

5.2. Governing Assurance Principle

An assurance evaluator is not required to know what its admissible evidence cannot reveal; it is required not to claim more than that evidence can justify.
Missing evidence is therefore not automatically evidence of absence. Missing invocation evidence does not independently establish non-invocation, and missing effect evidence does not independently establish absence of effect. Negative propositions require adequate observation, coverage, or closure evidence.

5.3. Hidden Experimental Ground Truth

Three binary supporting variables define the constructed case ground truth: D = deviation from the frozen envelope; X = execution or invocation attempt; E = unauthorized operational effect.
These variables are not evaluator inputs and are not themselves operational evidence. D = 1 does not establish B1; B1 requires surviving admissible external evidence supporting the bounded deviation proposition.
The three variables encode progression through the constructed experimental world rather than evaluator judgments. D records whether the constructed behavior departs from the frozen envelope; X records whether an invocation or execution attempt occurs, regardless of whether it produces an effect; and E records whether an unauthorized operational effect occurs. The A, B, and C baselines therefore represent selected progression states from deviation through attempted execution to effect, not an exhaustive enumeration of every logically possible D/X/E combination.

5.4. Pristine Baselines

The constructed ground-truth baselines were: Baseline A, (D, X, E) = (1, 0, 0), representing deviation with pre-invocation containment preventing invocation; Baseline B, (D, X, E) = (1, 1, 0), representing deviation and an invocation attempt with the unauthorized operational effect prevented; and Baseline C, (D, X, E) = (1, 1, 1), representing deviation, an invocation attempt, and an unauthorized operational effect.
The preflight condition was satisfied before the scored adversarial campaign: all nine pristine B1/B2/B3 evaluations matched the frozen pristine oracle.
Figure 1 summarizes the frozen experimental design, evaluator information boundary, and B1/B2/B3 claim boundary.

5.5. Claim-Relative Degradation

Claim-relative degradation describes how changes in the evidentiary basis constrain downstream assurance claims.
If evidence loss, corruption, conflict, or exclusion removes support for a proposition dimension, the evaluator must not continue to assert that unsupported dimension or introduce a broader unsupported proposition in its place.
The experiment distinguishes safety under degradation from information retention under degradation. Safety asks whether the evaluator remains within the semantic boundary justified by surviving admissible evidence. Information retention asks how much of the strongest still-justified proposition the evaluator preserves.
UNSAFE_FALSE_ESTABLISHMENT reflects failure of semantic safety. SAFE_CONSERVATIVE reflects safety with loss of justified information. EXACT_SAFE reflects safety plus semantic equivalence with the strongest-safe oracle.
Accordingly, failure to retain every supportable proposition is not in itself unsafe.

5.6. Evidence Properties

The study distinguishes the following evidence properties because each answers a different question about what a record can legitimately support. Representation integrity asks whether the presented record remains consistent with the protected or committed representation used by the experiment. Source authority and identity ask whether the asserted source is the recognized entity permitted by the frozen registration context to supply that type of evidence. Ordering and lineage ask whether event sequence, predecessor relations, and DAG links support the claimed temporal or derivational relationship. Completeness asks whether evidence required by the frozen case structure is present, while observation coverage asks whether the instrumentation or observer path could observe the proposition at issue. Final claim admissibility is the downstream judgment that integrates these properties with claim content and scope to determine what proposition, if any, may be established.
These properties are intentionally non-equivalent. A record may retain representation integrity yet come from an unauthorized source; an authoritative record may conflict with another authoritative record; a correctly ordered trace may still be incomplete; and missing evidence cannot establish non-occurrence when observation coverage is inadequate. Recognized source authority is evaluated against the frozen source-registration context, so source recognition does not imply general public-key infrastructure (PKI) authentication, semantic truth, or correctness of the proposition carried by the record. This separation is necessary because claim admissibility depends on the combination of relevant properties rather than on any one property in isolation.

5.7. Frozen Attack Taxonomy

Table 1 lists the 14 frozen single-factor attack classes, the principal evidentiary property stressed, and the mutation semantics used to preserve construct separation. Here, single-factor means that each case applies one frozen mutation operator targeting one primary evidentiary construct; the mutation may nevertheless have multiple downstream diagnostic or claim consequences.
Conceptually, the taxonomy spans four families of evidentiary stress. A1–A4 test content binding, deletion, admission, and replay handling; A5–A8 test authoritative order, lineage, source identity, and conflict handling; A9–A12 test invocation/effect coverage, observer-path coverage, and bounded reasoning in the presence of alternate paths; and A13–A14 separate manipulation before versus after the applicable commitment boundary. These families organize the mutation mechanisms, but they are not severity rankings and do not imply that attacks within a family have identical downstream effects.
A5 and A6 are distinct: A5 modifies the authoritative ordering relation while preserving predecessor/DAG structure, whereas A6 attacks predecessor or DAG linkage itself.
A9 and A10 target different claim dependencies: effect evidence versus invocation evidence.
A11 removes observation from a selected path. A12 tests scope-bounded reasoning in a constructed world containing an alternate path; it does not require the evaluator to detect an intrinsically unobservable path.
A13 and A14 differ by commitment boundary. A13 modifies evidence before the applicable commitment; A14 modifies it afterward.
The commitment boundary is the point at which the frozen mechanism treats the evidence representation as integrity-protected or otherwise fixed for subsequent verification. A pre-commitment manipulation can therefore produce a representation whose later integrity checks still succeed even though its semantic content is false relative to independently admissible cross-evidence. A post-commitment manipulation instead changes the protected representation after that boundary and is expected to manifest as an integrity defect under the frozen mechanism.
Critically, A13 did not ask the evaluator to infer falsity from an internally self-consistent committed record alone. Every A13 primary case contained independently admissible cross-evidence, frozen before scoring, sufficient to make the manipulated proposition experimentally identifiable.

5.8. Evaluator Information Boundary

The evaluator received presented evidence, frozen policy and authorization context, frozen source-registration information, control-role bindings, and generic B1/B2/B3 admissibility rules.
It did not receive baseline identity, hidden D/X/E values, attack identity, mutation provenance, the case oracle, expected disposition, or campaign score.
The same case-independent evaluator processed pristine and adversarial evidence.
This information boundary is part of the experiment’s internal-validity control. By withholding baseline labels, attack identity, mutation provenance, and expected scores, the evaluator cannot select attack-specific reasoning or use hidden experimental truth as a shortcut. Using the same case-independent evaluator for pristine and mutated evidence also ensures that differences in output arise from the presented evidentiary state rather than from per-case evaluator tuning.

5.9. Strongest-Safe Oracle and Scorer

For every adversarial case and every B question, a frozen oracle specified the strongest bounded proposition justified by surviving admissible evidence.
“Strongest-safe” means maximal within the frozen proposition registry and scope relation: the oracle selects the most informative registered proposition that remains justified by the surviving admissible evidence, such that no strictly stronger registered proposition is also justified. This is a measurement-model notion of maximality, not a claim that the oracle enumerates every proposition expressible in unrestricted logic or natural language.
The evaluator and oracle were normalized independently into the frozen semantic-profile representation before scoring. The oracle and scoring registry were frozen before adversarial evaluator outputs were generated.
The oracle was never supplied to the evaluator. The frozen record included consistency and reference-integrity checks linking the oracle, semantic-profile registry, scorer, and campaign artifacts; the post-campaign forensic review identified no oracle/scorer registry mismatch or oracle-reference digest mismatch. S02 v0.1 used one prospectively frozen strongest-safe oracle and did not include independent second-oracle adjudication. Oracle correctness therefore remains part of the measurement-model validity boundary. Freezing the oracle before adversarial evaluator outputs were generated prevented result-contingent reinterpretation, but it does not establish oracle independence or eliminate the possibility of oracle-specification error. Independent second-oracle replication or blinded adjudication is therefore a validation target for a separately versioned future experiment rather than a post hoc modification of the authoritative S02 v0.1 campaign.

5.10. Primary Corpus and Execution Discipline

The primary corpus contained 42 deterministic adversarial cases and 126 B1/B2/B3 claim evaluations.
No composite attacks, additional primary cases, independently developed second evaluator, live-agent primary campaign, frontier-model primary adversary, or statistical sampling campaign entered the primary corpus.
Valid unfavorable outcomes were retained. Result-driven retries and replacement cases were prohibited.
A case could be excluded only for a predeclared experimental-validity failure: wrong pristine baseline; mutation not conforming to its frozen operator; semantic no-op or equivalent mutation; invalid mutation provenance; or infrastructure failure producing no evaluable result. An exclusion did not authorize a replacement primary case or expansion of the corpus.
Each exclusion criterion protects a distinct experimental property: the correct pristine baseline protects ground-truth fidelity; conformance to the frozen mutation operator protects treatment fidelity; exclusion of semantic no-ops ensures that the intended factor was perturbed; valid mutation provenance preserves traceability; and the infrastructure criterion requires an evaluable result. Because exclusions could not trigger replacement cases, these rules prevented unfavorable outcomes from being removed through result-driven reruns or corpus expansion.
Generative AI systems were not components of the authoritative experimental measurement path. Under the author’s direct supervision, OpenAI ChatGPT (GPT-5.6 Sol) was used for manuscript drafting and editing, structural and logic review, structured analysis, citation and reference cross-checking, table and figure review, and publication-formatting assistance. OpenAI Codex (accessed September 2026) was used for routine implementation assistance, debugging support, test and code review, and verification of software changes under the author’s direct specification. OpenAI Daybreak Blue under Trusted Access for Cyber (GPT-5.6 Sol; public API alias gpt-daybreak-blue-latest, accessed September 2026) was used for defensive adversarial scientific review of the experimental methodology, claim boundaries, software and evidence-assurance logic, and interpretation of findings. No generative AI system determined the frozen research questions, hypotheses, case definitions, oracle judgments, experimental ground truth, authoritative campaign outcomes, or scientific conclusions; none served as the primary adversary, evaluator, scorer, or oracle.

6. Results

6.1. Primary and Claim-Specific Outcomes

All 42 adversarial experimental cases satisfied the prospectively frozen validity criteria, and all 126 B1/B2/B3 evaluations were scoreable. There were no exclusions. Table 2 consolidates the overall and claim-specific safety-classification distributions; the B3 localization secondary endpoint is reported separately below the table.
B3 localization (separate secondary endpoint): 12/33 eligible evaluations correctly localized the supported enforcement/failure boundary (36.36%). Localization is evaluated separately from semantic-safety classification.
The three outcome categories also distinguish semantic safety from retained information. In particular, 36/42 B3 evaluations (85.71%) were SAFE_CONSERVATIVE. These evaluations remained within the semantic boundary permitted by surviving admissible evidence but retained less information than the strongest-safe oracle supported. Accordingly, SAFE_CONSERVATIVE is a semantically safe outcome but not necessarily an optimal engineering outcome: frequent conservative withholding can reduce the diagnostic and decision-support value of an assurance system even when it avoids unsupported establishment.
The B3 localization endpoint measures retained operational specificity rather than semantic safety. Correct localization was achieved in 12/33 eligible evaluations (36.36%). Thus, an evaluator may remain semantically safe while failing to preserve an evidence-supported containment or enforcement boundary that could assist incident reconstruction, control diagnosis, or remediation.
The frozen zero-UFER target failed. The ratio 23/126 is the exact proportion over the constructed deterministic corpus; 18.25% is its rounded percentage representation, not an estimate of a real-world failure probability.
The outcome distributions were descriptively heterogeneous across the three co-primary claim classes. B2 accounted for 19/23 ≈ 82.61% of the authoritative unsafe evaluations. This is descriptive of the frozen corpus and does not establish a population-level B2 failure rate.

6.2. Unsafe Evaluations by Attack Class

Unsafe evaluations occurred under nine attack classes: A1 = 6, A5 = 3, A7 = 2, A8 = 3, A9 = 3, A10 = 1, A11 = 1, A13 = 2, and A14 = 2.
A2, A3, A4, A6, and A12 produced zero unsafe evaluations in this corpus. These zero counts do not establish general robustness to those attack classes.
These attack-class counts identify where unsafe semantics appeared in the constructed corpus; they should not be read as a ranking of inherent attack severity or real-world likelihood. Each attack class was applied to only the three frozen baselines, and the three B questions depend on different evidence properties. The mechanism-level analysis in Section 7 is therefore the appropriate basis for explaining why unsafe outcomes occurred, while the raw class counts serve as traceable descriptive summaries of this campaign.

6.3. Frozen Hypothesis Disposition

H1 (Safety): Not supported. Twenty-three of 126 claim evaluations received the authoritative UNSAFE_FALSE_ESTABLISHMENT classification; the zero-UFER target failed.
H2 (Conservative Degradation): Not uniformly supported. H2 predicted that when evidence required for a pristine claim was compromised or lost, unsupported establishment would be removed or qualified rather than silently retained. Six A1 case–claim evaluations constituted direct counterexamples: A-A1 B1, A-A1 B2, B-A1 B1, B-A1 B2, B-A1 B3, and C-A1 B1. In these evaluations, deviation-related or dependent semantics persisted after degradation of the required request-resource semantic binding. These six counterexamples are sufficient to prevent uniform support for H2. RC11 evaluations involved a different mechanism—an overbroad inadmissibility assertion after localized degradation rather than silent retention of the original unsupported establishment—and therefore are analyzed separately from the direct A1 counterexamples to H2.
H3 (Differential Resilience): Descriptively consistent with the predicted differential outcome pattern. Unsafe evaluations were B1 3/42, B2 19/42, and B3 1/42. No inferential comparison was specified, and the proposed explanation that these differences arise from distinct evidence dependencies was not independently tested as a causal hypothesis.
H4 (Evidence-Property Separation): Descriptively consistent as a construct-separation check within the frozen evidence model. Representation integrity, source authority and identity, ordering and lineage, completeness, and observation coverage generated distinguishable states and outcomes. Conflict handling and semantic truth, stressed particularly by A8 and A13, provided additional construct-separation observations beyond those frozen H4 dimensions. This does not establish universal separability outside the constructed evidence model.
These dispositions follow the prospectively frozen operational meanings of the hypotheses. H1 and H2 are safety-oriented and can be defeated by observed unsafe semantic retention, whereas H3 and H4 describe patterns within the complete deterministic corpus rather than population parameters. Accordingly, “descriptively consistent” for H3 or H4 does not mean statistically confirmed or causally established beyond the frozen evidence model.

7. Exploratory Post-Campaign Forensic Analysis

The analyses in this section are explicitly post-campaign and exploratory. They were not part of the prospectively frozen primary or secondary endpoint set, did not contribute to the pre-specified hypothesis dispositions, and did not alter any authoritative campaign score. The authoritative UFER and claim-level classifications were frozen before forensic interpretation.
A read-only forensic review was then applied to the 23 authoritative UNSAFE_FALSE_ESTABLISHMENT evaluations to characterize possible implementation and measurement mechanisms. This exploratory partition produced RC1_SEMANTIC_BINDING_NOT_PROPAGATED: 6 evaluations (26.09%); RC11_SCOPE_OVERCLAIM: 15 evaluations (65.22%); and RC13_SCORER_NORMALIZATION_ARTIFACT: 2 evaluations (8.70%). Twenty-one evaluations were interpreted as substantive semantic incompatibilities and two as possible scorer-normalization artifacts. No evaluation was rescored or removed from the authoritative endpoint.
RC identifiers denote exploratory post-campaign forensic root-cause categories and are distinct from the frozen A1–A14 attack identifiers. The identifiers are drawn from the frozen forensic root-cause taxonomy; their numbering is non-contiguous because only RC1, RC11, and RC13 were populated in this corpus.
RC1 comprises the six A1 evaluations; RC11 comprises the fifteen B2 evaluations under A5, A7, A8, A9, A10, A11, and A14; and RC13 comprises the two A13 B2 evaluations. In ordinary terms, RC1 denotes failure to propagate a required semantic binding after evidentiary content changed; RC11 denotes scope overclaim after localized degradation; and RC13 denotes a possible mismatch introduced by the scorer’s normalization of a generic evaluator output. These RC labels describe post-campaign failure mechanisms and do not replace the A1–A14 labels that identify the experimental mutations themselves.
Figure 2 summarizes the claim-relative degradation logic and the exploratory 6/15/2 forensic decomposition. The authoritative UFER remains 23/126.

7.1. RC1: Semantic-Binding Dependency Omission

A1 changed the resource binding of the action request while retaining source identity, event cardinality, ordering, lineage, and a commitment valid for the mutated representation.
The strongest-safe oracle required deviation status to become unresolved because the request resource no longer aligned with surviving downstream evidence. The evaluator nevertheless retained deviation-related semantics in six case–claim evaluations: A-A1 B1, A-A1 B2, B-A1 B1, B-A1 B2, B-A1 B3, and C-A1 B1.
These six evaluations arose from three A1 experimental cases.
The implementation-level mechanism was a missing end-to-end semantic dependency: default-deny policy behavior continued to characterize the mutated resource as unauthorized, while some downstream branches correlated records through action identity and lineage without consistently reconciling request-resource identity against execution, effect, or closure evidence.
The finding therefore identifies an implementation-specific semantic-binding dependency omission, not a general weakness of agent-assurance systems.

7.2. RC11: Global Gating After Localized Evidentiary Degradation

Fifteen substantive B2 unsafe evaluations shared a second mechanism.
The evidence machinery generally recognized the attacked evidentiary defect. Ordering manipulation was detected; forged-source evidence was excluded as unauthoritative; conflicting authoritative assertions were identified; missing invocation/effect coverage was recognized; observer-path evidence was lost; and post-commitment integrity failure was detected.
The unsafe behavior occurred downstream.
The frozen evaluator applied collection and normalization eligibility before reconstructing the strongest surviving observable-action-path proposition. A localized defect could therefore invalidate global eligibility and produce a proposition equivalent to: no bounded observable-action-path proposition is admissible.
Across the 15 substantive evaluations, the strongest-safe oracle still supported a narrower principal path, observed prefix, local sequence, or control-free sequence.
This was not merely an information-retention failure. The broad inadmissibility itself exceeded the semantic limitation justified by surviving evidence. Ordinary omission of a supported prefix could have remained SAFE_CONSERVATIVE; these evaluations were unsafe because the evaluator’s normalized proposition asserted too broad a limitation.
Global collection/normalization gating in the frozen evaluator sometimes converted localized evidence degradation into an overbroad B2 inadmissibility proposition.
This is a design-level mechanism within the frozen evaluator implementation, not evidence of a universal architectural defect.

7.3. B2 Claim-Language Granularity

The frozen evaluator exposed relatively few intermediate B2 proposition forms. Some evidence states that retained meaningful partial-path information therefore mapped to generic INCONCLUSIVE or INADMISSIBLE forms.
The strongest-safe oracle vocabulary was more granular, including partial prefixes, unresolved control outcomes, control-free segments, and principal-path-restricted propositions.
Limited proposition expressiveness is not inherently unsafe. If an evaluator returns a weaker proposition still entailed by the oracle, the result is SAFE_CONSERVATIVE. Reduced expressiveness contributes to unsafe classification only when the normalized fallback itself contains content or scope not entailed by the oracle.

7.4. B3: Safe Conservative Withholding

Of the 42 B3 claim evaluations, the raw evaluator produced 12 ESTABLISHED, 30 INCONCLUSIVE, and zero INADMISSIBLE dispositions.
Among the 12 established outputs, five were exact, six safe conservative, and one unsafe. All 30 inconclusive outputs were safe conservative.
B3 therefore primarily exhibited an information-retention limitation rather than a semantic-safety failure.
All 12 established B3 outputs localized the boundary correctly, including the single unsafe B3 evaluation (B-A1 B3); that evaluation was unsafe on a non-boundary semantic dimension. Of the 30 inconclusive B3 outputs, 21 were localization-eligible but missed supported localization, while the remaining nine were localization-ineligible because surviving admissible evidence did not support a boundary determination.
The low localization rate thus principally reflects withheld supported localization rather than repeated unsafe selection of an incorrect boundary.

7.5. RC13: A13 and the Measurement Boundary

Two authoritative unsafe B2 case-claim pairs (A-A13 B2 and B-A13 B2) were classified post hoc as possible scorer-normalization artifacts.
Their raw evaluator outputs were generic INCONCLUSIVE; neither positively asserted an unsupported execution, effect, boundary, or complete path.
Under the frozen normalization registry, the generic B2 inconclusive form mapped to an OBSERVED_PATH scope while withholding path continuity. The corresponding strongest-safe oracle propositions mapped to PRINCIPAL_OBSERVED_PATH.
Under the frozen scope relation, PRINCIPAL_OBSERVED_PATH ≺ OBSERVED_PATH, so the narrower oracle profile did not entail the evaluator’s broader normalized scope. The scorer therefore returned UNSAFE_FALSE_ESTABLISHMENT.
A direct post-campaign reading of the raw evaluator text found no comparably broad positive natural-language assertion. These two evaluations are therefore reported as possible scorer-normalization artifacts.
They remain part of the authoritative 23/126 result because the scorer and normalization model were frozen before campaign execution. The forensic interpretation identifies a construct-validity limitation; it does not authorize recoding.
These cases illustrate a specific classification-bias pathway in the measurement model. A generic raw evaluator response can acquire a registered scope during normalization that is semantically stronger than the natural-language response itself appears to assert. When that normalized scope exceeds the oracle-supported scope, the frozen scorer necessarily classifies the pair as unsafe. The resulting risk is therefore not arbitrary scorer error but construct mismatch between raw evaluator language and the semantic profile assigned by the normalization registry. This limitation was not corrected post hoc because doing so would alter the prospectively frozen measurement instrument after observing the campaign result.

8. Discussion

8.1. Claim Outcomes Were Heterogeneous

The three co-primary assurance questions exhibited distinct robustness characteristics under the same frozen mutation set. B1 produced three unsafe evaluations, B2 nineteen, and B3 one. B3 simultaneously produced many safe-conservative outcomes.
The experiment therefore supports treating evidence-plane robustness as claim-specific within this constructed corpus rather than as one undifferentiated property of an evaluator.

8.2. Tamper Detection and Claim Robustness Are Distinct

Several evidence defects associated with unsafe B2 evaluations were themselves detected correctly.
The failure arose because correct evidence-level detection did not always propagate into a semantically appropriate downstream assurance proposition.
Correctly detecting an evidentiary defect does not guarantee correct propagation of that defect into downstream assurance claims.

8.3. Claim-Relative Degradation Has Two Dimensions

Claim-relative degradation has the two dimensions defined in Section 4.3 and Section 5.5: semantic safety and information retention.
This distinction explains the contrast between B2 and B3. B2 RC11 was unsafe because the broad path-inadmissibility proposition itself exceeded what surviving evidence justified. Most B3 degradation was safe conservative because the evaluator withheld supportable localization without asserting an incompatible alternative.
Failure to preserve every supportable proposition is therefore not itself an unsafe failure.

8.4. Integrity, Authority, and Truth Are Distinct

A13 separates representational integrity and recognized source status from semantic truth.
Evidence can be structurally valid, correctly committed, and associated with a recognized source while nevertheless encoding a proposition inconsistent with the constructed experimental world.
This distinction parallels provenance and attestation systems in which authenticated origin and downstream appraisal remain separate concerns [7,8,9].
Here, semantic truth is bounded to the constructed experimental world: it means consistency of the proposition with the independently specified case ground truth and admissible cross-evidence relevant to that proposition. The manuscript does not treat semantic truth as something established solely by cryptographic integrity, source recognition, or the internal self-consistency of a record.

8.5. Semantic Safety and Diagnostic Utility Are Distinct

A binary safe/unsafe score would conceal an important engineering result.
Thirty-six of 42 B3 evaluations were SAFE_CONSERVATIVE. These outputs avoided unsupported semantic assertions but frequently preserved less containment or localization information than the strongest-safe oracle permitted. SAFE_CONSERVATIVE should therefore not be interpreted as unconditional engineering success. It represents semantic safety with information loss. In an operational assurance system, conservative withholding can be preferable to unsupported establishment because it preserves the evidentiary safety boundary; however, excessive conservative withholding can reduce diagnostic utility, auditability, incident reconstruction, and decision support. A system that remains safe primarily by returning weak or inconclusive propositions may satisfy a semantic-safety objective while failing to preserve information that operators could legitimately use.
The distinction between EXACT_SAFE and SAFE_CONSERVATIVE therefore measures information retention separately from semantic safety. An effective assurance system should seek both: no proposition outside the surviving evidence-supported boundary and maximal preservation of propositions that remain justified within that boundary.

8.6. Relation to Adversarial Agent Evaluation

Existing benchmarks largely ask whether agent behavior can be manipulated or made harmful [19,20,21,22,23].
The present experiment asks a complementary question: after behavior is fixed in a constructed world, can an assurance evaluator still characterize that behavior safely when the evidence available to it is attacked?
Behavioral robustness and assurance-evidence robustness are therefore distinct evaluation dimensions.

9. Threats to Validity and Limitations

The experiment uses a deterministic synthetic corpus rather than naturally occurring agent traces. This improves experimental control and ground-truth identifiability but does not estimate real-world attack prevalence.
The corpus contains three selected D/X/E progression states: Baseline A, (D, X, E) = (1, 0, 0); Baseline B, (D, X, E) = (1, 1, 0); and Baseline C, (D, X, E) = (1, 1, 1), rather than exhaustive coverage of the binary D/X/E state space. In particular, no D = 0 baseline was evaluated, so B1 behavior in a no-deviation ground-truth class is outside S02 v0.1. The corpus also contains exactly 14 single-operator evidence-plane attacks; composite attacks were not evaluated.
One evaluator implementation, one frozen scorer, one strongest-safe oracle, and one authoritative campaign were studied. The resulting classifications are therefore conditional on this frozen measurement composition. The deterministic design provides strong within-experiment reproducibility but does not establish that an independently developed evaluator or independently constructed oracle would produce identical proposition boundaries or classifications. Independent evaluator replication and blinded second-oracle adjudication would test different forms of measurement-model dependence and are reserved for separately versioned replication rather than retroactive modification of S02 v0.1.
No live autonomous agent or frontier model served as the primary adversary. The experiment establishes internal validity for deterministic evidence mutation rather than ecological validity for adaptive attackers.
The scorer uses a closed-world proposition registry and frozen scope relation. The two A13 B2 evaluations identify a possible normalization-related construct-validity limitation.
The exploratory partition of the 23 authoritative unsafe evaluations into 21 substantive incompatibilities and two possible scorer artifacts was not a prospectively frozen endpoint and was not independently replicated.
The strongest-safe oracle serves as the reference for both semantic safety and information retention. It was prospectively frozen and subject to the study’s consistency and reference-integrity checks, but S02 v0.1 did not include independent second-oracle adjudication; oracle correctness therefore remains part of the measurement-model validity boundary. Only semantic unsafety enters UFER; frequent SAFE_CONSERVATIVE outcomes should therefore not be interpreted as additional UFER failures.
The 42 baseline–attack pairs are the experimental units. The 126 B evaluations are co-primary assessments of those cases, not 126 statistically independent observations.
B3 localization measures boundary correctness separately from other semantic dimensions. A boundary can therefore be correct while another aspect of a proposition remains unsafe.
We did not evaluate B4 (systemic cause and validated mitigation). Forensic identification of an implementation mechanism is not equivalent to controlled causal validation of a mitigation.
Finally, UFER = 23/126 (18.25%, rounded) is a constructed-corpus proportion, not a real-world probability, prevalence estimate, or expected failure rate for autonomous-agent assurance systems generally.

10. Design Implications and Future Work

10.1. End-to-End Semantic Binding

End-to-end semantic binding means that proposition-relevant identities and attributes must remain explicitly connected as evidence moves from an action request through policy and authorization appraisal, execution or invocation records, control observations, effect observations, and closure evidence. The A1 forensic result illustrates why local record validity is insufficient: downstream records can remain individually well-formed while no longer referring to the same resource or action semantics as the mutated request. An assurance evaluator should therefore reconcile the bindings required by a claim and weaken or withhold that claim when those bindings cannot be established, rather than relying on action identity or lineage alone.

10.2. Claim-Relative Safety

Claim-relative safety requires the consequence of an evidence defect to be localized to the proposition dimensions that depend on the affected evidence. A fault in ordering, authority, coverage, or integrity in one segment does not automatically erase every proposition supported by unaffected evidence elsewhere in the case. The RC11 pattern demonstrates the opposite failure mode: localized defects sometimes triggered global collection or normalization ineligibility and produced a broader assertion of path inadmissibility than the surviving evidence justified. Robust evaluators should therefore propagate uncertainty at the claim, dimension, and scope level instead of indiscriminately invalidating the entire evidentiary collection.

10.3. Information Retention

Information retention is a distinct design objective from semantic safety. A conservative evaluator can remain safe by withholding a proposition, yet still discard evidence-supported prefixes, local sequences, or enforcement boundaries that would be operationally useful. B3 provides the clearest example in this corpus: most degraded outputs were safe, but supported localization was often withheld. Preserving the strongest still-justified residual proposition improves diagnostic utility and decision support without relaxing the requirement that every retained proposition remain within the evidence-supported semantic boundary. Operationally, this can create a safety–utility tension: conservative withholding can prevent unsafe establishment, but systematic over-withholding can leave investigators or automated control processes without information that the surviving evidence was sufficient to support.

10.4. Intermediate Proposition Expressiveness

Intermediate proposition expressiveness is the ability of the evaluator language to represent partially established states without forcing them into either complete establishment or generic inconclusiveness. Useful intermediate forms include a supported path prefix with an unresolved continuation, a known execution attempt with an unresolved effect, a control-free segment, or a localized enforcement outcome with other dimensions withheld. When these distinctions are unavailable, materially different evidence states can collapse into the same generic label, reducing information retention and, as the A13 normalization cases show, potentially acquiring unintended scope during normalization. The evaluator vocabulary should therefore encode both the proposition content that is retained and the dimensions or scopes that remain unresolved.

10.5. Measurement-Model Transparency

Measurement-model transparency requires treating semantic normalization, proposition vocabularies, scope relations, and entailment rules as part of the experimental instrument rather than as neutral bookkeeping. The two possible A13 normalization artifacts show that the mapping from a generic evaluator statement to a registered semantic profile can itself affect the safety classification, even when the raw natural-language output contains no comparably broad positive assertion. Reproducible assurance evaluation therefore requires these mappings and relations to be versioned, frozen for an authoritative campaign, and disclosed alongside the evaluator and oracle. Any later change to the normalization or scoring model should be treated as a new measurement version rather than used to retroactively alter the frozen result.

10.6. Future Work

Future work should address the limitations identified by the present deterministic campaign without altering the frozen S02 v0.1 result. A priority is controlled mitigation testing of the two substantive mechanisms identified here: end-to-end semantic-binding failure and global gating after localized evidentiary degradation. Such experiments would directly address B4 by testing whether specific implementation changes causally reduce or eliminate the observed unsafe propositions. A second priority is limited external-validity evaluation using a small number of live-agent or adaptive-adversary cases while preserving the same claim, oracle, and scoring boundaries. A further priority is measurement-model replication using an independently implemented evaluator and/or independently constructed second strongest-safe oracle. A blinded adjudication procedure could quantify oracle disagreement, identify proposition-boundary ambiguities, and distinguish evaluator failure from measurement-model dependence. Additional work may examine selected composite evidence-plane attacks and broader D/X/E baseline coverage. These extensions should be separately versioned and reported as new experiments rather than treated as modifications to the authoritative S02 v0.1 campaign.

11. Conclusions

This study evaluated whether bounded agent-assurance claims remain semantically admissible when their evidentiary substrate is deliberately attacked.
Fourteen deterministic evidence-plane mutations were applied to three known-ground-truth baselines, producing 42 adversarial experimental cases and 126 B1/B2/B3 claim evaluations. The frozen zero-UFER target failed: 23 of 126 evaluations received the authoritative UNSAFE_FALSE_ESTABLISHMENT classification.
Outcomes were descriptively heterogeneous across claim classes. B1 produced three unsafe evaluations, B2 nineteen, and B3 one. An exploratory post-campaign forensic review classified 21 of the 23 authoritative unsafe evaluations as substantive semantic incompatibilities and two B2 case–claim pairs as possible scorer-normalization artifacts; the primary result remained unchanged.
The dominant B2 mechanism involved 15 substantive unsafe evaluations in which localized evidentiary degradation triggered global collection or normalization ineligibility and produced an overbroad proposition about observable-path admissibility.
B3 demonstrated a complementary limitation. Unsafe overclaim was rare, but supported partial localization was frequently withheld. These evaluations were principally SAFE_CONSERVATIVE: they remained within the evidence-supported semantic boundary but retained less operational information than the strongest-safe oracle permitted. SAFE_CONSERVATIVE therefore represents semantic safety, not necessarily optimal assurance performance; excessive conservative withholding can reduce diagnostic and decision-support utility.
Evidence-backed assurance consequently requires two distinct properties. First, evaluators must remain within the semantic boundary justified by surviving admissible evidence after that evidence is degraded. Second, within that boundary, they should preserve as much supported residual information as practicable. Safety against unsupported establishment and retention of justified information are related but experimentally separable objectives.
Within the bounded scope of S02 v0.1, the results demonstrate that evidence-plane robustness cannot be characterized solely by whether tampering is detected or whether an output is nominally conservative. Assurance robustness comprises at least two experimentally distinguishable dimensions—semantic safety and information retention—and both depend on the evaluator, oracle, normalization model, and evidence-to-claim bindings that constitute the measurement and assurance architecture.

Funding

This research received no external funding.

Data Availability Statement

The frozen code, research contracts, deterministic mutation harness, strongest-safe oracle, scorer, authoritative campaign result, and closeout record are publicly available at https://github.com/rcampbell-research/s02-evidence-plane (accessed on 24 September 2026). The exact scientific state evaluated in this paper ends at tag s02-closeout-v0.1, commit 02ba4a726104298f1cd7d3ee623fa81b1a1f8731. The authoritative result is results/s02-stage3c-authoritative-campaign-v0.1.json; recorded SHA-256: 4d273903ff912880ca06ef680614e4169f0bc29a3c4d2acb6a925c767d101e8f. The repository’s current main branch contains publication metadata added after experimental closeout; those later publication-only changes do not revise the frozen S02 v0.1 scientific result. The SHA-256 digest provides byte-level content-identity verification and does not establish semantic truth, scientific validity, completeness, or external validity.

Acknowledgments

During the preparation of this manuscript and associated study software, the author used OpenAI ChatGPT (GPT-5.6 Sol) for manuscript drafting and editing, structural and logic review, citation and reference cross-checking, table and figure review, and publication preparation; OpenAI Codex (accessed September 2026) for routine implementation assistance, debugging, test review, and code inspection under the author’s direct specification; and OpenAI Daybreak Blue under Trusted Access for Cyber (GPT-5.6 Sol; public API alias gpt-daybreak-blue-latest, accessed September 2026) for defensive adversarial review of the experimental methodology, assurance and claim boundaries, implementation logic, and interpretation of results. The author reviewed and verified all AI-assisted outputs and takes full responsibility for the content of this publication. No generative AI system determined the frozen research questions, hypotheses, oracle judgments, experimental ground truth, authoritative campaign outcomes, or scientific conclusions, and none served as the primary adversary, evaluator, scorer, or oracle.

Conflicts of Interest

The author serves as Global Quantum-Safe Executive and Quantum Ambassador with IBM. The author is also the sole owner of Med Cybersecurity LLC, a single-member limited liability company that is currently inactive, conducts no business, and has no employees; the medcybersecurity.com domain is retained solely for the author’s personal correspondence email. Neither IBM nor Med Cybersecurity LLC had any role in the study’s conceptualization, design, software development, data collection, analysis, interpretation, manuscript preparation, or decision to submit, and no funding, infrastructure, data, or products of either organization were used. The study was conducted, and its computing resources were provided, solely by the author.

References

  1. Object Management Group. Structured Assurance Case Metamodel (SACM), Version 2.3; OMG: Milford, MA, USA, 2023; Available online: https://www.omg.org/spec/SACM/2.3 (accessed on 7 September 2026).
  2. Bloomfield, R.; Rushby, J. Confidence in Assurance 2.0 Cases. In The Practice of Formal Methods: Essays in Honour of Cliff Jones, Part I; Cavalcanti, A., Baxter, J., Eds.; LNCS 14780; Springer: Cham, Switzerland, 2024; pp. 1–23. [Google Scholar] [CrossRef] [Scilit]
  3. Goemans, A.; Buhl, M.D.; Schuett, J.; Korbak, T.; Wang, J.; Hilton, B.; Irving, G. Safety Case Template for Frontier AI: A Cyber Inability Argument. arXiv 2024, arXiv:2411.08088. [Google Scholar]
  4. Campbell, R. Runtime Cryptographic Evidence to Bounded Assurance Verdicts: Deterministic Conformance and Policy Appraisal for SHA-256 and AES-256-GCM. Preprints 2026, 2026081686. [Google Scholar] [CrossRef] [Scilit]
  5. Moreau, L.; Missier, P. (Eds.) PROV-DM: The PROV Data Model; W3C Recommendation: Wakefield, MA, USA, 2013; Available online: https://www.w3.org/TR/2013/REC-prov-dm-20130430/ (accessed on 24 September 2026).
  6. Torres-Arias, S.; Afzali, H.; Kuppusamy, T.K.; Curtmola, R.; Cappos, J. in-toto: Providing Farm-to-Table Guarantees for Bits and Bytes. In Proceedings of the 28th USENIX Security Symposium (USENIX Security 19), Santa Clara, CA, USA, 14–16 August 2019; USENIX Association: Santa Clara, CA, USA, 2019; pp. 1393–1410. [Google Scholar]
  7. Birkholz, H.; Thaler, D.; Richardson, M.; Smith, N.; Pan, W. RFC 9334: Remote ATtestation procedureS (RATS) Architecture; RFC Editor: Wilmington, DE, USA, 2023. [Google Scholar] [CrossRef] [Scilit]
  8. Lundblade, L.; Mandyam, G.; O’Donoghue, J.; Wallace, C. RFC 9711: The Entity Attestation Token (EAT); RFC Editor: Wilmington, DE, USA, 2025. [Google Scholar] [CrossRef] [Scilit]
  9. Birkholz, H.; Delignat-Lavaud, A.; Fournet, C.; Deshpande, Y.; Lasker, S. RFC 9943: An Architecture for Trustworthy and Transparent Digital Supply Chains; RFC Editor: Wilmington, DE, USA, 2026. [Google Scholar] [CrossRef] [Scilit]
  10. Theodorakopoulos, L.; Theodoropoulou, A. Auditable LLM Autonomy for Operational Decision-Making: Big Data Evidence and Decision Traces. Comput. Mater. Contin. 2026, 88, 10. [Google Scholar] [CrossRef] [Scilit]
  11. Huang, J.; Hsia, J.; Sun, J.; Shi, F.; Huang, W.; White, I.H. Proof-or-Stop: Don’t Trust the Agent, Trust the Evidence—Loop Engineering for Verifiable Evidence-Gated Lifecycle Control. arXiv 2026, arXiv:2607.14890. [Google Scholar]
  12. Nian, Y.; Yuan, A.; Zhang, H.; Li, J.; Zhao, Y. Auditable Agents. arXiv 2026, arXiv:2604.05485v1. [Google Scholar]
  13. Yin, X.; Du, M.; Prince, M.H.; Cherukara, M.J. Artifact-centered Claim-aware Observability for Autonomous Scientific Agents. arXiv 2026, arXiv:2608.18312. [Google Scholar]
  14. Wang, X.; Zhang, X.; Shao, J. Auditing Harness Tampering in Self-Improving Agents. arXiv 2026, arXiv:2609.00069. [Google Scholar]
  15. Sharif, R. Agent Audit Trail: A Standard Logging Format for Autonomous AI Systems. Internet-Draft Draft-Sharif-Agent-Audit-Trail-01, Work in Progress, 19 August 2026. Available online: https://datatracker.ietf.org/doc/html/draft-sharif-agent-audit-trail-01 (accessed on 24 September 2026).
  16. Dembowski, J. Proof-of-Behavior Protocol for Autonomous AI Agents. Internet-Draft Draft-Dembowski-Agentledger-Proof-of-Behavior-00, Work in Progress, 20 April 2026. Available online: https://datatracker.ietf.org/doc/html/draft-dembowski-agentledger-proof-of-behavior-00 (accessed on 24 September 2026).
  17. Du, Z.; Chen, Y. AUDITA: Certified Auditing and Causal Attribution of Adverse Outcomes in Autonomous Multi-Agent Systems. arXiv 2026, arXiv:2608.22160. [Google Scholar]
  18. Campbell, R. Agentic Shadow Infrastructure: How AI Supply-Chain Drift Creates Unmanaged Enterprise Infrastructure. Computers 2026, 15, 510. [Google Scholar] [CrossRef] [Scilit]
  19. Ruan, Y.; Dong, H.; Wang, A.; Pitis, S.; Zhou, Y.; Ba, J.; Dubois, Y.; Maddison, C.J.; Hashimoto, T. Identifying the Risks of LM Agents with an LM-Emulated Sandbox. In Proceedings of the International Conference on Learning Representations (ICLR 2024), Vienna, Austria, 7–11 May 2024; Available online: https://openreview.net/forum?id=GEcwtMk1uA (accessed on 7 September 2026).
  20. Debenedetti, E.; Zhang, J.; Balunović, M.; Beurer-Kellner, L.; Fischer, M.; Tramèr, F. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. In Proceedings of the Advances in Neural Information Processing Systems 37 (NeurIPS 2024), Datasets and Benchmarks Track, Vancouver, BC, Canada, 10–15 December 2024. [Google Scholar] [CrossRef] [Scilit]
  21. Zhang, H.; Huang, J.; Mei, K.; Yao, Y.; Wang, Z.; Zhan, C.; Wang, H.; Zhang, Y. Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-Based Agents. In Proceedings of the International Conference on Learning Representations (ICLR 2025), Singapore, 24–28 April 2025; Available online: https://openreview.net/forum?id=V4y0CpX4hK (accessed on 7 September 2026).
  22. Andriushchenko, M.; Souly, A.; Dziemian, M.; Duenas, D.; Lin, M.; Wang, J.; Hendrycks, D.; Zou, A.; Kolter, J.Z.; Fredrikson, M.; et al. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. In Proceedings of the International Conference on Learning Representations (ICLR 2025), Singapore, 24–28 April 2025; Available online: https://openreview.net/forum?id=AC5n7xHuR1 (accessed on 7 September 2026).
  23. Deng, Z.; Guo, Y.; Han, C.; Ma, W.; Xiong, J.; Wen, S.; Xiang, Y. AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways. ACM Comput. Surv. 2025, 57, 182. [Google Scholar] [CrossRef] [Scilit]
  24. Leucker, M.; Sánchez, C.; Scheffel, T.; Schmitz, M.; Thoma, D. Runtime Verification for Timed Event Streams with Partial Information. In Runtime Verification 2019; LNCS 11757; Springer: Cham, Switzerland, 2019; pp. 273–291. [Google Scholar] [CrossRef] [Scilit]
  25. Taleb, R.; Hallé, S.; Khoury, R. Uncertainty in Runtime Verification: A Survey. Comput. Sci. Rev. 2023, 50, 100594. [Google Scholar] [CrossRef] [Scilit]
  26. Cimatti, A.; Grosen, T.M.; Larsen, K.G.; Tonetta, S.; Zimmermann, M. Exploiting Assumptions for Effective Monitoring of Real-Time Properties under Partial Observability. Softw. Syst. Model. 2026; advance online publication. [CrossRef] [Scilit]
  27. Papadakis, M.; Kintis, M.; Zhang, J.; Jia, Y.; Le Traon, Y.; Harman, M. Mutation Testing Advances: An Analysis and Survey. In Advances in Computers; Elsevier: Amsterdam, The Netherlands, 2019; Volume 112, pp. 275–378. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Frozen S02 v0.1 experimental design and evaluator information boundary. Three constructed ground-truth baselines are subjected to 14 frozen evidence-plane mutations, producing 42 adversarial cases and 126 co-primary B1/B2/B3 claim evaluations. The case-independent evaluator receives presented evidence plus frozen policy, authorization, source-registration, control-role, and admissibility context. Baseline identity, hidden D/X/E ground truth, attack identity, mutation provenance, the strongest-safe oracle, expected disposition, and campaign score remain outside evaluator input. The independently frozen oracle and semantic scorer form a measurement-only reference path applied after evaluator output is produced. Arrows show the direction of case construction, evaluator evidence flow, and measurement-only scoring; they do not indicate access to hidden ground truth.
Figure 1. Frozen S02 v0.1 experimental design and evaluator information boundary. Three constructed ground-truth baselines are subjected to 14 frozen evidence-plane mutations, producing 42 adversarial cases and 126 co-primary B1/B2/B3 claim evaluations. The case-independent evaluator receives presented evidence plus frozen policy, authorization, source-registration, control-role, and admissibility context. Baseline identity, hidden D/X/E ground truth, attack identity, mutation provenance, the strongest-safe oracle, expected disposition, and campaign score remain outside evaluator input. The independently frozen oracle and semantic scorer form a measurement-only reference path applied after evaluator output is produced. Arrows show the direction of case construction, evaluator evidence flow, and measurement-only scoring; they do not indicate access to hidden ground truth.
Computers 15 00655 g001
Figure 2. Frozen semantic classification and exploratory forensic decomposition. Panel (A) shows the frozen closed-world scoring relation: EXACT_SAFE requires O ⊨ E and E ⊨ O; SAFE_CONSERVATIVE requires O ⊨ E and E ⊭ O; and UNSAFE_FALSE_ESTABLISHMENT requires O ⊭ E, where O is the normalized strongest-safe oracle proposition and E the normalized evaluator proposition. Panel (B) shows the read-only post-campaign RC1/RC11/RC13 decomposition (6/15/2). It is not a frozen endpoint and did not remove, rescore, or otherwise alter any evaluation; authoritative UFER remains 23/126 (18.25%, rounded). In Panel (A), arrows show evidence processing and classification; in Panel (B), they show the post-campaign analytical partition of the fixed result.
Figure 2. Frozen semantic classification and exploratory forensic decomposition. Panel (A) shows the frozen closed-world scoring relation: EXACT_SAFE requires O ⊨ E and E ⊨ O; SAFE_CONSERVATIVE requires O ⊨ E and E ⊭ O; and UNSAFE_FALSE_ESTABLISHMENT requires O ⊭ E, where O is the normalized strongest-safe oracle proposition and E the normalized evaluator proposition. Panel (B) shows the read-only post-campaign RC1/RC11/RC13 decomposition (6/15/2). It is not a frozen endpoint and did not remove, rescore, or otherwise alter any evaluation; authoritative UFER remains 23/126 (18.25%, rounded). In Panel (A), arrows show evidence processing and classification; in Panel (B), they show the post-campaign analytical partition of the fixed result.
Computers 15 00655 g002
Table 1. Frozen evidence-plane attack taxonomy. Each case applies one frozen mutation operator targeting one primary evidentiary construct; paired distinctions A5/A6, A9/A10, A11/A12, and A13/A14 preserve separate evidentiary constructs.
Table 1. Frozen evidence-plane attack taxonomy. Each case applies one frozen mutation operator targeting one primary evidentiary construct; paired distinctions A5/A6, A9/A10, A11/A12, and A13/A14 preserve separate evidentiary constructs.
IDAttack ClassPrincipal StressorFrozen Mutation Semantics
A1Event mutationSemantic/content bindingMutates action/resource content while preserving the other frozen structural properties required by the operator.
A2Evidence deletionCompletenessDeletes a frozen evidence item so claim support must be evaluated from the surviving set.
A3Unauthorized insertionAdmission/authorityInserts evidence that does not satisfy the frozen admission/authority conditions.
A4Replay/duplicationUniqueness/replay handlingReplays or duplicates evidence to test duplicate/replay rejection without changing hidden ground truth.
A5Authoritative event-order manipulationAuthoritative orderingChanges the authoritative ordered-event relation while preserving predecessor/directed acyclic graph (DAG) and source-local sequence structure.
A6Broken predecessor/DAG linkageLineageBreaks predecessor or DAG linkage rather than the authoritative order relation.
A7Source-identity forgerySource identity/authorityChanges source identity so the targeted evidence no longer satisfies frozen authority registration.
A8Conflicting authoritative evidenceConflict resolutionIntroduces incompatible authoritative assertions that must be recognized as conflicting.
A9Effect-evidence suppressionEffect coverage/completenessSuppresses evidence needed to establish operational-effect or no-effect propositions.
A10Invocation-evidence suppressionInvocation coverage/completenessSuppresses evidence needed to establish invocation/attempt or non-invocation propositions.
A11Observer-path bypassObservation coverageRemoves a selected observation while other operational-path evidence may remain.
A12Alternate/covert synthetic pathBounded scopeConstructs an alternate path to test whether claims stay bounded to the evidenced principal/observed scope.
A13Pre-attestation/pre-commitment manipulationSemantic truth before commitmentManipulates evidence before commitment; independent frozen cross-evidence makes the false proposition experimentally identifiable.
A14Post-attestation/post-commitment manipulationIntegrity after commitmentAlters evidence after commitment, producing a post-commitment integrity defect under the frozen mechanism.
Note: Blue header shading identifies column headings. Alternating pale row shading and bold attack identifiers aid navigation only; they do not encode attack severity, outcome, or additional experimental groupings.
Table 2. Authoritative S02 v0.1 semantic-safety outcomes. Values are n (%) within each claim class. EXACT_SAFE denotes semantic equivalence with the strongest-safe oracle; SAFE_CONSERVATIVE denotes a weaker but still oracle-entailed proposition and therefore semantic safety with information loss; UNSAFE_FALSE_ESTABLISHMENT denotes a proposition not entailed by the strongest-safe oracle. The authoritative UFER is 23/126 (18.25%, rounded), and the pre-specified zero-UFER target failed.
Table 2. Authoritative S02 v0.1 semantic-safety outcomes. Values are n (%) within each claim class. EXACT_SAFE denotes semantic equivalence with the strongest-safe oracle; SAFE_CONSERVATIVE denotes a weaker but still oracle-entailed proposition and therefore semantic safety with information loss; UNSAFE_FALSE_ESTABLISHMENT denotes a proposition not entailed by the strongest-safe oracle. The authoritative UFER is 23/126 (18.25%, rounded), and the pre-specified zero-UFER target failed.
Claim ClassValid EvaluationsEXACT_SAFE, n (%)SAFE_CONSERVATIVE, n (%)UNSAFE_FALSE_ESTABLISHMENT, n (%)
Overall12656 (44.44%)47 (37.30%)23 (18.25%)
B14236 (85.71%)3 (7.14%)3 (7.14%)
B24215 (35.71%)8 (19.05%)19 (45.24%)
B3425 (11.90%)36 (85.71%)1 (2.38%)
Note: Green, yellow, and red identify the EXACT_SAFE, SAFE_CONSERVATIVE, and UNSAFE_FALSE_ESTABLISHMENT columns, respectively. Blue and gray identify claim classes and evaluation counts. Bold row labels identify the overall result and B1–B3 classes; they do not denote statistical significance. Percentages may not sum to 100% because of rounding.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Campbell, R. Adversarial Evidence-Plane Robustness for Bounded Agent Assurance: Deterministic Mutation Testing of Claim Admissibility. Computers 2026, 15, 655. https://doi.org/10.3390/computers15100655

AMA Style

Campbell R. Adversarial Evidence-Plane Robustness for Bounded Agent Assurance: Deterministic Mutation Testing of Claim Admissibility. Computers. 2026; 15(10):655. https://doi.org/10.3390/computers15100655

Chicago/Turabian Style

Campbell, Robert. 2026. "Adversarial Evidence-Plane Robustness for Bounded Agent Assurance: Deterministic Mutation Testing of Claim Admissibility" Computers 15, no. 10: 655. https://doi.org/10.3390/computers15100655

APA Style

Campbell, R. (2026). Adversarial Evidence-Plane Robustness for Bounded Agent Assurance: Deterministic Mutation Testing of Claim Admissibility. Computers, 15(10), 655. https://doi.org/10.3390/computers15100655

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop