Next Article in Journal
UAV Identification Under Low SNR via Multi-Resolution Analysis and Riemannian Structure Preservation
Previous Article in Journal
A Knowledge-Driven Informed Search Framework for 3D Multi-UAV Cooperative Path Planning in Nearshore Coastal Mountainous Environments Using Neural Spatial Priors
 
 
Article
Peer-Review Record

Evidence-Carrying Mission Admission Contracts for Natural-Language UAV Task Submission

by Zhiwei Huang 1,2, Gang Wei 2, Gang Wang 2,*, Xuan Liu 1,2, Haolun Sun 1,2, Xiaoyang Han 2 and Hui Yuan 2
Reviewer 1: Anonymous
Reviewer 2: Anonymous
Reviewer 3: Anonymous
Reviewer 4:
Submission received: 22 June 2026 / Revised: 17 August 2026 / Accepted: 23 August 2026 / Published: 16 September 2026
(This article belongs to the Section Artificial Intelligence in Drones (AID))

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

The manuscript introduces EAMSR, an evidence-carrying Mission Admission Contract framework intended to govern natural-language UAV task submission before tasks enter the downstream planning chain. The framework separates large-language-model-generated candidates from the final admission decision and combines clause-level evidence tracing, authority and mutability constraints, isolation of unsupported semantic increments, mission-consequence screening, and backend plan-witness validation. The final output is classified as ADMIT, CLARIFY, or REJECT.

Nevertheless, the present evidence does not yet support several of the manuscript’s strongest conclusions. The principal concern is that the benchmark, governance rules, human labels, backend model, and EAMSR decision logic appear to have been developed within the same controlled framework. This creates a substantial risk of circular validation: the method is evaluated against labels and constraints that may encode essentially the same admission logic that the method implements. In addition, the comparison methods are simplified configurations rather than reproduced state-of-the-art systems, the LLM implementations are insufficiently specified, statistical uncertainty is not reported, and validation remains predominantly synthetic. These issues are important because the paper presents EAMSR as a safety-relevant governance layer for UAV mission admission.

The manuscript contains a potentially strong contribution, but significant methodological clarification, stronger external validation, and more cautious interpretation are required before publication.

 

Comments:

 

  • The manuscript positions the Mission Admission Contract as a new decision object situated between natural-language interpretation and mission planning. This is potentially valuable. However, several components resemble existing concepts from assume–guarantee contracts, safety assurance cases, policy-enforced planning, provenance-aware specification, formal requirement refinement, and runtime-assurance architectures. The related-work section currently explains why individual prior approaches do not provide the complete combination proposed by EAMSR, but it does not sufficiently distinguish:
  • the MAC from an assurance case or proof-carrying artifact;
  • the proof-obligation gate from conventional policy-based authorization and constraint validation;
  • the backend witness from standard plan validation;
  • mission-consequence equivalence from semantic-preserving specification refinement;
  • evidence-carrying clauses from provenance-tracked formal requirements

The authors should add a structured comparison table covering at least the following dimensions: natural-language grounding, clause provenance, evidence status, authority, mutability, ambiguity handling, semantic-preservation criteria, backend executability, audit output, and UAV-specific validation. The table should include the most closely related methods rather than broad research categories.

The paper should then identify precisely which element constitutes the principal methodological novelty:

  • the MAC representation itself;
  • the conjunctive admission predicate;
  • the authority-bounded refinement relation;
  • the mission-consequence equivalence mechanism;
  • or the integration of these mechanisms.

At present, the novelty appears to reside mainly in system-level integration. That is still publishable, but it should be stated accurately and supported by direct comparisons.

2. The construction of EAMSR-Bench requires substantially greater clarification because the benchmark appears to have been developed within the same conceptual and technical framework as the proposed method. Each sample contains a natural-language instruction, trusted context, protected invariants, evidence base, governance policy, backend model, and human admission label. If these elements were designed by the same researchers who defined the EAMSR admission logic, the evaluation may become partially self-confirming, since the benchmark may encode the same assumptions, decision boundaries, and governance rules that the framework is intended to reproduce. The authors should explain in detail how the 120 instructions were generated, whether any were obtained or adapted from real UAV operators, who designed the six admission-risk categories, and whether the benchmark was created before or after the framework was finalized. The manuscript should also clarify whether the human annotators were independent of the developers, whether they had access to EAMSR outputs, and how borderline cases were handled. A stronger validation should include an independently constructed test subset prepared and annotated by UAV operators or mission-planning specialists who were not involved in developing the framework. Without such an independent test set, the reported results should be presented as performance under a controlled, internally designed benchmark rather than evidence of broad generalization.

3. The human-labeling procedure is currently insufficiently documented. The manuscript reports 42 ADMIT, 47 CLARIFY, and 31 REJECT cases, but it does not adequately describe the number of annotators, their expertise, the annotation instructions, or the adjudication procedure. This is particularly important because the distinction between CLARIFY and REJECT may depend on subjective judgments regarding whether missing evidence can realistically be supplied, whether authorization may be obtained, or whether a conflict is fundamentally irresolvable. The authors should report whether the samples were independently annotated, whether the annotators were blinded to system outputs, and how disagreements were resolved. An inter-rater reliability statistic, such as Fleiss’ kappa, Krippendorff’s alpha, or pairwise Cohen’s kappa, should be provided. Class-specific agreement would also be useful, particularly for CLARIFY and REJECT. Representative disagreement cases should be discussed to demonstrate whether the remaining classification errors reflect limitations of EAMSR or genuine ambiguity in the annotation process.

4. The comparison configurations are useful for illustrating the incremental contribution of individual admission mechanisms, but they are not sufficiently strong to support broad claims of superiority over existing methods. The manuscript explicitly states that Direct-LLM, LLM+Backend, Rule-Gate, and Greedy Relaxation are controlled configurations rather than complete implementations of systems such as SafeGate, DUPLEX, or UAV-NL-STL-Repair. Consequently, these comparisons should be framed as controlled architectural baselines rather than state-of-the-art benchmarks. The current configurations are also structurally disadvantaged because they omit precisely the evidence, authority, semantic-isolation, consequence-screening, and audit components that distinguish EAMSR. The authors should either implement one or more representative published methods under the same evaluation protocol or moderate the comparative claims accordingly. A particularly informative additional baseline would be a deterministic, non-LLM admission system using the same ontology, evidence base, governance policy, and backend model. This would help establish whether the LLM provides meaningful value or merely generates candidate clauses that are subsequently governed by conventional rule-based mechanisms. The manuscript should also confirm that all configurations used equivalent prompts, candidate budgets, backend access, retry limits, and computational resources.

5. The description of the language-model implementation is not sufficiently precise for scientific reproduction. Terms such as “GPT-4-class model” and “DeepSeek-class model” are too general, particularly because model behavior can vary substantially across versions, providers, decoding configurations, and structured-output modes. The authors should provide the exact model names, release or API versions, temperatures, top-p values, token limits, system prompts, user prompts, few-shot examples, retry procedures, and output-parsing methods. It should also be stated whether function calling, JSON schemas, constrained decoding, or post-processing rules were used. The manuscript should clarify whether benchmark examples were employed during prompt development, whether identical prompts were used across all model backbones, how malformed responses were treated, and whether candidate generation was deterministic. Since each task was reportedly repeated five times, the authors must explain whether the final reported decision was obtained from a single run, majority voting, mean performance, best-case selection, or another aggregation method. Exact model versions and collection dates should be recorded because proprietary models may change over time.

6. The statistical treatment is currently insufficient for a dataset of only 120 tasks, with several individual risk categories containing just 18 samples. Single percentage values do not adequately represent uncertainty, and an observed unwarranted admission rate of 0.0% should not be interpreted as evidence that the true error probability is zero. The authors should report 95% confidence intervals for admission accuracy, unwarranted admission rate, class-specific precision and recall, and other major metrics. A complete confusion matrix should be included, together with macro-averaged precision, recall, and F1-score for the three decision classes. Since the experiments include repeated runs, the manuscript should provide mean values and standard deviations, or medians and interquartile ranges, rather than only aggregate percentages. Paired bootstrap comparisons or another appropriate paired statistical test should be used to assess whether the improvements over the controlled baselines are statistically meaningful. For the zero observed unwarranted admissions, an exact binomial confidence interval should be reported, and the manuscript should state that no unwarranted admissions were observed in the evaluated test set rather than claiming that such decisions were eliminated generally.

7. Several metrics are closely aligned with the internal architecture of EAMSR and may therefore favor the proposed method by construction. Evidence coverage rate, untrusted semantic increment blocking rate, protected-invariant preservation rate, consequence clarification rate, and audit-trail completeness directly correspond to modules that are explicitly implemented within EAMSR. Comparison methods that do not maintain equivalent internal structures are consequently unlikely to perform well on these measures. The authors state that the metrics were evaluated post hoc using unified human annotations, but the corresponding protocol is not sufficiently described. The manuscript should formally define the unit of analysis for each metric, explain how hard clauses and unsupported increments were counted, specify how partial evidence was treated, and clarify how semantic preservation was assessed after task reformulation. The assessors should ideally be blinded to the configuration that generated each result. Audit-trail completeness should be presented primarily as a system-capability measure rather than as an independent indicator of decision quality, since EAMSR is explicitly designed to generate such an audit trail.

8. The formal properties presented in the manuscript are useful architectural statements, but their status should be described more carefully. Protected-invariant preservation, preservation of user hard intent, exclusion of unsupported assumptions, and bounded closure largely follow from the definitions of the proof obligations, admissible-refinement relation, and finite candidate budget. These properties therefore appear to be guarantees of the abstract specification under assumed-correct implementation rather than independently verified guarantees of the complete operational system. The authors should distinguish clearly between properties guaranteed by construction in the formal model, properties empirically tested in the implementation, and assumptions that must remain valid for those properties to hold. The framework depends on the correctness and completeness of the trusted context, evidence base, governance policy, ontology, clause extraction process, authority assignments, and backend model. An incorrect or outdated evidence source, an improperly defined immutable constraint, or an incomplete backend model could permit a formally valid but operationally inappropriate admission. Unless the implementation has been machine-verified, the manuscript should avoid language that may be interpreted as formal certification of the complete system.

9. The mission-consequence mechanism is central to the paper, but the mathematical and operational definitions are currently incomplete. The manuscript introduces a consequence signature, a distance function, an equivalence threshold of 0.15, and a weighted refinement objective, yet it does not provide sufficient detail regarding how these quantities are calculated. The authors should define the exact form of the distance function and explain how heterogeneous components, including Boolean states, categorical mission goals, spatial restrictions, temporal margins, communication feasibility, and energy reserves, are normalized and combined. The values of all weighting coefficients in the refinement objective should be reported and justified. Sensitivity studies should be performed for both the consequence threshold and the optimization weights. The authors should also provide complete examples of two contracts considered consequence-equivalent and two contracts that trigger clarification because they belong to different classes. A single global threshold may not be operationally meaningful across all mission types; for example, a small change in return-energy reserve may be more critical than a similar normalized change in an optional observation preference. Dimension-specific or governance-defined equivalence criteria may therefore be more appropriate.

10. The backend witness is a critical component of the admission process, but the backend itself is insufficiently specified. The manuscript should describe the planning formalism, state representation, action model, search algorithm, temporal discretization, energy-consumption model, communication map, payload constraints, weather assumptions, infeasibility criteria, and termination rules. It should also explain how return-to-home energy is estimated and whether reserve calculations account for distance, altitude, wind, payload, and vehicle performance. The relationship between the backend predicates and the full witness condition should be simplified or clarified, since the current formulation appears to restate several constraints both individually and through the general task-model satisfaction condition. More importantly, the authors should establish whether the same backend logic was used to construct benchmark labels and validate admitted tasks. If the benchmark, human labels, and witness generation rely on the same constraint implementation, the evaluation may not detect systematic errors in the backend model. Some form of independent validation, cross-checking, or alternative planner implementation would strengthen the evidence considerably.

11. The evaluation remains predominantly synthetic and task-level, which limits the conclusions that can be drawn for real UAV mission operations. The authors appropriately state that EAMSR is a pre-execution admission mechanism rather than a flight-safety system, and a full flight campaign is therefore not strictly necessary. Nevertheless, the paper would benefit from stronger integration with an actual UAV mission-management environment. Suitable options include software-in-the-loop integration with PX4 or ArduPilot, hardware-in-the-loop validation, use of recorded real-flight telemetry, operator-in-the-loop experiments, or a larger AirSim campaign with randomized disturbances and operational uncertainty. The existing simulation cases appear primarily illustrative and do not demonstrate that admitted contracts can be reliably converted into mission artifacts accepted by standard autopilot systems. A practical interface showing how the admitted obligations and backend witness are translated into waypoints, mission commands, geofences, payload actions, and abort conditions would significantly improve the relevance of the contribution to the readership of the journal.

12. Authority and mutability are presented as core elements of the framework, but the operational policy model remains abstract. In real UAV operations, authorization may depend on operator role, mission category, jurisdiction, airspace approval, emergency status, payload type, organization, and temporal validity. The authors should provide at least one complete policy example containing defined authority levels, source classes, permitted modifications, immutable constraints, delegations, conflict-resolution rules, and expired or contradictory permissions. A worked example should distinguish clearly among operator hard requirements, regulatory constraints, vehicle limitations, trusted defaults, contextual facts, and LLM-generated hypotheses. The manuscript should also explain how conflicting authorized sources are prioritized and whether an operator can override a system default, a mission supervisor can override an operator preference, or an emergency rule can override a normal geofence policy. Without such concrete examples, authority-bounded refinement remains conceptually appealing but difficult to assess operationally.

13. The residual errors should be analyzed at case level rather than only through aggregated percentages. Five of the eight reported errors occur in the LLM-assumption contamination category, two occur in consequence ambiguity, and one occurs in authority or protected-invariant handling. This distribution suggests that some of the most important limitations are concentrated in semantic isolation and classification of uncertain content. For each erroneous case, the authors should provide the original natural-language instruction, relevant trusted context, human label, system label, extracted anchors, generated clauses, failed or passed proof obligations, and the final reason for disagreement. The analysis should identify whether the error originated from anchor extraction, clause generation, evidence matching, authority classification, consequence-distance calculation, refinement selection, or backend diagnosis. This would allow readers to determine whether the system fails conservatively, whether errors are caused by genuine ambiguity, and which modules require further development.

14. Several claims should be moderated to reflect the limited scale and controlled nature of the evaluation. In particular, the statement that unwarranted ADMIT decisions are eliminated should be replaced by a formulation indicating that no such decisions were observed on EAMSR-Bench under the evaluated settings. Similarly, claims that EAMSR ensures or guarantees admission integrity should be explicitly conditioned on the correctness and completeness of the trusted context, evidence base, governance rules, ontology, semantic extraction, and backend task model. The manuscript already acknowledges some of these limitations in later sections, but the abstract and highlights present the findings more generally. The distinction between demonstrated benchmark performance, formal properties of the abstract architecture, and expected operational behavior should be maintained consistently throughout the manuscript.

15. The availability of code, benchmark data, AirSim configurations, and validation records is a positive aspect of the manuscript, but the reproducibility package should be more explicitly documented and permanently archived. The repository should contain all benchmark instructions, trusted contexts, governance policies, protected invariants, evidence sources, backend models, human labels, annotation guidelines, exact prompts, model identifiers, decoding parameters, raw outputs, random seeds, repeated-run results, error cases, and scripts used to generate every table and figure. Dependency versions and installation instructions should also be provided. Since GitHub content may change after publication, the authors should archive a tagged release using a permanent repository such as Zenodo and provide a DOI corresponding exactly to the submitted or accepted manuscript version.

 

The manuscript would benefit from substantial editorial refinement and improved consistency. Algorithm 1 contains apparently incomplete steps and should be corrected so that candidate generation, consequence evaluation, clarification, refinement, and audit closure are fully represented. The notation should be revised because the symbol KK appears to denote both candidate clauses and the conflict core, while the audit-trail equation appears to contain a duplicated term. ADMIT and ACCEPT should not be used interchangeably, and the terminology “unsafe admission rate” and “unwarranted admission rate” should also be standardized, with the latter being more appropriate because the framework does not establish complete flight safety. The authors should provide tables defining all clause source categories, authority levels, mutability modes, trust labels, and possible evidence states. The ternary support classification should be expanded or discussed to distinguish unsupported, contradicted, missing, stale, and uncertain information. The manuscript should explain how evidence freshness is verified, especially for weather, communication coverage, airspace restrictions, battery state, and payload availability, and how contradictory trusted-context facts are resolved. The UAV ontology should be described quantitatively and structurally, including its vocabulary, relations, clause families, and availability. The imbalance between the 30 T1 samples and the 18 samples assigned to each other risk type should be justified, and the meaning of T1 or “normal-risk” tasks should be introduced before use. Class distributions should also be reported for each mission scenario. Metric denominators should be stated clearly in every relevant table note. Confidence intervals or run-to-run variability should be added to the principal figures and robustness tables. Figure 2 is visually overloaded and contains text that is difficult to read; it should be simplified or divided into separate architecture and decision-flow figures. Figure 3 should avoid switching between UAR and 100−UAR100-\mathrm{UAR}, or the transformation should be made unmistakably clear. The exact definition of decision consistency should be provided, and the number and generation procedure of the scenario-variation tasks should be reported. Runtime measurements should state whether external API latency was included and should provide variation rather than only a mean value. The practical significance of the reported admission time should be discussed for both routine preflight planning and emergency-response missions. AirSim parameters should be presented in a dedicated table, including the UAV model, map, speed assumptions, energy representation, communication modeling, and failure conditions. The manuscript should consistently distinguish task feasibility, task-level executability, mission admissibility, operational safety, and flight safety. The language is generally understandable but frequently repetitive and overly dense, particularly in the Introduction, Related Work, and Discussion, and professional English editing is recommended. Finally, broad citation ranges should be checked to ensure that every cited reference directly supports the associated statement, and all placeholder DOI information, repository links, and journal metadata should be updated before publication.

Author Response

Response to Reviewer 1

We sincerely thank the Reviewer for the careful, technically detailed, and constructive assessment of our manuscript. The comments identified several important issues concerning the positioning of the methodological novelty, the risk of circular validation, the independence and reliability of the human labels, the strength and fairness of the comparison configurations, reproducibility of the LLM implementation, statistical uncertainty, interpretation of architecture-specific metrics, the status of the formal properties, the definition of mission-consequence compatibility, backend fidelity, and the scope of operational validation.

We have revised the manuscript substantially in response. In particular, we have clarified the principal novelty and its relationship to prior concepts; explicitly characterized EAMSR-Bench as an internally constructed controlled benchmark; added an independently authored external test set; strengthened the annotation protocol and reported inter-annotator agreement; introduced a deterministic non-LLM attribution baseline; provided substantially more implementation detail; added confidence intervals, confusion matrices, repeated-run statistics, and paired significance tests; formalized the metric-auditing protocol; clarified the assumption boundary of the formal properties; expanded the consequence and backend models; added case-level error attribution; moderated the claims throughout the manuscript; and strengthened the permanent reproducibility package.

We are grateful for the opportunity to improve both the technical rigor and the presentation of the work. Our point-by-point responses are provided below.

Comment 1 — Positioning of the methodological novelty and relation to prior work

Reviewer’s comment: The Mission Admission Contract and several associated mechanisms resemble existing concepts from assurance cases, proof-carrying artifacts, policy-based authorization, plan validation, semantic-preserving refinement, and provenance-tracked requirements. The manuscript should provide direct structured comparisons and identify precisely where the principal methodological novelty lies.

Response:
We thank the Reviewer for this important observation. We agree that the previous manuscript did not draw the conceptual boundaries sharply enough and, as written, could give the impression that the individual ingredients of EAMSR were being claimed as independently new.

We have therefore substantially revised the Related Work section and added two structured comparison tables. Table 1 now compares EAMSR directly with closely related systems, including nl2spec, SafeGate, DUPLEX, and CROME, along the requested dimensions of natural-language grounding, clause provenance, evidence status, authority, mutability, ambiguity handling, semantic preservation, backend executability, audit output, and UAV-specific validation. Table 2 separately delineates the conceptual boundaries between the principal EAMSR components and their closest established counterparts.

More importantly, we now state explicitly that the principal methodological novelty of EAMSR is the integration of these mechanisms into a single conjunctive mission-admission predicate for natural-language UAV task submission, rather than any claim that the individual concepts have no precedent. In the revised text, we explicitly acknowledge that the MAC adapts ideas from provenance-tracked specifications; the proof-obligation gate extends policy-based authorization with evidence, mutability, and untrusted-semantic-isolation requirements; authority-bounded refinement builds on contract-refinement ideas; mission-consequence compatibility is related to semantic-preserving refinement; and the backend witness is related to conventional plan validation. What is new is their joint use as first-class admission variables before a natural-language mission enters the planning chain.

We believe this revised formulation is both more precise and more appropriately restrained. The corresponding revisions appear in Section 2 and Tables 1–2.

Comment 2 — Construction of EAMSR-Bench and risk of circular validation

Reviewer’s comment: Because the benchmark, governance rules, backend model, and admission framework appear to have been developed within the same conceptual framework, the evaluation may be partially self-confirming. The manuscript should document benchmark construction and ideally include an independently constructed test subset.

Response:
We fully agree with the Reviewer that this is a central validity concern. Rather than attempting to dismiss it, we have made the co-development issue explicit in the revised manuscript and changed the interpretation of the internal benchmark accordingly.

Section 4.1.1 now characterizes EAMSR-Bench as a controlled, internally constructed architecture-stress benchmark. We explicitly state that the 120 Chinese-language mission instructions were authored by the research team using Chinese-language operator manuals, civil-aviation regulations, and mission-planning guidelines; that no instruction was copied from operational flight logs or supplied by external UAV operators; that the six admission-risk types were defined by the research team and were motivated by the EAMSR proof-obligation structure; and that benchmark and framework development occurred concurrently. We therefore now state directly that this design introduces a potential circular-validation risk and that performance on EAMSR-Bench should be interpreted as controlled internal-benchmark performance rather than broad evidence of operational generalization.

To provide an additional layer of independent-authoring validation, we also constructed a new external test set containing 48 natural-language mission requests. These requests were prepared by UAV mission-planning practitioners who were not involved in developing EAMSR, defining the risk taxonomy, designing the prompts, or constructing the admission rules, and who did not have access to the EAMSR-Bench instructions. The external set was used exclusively for evaluation; no prompt, threshold, or model parameter was tuned on it.

We have deliberately retained an important qualification: the external requests still operate within the same admission ontology and decision semantics. We therefore describe this experiment as independent-authoring robustness, not as full operational external validation. On this set, EAMSR achieves 89.6% three-class accuracy and 100.0% binary admission accuracy, with no unwarranted ADMIT decision observed among the 32 reference non-admissible cases.

These changes are reported in Sections 4.1.1, 4.5.1, and 4.6.

Comment 3 — Human-labeling procedure and inter-rater reliability

Reviewer’s comment: The number and expertise of annotators, annotation instructions, blinding, adjudication procedure, and inter-rater reliability should be documented, particularly for the CLARIFY/REJECT distinction.

Response:
We thank the Reviewer for identifying this omission. We agree that the validity of the three-class evaluation depends critically on a transparent reference-labeling protocol, particularly because CLARIFY and REJECT encode a remediability judgment rather than a purely mechanical distinction.

The revised Section 4.1.1 now describes the annotation protocol in detail. Each benchmark sample was independently labeled by two non-developer domain experts, each with more than three years of experience in UAV mission-planning and autonomous-system research. They received the same raw task information—mission instruction, trusted context, protected invariants, evidence base, governance policy, and backend model—but were blinded to EAMSR outputs and audit trails. The initial annotators were not involved in designing the EAMSR admission logic or the benchmark risk taxonomy.

We also formalized the CLARIFY/REJECT boundary through a remediability criterion. A case is labeled CLARIFY when the blocking condition can potentially be resolved through additional evidence, authorization, or clarification of a consequence-sensitive ambiguity. REJECT is assigned when the request conflicts with an immutable/protected admission condition or is conclusively infeasible within the declared context, governance policy, and admissible refinement space.

Disagreements were adjudicated by a senior annotator, with novel cases escalated to a three-member domain-expert panel. We now report pairwise Cohen’s κ = 0.86 for the three-class labels. Class-specific agreement is 1.00 for ADMIT, 0.83 for CLARIFY, and 0.90 for REJECT. For the 78 non-ADMIT cases, initial agreement on the CLARIFY/REJECT distinction was 85.9% (67/78). Eleven initial disagreements occurred among the 120 benchmark cases, all on the CLARIFY–REJECT boundary; eight were resolved by the senior annotator and three by the expert panel.

The complete disagreement records and adjudication outcomes are included in the reproducibility package. We believe these additions make both the reference-label construction and its remaining uncertainty considerably more transparent.

Comment 4 — Strength and fairness of the comparison configurations

Reviewer’s comment: Direct-LLM, LLM+Backend, Rule-Gate, and Greedy Relaxation should not be presented as reproduced state-of-the-art methods. A deterministic non-LLM admission baseline using the same ontology/governance/backend would also be informative.

Response:
We agree with the Reviewer and have revised both the terminology and the experimental design.

First, the revised manuscript explicitly describes Direct-LLM, LLM+Backend, Rule-Gate, and Greedy Relaxation as controlled architectural comparison configurations, not as complete reproductions of SafeGate, DUPLEX, UAV-NL-STL-Repair, or other published systems. We have correspondingly moderated comparative statements throughout the manuscript. The purpose of these configurations is now stated narrowly: to isolate the effect of adding specific admission mechanisms under a common task-level protocol.

Second, following the Reviewer’s suggestion, we added Deterministic-MAC, a deterministic non-LLM admission system. It uses the same ontology, evidence base, governance policy, proof-obligation gate, and task-level backend as EAMSR, but replaces LLM candidate generation with a fixed ontology/template mapping procedure and deterministic refinement enumeration. This provides a more direct attribution test of whether the LLM contributes to admission integrity or mainly to semantic proposal and coverage.

The results are informative in this regard. Deterministic-MAC also produces zero observed unwarranted admissions on the internal and independently authored external sets, but its three-class accuracy is lower than that of EAMSR. Its clause recall, governance-valid candidate availability, and successful-refinement rate are also lower. We therefore revised our interpretation: the LLM primarily improves semantic coverage and refinement flexibility, whereas the admission boundary itself is enforced by the governance and backend layers.

Finally, the LLM-based configurations now use the same task texts, trusted context, ontology, backend model, LLM backbone, candidate budget, refinement budget, malformed-response treatment, and computational resources. No configuration is given a larger search budget or a more favorable parsing setup.

These revisions are described in Section 4.1.2 and Tables 13, 28, and 29.

Comment 5 — Reproducibility of the LLM implementation

Reviewer’s comment: Exact model identifiers, decoding settings, prompts, token limits, retry rules, structured-output handling, prompt-development procedure, repeated-run aggregation, and collection dates should be reported.

Response:
We agree. The previous use of labels such as “GPT-4-class” and “DeepSeek-class” was insufficient for scientific reproduction.

Section 4.1.4 and Table 14 now provide the concrete model identifiers used in the experiments: gpt-4-turbo-preview, Qwen2.5-72B-Instruct, and deepseek-v3, together with provider, temperature, top-p, maximum token budget, and seed settings. The evaluated configuration uses temperature 0, top-p 1.0, and a 4096-token output limit. API responses were collected between January and March 2025. We also explicitly note that the proprietary providers did not expose a more immutable snapshot identifier for the endpoints used; consequently, exact future reproduction of proprietary-model outputs cannot be guaranteed.

The manuscript now states that no function calling or JSON-mode constraints were used. The system prompt specifies the UAV ontology slots and requests a delimited textual clause format, which is parsed by a deterministic regular-expression-based extractor. No few-shot examples were included, and no EAMSR-Bench instructions were used during prompt development. The ontology-based system prompt was prepared before benchmark construction. Exact system and user prompts are included in the archived reproducibility package.

Malformed responses consume a candidate-generation attempt, with at most two retries subject to the remaining candidate budget. Exhaustion without a parseable response is recorded explicitly as BUDGET_EXCEEDED.

We have also clarified the role of the five repeated runs. The principal reported result is the first predefined run (seed 42); it is not obtained by majority voting, best-case selection, or selecting the most favorable run. The five runs are used to characterize run-to-run serving-side variability through mean/standard deviation and the decision-consistency metric.

These details are now provided in Section 4.1.4, Table 14, and the reproducibility package.

Comment 6 — Statistical uncertainty and significance

Reviewer’s comment: Point percentages are insufficient for a dataset of 120 tasks. Confidence intervals, a confusion matrix, class-specific metrics, repeated-run variability, paired statistical tests, and an exact interval for zero observed unwarranted admissions should be provided.

Response:
We agree with the Reviewer and have substantially expanded the statistical analysis.

For the prespecified primary accuracy measures, the revised manuscript reports Wilson 95% confidence intervals. EAMSR’s three-class admission accuracy is 93.3% with a Wilson 95% CI of [87.4%, 96.6%], and binary admission accuracy is 100.0% with its corresponding Wilson interval. Because no unwarranted ADMIT decisions were observed among the 78 reference non-admissible cases, we now report the result as 0/78 observed unwarranted admissions, together with a one-sided 95% Clopper–Pearson upper bound of 3.8%. We no longer interpret 0.0% as evidence that the underlying error probability is literally zero.

A complete three-class confusion matrix has been added, together with precision, recall, and F1 for ADMIT, CLARIFY, and REJECT and the corresponding macro averages. Because the class-specific measures are secondary/descriptive quantities with relatively small class denominators, we distinguish them from the primary inferential measures rather than treating every derived diagnostic quantity as an independent hypothesis test.

We also added paired statistical comparisons. Exact two-sided McNemar tests are used for the binary admission outcomes and UAR comparisons, while a paired bootstrap analysis with 10,000 resamples is used for the three-class outcomes. Improvements over all four controlled comparison configurations remain statistically significant after Holm–Bonferroni correction.

Finally, Table 25 reports mean ± sample standard deviation over five repeated runs for each LLM backbone, rather than presenting only a single aggregate percentage. Complete per-run results and statistical test records are included in the reproducibility archive.

These additions appear in Sections 4.1.3, 4.2, and 4.4 and in Tables 15–17 and 25.

Comment 7 — Architecture-aligned evaluation metrics

Reviewer’s comment: Several metrics correspond directly to EAMSR modules and may favor the proposed architecture. Their units and auditing protocol should be defined, assessors should ideally be blinded, and audit-trail completeness should be interpreted as a system-capability measure rather than independent evidence of quality.

Response:
We agree with this concern and have revised the evaluation framework to prevent the diagnostic metrics from being interpreted as independent evidence of overall superiority.

The revised manuscript now organizes the metrics into three tiers. Binary admission correctness and UAR form the primary admission tier; three-class admission accuracy and per-class precision/recall/F1 form a secondary decision tier; and ECR, UBR, PIR, WSR, CNR, and ATR are explicitly designated as diagnostic metrics that explain which mechanism succeeds or fails. We no longer use the diagnostic metrics as interchangeable evidence of general decision superiority.

We have also added a metric-auditing protocol. For example, ECR is evaluated at the unique atomic hard-clause level, and only Sup(κ)=Y counts as evidenced; contradicted, invalid-source, or insufficient-evidence clauses do not. UBR is evaluated over unique LLM-introduced hypothetical atomic claims rather than all generated text. PIR is evaluated over the protected invariants applicable to the task. ATR is complete only when all mandatory fields relevant to the corresponding decision path are populated, while structurally non-applicable fields must be recorded explicitly as such. Denominators and the treatment of non-applicable cases are now stated.

We also agree with the Reviewer that assessor blinding would have been preferable for the post-hoc diagnostic re-audit. In the present revision, the diagnostic assessors had access to configuration identities; we therefore do not claim that this audit was blinded, and this limitation is now stated explicitly.

Similarly, metrics that follow directly from an admission rule are now interpreted cautiously. In particular, the 100% witness-success rate among ADMIT decisions is identified as a conformance property of the admission rule, rather than an independent performance result. ATR is likewise treated principally as an auditability/system-capability measure.

These clarifications are included in Section 4.1.3 and the corresponding results discussion.

Comment 8 — Status and interpretation of the formal properties

Reviewer’s comment: The formal properties largely follow from the definitions and should be distinguished from machine-verified guarantees of the complete operational implementation.

Response:
We agree with the Reviewer’s interpretation and have revised the manuscript accordingly.

The properties in Section 3.5 are now explicitly introduced as specification-level properties of the abstract admission model. The text states that they hold provided that clause extraction, evidence status, authority assignment, governance policy, and backend predicates are implemented consistently with their definitions. We explicitly state that these properties do not constitute machine-verified certification or formal assurance of the complete operational UAV system.

We also added Table 10 to make the guarantee boundary explicit. The assumptions cover: correct representation of user hard intent by validated anchors; correctness, currency, and prioritization of evidence; fidelity of the authority and mutability policy; adequacy of the task-level backend abstraction; and the semantics of positive and absent witnesses. In particular, an absent witness is not treated as proof of infeasibility unless the backend can establish infeasibility conclusively.

The formulation of the user-intent property has also been qualified as preservation of represented user hard intent, because an anchor omitted by the extraction stage cannot subsequently be protected by the governance layer.

We believe these revisions more clearly separate guarantees by construction in the abstract specification, empirically evaluated behavior of the implementation, and assumptions on which both depend.

Comment 9 — Definition of mission-consequence compatibility

Reviewer’s comment: The consequence signature, distance function, threshold, weighting coefficients, normalization, sensitivity, and examples should be specified. A single global threshold may not be sufficient across heterogeneous mission dimensions.

Response:
We thank the Reviewer for this comment. We agree that the previous formulation was too compact for reproducibility and that a single averaged threshold could conceal a large change in a safety-critical dimension.

Section 3.4 has therefore been expanded substantially. The mission-consequence signature is now explicitly defined over ten components: executability, hard goals, optional tasks, ordering, return-energy margin, time-window margin, airspace compliance, communication feasibility, payload satisfaction, and weather/observation-condition satisfaction. Table 8 defines the normalization and component-level distance used for each family. Boolean/numerical backend diagnostic quantities are normalized to [0,1], set-valued components use Jaccard distance, and return-energy and time-window margins use truncated normalized margins.

The aggregate consequence distance is now defined as a normalized weighted average. The evaluated setting uses uniform component weights. We additionally recognized the Reviewer’s concern that averaging alone can mask a critical change: the revised compatibility relation therefore includes a per-dimension veto for critical dimensions—hard goals, executability, airspace, and return-energy margin—in addition to the global distance threshold. Thus, a sufficiently large change in a critical component can make two candidates incompatible even when the average distance is small.

We also corrected the terminology. Because this relation is reflexive and symmetric but generally non-transitive, the revised text describes it as a tolerance/compatibility relation rather than a strict mathematical equivalence relation.

Four worked examples have been added, including two compatible contract pairs and two pairs that trigger clarification. In addition, Table 23 reports sensitivity to the global consequence threshold, and Table 24 reports sensitivity to the refinement-cost weights. The evaluated threshold and weight profile were fixed using a separate 20-task development subset rather than recalibrated on the 120-task evaluation set.

Finally, we now explicitly note that mission-specific or governance-defined dimension thresholds may be preferable in operational deployment and identify stakeholder-calibrated, dimension-specific criteria as future work.

Comment 10 — Specification and independence of the backend witness

Reviewer’s comment: The backend planning formalism, state/action models, search procedure, temporal resolution, energy and communication models, termination and infeasibility criteria should be specified. The use of the same backend in label construction and admission also raises a circularity concern.

Response:
We agree that the backend is sufficiently central to the admission decision that it must be described at a reproducible level.

Section 3.5 has been expanded to specify the task-level state representation, action vocabulary, witness-search procedure, return-to-home safety condition, airspace/geofence checks, communication model, payload and time-window conditions, weather checks, search depth, temporal discretization, and termination criteria. The backend state tracks local position, remaining energy, communication coverage, payload status, and elapsed mission time. Witness generation uses greedy forward search with bounded backtracking at a 1 s timestep, a depth limit of 20, and a 5 s per-contract wall-clock budget.

The return-to-home model accounts for horizontal distance, vertical displacement, landing energy, a pre-landing hover component, and a return-energy safety buffer. We have also made the model’s limitations explicit: it remains a task-level energy proxy and does not model continuous wind- or payload-dependent aerodynamic power consumption. Wind is handled through a task-level threshold, and payload affects capability availability rather than aerodynamic power.

Importantly, we have also revised the semantics of backend failure. The backend now exposes WITNESS_FOUND, CONCLUSIVE_INFEASIBLE, and INCONCLUSIVE. Failure to find a plan within a bounded search does not by itself justify REJECT; an inconclusive result produces CLARIFY. Only a conclusive incompatibility with the declared finite backend model can support REJECT.

We fully agree with the Reviewer that using the same L1 backend in benchmark construction and witness generation creates a shared-dependency concern. This is now acknowledged explicitly. In addition to a conservative L2 parameter cross-check, we therefore implemented an independent task-level Backend-B from the written constraint definitions, using independently written code, different discretization choices, and independently implemented energy and communication models. On 60 sampled MACs, L1 and Backend-B agree on 56/60 decisions (93.3%; Cohen’s κ = 0.867). The remaining disagreements occur near energy, communication, and time-window thresholds.

We present this as an independent consistency cross-check—not as absolute validation of backend correctness—and retain backend fidelity as an explicit assumption of the framework.

Comment 11 — Predominantly synthetic validation and practical UAV integration

Reviewer’s comment: Stronger integration with an actual UAV mission-management environment, such as PX4/ArduPilot, HIL, real telemetry, operator-in-the-loop testing, or a larger AirSim campaign, would strengthen operational relevance. A practical mapping to standard mission artifacts would also be valuable.

Response:
We agree with the Reviewer that an autopilot-level or operational validation campaign would provide substantially stronger evidence than task-level simulation.

We have strengthened the simulation and interface description in the present revision, but we would like to be precise about what has and has not been established. A full PX4/ArduPilot integration, HIL campaign, real-flight telemetry study, or operational user trial was not performed in this revision. Such experiments require a separate operational validation effort and would extend beyond the current pre-execution admission boundary. We have therefore chosen not to imply that the present results constitute such validation.

Instead, Section 4.4 now provides a dedicated AirSim configuration table and three end-to-end cases corresponding to backend non-closure, operator-authorized refinement, and direct admission. The examples expose the task-level mission artifacts more concretely, including waypoints, restricted-region margins, return-energy estimates, an authorized bypass route, and final witness outcomes. The AirSim parameters, vehicle model, speed, energy representation, communication representation, and modeled failure conditions are now explicitly documented.

At the same time, we now state that these cases provide qualitative end-to-end execution traceability only and do not extend the statistical conclusions of the benchmark. The Discussion explicitly notes the absence of realistic wind fields, GNSS degradation, sensor errors, and flight-control disturbances.

Finally, autopilot-level mission-artifact integration, including interfaces such as PX4/MAVLink, is now identified explicitly as future work. We believe this more conservative presentation better reflects the evidence currently available.

Comment 12 — Operationalization of authority and mutability

Reviewer’s comment: Authority and mutability require a more concrete policy example, including authority levels, permitted modifications, immutable constraints, conflict resolution, overrides, and distinctions between user requirements, regulatory constraints, defaults, contextual facts, and LLM hypotheses.

Response:
We agree and have substantially expanded this portion of the method.

The revised manuscript now defines five source categories and a corresponding authority/mutability structure. Tables 3, 6, and 7 specify clause source categories, the authority hierarchy, mutability classes, and allowed transformations. The governance policy now explicitly defines how conflicts are resolved: a higher-authority source prevails; at the same authority level, the more recent source prevails, followed by a governance-defined tie-breaking rule when required. Regulatory constraints are non-overridable. An operator may override eligible system defaults, whereas an emergency override is represented as a temporary, context-dependent authorization affecting specified non-regulatory constraints rather than as an unrestricted standing authority.

We also added a complete worked powerline-inspection example. In this example, an applicable regulatory altitude constraint is immutable; the operator’s inspection objective is protected; the operational return-reserve default may be increased by an authorized operator but not reduced below a regulatory minimum; an unavailable thermal-camera clause introduced by the LLM is revocable; and an LLM proposal to reduce the return reserve below the protected minimum is rejected because it fails both mutability and authority checks.

Evidence recency and contradictory trusted facts are now also handled explicitly. Weather, communication coverage, and airspace evidence are checked against context-defined validity windows, while battery and payload state are treated as current context fields. Conflicting trusted sources are resolved according to authority and recency.

We acknowledge that a complete jurisdiction-specific organizational authorization model—including every possible role, delegation, and temporal permission regime—would be deployment-specific. The revised policy model is therefore intended as an executable example of the framework rather than a universal UAV authorization standard.

Comment 13 — Case-level analysis of the residual errors

Reviewer’s comment: The eight residual errors should be analyzed individually to identify the failing stage and distinguish system limitations from genuine annotation ambiguity.

Response:
We agree. Aggregate error percentages alone were insufficient to determine where the remaining failures originate.

Table 15 now reports all eight residual errors individually, including an abbreviated form of the mission instruction, risk category, adjudicated reference label, EAMSR decision, failed processing stage, and root cause. Five errors are CLARIFY→REJECT over-strict decisions, primarily associated with conditional/adaptive language that the untrusted-semantic-isolation or mission-consequence mechanism interprets too conservatively. Three errors are REJECT→CLARIFY decisions, including one authority/mutability error and two cases in which untrusted semantic increments were insufficiently distinguished from remediable ambiguity.

This analysis shows that all residual errors remain on the CLARIFY–REJECT boundary; none crosses the ADMIT/non-ADMIT boundary. It also localizes the errors to USI screening, mission-consequence screening, and authority/mutability handling rather than treating them as undifferentiated classification mistakes.

To keep the manuscript readable, Table 15 contains the case-level summary rather than reproducing every complete contract trace. The full records—including original instructions, trusted context, anchors, generated clauses, proof-obligation outcomes, audit information, and adjudication records—are included in the permanent reproducibility package.

We have also connected this error topology to the human-labeling analysis: the CLARIFY/REJECT boundary is the same boundary on which inter-annotator agreement is lowest. We therefore discuss the residual errors as a combination of implementation limitations and genuine remediability ambiguity rather than attributing all discrepancies solely to model failure.

Comment 14 — Moderation and conditioning of the claims

Reviewer’s comment: Statements such as “eliminates unwarranted ADMIT decisions” and broad claims of guarantees should be moderated and conditioned on the controlled benchmark and validity of the trusted context, governance policy, ontology, and backend model.

Response:
We agree and have revised the wording throughout the Abstract, Introduction, Results, Discussion, and Conclusion.

The revised Abstract now reports that no unwarranted ADMIT decisions were observed among 78 non-admissible EAMSR-Bench cases (0/78) and provides the 95% Clopper–Pearson upper bound of 3.8%. We no longer interpret the observed 0.0% as evidence that unwarranted admission has been eliminated in general. The same conservative language is used for the independently authored external set.

Similarly, the formal-property statements are now explicitly conditional on the assumptions summarized in Table 10, including anchor representation, evidence fidelity, governance fidelity, backend fidelity, and witness semantics. We no longer present specification-level properties as certification of the complete operational system.

The Discussion now states explicitly that the benchmark, governance rules, labels, backend, and framework were co-developed; that the independently authored external set remains small and uses the same admission ontology; that the AirSim cases are task-level rather than flight-safety validation; and that the experiments establish controlled-benchmark performance and first-level external-authoring robustness rather than general operational assurance.

Finally, the manuscript consistently distinguishes mission admissibility and task-level executability from operational safety and flight safety. The concluding claim is now limited to pre-execution mission admission under the stated assumptions.

Comment 15 — Reproducibility package and permanent archiving

Reviewer’s comment: The reproducibility package should include all benchmark, model, annotation, prompt, raw-output, statistical, and configuration artifacts and should be permanently archived with a DOI.

Response:
We agree and have strengthened the reproducibility package accordingly.

The Data Availability Statement now identifies both the public development repository and a tagged archival release corresponding to the revised manuscript. The archived release is version v2.0.0-r1 (commit f42cc53) and has been deposited in Mendeley Data with the permanent DOI 10.17632/br4jbgyh6c.1. Although the Reviewer mentioned Zenodo as an example, we used Mendeley Data to achieve the same objective of a versioned, DOI-addressable permanent research-data record.

The archived release includes the benchmark instructions, trusted contexts, governance policies, protected invariants, evidence sources, backend models, human labels, annotation guidelines, exact prompts, model identifiers, decoding parameters, raw outputs, random seeds, repeated-run results, error cases, independent external test set, Deterministic-MAC implementation, Backend-B cross-check artifacts, AirSim configurations, and scripts used to generate the reported tables and figures.

This information is now stated explicitly in the Data Availability Statement.

Editorial, notation, presentation, and language comments

Reviewer’s comment: The manuscript also requires several editorial corrections concerning Algorithm 1, notation, terminology, evidence categories, ontology description, benchmark composition, metric denominators, figures, decision consistency, scenario perturbations, runtime reporting, AirSim parameters, safety terminology, English expression, citations, and metadata.

Response:
We thank the Reviewer for this detailed editorial review. We have systematically revised these issues throughout the manuscript.

Algorithm 1 has been completed so that candidate generation, parsing, duplicate handling, governance checking, consequence screening, bounded refinement, backend closure, clarification/rejection handling, and audit output are represented without the previous incomplete step. The notation has been disambiguated: candidate clause sets use (K^{(b)}), whereas the backend conflict core is denoted by (Q), and the audit-trail expression has been corrected. We now use ADMIT, CLARIFY, and REJECT consistently as the decision terminology and use unwarranted admission rate (UAR) rather than terminology that could be interpreted as a direct flight-safety metric.

Tables have been added for clause source categories, clause families, authority levels, mutability modes, and trust states. The evaluated UAV ontology is now described quantitatively as comprising eight clause families, 42 normalized predicate templates, 15 relation types, six entity types, and 28 deterministic ontology rules. Evidence semantics have also been clarified: Y denotes support from the resolved admissible fresh source, N denotes contradiction or an invalidated source, and U denotes insufficient evidence. Evidence freshness and conflict resolution are now described explicitly.

The benchmark section now introduces T1 (“normal-risk”) before its use, explains all six risk categories, and reports both the risk-type distribution and the ADMIT/CLARIFY/REJECT distribution for every mission scenario. Table notes now identify metric denominators and the handling of non-applicable cases where relevant.

The graphical presentation has also been reorganized. The framework figure is now presented as a clearer two-panel architecture and admission-decision flow, while the overall comparison results are shown separately. The former switching between UAR and (100-\mathrm{UAR}) has been removed; Figure 3 now reports backend-witness margins directly. Decision consistency is formally defined in Eq. (55), and the scenario-perturbation procedure now specifies the number of tasks and the generation distributions for target layouts, no-fly-zone layouts, communication maps, time windows, and objective augmentation.

Runtime is now reported with variability: the full EAMSR configuration requires 8.7 ± 2.3 s on average in the reported experiment, including external LLM API latency, which accounts for approximately 60% of the measured wall-clock time. The Discussion now distinguishes the likely acceptability of this latency for routine preflight admission from its potential limitation in time-critical emergency missions. AirSim parameters are reported in a dedicated table, including simulator version, scene, vehicle model, flight altitude, cruise speed, battery capacity, reserve threshold, energy representation, communication model, disturbance assumptions, and failure conditions.

Finally, we revised the manuscript for terminology consistency and reduced repetitive or overly dense passages, particularly in the Introduction, Related Work, and Discussion. Citation ranges, repository information, and manuscript metadata were also rechecked. Throughout the revised text, we now distinguish task feasibility, task-level executability, mission admissibility, operational safety, and flight safety more carefully.

We sincerely thank the Reviewer again for the detailed and constructive comments. They led us to reconsider not only the presentation of EAMSR but also the validity boundaries of the evaluation. We believe the revised manuscript is substantially more transparent regarding what is demonstrated by the present experiments, what follows by construction under explicit assumptions, and what remains to be established through future operational validation.

Author Response File: Author Response.pdf

Reviewer 2 Report

Comments and Suggestions for Authors

This paper presents an evidence-carrying Mission Admission Contract (MAC) framework for missions generated by means of operator interaction with LLMs. The proposed scheme is different from other state-of-the-art solutions, such as safety gates, pre-execution safety checks. The authors present MAC logical framework in detail, which seamlessly connect to the metrics defined for evaluation of the solution. The paper has interest to the community, but can be improved:

1) Eq. (30): the lambdas are not formally defined in the text.

2) Page 12: candidate-generation budget B_c and refinement budget B_r are not formally defined.

3) Page 13, §4.1.4: The authors employ different LLMs. Explanation should be given why the choice of LLM, or the use of different LLMs for different aspects of evaluation are not expected to have implication in the results.

4) Page 13, §4.1.4: "The total budget for candidate generation and refinement is set to 8. The 36
mission-consequence distance threshold is set to 0.15. Each task is repeated five times." These magic numbers should be justified. Are they empirical? Are five times repetitions enough?

5) Page 16: "The two configurations achieve similar feasibility-restoration rates, but their refinement boundaries differ." The two configurations Greedy Relaxation and EAMSR should be explicitly mentioned in the sentence.

 

Some editing comments:

1) In Algorithm 1, line 9 is empty, and it is the only empty line.

2) Section 4.1 should have a small paragraph of introductory text explaining the rationale. Check for other empty sections.

3) Tables and figures must all be referred in the text.

 

 

 

Author Response

We sincerely thank the Reviewer for the careful reading of our manuscript and for the constructive and specific comments. We greatly appreciate the Reviewer’s recognition of the potential interest of the proposed evidence-carrying Mission Admission Contract framework, as well as the suggestions concerning the mathematical definitions, implementation rationale, parameter settings, experimental design, and manuscript presentation.

We have carefully revised the manuscript in response to all of these comments. In particular, we have clarified the definitions and roles of the refinement-cost coefficients and the candidate-generation and refinement budgets; expanded the explanation of the LLM-backbone selection and its role in the evaluation; documented the empirical basis of the principal implementation settings; added sensitivity and repeated-run analyses; clarified the comparison between Greedy Relaxation and EAMSR; revised and completed Algorithm 1; added introductory text to Section 4.1; and systematically checked the references to tables and figures throughout the manuscript.

We are grateful for these comments, which helped us improve the clarity, reproducibility, and presentation of the revised manuscript. Our point-by-point responses are provided below.

Comment 1

Reviewer’s comment:
“Eq. (30): the lambdas are not formally defined in the text.”

Response:
We thank the Reviewer for identifying this lack of definition. We agree that, in the previous version, the weighting coefficients were introduced too abruptly and their individual roles were not explained with sufficient precision.

We have therefore revised the corresponding part of Section 3.4. Because additional material has been introduced during revision, the relevant refinement-cost formulation has been renumbered in the revised manuscript.

The revised text now explicitly explains that the four weighting coefficients correspond respectively to semantic displacement, the amount of editing introduced during refinement, loss of soft-preference satisfaction, and the diagnostic mission cost associated with the refined candidate. We also clarify that these terms are normalized before being combined, so that the different quantities can be compared on a common scale. The coefficients are constrained to form a normalized weighting profile, and the evaluated configuration uses equal weights for the four components.

We have further clarified that equal weighting is only the configuration used in the present evaluation and is not a requirement of the framework. Different weighting profiles may be adopted when a governance policy gives greater priority to, for example, minimizing semantic change or preserving operator preferences. To make this design choice more transparent, the revised manuscript also includes a sensitivity analysis over several alternative weighting profiles.

We appreciate this comment, as it prompted us to make both the mathematical meaning and the practical interpretation of the refinement-cost coefficients substantially clearer.

Comment 2

Reviewer’s comment:
“Page 12: candidate-generation budget B_c and refinement budget B_r are not formally defined.”

Response:
We agree with the Reviewer and have now clarified these two quantities explicitly in both the method description and Algorithm 1.

In the revised manuscript, the candidate-generation budget is defined as the maximum number of candidate-generation attempts permitted for a mission request. Each attempt produces a complete candidate clause set from the same validated mission anchors under the same decoding configuration. If two generated candidates are identical, the duplicate is discarded from the candidate set but the attempt is still recorded in the audit trail and consumes part of the generation budget. This makes the maximum number of generation attempts explicit and prevents unbounded candidate generation.

The refinement budget is correspondingly defined as the maximum number of authority-bounded refinement iterations that may be performed when direct admission cannot be completed. A refinement may be initiated after a remediable governance conflict, a mission-consequence incompatibility, or a remediable backend conflict. The revised Algorithm 1 now explicitly shows the refinement counter and the conditions under which another refinement is permitted.

For the evaluated configuration, the manuscript now states clearly that up to five candidate-generation attempts and up to three refinement attempts are allowed. The combined procedure is therefore bounded rather than open-ended.

We thank the Reviewer for this comment, which helped us make the computational and procedural meaning of these budgets much more explicit.

Comment 3

Reviewer’s comment:
“Page 13, §4.1.4: The authors employ different LLMs. Explanation should be given why the choice of LLM, or the use of different LLMs for different aspects of evaluation, are not expected to have implication in the results.”

Response:
We thank the Reviewer for this important comment. We agree that the role of the different LLM backbones was not sufficiently clear in the previous version.

We have revised Section 4.1.4 to clarify that different LLMs are not assigned to different functional stages of the EAMSR admission process. The principal experiments use one model as the primary candidate-generation backbone, while two additional models are introduced specifically to examine backbone sensitivity. Across these experiments, the task inputs, ontology, prompts, decoding settings, evidence base, governance policy, consequence-screening procedure, and backend task model are kept fixed.

We have also clarified an important architectural point: the LLM is not responsible for the final admission decision. Its role is limited to generating candidate clauses and proposing bounded refinements. Evidence verification, authority and mutability checking, isolation of unsupported semantic increments, mission-consequence screening, and backend-witness verification are performed outside the LLM. Thus, changing the backbone may affect how candidate clauses are expressed or how completely the model covers the intended semantics, but it does not remove or replace the external admission checks.

At the same time, we agree with the Reviewer that it would be inappropriate to assume that the choice of LLM has no effect on the results. We therefore evaluate this issue empirically rather than claiming complete model independence. The backbone-sensitivity experiment shows that the three evaluated models produce similar, but not identical, results. Three-class accuracy remains within a relatively narrow range, and decision consistency is high across repeated runs, although modest model-dependent variation remains.

Accordingly, the revised manuscript now makes a more restrained claim: the external governance and backend layers reduce dependence on the particular LLM used for candidate generation, but they do not make the complete system entirely backbone-independent.

We thank the Reviewer for encouraging us to clarify this distinction.

Comment 4

Reviewer’s comment:
“Page 13, §4.1.4: ‘The total budget for candidate generation and refinement is set to 8. The mission-consequence distance threshold is set to 0.15. Each task is repeated five times.’ These magic numbers should be justified. Are they empirical? Are five times repetitions enough?”

Response:
We fully agree with the Reviewer that these values should not appear as unexplained implementation constants. We have therefore substantially expanded Section 4.1.4 and added sensitivity and repeated-run analyses to explain how these settings were selected and how strongly the results depend on them.

First, the generation and refinement budgets were selected empirically using pilot experiments on a separate development subset of 20 tasks that was not included in the 120-task evaluation set. In these pilot experiments, successful cases typically converged within three to five candidate-generation attempts and one to two refinement steps. Based on these observations, we selected a maximum of five generation attempts and three refinement attempts to provide some additional headroom while keeping the procedure bounded and avoiding unnecessary computational and API cost. The revised manuscript also reports that increasing the candidate-generation budget beyond five produced only a small additional gain in governance-admissible candidate coverage on the development subset.

Second, the mission-consequence compatibility threshold was calibrated on the same development subset rather than on the 120-task evaluation set. Its purpose is to distinguish materially different mission outcomes from minor variations in non-critical preferences. To make the effect of this parameter transparent, we added a dedicated sensitivity analysis across a range of more conservative and more permissive threshold values. The results show the expected trade-off: smaller thresholds increase clarification and make the system more conservative, while larger thresholds eventually permit some non-admissible cases to cross the admission boundary. The evaluated value was therefore retained as a conservative operating point selected independently of the final evaluation set.

Third, we agree with the Reviewer that five repetitions should not be interpreted as an exhaustive statistical characterization of all possible serving-side variability. We have therefore revised the wording accordingly. The purpose of the five runs is limited to characterizing observed run-to-run stability under the available API and computational budget. The revised manuscript reports the predefined seeds, the mean and sample standard deviation across the runs, and a decision-consistency measure for each backbone. Across the evaluated models, the observed variability remains relatively small, with the largest sample standard deviation in three-class accuracy being 2.3 percentage points.

The main experimental result is not selected from the best of the five runs and is not obtained through majority voting. The principal result uses the predefined first run, while the five-run analysis is reported separately as a robustness characterization.

Finally, the revised manuscript now makes clear that the principal configurable parameters were calibrated on the separate 20-task development subset rather than on the final 120-task evaluation set.

We sincerely thank the Reviewer for this comment. It prompted us to replace several previously unexplained settings with a much more transparent development-set, sensitivity, and repeated-run justification.

Comment 5

Reviewer’s comment:
“Page 16: ‘The two configurations achieve similar feasibility-restoration rates, but their refinement boundaries differ.’ The two configurations Greedy Relaxation and EAMSR should be explicitly mentioned in the sentence.”

Response:
We agree and thank the Reviewer for pointing out this ambiguity.

The sentence has been revised to explicitly identify the two configurations as Greedy Relaxation and EAMSR. The revised text now states that Greedy Relaxation and EAMSR achieve comparable feasibility-restoration rates, but their refinement boundaries differ.

The subsequent explanation has also been retained and clarified. Greedy Relaxation treats constraint relaxation primarily as an optimization procedure for restoring feasibility, whereas EAMSR accepts only refinements that preserve protected constraints and explicitly anchored hard clauses and that remain within the predefined mission-consequence compatibility conditions.

This revision removes the ambiguous pronoun reference and makes the intended comparison immediately clear.

Editing Comments

Editing Comment 1

Reviewer’s comment:
“In Algorithm 1, line 9 is empty, and it is the only empty line.”

Response:
We thank the Reviewer for noticing this formatting inconsistency.

Algorithm 1 has been revised and reformatted, and the previously empty line has been removed. We also took this opportunity to re-examine the complete pseudocode and ensure that candidate generation, response parsing, duplicate detection, governance checking, consequence screening, bounded refinement, backend-witness evaluation, clarification or rejection handling, and final audit-trail assembly are represented consistently.

The line numbering and indentation of Algorithm 1 have also been rechecked throughout the revised manuscript. The current version no longer contains the isolated empty line identified by the Reviewer.

We appreciate the Reviewer’s careful attention to this presentation detail.

Editing Comment 2

Reviewer’s comment:
“Section 4.1 should have a small paragraph of introductory text explaining the rationale. Check for other empty sections.”

Response:
We agree with the Reviewer and have added a short introductory paragraph immediately after the heading of Section 4.1.

The new paragraph explains the rationale and organization of the evaluation protocol before the individual subsections are introduced. Specifically, it states that the evaluation consists of the benchmark and human reference labels, the controlled comparison configurations and their coverage boundaries, the metric definitions with uncertainty analysis, and the implementation settings.

We also reviewed the remaining manuscript structure for section headings that were followed immediately by subsections or otherwise lacked sufficient introductory context, and revised the relevant locations where appropriate.

We thank the Reviewer for this suggestion, which improves the continuity and readability of the experimental section.

Editing Comment 3

Reviewer’s comment:
“Tables and figures must all be referred in the text.”

Response:
We agree with the Reviewer.

We have systematically rechecked all tables and figures in the revised manuscript and ensured that each is explicitly introduced, cited, or discussed in the main text before or near its appearance. We also rechecked the numbering and internal cross-references after adding the new tables and figures introduced during revision.

In particular, the newly added tables concerning model configurations, parameter sensitivity, repeated-run statistics, deterministic-baseline comparisons, backend cross-checks, and AirSim settings are now explicitly referred to and interpreted in the corresponding sections.

We thank the Reviewer for drawing our attention to this presentation requirement.

We sincerely thank the Reviewer again for the careful and constructive comments. The suggestions concerning parameter definitions, the role of the LLM backbones, the justification of empirical settings, and the organization of the experimental section were particularly helpful.

In response, we have made a clearer distinction among quantities that are defined by the framework, quantities that were calibrated on a separate development subset, and quantities that are used only to characterize robustness and run-to-run variability. We have also revised the manuscript to avoid overstating either LLM-backbone independence or the statistical meaning of the five repeated runs.

We believe that these revisions have substantially improved the methodological clarity, transparency, reproducibility, and overall presentation of the manuscript, and we are grateful to the Reviewer for helping us strengthen the work.

Author Response File: Author Response.pdf

Reviewer 3 Report

Comments and Suggestions for Authors

A novel framework, named EAMSR, that reformulates natural-language UAV task submission as a Mission Admission Contract (MAC) rather than a simple text translation, is proposed in this paper. The main objective is to prevent unsupported assumptions, constraint relaxations, and changing task repairs from entering the UAV planning chain. EAMSR uses LLMs as an untrusted generator of candidate clauses, which are subsequently passed through a rigorous validation gate that checks evidence, authority, mutability, untrusted semantic increments, and backend executability. EAMSR shows 93.3% admission accuracy and eliminates unwarranted admissions over 120 tasks. The main strengths of this work lie in the strong formalization of the admission boundary and ablation studies demonstrating the necessity of its checks.

Comments regarding general concepts:

  • Isolating the admission process into a contract that carries evidence and defines explicit proof obligations is an original and valuable conceptualization.
  • The methodology is generally well founded, combining formal logic constraints with a dedicated backend task-level witness. However, some implementation details that are important for understanding the approach, such as the definition of the cost function weights and the anchors, are not described.
  • The experimental setup is robust but not large, utilizing a custom 120 task EAMSR-Bench across six UAV scenarios. The inclusion of AirSim simulations provides good qualitative mapping to real word operations.
  • The authors provide a GitHub repository containing the dataset, code, and validation artifacts, which theoretically ensures high reproducibility.

Specific comments, in order of importance:

  • Section 3.2, Equation 6 line 213: the mechanism for extracting requirement anchors from the natural-language instruction is undefined. Please specify how this is achieved.
  • Section 4.1.1: the generation process of the instructions is not clear. The authors should write how the natural-language instructions were generated.
  • Section 3.4, Equation 30 line 311: the cost weights lack explicit definitions.
  • GitHub repository (line 631 in the text): to ensure global accessibility and reproducibility, please ensure all documentation, variable names, and data files within the provided GitHub repository are translated into English.
  • GitHub repository (line 631 in the text): reviewing the file audit_p2.json in the repository inconsistencies seem to appear between the “instruction_preview” fields and the “anchors”. Please verify the dataset for translation artifacts or extraction errors to ensure full agreement with the claims on the paper.
  • Section 1, Figure 1: This figure provides limited technical value. Consider replacing it with an explicit system architecture block diagram or removing it and relying on the detailed workflow in Figure 2.

Author Response

We sincerely thank the Reviewer for the careful, constructive, and technically focused assessment of our manuscript. We greatly appreciate the Reviewer’s positive evaluation of the central conceptualization of EAMSR, particularly the formulation of mission admission as an evidence-carrying contract with explicit proof obligations. We also appreciate the comments identifying several implementation and reproducibility details that were insufficiently explained in the previous version.

In response, we have substantially revised the manuscript and the associated reproducibility materials. The principal revisions include a detailed specification of the requirement-anchor extraction procedure, a clearer description of how EAMSR-Bench instructions were constructed, explicit definitions and sensitivity analysis for the refinement-cost weights, a systematic review of the repository organization and audit artifacts, and a redesign of the framework figure to communicate the system architecture and admission flow more clearly.

We are grateful for these comments, which have helped us improve both the technical transparency and reproducibility of the work. Our point-by-point responses are provided below.

General Comment 1

Reviewer’s comment:
“Isolating the admission process into a contract that carries evidence and defines explicit proof obligations is an original and valuable conceptualization. The methodology is generally well founded, combining formal logic constraints with a dedicated backend task-level witness. However, some implementation details that are important for understanding the approach, such as the definition of the cost function weights and the anchors, are not described.”

Response:
We sincerely thank the Reviewer for this positive assessment and for identifying the need to make the implementation-level definitions more explicit.

We agree that the previous version placed greater emphasis on the formal admission logic than on some of the preprocessing and configuration details needed to reproduce the complete pipeline. We have therefore expanded Sections 3.2 and 3.4 substantially.

Section 3.2 now explains how requirement anchors are extracted from a natural-language instruction, including span detection, ontology-slot assignment, modality identification, and validation against the trusted context and deterministic ontology relations. We further clarify that anchor extraction is deterministic and is performed before LLM-based candidate generation.

Section 3.4 now defines the role of each refinement-cost component and explains how the corresponding weights are normalized and combined. The revised manuscript also states the weighting profile used in the evaluation and includes sensitivity results for alternative profiles.

These revisions are intended to make the transition from natural-language input to the formal admission process substantially more transparent.

General Comment 2

Reviewer’s comment:
“The experimental setup is robust but not large, utilizing a custom 120 task EAMSR-Bench across six UAV scenarios. The inclusion of AirSim simulations provides good qualitative mapping to real world operations.”

Response:
We thank the Reviewer for this balanced assessment. We agree that the original 120-task benchmark provides controlled coverage across multiple scenarios and admission-risk types, but does not by itself justify broad claims of operational generalization.

We have therefore revised the manuscript to characterize EAMSR-Bench explicitly as a controlled, internally constructed architecture-stress benchmark. We also added an independently authored external test set containing 48 additional natural-language mission requests across the same six mission scenarios. These requests were prepared independently of the EAMSR development team and were used exclusively for evaluation, without tuning prompts, thresholds, or model parameters on them.

At the same time, we have deliberately retained a conservative interpretation of the AirSim results. The AirSim cases are now described as qualitative end-to-end execution traceability rather than as statistical or real-flight validation. The Discussion also explicitly acknowledges that the present study does not cover real wind fields, sensor errors, GNSS degradation, flight-control disturbances, or full operational deployment.

We appreciate this comment because it encouraged us to distinguish more clearly among controlled benchmark evaluation, independently authored robustness testing, simulation-based qualitative validation, and future operational validation.

General Comment 3

Reviewer’s comment:
“The authors provide a GitHub repository containing the dataset, code, and validation artifacts, which theoretically ensures high reproducibility.”

Response:
We thank the Reviewer for recognizing the importance of the reproducibility materials. We agree that simply providing a repository is not sufficient unless the repository is internally consistent, understandable to external researchers, and tied clearly to the reported manuscript version.

We have therefore reviewed the reproducibility package together with the manuscript revision. In addition to the public development repository, the revised Data Availability Statement now identifies a tagged and permanently archived release corresponding to the revised manuscript. The release contains the benchmark data, external test set, human labels, annotation materials, exact prompts, model configurations, random seeds, repeated-run results, error records, backend cross-check materials, AirSim configurations, and supporting code and validation artifacts.

We have also responded specifically to the Reviewer’s repository-language and audit-file concerns below.

Specific Comments

Specific Comment 1

Reviewer’s comment:
“Section 3.2, Equation 6 line 213: the mechanism for extracting requirement anchors from the natural-language instruction is undefined. Please specify how this is achieved.”

Response:
We fully agree with the Reviewer. This was an important omission in the previous version because requirement anchors define the boundary between user-provided content and later semantic processing.

We have substantially expanded Section 3.2 to describe the anchor-extraction mechanism explicitly. In the revised implementation, anchor extraction is a deterministic preprocessing procedure and does not use the LLM.

The procedure consists of four stages. First, the instruction is segmented at linguistic boundaries such as clauses, conjunctions, and conditional constructions to identify candidate spans with independent mission meaning. Second, each candidate span is associated with an appropriate UAV ontology category, such as objective, spatial, temporal, payload, communication, energy, contingency, or preference information. Third, the modality of the span is identified so that mandatory, preferred, optional, and conditional content can be distinguished. Fourth, each candidate anchor is validated against the trusted context and deterministic ontology relations before it is allowed to enter candidate-clause generation.

We have also clarified the treatment of uncertain extraction cases. Spans with unresolved references, ambiguous negation scope, or unresolved cross-clause dependencies are retained as pending rather than being silently interpreted by the LLM. This design is intended to prevent a model-generated completion from subsequently being treated as if it had been explicitly provided by the operator. Anchor extraction is treated as trusted preprocessing in the current implementation, and the correctness of this preprocessing step is now stated explicitly as an assumption of the method.

The evaluated UAV ontology and its deterministic rule base are also described more concretely in the revised manuscript, with the complete ontology specification included in the reproducibility package.

We sincerely thank the Reviewer for this comment. The revised explanation makes clear where the trusted user representation comes from before LLM candidate generation begins.

Specific Comment 2

Reviewer’s comment:
“Section 4.1.1: the generation process of the instructions is not clear. The authors should write how the natural-language instructions were generated.”

Response:
We agree and have substantially expanded Section 4.1.1 to describe the benchmark-construction process.

The revised manuscript now states explicitly that the 120 EAMSR-Bench instructions are Chinese-language natural-language UAV mission requests authored by the research team using Chinese-language operator manuals, civil-aviation regulations, and mission-planning guidelines as domain references. They were not copied from operational flight logs and were not collected from external UAV operators.

The benchmark crosses six mission scenarios with six admission-risk categories. Each benchmark instance is constructed by combining scenario-specific mission content—such as objectives, target areas, and operating conditions—with a predefined risk-type template. Depending on the risk category, the template introduces a particular admission challenge through the instruction, evidence base, or trusted context; the normal-risk category leaves the admission obligations satisfiable. The manuscript now explains this construction procedure directly rather than presenting only the final benchmark composition.

We have also explicitly acknowledged that the benchmark and the framework were developed concurrently. Because the risk categories were motivated by the EAMSR proof-obligation structure, we now identify this as a potential source of circular validation and interpret EAMSR-Bench as a controlled internally designed benchmark rather than as independent operational evidence.

To partially address this limitation, the revised evaluation additionally includes 48 independently authored mission requests prepared by UAV mission-planning practitioners who were not involved in designing EAMSR, the risk taxonomy, the prompts, or the admission rules.

We thank the Reviewer for requesting this clarification. We believe the revised description now allows readers to understand both how the benchmark was constructed and the validity boundary associated with that construction process.

Specific Comment 3

Reviewer’s comment:
“Section 3.4, Equation 30 line 311: the cost weights lack explicit definitions.”

Response:
We agree with the Reviewer. The previous manuscript introduced the refinement objective without sufficiently explaining the individual weighting terms.

We have revised Section 3.4 so that each cost component is now described explicitly in the text. The four components account respectively for semantic displacement introduced by a refinement, the amount of editing required, loss of soft-preference satisfaction, and the estimated mission cost associated with the candidate.

Because these quantities have different native units, the revised manuscript also explains how they are normalized before combination. The corresponding coefficients are defined as non-negative normalized weights. The evaluated configuration uses equal weighting across the four components, while the framework permits alternative profiles when a governance policy assigns different priorities.

In addition, we have included a dedicated sensitivity analysis comparing several alternative weighting profiles. The results indicate that the admission decisions are less sensitive to these refinement-cost weights than to the mission-consequence compatibility threshold.

Because the manuscript has been substantially expanded, the equation numbering has changed in the revised version. We have therefore also rechecked all corresponding equation references.

We sincerely thank the Reviewer for identifying this missing definition.

Specific Comment 4

Reviewer’s comment:
“GitHub repository (line 631 in the text): to ensure global accessibility and reproducibility, please ensure all documentation, variable names, and data files within the provided GitHub repository are translated into English.”

Response:
We thank the Reviewer for this important reproducibility suggestion and agree that repository-level accessibility is essential for international readers.

We have reviewed the repository organization and standardized the user-facing documentation, schema descriptions, field descriptions, variable naming, and reproducibility instructions in English in the revised release.

We would, however, like to retain one scientifically important distinction. The original EAMSR-Bench mission instructions are Chinese because Chinese natural-language task submission is the experimental input studied in the current benchmark. Replacing these original instructions with English translations would change the source experimental data. We therefore preserve the original Chinese instructions as canonical source records while providing English renderings or explanatory fields where appropriate. The manuscript now also states explicitly that English mission instructions displayed in tables and figures are renderings of the Chinese originals rather than independent source instructions.

Thus, the revised repository is intended to make the dataset structure and experimental procedure understandable in English while preserving the original language data required for faithful reproduction of the experiment.

We appreciate the Reviewer’s suggestion, which has improved the accessibility of the reproducibility materials.

Specific Comment 5

Reviewer’s comment:
“GitHub repository (line 631 in the text): reviewing the file audit_p2.json in the repository inconsistencies seem to appear between the ‘instruction_preview’ fields and the ‘anchors’. Please verify the dataset for translation artifacts or extraction errors to ensure full agreement with the claims on the paper.”

Response:
We sincerely thank the Reviewer for examining the released audit artifacts at this level of detail. We consider this an important comment because the correspondence between the source instruction and the extracted anchors is central to the admission-traceability claim.

In response, we performed an additional record-level consistency review of the audit artifacts, including the correspondence among the canonical instruction, the instruction-preview field, and the extracted anchors. We also reviewed the relevant records for possible truncation, translation/rendering artifacts, and extraction inconsistencies.

In the revised release, the full source instruction is treated as the canonical record for anchor consistency checking; the preview field is retained only as a convenience field for inspection and is not treated as a substitute for the complete source instruction. Records identified during this review as having inconsistent preview or rendering information were corrected or regenerated against the canonical source record.

We have also strengthened the manuscript itself so that the anchor-extraction procedure is no longer implicit. The revised Section 3.2 now specifies that anchors are extracted deterministically from the source instruction, validated before clause generation, and that uncertain spans are retained as pending rather than silently completed.

In addition, the revised reproducibility package preserves the original benchmark instructions together with the relevant validation and audit artifacts, allowing the correspondence between source text, anchor records, and subsequent candidate clauses to be checked directly. The archived release information and included artifacts are now stated explicitly in the Data Availability Statement.

We are particularly grateful for this comment because it led us to recheck the audit data independently of the manuscript-level evaluation and to make the distinction between canonical source fields and convenience/display fields clearer.

Specific Comment 6

Reviewer’s comment:
“Section 1, Figure 1: This figure provides limited technical value. Consider replacing it with an explicitly system architecture block diagram or removing it and relying on the detailed workflow in Figure 2.”

Response:
We agree with the Reviewer that the previous Figure 1 did not communicate enough technical information relative to the space it occupied.

Rather than retaining the previous conceptual illustration, we have substantially redesigned Figure 1 as a technical overview of the EAMSR architecture and admission process.

The revised figure now contains two complementary views. The architecture view shows the progression from the operator mission request through anchor extraction and untrusted candidate construction to the evidence-carrying MAC, governance verification, mission-consequence screening, backend witness, and audit trail. The decision-flow view explicitly shows the transitions among candidate MAC generation, governance gating, consequence screening, backend closure, authority-bounded refinement, and the three possible final outcomes: ADMIT, CLARIFY, and REJECT.

The corresponding caption has also been rewritten to explain that final admission requires governance admissibility, consequence consistency, and backend-witness closure rather than only a successful LLM interpretation.

We believe this redesigned figure now serves the technical role suggested by the Reviewer and provides a concise architecture-level reference for the more detailed formal descriptions that follow.

We thank the Reviewer for this valuable presentation suggestion.

 

Closing Response

We sincerely thank the Reviewer again for the careful and constructive review. The comments were particularly valuable in identifying places where the conceptual framework was more complete than its implementation-level explanation.

In response, we have made the requirement-anchor extraction process explicit, documented the benchmark instruction-generation procedure, clarified the refinement-cost parameters, strengthened the repository documentation and audit consistency, and redesigned the framework figure to provide greater technical value. We have also further clarified the limitations of the internally constructed benchmark and supplemented it with independently authored evaluation data.

We believe that these revisions have substantially improved the transparency, reproducibility, and technical presentation of the manuscript, and we are grateful to the Reviewer for helping us strengthen the work.

Author Response File: Author Response.pdf

Reviewer 4 Report

Comments and Suggestions for Authors

The manuscript presents a solid and high-quality contribution in terms of methodology, results, and scientific soundness. The work is well-written and the level of English is fine.

However, to improve the final version, I suggest addressing the following minor points:

1.- Introduction: slightly expand the background context and include recent relevant references to complement the theoretical framework.

2.- Figures and tables: improve the visual quality, formatting, and clarity of the captions to facilitate interpretation.

Author Response

We sincerely thank the Reviewer for the positive and encouraging assessment of our manuscript. We greatly appreciate the Reviewer’s recognition of the methodological contribution, scientific soundness, quality of the results, and overall clarity of the presentation. We are also grateful for the two constructive suggestions concerning the background and references in the Introduction and the presentation quality of the figures and tables.

We have carefully revised the manuscript in response to both comments. In particular, we have strengthened the contextual positioning of the work by expanding the discussion of recent developments in natural-language UAV interaction, language-to-formal-specification methods, LLM-assisted robotic planning, specification repair, and pre-execution verification. We have also comprehensively reviewed the figures and tables, improving their organization, readability, terminology, captions, notes, and cross-references.

Our point-by-point responses are provided below.

Comment 1

Reviewer’s comment:
“Introduction: slightly expand the background context and include recent relevant references to complement the theoretical framework.”

Response:
We sincerely thank the Reviewer for this helpful suggestion. We agree that the original version could provide a broader and more up-to-date context for motivating the proposed mission-admission perspective.

We have therefore expanded the background discussion in the Introduction and the closely connected Related Work section. The revised manuscript now more clearly positions natural-language UAV task submission within the broader progression from language-to-specification translation, through planning and specification repair, to pre-execution verification and safety gating.

In particular, we now explain more explicitly that existing natural-language interfaces can transform operator instructions into structured task specifications, temporal-logic expressions, planning domains, waypoints, or constrained plans. However, structural correctness or planning feasibility alone does not establish that every resulting mission clause is supported by appropriate evidence, introduced by an authorized source, or preserved without an unintended change in mission consequences. This distinction provides the motivation for treating mission admission as a separate decision layer rather than as another text-translation or planning stage.

We have also incorporated and discussed recent representative studies to strengthen the connection with current developments in the field. These include recent work on interactive natural-language-to-temporal-logic translation and clarification, scalable LLM-based planning-domain generation for aerial robotics, LLM-enabled robotic planning, specification repair, formal verification, and safety guardrails. The revised discussion now places these studies alongside established work on natural-language UAV task specification, planner-based repair, policy-based verification, and contract-based refinement.

At the same time, we have been careful not simply to increase the number of citations. The revised text now organizes the literature according to the specific decision problem addressed by each family of methods: translation correctness, executability restoration, operational or pre-execution safety checking, and admission integrity. This organization also allows us to state the scope of our contribution more precisely.

To further clarify the relationship with prior work, the revised manuscript includes structured comparisons of closely related methods and explicitly explains that the principal contribution of EAMSR is not that every individual mechanism is unprecedented, but that evidence, authority, mutability, untrusted-semantic isolation, consequence compatibility, and backend executability are integrated into a single conjunctive admission process for natural-language UAV task submission.

We believe these revisions provide a stronger and more current background while also making the theoretical motivation and methodological boundary of EAMSR clearer. We thank the Reviewer for this valuable suggestion.

Comment 2

Reviewer’s comment:
“Figures and tables: improve the visual quality, formatting, and clarity of the captions to facilitate interpretation.”

Response:
We agree with the Reviewer and have comprehensively reviewed the figures and tables throughout the manuscript.

First, we revised the organization and visual presentation of the principal figures so that each figure communicates a more focused technical message. In particular, the framework overview has been redesigned to distinguish the system architecture from the admission-decision flow. The revised figure now shows more clearly how the operator request proceeds through anchor extraction, candidate construction, governance verification, mission-consequence screening, backend-witness checking, refinement, and the final ADMIT, CLARIFY, or REJECT decision.

Second, we revised the experimental figures to improve readability and reduce unnecessary visual complexity. Labels, axis descriptions, abbreviations, and annotations were rechecked for consistency. Where a figure contains multiple panels, the captions now explain the role of each panel and the interpretation of the reported quantities more explicitly.

Third, we revised the table formatting and table notes throughout the manuscript. Abbreviations and evaluation metrics are now defined more consistently, denominators and units are stated where necessary, and notes clarify the interpretation of quantities that could otherwise be ambiguous. Newly added tables also separate information that was previously embedded densely in the main text, including model configurations, benchmark composition, authority and mutability definitions, parameter sensitivity, repeated-run statistics, backend cross-checks, and AirSim settings.

We have also expanded the figure and table captions so that they are more self-contained. In particular, the revised captions now explain the relevant experimental setting, the meaning of key abbreviations or comparison groups, and, where appropriate, the limitations of the corresponding result. For example, controlled sensitivity analyses are explicitly distinguished from independent accuracy evaluations, and simulation examples are described as qualitative execution traceability rather than statistical validation.

Finally, we systematically checked that every figure and table is explicitly referred to and interpreted in the main text and that the numbering and cross-references remain consistent after the revisions.

We sincerely thank the Reviewer for this suggestion. We believe that the revised visual presentation and more informative captions substantially improve the accessibility and interpretability of the manuscript.

Closing Response

We sincerely thank the Reviewer again for the positive evaluation of our work and for these constructive suggestions. Although both comments were relatively minor, they prompted us to improve two important aspects of the manuscript: the broader positioning of the proposed framework within recent research and the clarity with which the technical and experimental results are communicated.

We believe that the expanded background and updated literature discussion, together with the revised figures, tables, captions, and formatting, have improved the completeness and readability of the manuscript. We are grateful to the Reviewer for helping us strengthen the final version.

Author Response File: Author Response.pdf

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

The authors have answered all my concerns. I recommend it for publishing. 

Back to TopTop