Next Article in Journal
Correction: Case, R.P.; Hupy, J.P. Methods for GIS-Driven Airspace Management: Integrating Unmanned Aircraft Systems (UASs), Advanced Air Mobility (AAM), and Crewed Aircraft in the NAS. Drones 2026, 10, 82
Previous Article in Journal
AQGTO: Adaptive Q-Learning-Guided Gorilla Troops Optimizer for 3D UAV Path Planning in Precision Agriculture
 
 
Article
Peer-Review Record

Formalizing the Implicit Mechanisms in UAV Energy Model Selection Through Decision Tree and Analytic Hierarchy Process

by Israel Kolaïgué Bayaola 1, Jean Louis Ebongué Kedieng Fendji 2,*, Blaise Omer Yenke 2, Marcellin Atemkeng 3 and Christiana Ibidun Obagbuwa 4
Reviewer 1: Anonymous
Reviewer 2:
Reviewer 3:
Reviewer 4: Anonymous
Submission received: 7 March 2026 / Revised: 1 May 2026 / Accepted: 3 May 2026 / Published: 8 May 2026

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

The manuscript addresses the lack of a standardized methodology for selecting energy-consumption models within the Unmanned Aerial Vehicle (UAV) research community. The authors propose a rigorous two-stage framework that effectively bridges qualitative feasibility filtering with quantitative multi-criteria prioritization through the Analytic Hierarchy Process. A significant strength of the work is the inductive derivation of the decision tree from a systematic literature review, which was subsequently validated against an independent 20% holdout set with a reported 100% predictive match.  The inclusion of an open-source Python-based GUI further enhances the paper's impact by providing a tangible tool for the community to improve methodological reproducibility. However, I have the following concerns:
1.    Fig. 3 is dense and hard to follow. A simpler schematic would enhance clarity, and illustrating the multi-stage hierarchy could better explain the proposed decision tree.
2.    The framework’s validation relies on replicating modeling choices from existing literature. Since these choices are acknowledged as largely ad-hoc, a 100% match demonstrates mimicry of current practices rather than proof of optimality. The authors should discuss how the framework can correct or improve upon suboptimal ad-hoc selections to ensure genuine predictive accuracy and methodological advancement.
3.    Also, the current validation is entirely retrospective based on published papers. To demonstrate broader applicability, the authors should consider a prospective case study in which the framework is applied to select a model for a new, unpublished mission.
4.    A sensitivity analysis of the values in matrices A^(k) should be conducted to ensure the accuracy of the predictions and to assess the impact of different parameter choices. This would strengthen the framework's robustness and clarify the influence of different assumptions on the results. 
5.    The results presented in Table 1 validate the match outcomes for the 20% holdout set; however, this procedure could be extended to incorporate additional energy models. I recommend that the authors discuss the principal challenges of such extensions and the practical implications for the framework's broader applicability.
6.    Minor concerns:
-    Table 1 should be reformatted to match the style of the other tables in the manuscript to ensure consistency and improve clarity.
-    The GitHub repository referenced in the manuscript is currently inaccessible. The authors should verify and update the link to ensure long-term availability. 

Author Response

We thank Reviewer 1 for the thorough and constructive assessment. We have addressed each concern in the revised manuscript, as detailed below.

Comment 1.1.  Fig. 3 is dense and hard to follow. A simpler schematic would enhance clarity, and illustrating the multi-stage hierarchy could better explain the proposed decision tree.

We acknowledge this concern. In the revised manuscript, Fig. 3 has been restructured to present the decision hierarchy in a more readable, layered format. The entry point, branching conditions, and terminal outputs are now visually distinguished through consistent color coding and spatial grouping by stage. Each stage is now labelled explicitly (Stage 1 through Stage 5) so that a reader can trace any decision path without simultaneously processing the full tree. We believe the revised figure substantially reduces cognitive load while preserving the full logical content of the framework.

Comment 1.2. The 100% match demonstrates mimicry of current practices rather than proof of optimality. The authors should discuss how the framework can correct or improve upon suboptimal ad-hoc selections.

This is a legitimate and important epistemological distinction. We thank the reviewer for raising it explicitly.

We fully concede that the framework was designed to formalize the implicit logic already present in the literature not to contradict it. The 100% predictive alignment therefore confirms that the framework accurately captures the dominant decision-making patterns in the field, which is the stated contribution. The claim is descriptive fidelity, not prescriptive optimality.

However, the reviewer is correct that the paper did not sufficiently discuss how the framework adds normative value beyond matching current practice. We have added a paragraph in Section 5.1 of the revised manuscript addressing this directly. Specifically, the framework adds value in three ways that go beyond mimicry: (i) it makes previously implicit constraints explicit and auditable, which enables identification of cases where a researcher's actual constraints do not support their chosen model, a situation that mere replication of literature habits obscures; (ii) the AHP stage introduces a principled ranking among feasible alternatives, which is absent from ad-hoc selection; and (iii) the fallback flexibility mechanism provides a structured response to experimental failures, which has no equivalent in current practice. We have moderated the language around "predictive accuracy" in the highlights and abstract to more precisely reflect that the framework formalizes observed patterns rather than certifying them as optimal.

Comment 1.3. A prospective case study, applying the framework to a new, unpublished mission, would demonstrate broader applicability.

We agree in principle that a prospective case study would strengthen the demonstration of applicability. In the revised manuscript, we have added a brief illustrative prospective application in Section 5.1, in which the framework is applied to a notional but realistic mission scenario. Specifically, a logistics UAV deployment in a resource-constrained environment without wind tunnel access or large telemetry datasets. This example walks through both the Boolean feasibility filtering and the AHP weighting under a specific constraint profile, demonstrating how the framework would guide a researcher facing a genuinely novel decision rather than replicating a published choice.

We note, however, that a fully experimental prospective study (i.e., applying the framework, selecting a model, collecting data, and measuring accuracy) lies beyond the scope of the current paper, which is positioned as a methodological framework contribution. We have acknowledged this explicitly in the limitations section and proposed it as a high-priority direction for future work.

Comment 1.4. A sensitivity analysis of the matrices A^(k) should be conducted.

Response: This concern is directly addressed in the revised manuscript. Section 4.7 now presents a formal sensitivity analysis evaluating the impact of contextual criteria weight perturbations on the global priority score vector S. We demonstrate that under equal weighting (w_base = [0.25, 0.25, 0.25, 0.25]), O1 dominates due to its aggregate superiority across accuracy, interpretability, and customization. When the weight vector is shifted to heavily favor development cost (w_cost = [0.10, 0.10, 0.70, 0.10]), the ranking inverts, with O5 and O4 rising to the top — a result that is mathematically consistent with their known resource efficiency. This analysis confirms that the AHP stage is not statically biased toward any single paradigm and responds predictably to changes in researcher priorities. The analysis also confirms that the framework's recommendations remain consistent for modest perturbations of up to ±0.10 in any individual weight component, establishing a reasonable robustness threshold.

Comment 1.5. The validation could be extended to incorporate additional energy models, and the principal challenges of such extensions should be discussed.

Response: In the revised manuscript, the validation corpus has been substantially expanded, from the original 7-paper holdout to an independent holdout set of 24 studies (Section 3.6 and Section 4.9). This expansion activates a broader range of terminal nodes, including O3 (transparent regression), and provides a more representative test of the framework's generalizability across diverse modeling contexts. We have also added a dedicated discussion in Section 5.3 addressing the practical challenges of further extending the validation corpus. These include: the difficulty of systematically inferring feasibility indicators (F1–F6) from papers that do not report them explicitly; the risk of O5 (reuse of black-box models) being underrepresented in the recent literature due to platform-specificity of empirical models; and the challenge of incorporating genuinely hybrid paradigms, such as PINNs, which do not fit cleanly into the current binary white-box/black-box taxonomy.

Comment 1.6 (Minor). Table 1 formatting inconsistency; GitHub link inaccessible.

Response: The formatting of Table 1 has been revised to align with the typographic conventions used throughout the manuscript. Regarding the GitHub repository, we have verified that the link (https://github.com/Bayaola/UAV-Energy-Model-Selection.git) is accessible and that the repository contains the complete Python GUI source code. We have additionally included a direct reference to the repository in the Data Availability Statement to ensure long-term discoverability.

Reviewer 2 Report

Comments and Suggestions for Authors

This paper makes a valuable contribution by formalizing UAV energy model selection through a principled two-stage framework. The decision tree derivation is sound, and the 100% validation match on the holdout set is compelling. Strengths include practical GUI implementation and clear presentation.

For improvement: (1) Add guidance for practitioners facing partial resource constraints—binary feasibility may oversimplify real scenarios. (2) Include AHP sensitivity analysis to demonstrate robustness. (3) Slightly moderate generalization claims given the modest holdout sample size. (4) Expand discussion of emerging paradigms like PINNs

Author Response

We thank Reviewer 2 for the positive overall assessment and for four targeted recommendations, each of which has been addressed in the revised manuscript.

Comment 2.1. Add guidance for practitioners facing partial resource constraints, as binary feasibility may oversimplify real scenarios.

Response: This concern is addressed directly in Section 5.2 of the revised manuscript, which is a new section titled "Addressing Partial Resource Constraints and Continuous Conditions." We explain that the AHP stage functions as a continuous absorption mechanism for Stage 1 ambiguities: a researcher who has only partial access to experimental infrastructure can conditionally set the relevant Boolean flag (e.g., F4 = 1 for limited wind tunnel access) while penalizing the associated AHP criterion (e.g., reducing the weight on Predictive Accuracy, C1, or increasing Development Cost, C3). This causes the global priority score for the resource-intensive option to be mathematically depressed without triggering a hard rejection. We also note that the fallback flexibility vector provides a natural mitigation path when a partially feasible option fails during execution. We further identify this limitation as a primary motivation for the planned integration of Fuzzy AHP in future work.

Comment 2.2. Include AHP sensitivity analysis to demonstrate robustness.

Response: This recommendation has been addressed in full through the addition of Section 4.7, which presents a formal sensitivity analysis evaluating the impact of contextual criteria weight perturbations on the global priority score vector S.

The analysis compares a baseline equal-weight scenario (w_base = [0.25, 0.25, 0.25, 0.25]), where O1 (novel white-box derivation) dominates due to its aggregate superiority in accuracy and interpretability, against extreme perturbations. For instance, when the weight vector is shifted to heavily favor development cost (w_cost = [0.10, 0.10, 0.70, 0.10]), the ranking inverts as expected, with O5 and O4 (reused or data-driven models) rising to the top.

These results confirm that the AHP stage is not statically biased toward any single paradigm and responds predictably to changes in researcher priorities. Furthermore, the analysis establishes a robustness threshold, demonstrating that the framework's recommendations remains stable for modest perturbations of up to pm 0.10 in any individual weight component. This confirms the structural stability of the framework for realistic weight profiles.

Comment 2.3. Slightly moderate generalization claims given the modest holdout sample size.

Response: We fully agree. The revised manuscript consistently uses measured language when characterizing the validation results. Phrases such as "rigorously validated," "100% predictive match proves its reliability," and similar formulations have been revised throughout. The abstract and conclusion now state that the framework "demonstrates a high degree of consistency between the framework's recommendations and the methodologies employed in the literature," which is an accurate characterization of what the expanded 24-study holdout demonstrates. We have also added an explicit caveat in Section 3.6.2 acknowledging that while the validation corpus is now substantially larger, the framework's generalizability to domains outside the surveyed literature (particularly emerging or hybrid UAV platforms) should be assessed through future prospective application.

Comment 2.4. Expand discussion of emerging paradigms like PINNs.

Response: The discussion of Physics-Informed Neural Networks has been expanded in Section 5.3.2. We now provide a substantive explanation of why PINNs represent a methodologically significant gap in the current framework: they occupy an intermediate position between white-box and black-box paradigms by embedding differential equation constraints into the neural network training objective, rendering them neither purely physics-derived nor purely data-driven. We discuss the specific challenges of incorporating PINNs as a terminal node. Particularly the difficulty of assigning Boolean feasibility conditions (F4 vs. F5 vs. F6) and the absence of consensus on how to assess their interpretability relative to classical white-box models. We identify these as the primary research questions to be resolved before PINNs can be formally integrated into the framework's decision logic.

Reviewer 3 Report

Comments and Suggestions for Authors

Having read and reviewed the submitted work in its entirety, I suggest the following improvements:

  1. In Section I, Introduction, I suggest clarifying more precisely which elements constitute the main novelty of this work compared to previous work.
  2. In Section I, Introduction, I suggest explaining the rationale for selecting 80% of the literature for training and 20% for validation.
  3. In section 2.4.1, I suggest explaining the main differences between white-box and black-box models.
  4. In section 3.4, I suggest defining and justifying the five outputs (O1–O5) and indicating whether these fully cover the space of real solutions.
  5. I suggest justifying in the work the reason for choosing to model variables such as infrastructure or data with binary values ​​instead of continuous variables.
  6. In section 3.5, I suggest considering the inclusion of more studies, as validating based on only 7 studies does not always allow for statistically significant conclusions.
  7. I suggest that section 6 be titled "Conclusions and Future Work."
  8. In the "Conclusions and Future Work" section, it is suggested that the conclusions be highlighted quantitatively according to the results obtained.

Author Response

We thank Reviewer 3 for a detailed and methodologically focused set of recommendations. Each has been incorporated into the revised manuscript.

Comment 3.1. Clarify more precisely which elements constitute the main novelty compared to previous work.

Response: The introduction of the revised manuscript has been restructured to more sharply distinguish the paper's contribution from adjacent work. The key novelty is the formalization of the implicit selection logic that already exists (undocumented) within the UAV energy modeling literature. Prior work either surveys modeling strategies or applies them; none has extracted and codified the decision-making structure underlying those choices, nor integrated it with a quantitative multi-criteria ranking mechanism. The introduction now explicitly contrasts this contribution with survey papers (which classify without prescribing) and with optimization studies (which assume a model without justifying the selection). The new title "Formalizing the Implicit Mechanisms in UAV Energy Model Selection Through Decision Tree and Analytic Hierarchy Process" was adopted precisely to reflect this positioning accurately.

Comment 3.2 Explain the rationale for the 80/20 training/validation split.

Response: Section 3.1 of the revised manuscript now includes an explicit justification. The 80/20 partitioning follows an established convention in machine learning and empirical research design for maximizing the training corpus while retaining a meaningful holdout for independent evaluation. In our context, the 23-study training corpus was designed to be large enough to capture the full diversity of operational scenarios (standard platforms, custom platforms, data-rich environments, infrastructure-rich environments), while the 24-study holdout (which in the revised version is substantially larger than the original 7-study set) provides a robust independent test. We also clarify in this section that the partition was performed chronologically (more recent papers prioritized for the holdout) to simulate a prospective evaluation scenario and to mitigate the risk of temporal leakage.

Comment 3.3. Explain the main differences between white-box and black-box models in Section 2.4.1.

Response: Section 2.4.1 has been revised to provide a more structured and didactically complete contrast between the two paradigms. The revised section now explicitly defines white-box models as those grounded in first-principles aerodynamic theory (momentum theory, blade element theory, propulsion mechanics), whose parameters have direct physical interpretations and whose structure is fixed by the governing physics. Black-box models are defined as those that map observable inputs to energy outputs through statistical learning, with internal parameters that do not correspond to physical quantities. The practical implications of this distinction (for interpretability, extrapolation, data requirements, and computational cost) are now laid out in a structured comparative manner before Table 1 summarizes them.

Comment 3.4. Define and justify the five outputs (O1–O5) and indicate whether they fully cover the space of real solutions.

Response: Section 3.5 of the revised manuscript now contains an explicit definition and justification for each terminal output, grounded in the training corpus analysis. The five outputs correspond to the five distinct modeling strategies empirically observed across the 23 training studies: novel white-box derivation (O1), reuse of established white-box models (O2), limited-data regression (O3), large-scale machine learning (O4), and reuse of existing empirical models (O5). We also address the coverage question directly: we acknowledge that this taxonomy does not formally exhaust the space of conceivable modeling approaches. Most notably, hybrid paradigms such as PINNs are not yet represented. We argue that the current five outputs cover the dominant strategies in contemporary UAV energy research (as evidenced by their activation across the 24-study holdout) and commit to extending the taxonomy as hybrid approaches gain prevalence.

Comment 3.5. Justify the use of binary values for infrastructure and data variables instead of continuous variables.

Response: This justification is now provided in Section 4.1 of the revised manuscript and elaborated in Section 5.2. The binary formulation was adopted for two reasons. First, it reflects the decision-theoretic reality that certain resources are either sufficient for a given modeling approach or they are not. A wind tunnel either permits aerodynamic parameter extraction at the required fidelity or it does not. Second, binary constraints allow the feasibility stage to function as a non-compensatory filter, which prevents a high score on one criterion from overriding a fundamental resource constraint. We acknowledge, however, that resource availability is frequently a matter of degree in practice, and we explain how the AHP stage can absorb this continuous nuance through criterion weighting without requiring a modification to the binary architecture. The planned Fuzzy AHP extension is identified as the formal resolution of this limitation.

Comment 3.6. Consider the inclusion of more studies, as validating based on only 7 studies does not always allow for statistically significant conclusions.

Response: This concern has been fully addressed. The revised manuscript expands the holdout validation set from 7 to 24 independent studies, drawn from recent literature published between 2023 and 2026. The expanded set covers a wider range of output categories (O1, O2, O3, O4) and a more diverse range of operational contexts, including swarm optimization, solar-powered UAVs, edge-based object detection, truck-drone collaborative delivery, and battery degradation prediction. The complete validation results are reported in Tables 4 and 5 of the revised manuscript.

Comment 3.7. Section 6 should be titled "Conclusions and Future Work."

Response: The section title has been updated to "Conclusions and Future Work" in the revised manuscript.

Comment 3.8. Conclusions should be highlighted quantitatively according to the results obtained.

Response: The conclusions section has been revised to include specific quantitative statements. These include: the training corpus size (23 studies), the holdout set size (24 studies), the distribution of framework outputs across the holdout (18 O2, 4 O4, 1 O1, 1 O3, 0 O5), the consistency ratio bounds achieved across all AHP matrices (CR < 0.1 in all cases), and the sensitivity analysis boundary conditions under which the top-ranked output changes.

Reviewer 4 Report

Comments and Suggestions for Authors

Major comments

- The core contribution of the paper is a decision framework for selecting UAV energy models, but the validation is too weak to support the strength of its claims. The framework is derived from the same literature corpus that is later used for validation, with only a holdout of 20 percent of seven articles, and the manuscript repeatedly emphasizes a "100 predictive match.” This is not convincing evidence of generalizability, especially given the very small sample and the fact that the categories and rules appear to be constructed manually from the same review process. The contribution therefore feels closer to a structured survey plus decision aid than a rigorously validated research framework.

- The methodological novelty is limited. The manuscript combines a literature-derived decision tree with a standard AHP ranking layer, but there is no strong theoretical development, no new modeling principle for UAV energy estimation itself, and no experimental comparison showing that this framework materially improves research results, model accuracy, time, or cost. For a journal like Drones, this may be considered too managerial or procedural unless supported by stronger technical evidence.

- The AHP stage is not adequately grounded and appears highly subjective. The pairwise matrices for precision, interpretability, cost, and adaptability are manually assigned using qualitative judgments, yet the manuscript does not explain who elicited these values, how expert bias was controlled, or whether alternative weight settings would change the ranking materially. Reporting acceptable consistency ratios is not enough; the real issue is the external validity of the judgments themselves.

- The “quantitative re-validation” is not fully independent of the framework design. The criteria elicitation protocol deterministically maps textual emphasis scores to Saaty intensities, but these emphasis scores themselves are still assigned by the author and potentially influenced by knowledge of the chosen methodology of the target paper. This raises a real risk of circular reasoning: the framework may be reproducing the authors’ interpretation of the papers rather than objectively predicting model choices.

- The feasibility layer is oversimplified for the UAV energy-modeling domain. Key factors are reduced to binary indicators such as whether infrastructure exists, whether telemetry can be collected, and whether a transferable model exists. In practice, these are continuous and nuanced conditions, as the authors themselves later acknowledge. This simplification weakens the practical usefulness of the framework and makes the final recommendations somewhat artificial.

- The paper does not evaluate whether the recommended model choices actually produce better downstream outcomes. Since the stated motivation is to guide energy modeling for optimization, planning, and operations, the study should ideally show at least one realistic case study where following the framework leads to better model selection, better prediction accuracy, or better mission-planning performance than an ad-hoc choice. Without that, the practical impact remains asserted rather than demonstrated.

- The manuscript is written in an overly promotional style, with repeated strong claims such as “rigorously validated,” “100percent predictive matches",” “highly reliable", and “mathematically substantiates operational validity.” This tone overstates what evidence can support and should be substantially moderated. The discussion and conclusion are currently read more as advocacy than as a critical scientific assessment.

- The fit to Drones is only moderate. The topic is related to UAV, but the work focuses primarily on advancing drone design, control, sensing, autonomy, experiments, or mission-level validation. It is instead a meta-level methodological selection tool based on literature classification. That can still be publishable, but only if the validation and practical demonstration are much stronger than they are now.

Minor comments

- The prose is often too verbose and repetitive. Many sections could be considerably tightened, especially the introduction, validation interpretation, and discussion.

- Some terminology is not fully standardized. The paper alternates between “black-box,” “empirical,” “ML,” “deep learning,” and “regression” in ways that sometimes blur category boundaries.

- The table presentation could be improved. The validation tables are useful, but several entries look crowded and typographically awkward in PDF extraction, reducing readability. 

- The paper would benefit from a clearer protocol for literature selection: search strings, databases, inclusion/exclusion criteria, and potential review bias are not sufficiently transparent from the current text.

- The GUI is a nice implementation detail, but at present it feels more like a software wrapper around the paper’s logic than evidence of scientific contribution. Its role should be secondary unless experimentally assessed.

- Several claims in the conclusions should be softened to match the actual evidence base, especially where broad statements are made from a very small holdout sample.

In its current form, I do not find the paper strong enough for acceptance in Drones. The idea is organized and potentially useful, but the validation is too limited, the AHP layer is too subjective, and the practical impact is not demonstrated convincingly. A publishable revision would need a much stronger and more independent validation design, a clearer review methodology, sensitivity analysis for AHP judgments, and at least one substantive case study showing real decision benefit.

Comments on the Quality of English Language

The manuscript is generally understandable, but the English is often overly verbose and uses an excessively promotional tone that overstates the strength of the evidence. The paper would benefit from careful language editing to improve concision, precision, and scientific neutrality throughout.

Author Response

We thank Reviewer 4 for the most extensive and critically rigorous review. The concerns raised are substantive and have directly shaped the major structural revisions of the manuscript. We address each in turn, and where we disagree with the reviewer's framing, we do so explicitly and with supporting argument.

Major Comment 4.1. Validation is too weak; the framework is derived from the same corpus later used for validation; 100% match on 7 papers is not convincing.

Response: We accept that the original holdout of 7 papers was insufficient to support the strength of the claims made. This has been corrected in the revised manuscript: the holdout set has been expanded to 24 independent studies, strictly isolated from the 23-study training corpus. The expanded validation activates four of the five terminal outputs and covers a substantially broader range of mission types, platform configurations, and modeling paradigms.

On the reviewer's deeper concern ( that the categories and rules were constructed from the same review process ) we offer the following clarification. The decision tree was derived inductively from the training corpus, and then applied to the holdout corpus without any post-hoc adjustment. The categories are not fitted to the holdout; they are inferred from one set and tested on another. This is the standard architecture of inductive validation. We acknowledge that the number of categories is modest (five) and that the decision rules are not learned via a statistical algorithm, which reduces the risk of overfitting relative to a parametric model. We have clarified this in Section 3.6.2 and moderated the associated claims.

Regarding the reviewer's characterization of the work as "a structured survey plus decision aid": we respectfully consider this a fair description of the contribution's nature, but not a reason for rejection. The paper does not claim to introduce a new energy model or optimization algorithm. It claims to formalize an implicit selection process. The appropriate benchmark for this claim is not a new modeling result, but rather a demonstration that the formalized logic is consistent with how the field actually makes decisions which the expanded validation addresses.

Major Comment 4.2. Methodological novelty is limited; no new modeling principle or experimental comparison is presented.

Response: We acknowledge that the paper does not introduce a new UAV energy model, which we consider a feature of its scope rather than a deficiency. The contribution is explicitly at the meta-level: it is a decision framework for selecting among existing models, not a new model itself. This is a legitimate and practically important class of contribution, particularly in fields where model selection is acknowledged to be implicit and unreproducible, as documented across our entire training corpus.

We do not claim that applying a decision tree and AHP constitutes methodological novelty in isolation. The novelty lies in the application of this combination to the specific problem of UAV energy model selection, grounded in systematic literature analysis, and operationalized through a community-facing tool. We have revised the introduction and novelty statement to express this more precisely.

Major Comment 4.3. The AHP pairwise matrices are highly subjective; who elicited the values? How was expert bias controlled?

Response: This is a valid and important concern. The revised manuscript addresses it directly in Section 4.3, which now includes a dedicated subsection on "Bias Control and Literature-Grounded Elicitation". We address this through a systematic elicitation protocol where pairwise comparison intensities (1–9) are not arbitrary but are mapped to verifiable benchmarks in the training corpus (e.g., the presence of physical differential operators vs. black-box weights). .

Major Comment 4.4. The emphasis scores E^(s) are still assigned by the author, creating a risk of circular reasoning.

Response: This concern identifies a genuine limitation of the elicitation protocol, and we do not dismiss it. The emphasis scores are derived from the textual content of each holdout paper, parsed for the core criteria. To the extent that this parsing is subjective, a risk of author-influenced alignment exists.

We have addressed this in two ways in the revised manuscript. First, we have made the elicitation mapping deterministic and fully explicit: the conversion from ordinal emphasis to Saaty intensity is defined by a fixed algebraic rule (Table in Section 4.9.1), leaving no degrees of freedom for post-hoc adjustment. Second, we note that the 24-study holdout now includes studies in which the framework's feasibility filtering is the operative mechanism, the emphasis scores are in many cases irrelevant because only one output survives Stage 1 filtering (e.g., Suo et al., where O3 is the sole feasible option). In these cases, the match is determined by the Boolean constraints alone, independent of the E^(s) assignment. We have restructured Section 4.9 to make this distinction explicit.

Major Comment 4.5. The feasibility layer is oversimplified; key factors are reduced to binary indicators.

Response: We acknowledge that reducing complex resource environments to binary indicators is a simplification. However, this formulation was adopted because it reflects the decision-theoretic reality that certain resources are either sufficient for a specific modeling approach or they are not; for instance, a wind tunnel either permits aerodynamic parameter extraction at the required fidelity or it does not. Furthermore, binary constraints allow the feasibility stage to function as a non-compensatory filter, preventing a high performance score on one criterion from overriding a fundamental lack of necessary infrastructure.

To address scenarios where resource availability is a matter of degree, we have added Section 5.2, titled "Addressing Partial Resource Constraints and Continuous Conditions." This section details how the AHP stage functions as a continuous absorption mechanism for Stage 1 ambiguities: a researcher with partial access to infrastructure can conditionally set the relevant Boolean flag while penalizing the associated AHP criterion (e.g., reducing the weight on Predictive Accuracy, C1, or increasing Development Cost, C3). This causes the global priority score for the resource-intensive option to be mathematically depressed without triggering a hard rejection. We also note that the fallback flexibility vector provides a natural mitigation path when a partially feasible option fails during execution. We identify this current limitation as a primary motivation for the planned integration of Fuzzy AHP in future work.

Major Comment 4.6. The paper does not evaluate whether the recommended model choices actually produce better downstream outcomes.

Response: We acknowledge this limitation explicitly in Section 5.3.1 of the revised manuscript. Demonstrating that framework-guided selection leads to measurably better model accuracy or mission outcomes than ad-hoc selection would require a controlled experiment in which the same research problem is addressed under both conditions — a study design that is both logistically demanding and outside the scope of the current paper. We have added an illustrative prospective application scenario in Section 5.1 (responding to Reviewer 1, Comment 1.3) as a partial demonstration, while committing to a full prospective validation study as the primary direction of future work. We consider this the most significant remaining limitation of the current manuscript.

Major Comment 4.7. The manuscript is written in an overly promotional style.

Response: We agree. The revised manuscript has been systematically edited to remove or moderate overclaiming language throughout. Key revisions include: replacing "rigorously validated" with "evaluated against an independent holdout set"; replacing "100% predictive match proves reliability" with "a high degree of consistency between the framework's recommendations and the methodologies employed in the literature"; removing phrases such as "mathematically substantiates operational validity" and "highly reliable prescriptive guide" except where qualified by appropriate caveats. The abstract, highlights, introduction, and conclusion have all been revised with this standard applied consistently.

Major Comment 4.8. The fit to the journal Drones is only moderate.

Response: We respectfully maintain that the manuscript is well-suited for Drones, as it addresses a fundamental challenge in the design and application of UAV systems. The journal’s scope emphasizes the development and operational use of unmanned vehicles; our framework provides the necessary methodological rigor for the critical "pre-design" phase selecting an energy consumption model that is technically appropriate for the specific mission constraints.

By formalizing this selection process, the paper offers a reproducible bridge between theoretical aerodynamic modeling and the practical deployment of UAVs in energy-constrained scenarios. We believe this contribution directly supports the journal's mission to advance the design and application of drone technologies by ensuring the underlying energy estimations are both auditable and consistent with current research standards.

.

Minor Comments. Prose verbosity, terminology standardization, table presentation, literature selection protocol, GUI positioning, conclusion moderation.

Response: All minor comments have been addressed in the revised manuscript. The introduction, validation interpretation, and discussion sections have been tightened to eliminate redundancy. Terminology is now standardized throughout: the paper consistently uses "white-box (physics-based)" and "black-box (data-driven/empirical/ML)" as the primary categorical labels, with more specific terms (deep learning, regression, polynomial fitting) used only at the level of individual outputs. Table presentation has been improved with consistent column widths and typographic alignment. The literature selection protocol is now documented in detail in Section 3.1, including search strings, databases, and inclusion/exclusion criteria. The GUI is repositioned in Section 4.10 as an operationalization tool rather than a scientific contribution per se. All broad claims in the conclusions have been revised to match the evidence base, as described under Major Comment 4.7 above.

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

The authors have successfully addressed all the comments.

Author Response

We thank the reviewer

Reviewer 4 Report

Comments and Suggestions for Authors

Major comment 1. Validation improved but remains not fully convincing. The revised manuscript substantially improves the validation by expanding the holdout set to 24 studies and separating it from the 23-study training corpus. This is a meaningful response to the previous major revision. Nevertheless, the validation remains retrospective and interpretive rather than prospective or outcome-based. The reported perfect alignment across all 24 studies should be treated cautiously, since the framework appears to reproduce already visible methodological choices in the literature rather than predict independently observable performance. The authors themselves acknowledge that downstream performance comparison is outside the current scope .

Major comment 2. AHP subjectivity is better documented but not solved. The manuscript now includes a bias-control discussion and claims that pairwise intensities are grounded in literature features such as interpretability and physical transparency . This is an improvement, but the actual numerical matrices still appear to be author-defined. Consistency ratios below 0.1 show internal mathematical consistency, not external validity. Therefore, the AHP component is now more transparent, but it remains subjective.

Major comment 3. The binary feasibility layer remains oversimplified. The authors now explicitly define Boolean feasibility indicators F1–F6 and map them to admissible outputs O1–O5 . They also added a discussion of partial resource constraints and future fuzzy extensions . This is a good correction, but the limitation remains central: real UAV modeling decisions are rarely binary, and the proposed workaround of manually adjusting AHP weights does not fully resolve the issue.

Major comment 4. Novelty is still moderate. The authors are now more honest that the paper does not introduce a new UAV energy model or new optimization method, but rather a structured decision aid. That framing is more defensible. Still, the contribution remains methodological-organizational rather than technically strong. For Drones, this may be acceptable if the journal values survey/framework papers, but it is not a strong technical contribution.

Major comment 5. Some promotional language remains. Although the response claims that overclaiming was reduced, the manuscript still contains strong statements such as “100\% predictive match,” “proving reliability,” and “mathematically optimal” fallback selection in some places, especially in the highlights and validation synthesis . These should be moderated further. A safer wording would be “high consistency with observed literature choices,” not “predictive accuracy” or “optimal modeling strategy.”

Minor comments. The literature selection protocol is clearer, but Google Scholar, ResearchGate, and Perplexity are not sufficient for a fully systematic review unless complemented by Scopus, Web of Science, or IEEE Xplore. The GUI figures are useful but should remain secondary, as the authors now state. Some tables are still dense, especially the validation and AHP tables. The English has generally improved but is still verbose and somewhat repetitive. The conclusion is acceptable but should emphasize limitations more strongly.

Decision: Minor Revision / Borderline Accept after revision.
The authors have partially but substantially addressed the major-revision comments. I would not reject the paper at this stage because the revised version is clearly stronger and more transparent. However, I would still request a minor revision before acceptance, mainly to further moderate the '100\%' predictive” language, clarify the subjectivity of the AHP scoring, and explicitly state that the framework has not yet been validated through downstream UAV performance experiments.

Comments on the Quality of English Language

The English has improved compared with the previous version, but the manuscript remains verbose and occasionally promotional. Further editing is recommended to improve conciseness, precision, and the moderation of strong claims.

Author Response

We thank the reviewer for the constructive and rigorous feedback. The push to moderate our language and explicitly define the boundaries of our validation has significantly strengthened the manuscript. We have implemented all requested minor revisions, focusing specifically on scaling back promotional phrasing and clarifying the epistemological limits of our AHP scoring and validation methods.

Response to Major Comment 1 (Validation limitations)

We fully agree with this assessment. The validation demonstrates that the framework accurately mirrors the state-of-the-art decision logic, but it does not prove that this logic translates to better physical flight performance. To make this limitation unambiguous, we have added an explicit disclaimer in the Conclusions and Limitations sections stating that the framework has not yet been validated through prospective, downstream UAV performance experiments, identifying this as the critical next step for future empirical research.

Response to Major Comment 2 (AHP Subjectivity)

The reviewer makes a vital distinction here between internal consistency (CR < 0.1) and external validity. We have revised Section 4.3 to explicitly concede this point. The text now clearly states that while the intensities are anchored to literature features, the final matrices remain an author-defined synthesis. We have removed any implication that the AHP scoring is entirely objective, framing it instead as a transparent but subjective baseline that requires future multi-expert consensus.

Response to Major Comment 3 (Binary Feasibility)

We appreciate the acknowledgment of our additions. We concur that the manual AHP weight adjustment is a workaround rather than a fundamental solution to the non-binary nature of real-world resource constraints. We have ensured this limitation remains central in our discussion (Section 5.3), firmly pointing to Fuzzy AHP as the necessary architectural upgrade for future iterations.

Response to Major Comment 4 (Novelty)

We accept this framing completely. We have ensured that the introduction and conclusion position the paper strictly as a methodological-organizational contribution (a structured decision aid) rather than a novel technical optimization algorithm.

Response to Major Comment 5 (Promotional Language)

We have conducted a thorough review of the manuscript, to remove overclaiming language. Phrases such as "100% predictive match," "proving reliability," and "mathematically optimal" have been systematically replaced with phrases like "high degree of consistency between the framework's recommendations and the methodologies deployed in state-of-the-art literature," "demonstrating structural alignment," and "highest-ranked feasible alternative."

Response to Minor Comments

All minor comments have been addressed:

  1. Databases: We have added a caveat in Section 3.1.1 acknowledging that relying on Google Scholar and Perplexity, without strict filtering through Web of Science or Scopus, limits the systematic rigidity of the review protocol.
  2. Tables and Verbosity: We have condensed the validation and AHP tables to improve readability and conducted another proofreading pass to reduce prose verbosity.

Conclusion: The conclusion has been restructured to emphasize the framework's limitations (particularly the lack of downstream experimental validation and AHP subjectivity) much more strongly.

Back to TopTop