A Hybrid Quantum-Classical Framework for Saliency-Aware Medical Image Encoding
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThis paper introduces a saliency-aware hybrid quantum image representation, where important image regions are encoded with higher precision and less important regions with lower precision. The general idea is interesting, and bringing content awareness into quantum image representation is a reasonable direction. That said, in its current form, the paper has major conceptual, methodological, and presentation problems, and I do not find the evidence strong enough to support publication as it stands.
My main concern is that the claims are too strong for what the results actually show. The paper repeatedly suggests better compression and clear promise for medical imaging, but the data do not really support those broader conclusions. Based on Table 6, SAHQR does not clearly outperform the main baseline methods on the most important resource measures. It uses more qubits than several alternatives, and its circuit depth and gate count are higher than some simpler methods. Its compression ratio is also much lower than that of amplitude-encoding approaches. From the results shown, the clearest point is that SAHQR is content-aware, not that it is broadly better overall. The paper should present the contribution in that more careful and limited way.
A second issue is the comparison framework. The manuscript compares several very different quantum image representation methods using the same small group of metrics, such as qubit count, circuit depth, gate count, and wall-clock encoding time. But these methods are based on different encoding ideas and do not have the same assumptions, state-preparation costs, or readout behavior. Because of that, the comparison does not feel fully fair or consistent. In particular, the reported gate counts and depths for the amplitude-encoding methods look unrealistically small for 256-pixel image preparation. The paper needs to explain the cost model much more clearly. The authors should say exactly what they count as a gate, at what level the circuit is being counted, whether multi-controlled operations are treated as one gate or decomposed, and whether the same optimization procedure was used for all methods.
There is also a major inconsistency in the SAHQR resource estimates. The manuscript gives a theoretical average gate count of about 1329.6, but later reports a measured mean gate count of 346.12 and says the difference comes from optimization and decomposition. That explanation is not convincing in its present form. The gap is too large and needs a much more precise technical explanation. Without that, it is hard to have confidence in the comparison tables.
The saliency part is also overstated, especially in the discussion of medical imaging. The paper uses Sobel gradient magnitude as the saliency detector. That can highlight edges and local contrast, but it does not by itself show that the method is preserving diagnostically important regions. The manuscript should be much more careful here. If the authors want to make medical claims, they need stronger validation, such as comparison with expert-labeled regions, tumor masks, or at least region-based fidelity measurements.
Another important limitation is the image size. Everything is resized to 16×16, which makes the study very abstract. At that resolution, a lot of meaningful structure in MRI and SAR images is already lost before the quantum encoding even begins. Because of that, the current results are better understood as a toy-scale proof of concept, not as a convincing demonstration for realistic medical or remote-sensing data. This limitation should be stated clearly, and the claims should be scaled back.
The statistical analysis also needs work. With datasets this large, even very small differences can produce highly significant p-values. Reporting p-values alone is not enough. The paper would be stronger if it also reported effect sizes, confidence intervals, and some discussion of practical significance. Table 7 is also confusing, because its title suggests one comparison, while the content includes several different method pairs.
The manuscript also needs substantial editing. The writing is repetitive, some sections restate the same points more than once, and a number of claims come across as too promotional. The introduction is longer and more tutorial than necessary for a research paper. There are also many figures, but not all of them add real value. In particular, the Q-sphere plots are visually interesting, but they do not provide much direct support for the main technical claims. I also noticed problems in the data-availability section and the author-contribution section, including wording errors and an apparent inconsistency in the listed contributors. These should be corrected carefully.
The references also need another pass. Some citations seem mismatched or only loosely connected to the point being made, and some appear to have incorrect publication details. The bibliography should be checked carefully and focused more tightly on the most relevant literature.
Overall, I think there is an interesting idea here, but the paper in its current form does not yet provide a rigorous or convincing case. The framing should be more modest, the circuit-cost analysis needs to be clarified, the medical relevance claims should be toned down unless supported more directly, and the fidelity-based evaluation should be strengthened.
Comments on the Quality of English LanguageThis paper introduces a saliency-aware hybrid quantum image representation, where important image regions are encoded with higher precision and less important regions with lower precision. The general idea is interesting, and bringing content awareness into quantum image representation is a reasonable direction. That said, in its current form, the paper has major conceptual, methodological, and presentation problems, and I do not find the evidence strong enough to support publication as it stands.
My main concern is that the claims are too strong for what the results actually show. The paper repeatedly suggests better compression and clear promise for medical imaging, but the data do not really support those broader conclusions. Based on Table 6, SAHQR does not clearly outperform the main baseline methods on the most important resource measures. It uses more qubits than several alternatives, and its circuit depth and gate count are higher than some simpler methods. Its compression ratio is also much lower than that of amplitude-encoding approaches. From the results shown, the clearest point is that SAHQR is content-aware, not that it is broadly better overall. The paper should present the contribution in that more careful and limited way.
A second issue is the comparison framework. The manuscript compares several very different quantum image representation methods using the same small group of metrics, such as qubit count, circuit depth, gate count, and wall-clock encoding time. But these methods are based on different encoding ideas and do not have the same assumptions, state-preparation costs, or readout behavior. Because of that, the comparison does not feel fully fair or consistent. In particular, the reported gate counts and depths for the amplitude-encoding methods look unrealistically small for 256-pixel image preparation. The paper needs to explain the cost model much more clearly. The authors should say exactly what they count as a gate, at what level the circuit is being counted, whether multi-controlled operations are treated as one gate or decomposed, and whether the same optimization procedure was used for all methods.
There is also a major inconsistency in the SAHQR resource estimates. The manuscript gives a theoretical average gate count of about 1329.6, but later reports a measured mean gate count of 346.12 and says the difference comes from optimization and decomposition. That explanation is not convincing in its present form. The gap is too large and needs a much more precise technical explanation. Without that, it is hard to have confidence in the comparison tables.
The saliency part is also overstated, especially in the discussion of medical imaging. The paper uses Sobel gradient magnitude as the saliency detector. That can highlight edges and local contrast, but it does not by itself show that the method is preserving diagnostically important regions. The manuscript should be much more careful here. If the authors want to make medical claims, they need stronger validation, such as comparison with expert-labeled regions, tumor masks, or at least region-based fidelity measurements.
Another important limitation is the image size. Everything is resized to 16×16, which makes the study very abstract. At that resolution, much of the meaningful structure in MRI and SAR images is already lost before quantum encoding even begins. Because of that, the current results are better understood as a toy-scale proof of concept, not as a convincing demonstration for realistic medical or remote-sensing data. This limitation should be stated clearly, and the claims should be scaled back.
The statistical analysis also needs work. With datasets this large, even very small differences can produce highly significant p-values. Reporting p-values alone is not enough. The paper would be stronger if it also reported effect sizes, confidence intervals, and discussed practical significance. Table 7 is also confusing because its title suggests a single comparison, while the content includes several method pairs.
The manuscript also needs substantial editing. The writing is repetitive; some sections restate the same points multiple times, and a number of claims come across as overly promotional. The introduction is longer and more tutorial than necessary for a research paper. There are also many figures, but not all of them add real value. In particular, the Q-sphere plots are visually interesting, but they do not provide much direct support for the main technical claims. I also noticed problems in the data availability and author contribution sections, including wording errors and an apparent inconsistency in the list of contributors. These should be corrected carefully.
The references also need another pass. Some citations seem mismatched or only loosely connected to the point being made, and some appear to have incorrect publication details. The bibliography should be checked carefully and focused more tightly on the most relevant literature.
I think there is an interesting idea here, but the paper, in its current form, does not yet present a rigorous or convincing case. The framing should be more modest; the circuit-cost analysis needs clarification; the medical-relevance claims should be toned down unless supported more directly; and the fidelity-based evaluation should be strengthened.
Author Response
We sincerely thank the reviewer for their valuable comments and insightful suggestions, which have helped us significantly improve the quality and clarity of the manuscript.
Reviewer 1
Comment 1 : This paper introduces a saliency-aware hybrid quantum image representation, where important image regions are encoded with higher precision and less important regions with lower precision. The general idea is interesting, and bringing content awareness into quantum image representation is a reasonable direction. That said, in its current form, the paper has major conceptual, methodological, and presentation problems, and I do not find the evidence strong enough to support publication as it stands.
Response 1 :We thank the reviewer for this important observation. We agree that the original manuscript overstated the performance advantages of SAHQR. The primary contribution of our work is content-aware adaptive precision encoding, rather than universal superiority across all resource metrics. We have revised the manuscript to reflect a more balanced and evidence-supported interpretation of results.(Page No 36 - conclusion and Abstract )
Comment 2 : My main concern is that the claims are too strong for what the results actually show. The paper repeatedly suggests better compression and clear promise for medical imaging, but the data do not really support those broader conclusions. Based on Table 6, SAHQR does not clearly outperform the main baseline methods on the most important resource measures. It uses more qubits than several alternatives, and its circuit depth and gate count are higher than some simpler methods. Its compression ratio is also much lower than that of amplitude-encoding approaches. From the results shown, the clearest point is that SAHQR is content-aware, not that it is broadly better overall. The paper should present the contribution in that more careful and limited way.
Response 2 :We agree that comparing heterogeneous quantum image representations using identical metrics without clarifying assumptions may lead to ambiguity. We have now explicitly defined the cost model and ensured consistent evaluation criteria across all methods. We have included the sub section Cost Model and Evaluation Assumptions (page no .24) and also referred to Table 6
Comment 3 :A second issue is the comparison framework. The manuscript compares several very different quantum image representation methods using the same small group of metrics, such as qubit count, circuit depth, gate count, and wall-clock encoding time. But these methods are based on different encoding ideas and do not have the same assumptions, state-preparation costs, or readout behavior. Because of that, the comparison does not feel fully fair or consistent. In particular, the reported gate counts and depths for the amplitude-encoding methods look unrealistically small for 256-pixel image preparation. The paper needs to explain the cost model much more clearly. The authors should say exactly what they count as a gate, at what level the circuit is being counted, whether multi-controlled operations are treated as one gate or decomposed, and whether the same optimization procedure was used for all methods.
Response 3 :We agree that comparing heterogeneous quantum image representations using identical metrics without clarifying assumptions may lead to ambiguity. We have now explicitly defined the cost model and ensured consistent evaluation criteria across all methods. We have included the sub section Cost Model and Evaluation Assumptions (page no .24) and also referred to Table 6
Comment 4 :There is also a major inconsistency in the SAHQR resource estimates. The manuscript gives a theoretical average gate count of about 1329.6, but later reports a measured mean gate count of 346.12 and says the difference comes from optimization and decomposition. That explanation is not convincing in its present form. The gap is too large and needs a much more precise technical explanation. Without that, it is hard to have confidence in the comparison tables.
Response 4 :We appreciate the reviewer’s concern regarding the discrepancy between theoretical and measured gate counts. Explanation of Theoretical vs Empirical Gate Counts and table 7 added on page no 27,28
Comment 5 :The saliency part is also overstated, especially in the discussion of medical imaging. The paper uses Sobel gradient magnitude as the saliency detector. That can highlight edges and local contrast, but it does not by itself show that the method is preserving diagnostically important regions. The manuscript should be much more careful here. If the authors want to make medical claims, they need stronger validation, such as comparison with expert-labeled regions, tumor masks, or at least region-based fidelity measurements.
Response 5 : We agree with the reviewer that the saliency discussion was overstated. Replaced strong claims such as “preserves diagnostically important regions” with “enhances edge and gradient-based structural features” and Clinical relevance requires validation against expert annotations or segmentation masks, which is beyond the scope of this study has been added
Comment 6 :Another important limitation is the image size. Everything is resized to 16×16, which makes the study very abstract. At that resolution, a lot of meaningful structure in MRI and SAR images is already lost before the quantum encoding even begins. Because of that, the current results are better understood as a toy-scale proof of concept, not as a convincing demonstration for realistic medical or remote-sensing data. This limitation should be stated clearly, and the claims should be scaled back.
Response 6 : We acknowledge this important limitation.subsection added Limitations and Scalability Considerations (Page No 39)
Comment 7 :The statistical analysis also needs work. With datasets this large, even very small differences can produce highly significant p-values. Reporting p-values alone is not enough. The paper would be stronger if it also reported effect sizes, confidence intervals, and some discussion of practical significance. Table 7 is also confusing, because its title suggests one comparison, while the content includes several different method pairs.
Response 7 : We thank the reviewer for this valuable suggestion. Table 8 is updated and caption also modified (Page no 32)
Comment 8 :The manuscript also needs substantial editing. The writing is repetitive, some sections restate the same points more than once, and a number of claims come across as too promotional. The introduction is longer and more tutorial than necessary for a research paper. There are also many figures, but not all of them add real value. In particular, the Q-sphere plots are visually interesting, but they do not provide much direct support for the main technical claims. I also noticed problems in the data-availability section and the author-contribution section, including wording errors and an apparent inconsistency in the listed contributors. These should be corrected carefully.
Response 8 :We have performed a comprehensive revision of the manuscript. The data-availability section and the author-contribution section have been revised in the manuscript .Q-sphere plots are removed.
Comment 9 :The references also need another pass. Some citations seem mismatched or only loosely connected to the point being made, and some appear to have incorrect publication details. The bibliography should be checked carefully and focused more tightly on the most relevant literature.
Response 9 : We have carefully reviewed and corrected the bibliography:
- Fixed incorrect publication details
- Removed weak or loosely related citations
- Added more relevant and recent works in quantum image processing
- Ensured all references directly support the associated claims
Comment 10 :Overall, I think there is an interesting idea here, but the paper in its current form does not yet provide a rigorous or convincing case. The framing should be more modest, the circuit-cost analysis needs to be clarified, the medical relevance claims should be toned down unless supported more directly, and the fidelity-based evaluation should be strengthened.
Response 10 :
We appreciate the reviewer’s constructive summary.
We have made the following major improvements:
- More Modest Framing
- Replaced strong claims with: “proof-of-concept”, “exploratory study”, “preliminary validation”
- Improved Cost Analysis
- Added detailed cost model (Comment 3 fix)
- Ensured fair comparison across methods
- Reduced Medical Claims
- Removed unsupported clinical interpretations
- Clearly stated limitations
- Strengthened Evaluation
- Added statistical rigor (effect size, CI)
- Clarified interpretation of results
- Clear Positioning
“The proposed method demonstrates a structurally efficient encoding strategy under constrained quantum resources, but further validation is required for real-world applicability.”
Reviewer 2 Report
Comments and Suggestions for AuthorsThe manuscript submitted by the authors addresses a relevant and timely problem in quantum image processing, particularly in the context of efficient encoding for large-scale medical imaging data. The idea of introducing content-awareness through saliency detection represents a meaningful conceptual advancement over traditional uniform encoding schemes. Moreover, the work is well-structured, methodologically sound, and demonstrates a commendable level of experimental effort. However, several important issues related to practical relevance, scalability, and positioning within the broader field should be addressed prior to publication.
These specific concerns are listed below
The experiments conducted by the authors were performed using classical simulation (Qiskit statevector simulator). While this is common in early-stage QIP research, the manuscript should include simulation using Performance under noise models, or Feasibility on NISQ hardware. Given that circuit depth and gate count are critical in noisy environments, the absence of such analysis limits the practical significance of the work.
The proposed method presented here relies heavily on classical saliency detection prior to quantum encoding, raising if SAHQR is a quantum method or a classical preprocessing + quantum encoding pipeline. Therefore, the authors should clarify the novelty in terms of quantum contribution, and discuss limitations of hybrid approaches.
The manuscript claims improved compression efficiency but only compares against other quantum representations. Here, It would be adequate to include at least a qualitative (preferably quantitative) comparison with classical baselines.
All experiments conducted here were performed on 16 × 16 images, which is significantly below real-world medical imaging resolutions. Although justified by simulation constraints, this raises concerns about scalability of the approach and applicability to real datasets. In addition, Q-sphere visualizations are visually informative but lack quantitative interpretation.
Finally, while SAHQR improves compression, also introduces additional preprocessing (saliency detection), and additional circuit components (saliency qubit, conditional encoding). Nonetheless, the manuscript does not clearly quantify whether the benefits of theses analysis outweigh the added costs.
Author Response
We appreciate the reviewer’s careful evaluation of our work and have addressed all the comments in detail, incorporating the suggested revisions to strengthen the manuscript.
Reviewer 2
Comment 1 :The experiments conducted by the authors were performed using classical simulation (Qiskit state vector simulator). While this is common in early-stage QIP research, the manuscript should include simulation using Performance under noise models, or Feasibility on NISQ hardware. Given that circuit depth and gate count are critical in noisy environments, the absence of such analysis limits the practical significance of the work.
Response 1 : We thank the reviewer for highlighting the importance of evaluating quantum circuits under realistic noisy conditions. We agree that performance under noise and feasibility on NISQ hardware are critical for assessing practical applicability. The subsection has been added as Performance Under Noise Models and NISQ Feasibility on page no 24
Comment 2 : The proposed method presented here relies heavily on classical saliency detection prior to quantum encoding, raising if SAHQR is a quantum method or a classical preprocessing + quantum encoding pipeline. Therefore, the authors should clarify the novelty in terms of quantum contribution, and discuss limitations of hybrid approaches.
Response 2 : We appreciate the reviewer’s insightful observation regarding the hybrid nature of the proposed SAHQR framework. In revised manuscript the subsection has been added as Hybrid Nature and Quantum Contribution on page no 23
Comment 3 :The manuscript claims improved compression efficiency but only compares against other quantum representations. Here, It would be adequate to include at least a qualitative (preferably quantitative) comparison with classical baselines.
Response 3 :We thank the reviewer for this important suggestion. We agree that comparing with classical baselines strengthens the claims of compression efficiency. In the revised manuscript, we have included comparisons with classical image compression methods. The point Comparison with Classical Compression Methods is added in revised manuscript on page no 33 and 34
Comment 4 :All experiments conducted here were performed on 16 × 16 images, which is significantly below real-world medical imaging resolutions. Although justified by simulation constraints, this raises concerns about scalability of the approach and applicability to real datasets. In addition, Q-sphere visualizations are visually informative but lack quantitative interpretation.
Response 4: We acknowledge the reviewer’s concern regarding the use of 16×16 images. This limitation arises due to the exponential growth of quantum state space and simulator constraints. We have added it in future scope and limitations .
Comment 5 :Finally, while SAHQR improves compression, also introduces additional preprocessing (saliency detection), and additional circuit components (saliency qubit, conditional encoding). Nonetheless, the manuscript does not clearly quantify whether the benefits of theses analysis outweigh the added costs.
Response 5:
We thank the reviewer for pointing out the need for a clear trade-off analysis.We agree that SAHQR introduces additional overhead due to:
- Saliency detection preprocessing
- Additional qubit (saliency qubit)
- Controlled operations
To address this, we have introduced a comprehensive cost-benefit analysis included in Cost-Benefit Analysis of SAHQR page no 38
Round 2
Reviewer 2 Report
Comments and Suggestions for AuthorsThe authors have correctly addressed the requested comments and now the manuscript is suitable for publication in its present form

