Benchmarking Barren Plateau Mitigation Strategies in Quantum Neural Networks on Standard and Medical Image Datasets
Abstract
1. Introduction
1.1. Motivation and Objectives
1.2. Contributions
1.3. Paper Organization
2. Preliminary Background
2.1. Parameterized Quantum Circuits (PQCs) and Cost Function Landscape
2.2. Manifestation of Barren Plateaus
2.2.1. Vanishing Gradients
2.2.2. Flat Optimization Landscapes
2.3. Relevance to Quantum Neural Networks
2.4. Theoretical Foundations and Indicators
2.5. Key Factors Contributing to Barren Plateaus
- 1
- Expressibility of Parameterized Quantum Circuits (PQCs): Highly expressive PQCs, which can approximate arbitrary unitary transformations, tend to create more entanglement among qubits. While expressibility is a desirable feature for capturing complex correlations, it also increases the likelihood of gradients vanishing. The variance of gradients decreases as circuits become more expressive, as shown by the relationship in Equation (5)where and are density matrices, and represents the Hilbert space dimension [16].
- 2
- Entanglement-Induced BPs: Marrero et al. demonstrated that BPs can also arise from excessive entanglement in the circuit. Deeply entangled states lead to near-uniform sampling in the Hilbert space, which further diminishes gradient variance. This effect is exacerbated in circuits with global cost functions, as they involve measurements over all qubits, amplifying the gradient vanishing effect [17].
- 3
- Cost Function Locality: Cerezo et al. showed that the structure of the cost function significantly impacts the emergence of BPs. Local cost functions, which involve a small subset of qubits, tend to mitigate gradient decay compared to global cost functions. This is formalized as follows in Equation (6)where the polynomial factor depends on the locality of the cost function [14].
- 4
- Initialization and Parameter Scaling: The choice of parameter initialization plays a crucial role in determining the onset of BPs. Improper initialization can cause the circuit to operate in regions of parameter space where gradients are uniformly small. Techniques like Beta Initialization and Fourier-based parameterization have been proposed to address this issue by ensuring the variance of initial parameters is distributed optimally across the circuit layers [17].
3. Methodology
3.1. Scope of Noise Modeling
3.2. Datasets
- 1
- Iris: A small dataset with 150 samples and four features across three classes, commonly used for evaluating the basic performance of classification algorithms. It serves as a benchmark for assessing BP mitigation in simpler QNN configurations.
- 2
- MNIST: A dataset containing 70,000 grayscale images of handwritten digits (0–9), each 28 × 28 pixels. This dataset offers moderate complexity and is frequently used to evaluate model performance in classification tasks, including QNNs.
- 3
- MedMNIST: A high-dimensional dataset composed of medical images from various categories, with greater complexity and dimensionality. This dataset is ideal for evaluating BP mitigation strategies in large-scale QNN applications, testing scalability and the ability to handle intricate data structures.
3.3. Preprocessing and Quantum Feature Encoding
3.4. Experimental Setup
3.4.1. Quantum Neural Network Setup
- Number of qubits: Experiments are conducted for 2, 4, 8, 12, 16, and 20 qubits.
- Circuit layers: Each circuit includes two parameterized layers with single-qubit rotations (RX, RY) and controlled-Z (CZ) entangling gates.
- Measurement: The expectation value of the Pauli-Z observable is computed for the final state.
3.4.2. Hyperparameters
- Learning rate: 0.001 (scaled for specific initialization strategies when required).
- Batch size: 16 samples per batch for all datasets.
- Optimizer: Adam optimizer is employed for all experiments, with adaptive learning based on initialization.
- Epochs: Each experiment is run for 30 epochs.
3.5. BP Mitigation Techniques
- Initialization-Based Strategy
- 1
- Beta Initialization: Kulshrestha et al. [18] propose initializing model weights using a Beta distribution fitted to the normalized input data. The input is normalized to as in Equation (7)and the Beta distribution, as follows in Equation (8),is used to determine shape parameters and . Weights are then initialized as , promoting stable gradient propagation. In related work, small stochastic perturbations may be used to improve exploration of the parameter landscape; however, the present benchmark does not use this mechanism as a hardware-noise model. This approach mitigates the rapid gradient variance decay in large models and improves optimization efficiency, particularly in binary classification tasks [19,20].
- 2
- Uniform Norm Initialization: This strategy leverages a uniform distribution aligned with data-specific statistical properties. Weights W are flattened and fitted to a uniform distribution with bounds a and b, and normalized as in the below Equation (9)to fall within . Throughout the training, gradient norms are monitored, with reinitialization or learning rate adjustments if norms drop below , helping maintain stable gradient flow and preventing BPs [21,22].
- 3
- Gaussian Initialization: In Equation (10) parameters are initialized from a Gaussian distribution , with scaled to the number of layers.This approach is known to improve gradient retention in deep QNNs by minimizing rapid gradient decay. QNNs were implemented with parameterized and rotation layers and controlled-Z gates to enhance model expressibility. Gradient norms are monitored, with reinitialization triggered if norms drop below a threshold, helping mitigate BPs [23].
- 4
- He Initialization: The He_normal initialization sets the weights of a layer with input units according to a Gaussian (normal) distribution with mean 0 and variance, which can be defined as in Equation (11)Mathematically, the weights can be initialized as follows in Equation (12):where represents the number of input units in the layer.In He_uniform initialization, the weights are sampled from a uniform distribution within the range , where it can be defined as in Equation (13)Mathematically, the weights can be initialized as follows in Equation (14)where denotes the number of input units.
- 5
- Xavier Initialization: The Xavier_normal initialization (Equation (15)) sets the weights of a layer with input units and output units according to a Gaussian (normal) distribution with mean 0 and variance.Mathematically, the weights are initialized as follows in Equation (16)In Xavier_uniform initialization, the weights of a layer with input units and output units are sampled from a uniform distribution within the range , as in Equation (17)Mathematically, the weights can be initialized as in Equation (18):
3.5.1. CNN-Based Initialization
3.5.2. Model-Based Variational Encoder
3.5.3. Optimization-Based Time-Nonlocal Fourier Parameterization
3.6. Reference Initialization Configuration
3.7. Evaluation Metrics
3.7.1. Gradient Variance
- -
- represents the gradient vector of the cost function C at epoch i,
- -
- is the Euclidean norm (magnitude) of the gradient vector for epoch i,
- -
- is the mean gradient norm across N epochs,
- -
- N is the total number of epochs or data points considered.
3.7.2. Training Loss
- is the loss function (e.g., mean squared error, cross-entropy),
- and are the true and predicted labels, respectively,
- N is the number of data points.
3.8. Evaluation Scope and Downstream Classification Metrics
4. Results and Discussion
4.1. Variance Analysis: Mitigation of Barren Plateaus
4.2. Loss Curve Analysis: Evaluating Convergence and Stability
4.3. Benchmark Viewpoints
4.4. Circuit Depth and Scalability Considerations
4.5. Practical QNN Setbacks and Benchmark Implications
4.6. Discussion
5. Related Work
5.1. Initialization-Based Mitigation Strategies
5.2. Model-Based Mitigation Strategies
5.3. Optimization-Based Mitigation Strategies
5.4. Evaluation Tools and Theoretical Insights
6. Conclusions
Limitations and Future Scope
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Preskill, J. Quantum computing in the NISQ era and beyond. Quantum 2018, 2, 79. [Google Scholar] [CrossRef]
- Bharti, K.; Cervera-Lierta, A.; Kyaw, T.H.; Haug, T.; Alperin-Lea, S.; Anand, A.; Degroote, M.; Heimonen, H.; Kottmann, J.S.; Menke, T.; et al. Noisy intermediate-scale quantum algorithms. Rev. Mod. Phys. 2022, 94, 015004. [Google Scholar] [CrossRef]
- McClean, J.R.; Romero, J.; Babbush, R.; Aspuru-Guzik, A. The theory of variational hybrid quantum-classical algorithms. New J. Phys. 2016, 18, 023023. [Google Scholar] [CrossRef]
- Farhi, E.; Goldstone, J.; Gutmann, S. A quantum approximate optimization algorithm. arXiv 2014, arXiv:1411.4028. [Google Scholar]
- Peruzzo, A.; McClean, J.; Shadbolt, P.; Yung, M.H.; Zhou, X.Q.; Love, P.J.; Aspuru-Guzik, A.; O’brien, J.L. A variational eigenvalue solver on a photonic quantum processor. Nat. Commun. 2014, 5, 4213. [Google Scholar] [CrossRef] [PubMed]
- Dealing with Unreliable Annotations: A Noise-Robust Network for Semantic Segmentation through a Transformer-Improved Encoder and Convolution Decoder. Appl. Sci. 2023, 13, 7966. [CrossRef]
- Schuld, M.; Sinayskiy, I.; Petruccione, F. The quest for a quantum neural network. Quantum Inf. Process. 2014, 13, 2567–2586. [Google Scholar] [CrossRef]
- Biamonte, J.; Wittek, P.; Pancotti, N.; Rebentrost, P.; Wiebe, N.; Lloyd, S. Quantum machine learning. Nature 2017, 549, 195–202. [Google Scholar] [CrossRef] [PubMed]
- Havlíček, V.; Córcoles, A.D.; Temme, K.; Harrow, A.W.; Kandala, A.; Chow, J.M.; Gambetta, J.M. Supervised learning with quantum-enhanced feature spaces. Nature 2019, 567, 209–212. [Google Scholar] [CrossRef] [PubMed]
- Mitarai, K.; Negoro, M.; Kitagawa, M.; Fujii, K. Quantum circuit learning. Phys. Rev. A 2018, 98, 032309. [Google Scholar] [CrossRef]
- Kandala, A.; Mezzacapo, A.; Temme, K.; Takita, M.; Brink, M.; Chow, J.M.; Gambetta, J.M. Hardware-efficient variational quantum eigensolver for small molecules and quantum magnets. Nature 2017, 549, 242–246. [Google Scholar] [CrossRef] [PubMed]
- Zhou, L.; Wang, S.T.; Choi, S.; Pichler, H.; Lukin, M.D. Quantum approximate optimization algorithm: Performance, mechanism, and implementation on near-term devices. Phys. Rev. X 2020, 10, 021067. [Google Scholar] [CrossRef]
- Benedetti, M.; Lloyd, E.; Sack, S.; Fiorentini, M. Parameterized quantum circuits as machine learning models. Quantum Sci. Technol. 2019, 4, 043001. [Google Scholar] [CrossRef]
- Cerezo, M.; Sone, A.; Volkoff, T.; Cincio, L.; Coles, P.J. Cost-function-dependent barren plateaus in shallow quantum neural networks (2020). arXiv 2001, arXiv:2001.00550. [Google Scholar]
- Grant, E.; Wossnig, L.; Ostaszewski, M.; Benedetti, M. An initialization strategy for addressing barren plateaus in parametrized quantum circuits. Quantum 2019, 3, 214. [Google Scholar] [CrossRef]
- McArdle, S.; Endo, S.; Aspuru-Guzik, A.; Benjamin, S.C.; Yuan, X. Quantum computational chemistry. Rev. Mod. Phys. 2020, 92, 015003. [Google Scholar] [CrossRef]
- Ortiz Marrero, C.; Kieferová, M.; Wiebe, N. Entanglement-induced barren plateaus. PRX Quantum 2021, 2, 040316. [Google Scholar] [CrossRef]
- Kulshrestha, A.; Safro, I. Beinit: Avoiding barren plateaus in variational quantum algorithms. arXiv 2022, arXiv:2204.13751. [Google Scholar]
- Kaminishi, E.; Mori, T.; Sugawara, M.; Yamamoto, N. Impact of Measurement Noise on Escaping Saddles in Variational Quantum Algorithms. arXiv 2024, arXiv:2406.09780. [Google Scholar]
- Liu, J.; Wilde, F.; Mele, A.A.; Jiang, L.; Eisert, J. Stochastic noise can be helpful for variational quantum algorithms. arXiv 2022, arXiv:2210.06723. [Google Scholar]
- Kashif, M.; Rashid, M.; Al-Kuwari, S.; Shafique, M. Alleviating barren plateaus in parameterized quantum machine learning circuits: Investigating advanced parameter initialization strategies. In Proceedings of the 2024 Design, Automation & Test in Europe Conference & Exhibition (DATE); IEEE: New York, NY, USA, 2024; pp. 1–6. [Google Scholar]
- Zhang, K.; Liu, L.; Hsieh, M.H.; Tao, D. Escaping from the barren plateau via gaussian initializations in deep variational quantum circuits. Adv. Neural Inf. Process. Syst. 2022, 35, 18612–18627. [Google Scholar] [CrossRef]
- Shi, X.; Shang, Y. Avoiding barren plateaus via gaussian mixture model. arXiv 2024, arXiv:2402.13501. [Google Scholar]
- McClean, J.R.; Boixo, S.; Smelyanskiy, V.N.; Babbush, R.; Neven, H. Barren plateaus in quantum neural network training landscapes. Nat. Commun. 2018, 9, 4812. [Google Scholar] [CrossRef] [PubMed]
- Sauvage, F.; Sim, S.; Kunitsa, A.A.; Simon, W.A.; Mauri, M.; Perdomo-Ortiz, A. FLIP: A flexible initializer for arbitrarily-sized parametrized quantum circuits. arXiv 2021, arXiv:2103.08572. [Google Scholar]
- Sack, S.H.; Medina, R.A.; Michailidis, A.A.; Kueng, R.; Serbyn, M. Avoiding barren plateaus using classical shadows. PRX Quantum 2022, 3, 020365. [Google Scholar] [CrossRef]
- Rad, A.; Seif, A.; Linke, N.M. Surviving the barren plateau in variational quantum circuits with bayesian learning initialization. arXiv 2022, arXiv:2203.02464. [Google Scholar]
- Friedrich, L.; Maziero, J. Avoiding barren plateaus with classical deep neural networks. Phys. Rev. A 2022, 106, 042433. [Google Scholar] [CrossRef]
- Mele, A.A.; Mbeng, G.B.; Santoro, G.E.; Collura, M.; Torta, P. Avoiding barren plateaus via transferability of smooth solutions in a Hamiltonian variational ansatz. Phys. Rev. A 2022, 106, L060401. [Google Scholar] [CrossRef]
- Grimsley, H.R.; Barron, G.S.; Barnes, E.; Economou, S.E.; Mayhall, N.J. Adaptive, problem-tailored variational quantum eigensolver mitigates rough parameter landscapes and barren plateaus. npj Quantum Inf. 2023, 9, 19. [Google Scholar] [CrossRef]
- Liu, H.Y.; Sun, T.P.; Wu, Y.C.; Han, Y.J.; Guo, G.P. Mitigating barren plateaus with transfer-learning-inspired parameter initializations. New J. Phys. 2023, 25, 013039. [Google Scholar] [CrossRef]
- Park, C.Y.; Killoran, N. Hamiltonian variational ansatz without barren plateaus. Quantum 2024, 8, 1239. [Google Scholar] [CrossRef]
- Li, G.; Song, Z.; Wang, X. VSQL: Variational shadow quantum learning for classification. Proc. AAAI Conf. Artif. Intell. 2021, 35, 8357–8365. [Google Scholar] [CrossRef]
- Bharti, K.; Haug, T. Quantum-assisted simulator. Phys. Rev. A 2021, 104, 042418. [Google Scholar] [CrossRef]
- Du, Y.; Huang, T.; You, S.; Hsieh, M.H.; Tao, D. Quantum circuit architecture search for variational quantum algorithms. npj Quantum Inf. 2022, 8, 62. [Google Scholar] [CrossRef]
- Zhang, Z.; Chen, Z.; Huang, H.; Jia, Z. Quark: A Gradient-Free Quantum Learning Framework for Classification Tasks. arXiv 2022, arXiv:2210.01311. [Google Scholar]
- Selvarajan, R.; Sajjan, M.; Humble, T.S.; Kais, S. Dimensionality reduction with variational encoders based on subsystem purification. Mathematics 2023, 11, 4678. [Google Scholar] [CrossRef]
- Tüysüz, C.; Clemente, G.; Crippa, A.; Hartung, T.; Kühn, S.; Jansen, K. Classical splitting of parametrized quantum circuits. Quantum Mach. Intell. 2023, 5, 34. [Google Scholar] [CrossRef]
- Kashif, M.; Al-Kuwari, S. ResQNets: A residual approach for mitigating barren plateaus in quantum neural networks. EPJ Quantum Technol. 2024, 11, 4. [Google Scholar] [CrossRef]
- Shin, M.; Lee, S.; Lee, M.; Ji, D.; Yeo, H.; Lee, H.J.; Jeong, K. Layerwise Quantum Convolutional Neural Networks Provide a Unified Way for Estimating Fundamental Properties of Quantum Information Theory. arXiv 2024, arXiv:2401.07716. [Google Scholar]
- Zhang, H.K.; Liu, S.; Zhang, S.X. Absence of barren plateaus in finite local-depth circuits with long-range entanglement. Phys. Rev. Lett. 2024, 132, 150603. [Google Scholar] [CrossRef] [PubMed]
- Ostaszewski, M.; Grant, E.; Benedetti, M. Structure optimization for parameterized quantum circuits. Quantum 2021, 5, 391. [Google Scholar] [CrossRef]
- Skolik, A.; McClean, J.R.; Mohseni, M.; Van Der Smagt, P.; Leib, M. Layerwise learning for quantum neural networks. Quantum Mach. Intell. 2021, 3, 5. [Google Scholar] [CrossRef]
- Gharibyan, H.; Su, V.; Tepanyan, H. Hierarchical Learning for Quantum ML: Novel Training Technique for Large-Scale Variational Quantum Circuits. arXiv 2023, arXiv:2311.12929. [Google Scholar]
- Haug, T.; Kim, M. Optimal training of variational quantum algorithms without barren plateaus. arXiv 2021, arXiv:2104.14543. [Google Scholar]
- Mele, A.A.; Angrisani, A.; Ghosh, S.; Khatri, S.; Eisert, J.; França, D.S.; Quek, Y. Noise-induced shallow circuits and absence of barren plateaus. arXiv 2024, arXiv:2403.13927. [Google Scholar]
- Sannia, A.; Tacchino, F.; Tavernelli, I.; Giorgi, G.L.; Zambrini, R. Engineered dissipation to mitigate barren plateaus. npj Quantum Inf. 2024, 10, 81. [Google Scholar] [CrossRef]
- Zambrano, L.; Muñoz-Moller, A.D.; Muñoz, M.; Pereira, L.; Delgado, A. Avoiding barren plateaus in the variational determination of geometric entanglement. Quantum Sci. Technol. 2024, 9, 025016. [Google Scholar] [CrossRef]
- Broers, L.; Mathey, L. Mitigated barren plateaus in the time-nonlocal optimization of analog quantum-algorithm protocols. Phys. Rev. Res. 2024, 6, 013076. [Google Scholar]
- Heyraud, V.; Li, Z.; Donatella, K.; Le Boité, A.; Ciuti, C. Efficient estimation of trainability for variational quantum circuits. PRX Quantum 2023, 4, 040335. [Google Scholar] [CrossRef]
- Kieferova, M.; Carlos, O.M.; Wiebe, N. Quantum Generative Training Using R∖’enyi Divergences. arXiv 2021, arXiv:2106.09567. [Google Scholar]
- Sciorilli, M.; Borges, L.; Patti, T.L.; García-Martín, D.; Camilo, G.; Anandkumar, A.; Aolita, L. Towards large-scale quantum optimization solvers with few qubits. arXiv 2024, arXiv:2401.09421. [Google Scholar]
- Falla, J.; Langfitt, Q.; Alexeev, Y.; Safro, I. Graph representation learning for parameter transferability in quantum approximate optimization algorithm. Quantum Mach. Intell. 2024, 6, 46. [Google Scholar] [CrossRef]
- Patti, T.L.; Najafi, K.; Gao, X.; Yelin, S.F. Entanglement devised barren plateau mitigation. Phys. Rev. Res. 2021, 3, 033090. [Google Scholar] [CrossRef]
- Larocca, M.; Czarnik, P.; Sharma, K.; Muraleedharan, G.; Coles, P.J.; Cerezo, M. Diagnosing barren plateaus with tools from quantum optimal control. Quantum 2022, 6, 824. [Google Scholar] [CrossRef]
- Park, S.; Kim, J. Quantum Neural Network Software Testing, Analysis, and Code Optimization for Advanced IoT Systems: Design, Implementation, and Visualization. arXiv 2024, arXiv:2401.10914. [Google Scholar]
- Cerezo, M.; Sone, A.; Volkoff, T.; Cincio, L.; Coles, P.J. Cost function dependent barren plateaus in shallow parametrized quantum circuits. Nat. Commun. 2021, 12, 1791. [Google Scholar] [CrossRef] [PubMed]







| Method Group | Gradient Variance | Training Loss | Accuracy | Macro-F1/AUC |
|---|---|---|---|---|
| Stochastic reference initialization | Reported | Reported | Not reported | Not reported |
| Beta initialization | Reported | Reported | Not reported | Not reported |
| Gaussian initialization | Reported | Reported | Not reported | Not reported |
| Uniform Norm initialization | Reported | Reported | Not reported | Not reported |
| CNN-based initialization | Reported | Reported | Not reported | Not reported |
| He-normal/He-uniform | Reported | Reported | Not reported | Not reported |
| Xavier-normal/Xavier-uniform | Reported | Reported | Not reported | Not reported |
| Model-based variational encoder | Reported | Reported | Not reported | Not reported |
| Optimization-based Fourier parameterization | Reported | Reported | Not reported | Not reported |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Rahman, M.; Liu, R.; Majumder, A.; Paul, P.C.; Mo, K.; Begum, A.; Sultana, K.; Akter, N.; Wei, L.; Zhang, Y.; et al. Benchmarking Barren Plateau Mitigation Strategies in Quantum Neural Networks on Standard and Medical Image Datasets. J. Imaging 2026, 12, 275. https://doi.org/10.3390/jimaging12070275
Rahman M, Liu R, Majumder A, Paul PC, Mo K, Begum A, Sultana K, Akter N, Wei L, Zhang Y, et al. Benchmarking Barren Plateau Mitigation Strategies in Quantum Neural Networks on Standard and Medical Image Datasets. Journal of Imaging. 2026; 12(7):275. https://doi.org/10.3390/jimaging12070275
Chicago/Turabian StyleRahman, Maqsudur, Rui Liu, Anup Majumder, Pintu Chandra Paul, Kangtong Mo, Amena Begum, Kashmi Sultana, Nahida Akter, Lu Wei, Ye Zhang, and et al. 2026. "Benchmarking Barren Plateau Mitigation Strategies in Quantum Neural Networks on Standard and Medical Image Datasets" Journal of Imaging 12, no. 7: 275. https://doi.org/10.3390/jimaging12070275
APA StyleRahman, M., Liu, R., Majumder, A., Paul, P. C., Mo, K., Begum, A., Sultana, K., Akter, N., Wei, L., Zhang, Y., & Zhuang, J. (2026). Benchmarking Barren Plateau Mitigation Strategies in Quantum Neural Networks on Standard and Medical Image Datasets. Journal of Imaging, 12(7), 275. https://doi.org/10.3390/jimaging12070275

