Depth Fragility and Skeletal Universality: Decoupling Topology and Function in Deep Neural Networks
Abstract
1. Introduction
- 1.
- The Zombie Network State: Under random edge pruning, a network can remain almost entirely connected (, is ≈1.0) yet exhibit a complete functional collapse. This indicates that topological reachability does not imply computational capability, because the surviving random connections carry insufficient signal energy to propagate useful feature representations through the network.
- 2.
- Depth-Induced Fragility: We uncover a trade-off between expressivity and structural stability. While increasing network depth improves performance, it introduces a severe fragility to stochastic noise. Deep networks suffer a catastrophic “avalanche” collapse at , compared to the gradual decay observed in shallow architectures ().
- 3.
- Universality of the Skeleton in this Setting: Under magnitude pruning, functional robustness becomes nearly invariant to architecture. Networks of different configurations converge to a similar critical threshold (–), suggesting that the functional skeleton reflects the task’s complexity rather than the model’s capacity, for the architectures and dataset studied here.
2. Methods
2.1. Notation
2.2. Dataset
2.3. Network Modeling and Architectural Variants
- Shallow-MLP (): minimal-depth architecture with one hidden layer of 512 nodes.
- Deep-MLP (): increased depth leading to multiplicative signal attenuation and enhanced sensitivity to stochastic perturbations.
- Wide-MLP (): increased width at the same depth as the Deep-MLP, introducing topological redundancy that can temporarily buffer random damage.
2.4. Edge Pruning Protocol
- Stochastic pruning (random edge removal) simulates non-selective synaptic decay or hardware failure by removing connections with uniform probability. This corresponds to classical random edge percolation.
- Magnitude pruning (weak edge removal) simulates targeted skeletal extraction by prioritizing the removal of connections with the smallest absolute weight values .
2.5. Training Protocol
2.6. Order Parameters and Critical Thresholds
3. Results
3.1. Phase 1: Emergence of the Heavy-Tailed Skeleton
3.2. Phase 2: The Connectivity Plateau
3.3. Phase 3: The Zombie Network and Functional Decoupling
3.4. Impact of Structural Complexity: The Depth Fragility Paradox
4. Discussion
4.1. Scale-Free Universality and the Functional Skeleton
4.2. The Zombie Phase and Depth-Induced Fragility
4.3. Implications for Hardware Design and Limitations
- Experiments were conducted on a single dataset (Fashion-MNIST) and three MLP architectures. Modern architectures (e.g., Transformers, ResNets, convolutional networks with attention) may exhibit different behavior.
- The signal decay model ignored nonlinear activations.
- Biases and normalization parameters were excluded from pruning.
- Formal statistical tests of the power-law hypothesis for weight distributions are not provided.
- Critical exponents and finite-size scaling analysis, standard in statistical physics, were not derived.
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436–444. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zdeborová, L. Understanding deep learning is also a job for physicists. Nat. Phys. 2020, 16, 602–604. [Google Scholar] [CrossRef] [Scilit]
- Tumminello, M.; Aste, T.; Di Matteo, T.; Mantegna, R.N. A tool for filtering information in complex systems. Proc. Natl. Acad. Sci. USA 2005, 102, 10421–10426. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mantegna, R.N. Hierarchical structure in financial markets. Eur. Phys. J. B-Condens. Matter Complex Syst. 1999, 11, 193–197. [Google Scholar] [CrossRef] [Scilit]
- Bonanno, G.; Caldarelli, G.; Lillo, F.; Mantegna, R.N. Topology of correlation-based minimal spanning trees in real and model markets. Phys. Rev. E 2003, 68, 046130. [Google Scholar] [CrossRef] [Scilit]
- Garas, A.; Argyrakis, P.; Havlin, S. The structural role of weak and strong links in a financial market network. Eur. Phys. J. B 2008, 63, 265–271. [Google Scholar] [CrossRef] [Scilit]
- Watts, D.J.; Strogatz, S.H. Collective dynamics of ‘small-world’ networks. Nature 1998, 393, 440–442. [Google Scholar] [CrossRef] [Scilit]
- Barabási, A.L.; Albert, R. Emergence of scaling in random networks. Science 1999, 286, 509–512. [Google Scholar] [CrossRef] [Scilit]
- Newman, M.E. The structure and function of complex networks. SIAM Rev. 2003, 45, 167–256. [Google Scholar] [CrossRef] [Scilit]
- Onnela, J.P.; Saramäki, J.; Hyvönen, J.; Szabó, G.; Argollo de Menezes, M.; Kaski, K.; Kertész, J. Analysis of a large-scale weighted network of one-to-one human communication. New J. Phys. 2007, 9, 179. [Google Scholar] [CrossRef] [Scilit]
- Bellingeri, M.; Bevacqua, D.; Scotognella, F.; Alfieri, R.; Nguyen, Q.; Montepietra, D.; Cassi, D. Link and node removal in real social networks: A review. Front. Phys. 2020, 8, 228. [Google Scholar] [CrossRef] [Scilit]
- Newman, M. Networks; Oxford University Press: Oxford, UK, 2018. [Google Scholar]
- Albert, R.; Jeong, H.; Barabási, A.L. Error and attack tolerance of complex networks. Nature 2000, 406, 378–382. [Google Scholar] [CrossRef] [Scilit]
- Callaway, D.S.; Newman, M.E.; Strogatz, S.H.; Watts, D.J. Network robustness and fragility: Percolation on random graphs. Phys. Rev. Lett. 2000, 85, 5468. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kitsak, M.; Ganin, A.A.; Alderson, D.L.; Fowler, M.S.; Linkov, I. Stability of a giant connected component. Phys. Rev. E 2018, 97, 012309. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Stauffer, D.; Aharony, A. Introduction to Percolation Theory; CRC Press: Boca Raton, FL, USA, 2018. [Google Scholar]
- Cohen, R.; Erez, K.; Ben-Avraham, D.; Havlin, S. Resilience of the Internet to random breakdowns. Phys. Rev. Lett. 2000, 85, 4626. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pajevic, S.; Plenz, D. The organization of strong links in complex networks. Nat. Phys. 2012, 8, 429–436. [Google Scholar] [CrossRef] [Scilit]
- Nguyen, N.K.K.; Nguyen, Q.; Pham, H.; Tr, N.T.; Alfieri, R.; Cassi, D.; Bellingeri, M. Analytics solution for the robustness of incomplete information scale-free networks. Phys. Lett. A 2025, 565, 131123. [Google Scholar] [CrossRef] [Scilit]
- Barrat, A.; Barthelemy, M.; Pastor-Satorras, R.; Vespignani, A. The architecture of complex weighted networks. Proc. Natl. Acad. Sci. USA 2004, 101, 3747–3752. [Google Scholar] [CrossRef] [Scilit]
- Onnela, J.P.; Saramäki, J.; Hyvönen, J.; Szabó, G.; Lazer, D.; Kaski, K.; Kertész, J.; Barabási, A.L. Structure and tie strengths in mobile communication networks. Proc. Natl. Acad. Sci. USA 2007, 104, 7332–7336. [Google Scholar] [CrossRef] [Scilit]
- Bellingeri, M.; Bevacqua, D.; Sartori, F.; Turchetto, M.; Scotognella, F.; Alfieri, R.; Nguyen, N.; Le, T.; Nguyen, Q.; Cassi, D. Considering weights in real social networks: A review. Front. Phys. 2023, 11, 1152243. [Google Scholar] [CrossRef] [Scilit]
- Pastor-Satorras, R.; Vespignani, A. Epidemic spreading in scale-free networks. Phys. Rev. Lett. 2001, 86, 3200. [Google Scholar] [CrossRef] [Scilit]
- Newman, M.E. Spread of epidemic disease on networks. Phys. Rev. E 2002, 66, 016128. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, Y.; Paul, G.; Havlin, S.; Liljeros, F.; Stanley, H. Finding a Better Immunization Strategy. Phys. Rev. Lett. 2008, 101, 058701. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- LeCun, Y.; Denker, J.; Solla, S. Optimal brain damage. In Advances in Neural Information Processing Systems; MIT Press: Cambridge, MA, USA, 1989; Volume 2. [Google Scholar]
- Akhtar, N.; Mian, A. Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey. IEEE Access 2018, 6, 14410–14430. [Google Scholar] [CrossRef] [Scilit]
- Yuan, X.; He, P.; Zhu, Q.; Li, X. Adversarial Examples: Attacks and Defenses for Deep Learning. IEEE Trans. Neural Netw. Learn. Syst. 2019, 30, 2805–2824. [Google Scholar] [CrossRef] [Scilit]
- Han, S.; Pool, J.; Tran, J.; Dally, W. Learning both weights and connections for efficient neural network. In Advances in Neural Information Processing Systems (NeurIPS); MIT Press: Cambridge, MA, USA, 2015; Volume 28. [Google Scholar]
- Lazarevich, I.; Kozlov, A.; Malinin, N. Post-training deep neural network pruning via layer-wise calibration. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, BC, Canada, 11–17 October 2021. [Google Scholar]
- Frantar, E.; Alistarh, D. SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot. In Proceedings of the 40th International Conference on Machine Learning (ICML), PMLR 2023, Honolulu, HI, USA, 23–29 July 2023; Volume 202, pp. 10323–10337. [Google Scholar]
- Nguyen, N.K.K.; Nguyen, Q.; Pham, H.H.; Le, T.T.; Nguyen, T.M.; Cassi, D.; Scotognella, F.; Alfieri, R.; Bellingeri, M. Predicting the Robustness of Large Real-World Social Networks Using a Machine Learning Model. Complexity 2022, 2022, 3616163. [Google Scholar] [CrossRef] [Scilit]
- Frankle, J.; Carbin, M. The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks. In Proceedings of the International Conference on Learning Representations (ICLR), New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Tanaka, H.; Kunin, D.; Yamins, D.L.; Ganguli, S. Pruning neural networks without any data by iteratively conserving synaptic flow. In Advances in Neural Information Processing Systems (NeurIPS); MIT Press: Cambridge, MA, USA, 2020. [Google Scholar]
- Xiao, H.; Rasul, K.; Vollgraf, R. Fashion-MNIST: A Novel Image Dataset for Benchmarking Machine Learning Algorithms. arXiv 2017, arXiv:1708.07747. [Google Scholar] [CrossRef] [Scilit]
- Kingma, D.P.; Ba, J. Adam: A Method for Stochastic Optimization. arXiv 2014, arXiv:1412.6980. [Google Scholar]
- Goodfellow, I.; Bengio, Y.; Courville, A. Deep Learning; MIT Press: Cambridge, MA, USA, 2016. [Google Scholar]
- Albert, R.; Barabási, A.L. Statistical mechanics of complex networks. Rev. Mod. Phys. 2002, 74, 47. [Google Scholar] [CrossRef] [Scilit]
- Clauset, A.; Shalizi, C.R.; Newman, M.E. Power-law distributions in empirical data. SIAM Rev. 2009, 51, 661–703. [Google Scholar] [CrossRef] [Scilit]
- Li, H.; Xu, Z.; Taylor, G.; Studer, C.; Goldstein, T. Visualizing the loss landscape of neural nets. In Advances in Neural Information Processing Systems (NeurIPS); MIT Press: Cambridge, MA, USA, 2018. [Google Scholar]



| Symbol | Definition |
|---|---|
| Graph representation of a neural network with nodes V (neurons) and edges E (synaptic weights) | |
| W | Weight matrix of a fully connected layer |
| M | Binary pruning mask; |
| p | Edge pruning ratio (fraction of weights removed) |
| Normalized size of the giant connected component (GCC) | |
| Topological threshold: p at which drops below 0.9 | |
| Functional threshold: p at which test accuracy drops below 50% of baseline | |
| Test accuracy as a function of pruning ratio p | |
| L | Number of hidden layers (network depth) |
| Activation vector at layer l | |
| Expected squared signal norm at the output layer |
| Architecture | Config | Initial Acc. | Strategy | (Func.) | (Topo.) |
|---|---|---|---|---|---|
| Shallow-MLP | Narrow/Shallow | ∼89.0% | Magnitude | 0.95 | 0.94 |
| Random | 0.68 | 0.98 | |||
| Deep-MLP | Narrow/Deep | ∼89.0% | Magnitude | 0.85 | 0.94 |
| Random | 0.35 | 0.98 | |||
| Wide-MLP | Wide/Deep | ∼89.0% | Magnitude | 0.92 | 0.95 |
| Random | 0.60 | 0.97 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Nguyen, Q.; Pham, H.H.; Cassi, D.; Bellingeri, M. Depth Fragility and Skeletal Universality: Decoupling Topology and Function in Deep Neural Networks. Mathematics 2026, 14, 1438. https://doi.org/10.3390/math14091438
Nguyen Q, Pham HH, Cassi D, Bellingeri M. Depth Fragility and Skeletal Universality: Decoupling Topology and Function in Deep Neural Networks. Mathematics. 2026; 14(9):1438. https://doi.org/10.3390/math14091438
Chicago/Turabian StyleNguyen, Quang, Hai Ha Pham, Davide Cassi, and Michele Bellingeri. 2026. "Depth Fragility and Skeletal Universality: Decoupling Topology and Function in Deep Neural Networks" Mathematics 14, no. 9: 1438. https://doi.org/10.3390/math14091438
APA StyleNguyen, Q., Pham, H. H., Cassi, D., & Bellingeri, M. (2026). Depth Fragility and Skeletal Universality: Decoupling Topology and Function in Deep Neural Networks. Mathematics, 14(9), 1438. https://doi.org/10.3390/math14091438

