Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance
Abstract
1. Introduction
- We introduce a robust feature extraction pipeline based on the spectral analysis of layer activation covariance matrices, covering 14 key internal components of the AASIST3 architecture.
- We develop a meta-classification and attribution methodology using CatBoost v1.2.8 and TreeSHAP v0.3.1 [19] to quantify the contribution share of each processing branch (B0–B3) and global module (GAT-S, GAT-T).
- We define four operational archetypes—Effective Specialization, Effective Consensus, Ineffective Consensus, and Flawed Specialization—to classify the model’s behavior.
- We empirically demonstrate that AASIST3 dynamically adapts its strategy for different attacks (A07–A19) and identify a structural misalignment where the model confidently relies on an incorrect branch, leading to high error rates.
2. Materials and Methods
2.1. AASIST3 Architecture Components
- (1)
- Heterogeneous Stacking Graph Attention Layers (HSGAL): These layers form the computational core of the parallel branches. They employ graph attention mechanisms to capture complex, non-local spectro-temporal patterns within the audio data. We analyze the early-stage (HSGAL1) and late-stage (HSGAL2) layers across all four branches: B0-HSGAL1, B0-HSGAL2, B1-HSGAL1, B1-HSGAL2, B2-HSGAL1, B2-HSGAL2, and B3-HSGAL1, B3-HSGAL2.
- (2)
- Pooling Layers (Pool): Each branch includes a pooling operation for feature aggregation and dimensionality reduction. We analyze these as: B0-Pool, B1-Pool, B2-Pool, and B3-Pool.
- (3)
- Global Graph Attention Networks (GAT): Two global modules operate on the multidimensional features to capture holistic dependencies. GAT-S (Spectral) models relationships across frequency bins, while GAT-T (Temporal) models dependencies across time frames.
2.2. Methodology Pipeline
Confidence Interval Estimation for Reported Metrics
2.3. Justification of the Analysis Method
3. Results
3.1. Eigenvalue Count Ablation
3.2. Penalty Function Ablation
3.3. Overview of Model Performance and Internal Strategies
3.4. Correlation and Variance Analysis of Operational Archetypes
3.5. Detailed Per-Attack Analysis
3.6. Interpretation Framework for SHAP Analysis
3.6.1. Attacks A07 and A08 (Consensus Strategies)
3.6.2. Attacks A09 and A10 (Specialization vs. Complexity)
3.6.3. Attacks A11 and A12 (Spectral vs. Mixed)
3.6.4. Attacks A13 and A14 (Borderline vs. Distinct)
3.6.5. Attacks A15 and A16 (Weak vs. Robust Consensus)
3.6.6. Attacks A17 and A18 (Vulnerability and Failure)
3.6.7. Attack A19 (Global Spectral Detection)
3.7. Detailed Statistical Data
3.8. Single-Branch Retention Ablation
4. Discussion
4.1. Analysis of Operational Archetypes
4.2. Architectural Implications and Vulnerabilities
4.3. Connection to Broader Research
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| AASIST | Audio Anti-Spoofing using Integrated Spectro-Temporal GNNs |
| EER | Equal Error Rate |
| GAT | Graph Attention Network |
| HSGAL | Heterogeneous Stacking Graph Attention Layers |
| SHAP | SHapley Additive exPlanations |
References
- Yamagishi, J.; Wang, X.; Todisco, M.; Sahidullah, M.; Patino, J.; Nautsch, A.; Liu, X.; Lee, K.A.; Kinnunen, T.; Evans, N.; et al. ASVspoof 2021: Accelerating progress in spoofed and deepfake speech detection. In Proceedings of the 2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge, Online, 16 September 2021; pp. 47–54. [Google Scholar] [CrossRef]
- Wang, X.; Delgado, H.; Tak, H.; Jung, J.-W.; Shim, H.-J.; Todisco, M.; Kukanov, I.; Liu, X.; Sahidullah, M.; Kinnunen, T.H.; et al. ASVspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale. In Proceedings of the Automatic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024), Kos, Greece, 31 August 2024; pp. 1–8. [Google Scholar] [CrossRef]
- Wang, X.; Yamagishi, J.; Todisco, M.; Delgado, H.; Nautsch, A.; Evans, N.; Sahidullah, M.; Vestman, V.; Kinnunen, T.; Lee, K.A.; et al. ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech. Comput. Speech Lang. 2020, 64, 101114. [Google Scholar] [CrossRef]
- Borodin, K.; Kudryavtsev, V.; Mkrtchian, G.; Gorodnichev, M. Capsule-based and TCN-based Approaches for Spoofing Detection in Voice Biometry. Eng. Technol. Appl. Sci. Res. 2024, 14, 18409–18414. [Google Scholar] [CrossRef]
- Borodin, K.; Kudryavtsev, V.; Korzh, D.; Efimenko, A.; Mkrtchian, G.; Gorodnichev, M.; Rogov, O.Y. AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge. arXiv 2024, arXiv:2408.17352. [Google Scholar] [CrossRef]
- Kinnunen, T.; Lee, K.A.; Delgado, H.; Evans, N.; Todisco, M.; Sahidullah, M.; Yamagishi, J.; Reynolds, D.A. t-DCF: A Detection Cost Function for the Tandem Assessment of Spoofing Countermeasures and Automatic Speaker Verification. arXiv 2019, arXiv:1804.09618. [Google Scholar] [CrossRef]
- Sundararajan, M.; Taly, A.; Yan, Q. Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning, ICML’17, Sydney, Australia, 6–11 August 2017; Volume 70, pp. 3319–3328. [Google Scholar]
- Jung, J.W.; Heo, H.S.; Tak, H.; Shim, H.J.; Chung, J.S.; Lee, B.J.; Yu, H.J.; Evans, N. AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks. In Proceedings of the ICASSP 2022—2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, 23–27 May 2022; pp. 6367–6371. [Google Scholar] [CrossRef]
- Hu, Y.; Sompolinsky, H. The spectrum of covariance matrices of randomly connected recurrent neuronal networks with linear dynamics. PLoS Comput. Biol. 2022, 18, e1010327. [Google Scholar] [CrossRef] [PubMed]
- Sihag, S.; Mateos, G.; McMillan, C.; Ribeiro, A. coVariance Neural Networks. arXiv 2023, arXiv:2205.15856. [Google Scholar] [PubMed]
- Binkowski, J.; Janiak, D.; Sawczyn, A.; Gabrys, B.; Kajdanowicz, T. Hallucination Detection in LLMs Using Spectral Features of Attention Maps. arXiv 2025, arXiv:2502.17598. [Google Scholar] [CrossRef]
- Harzli, O.E.; Grau, B.C. Adversarial Attacks as Near-Zero Eigenvalues in the Empirical Kernel of Neural Networks. In Proceedings of the NeurIPS 2024 Workshop on Mathematics of Modern Machine Learning (M3L), Vancouver, BC, Canada, 14 December 2024. [Google Scholar]
- Lundberg, S.; Lee, S.I. A Unified Approach to Interpreting Model Predictions. arXiv 2017, arXiv:1705.07874. [Google Scholar] [CrossRef]
- Ge, W.; Patino, J.; Todisco, M.; Evans, N. Explaining deep learning models for spoofing and deepfake detection with SHapley Additive exPlanations. arXiv 2024, arXiv:2110.03309. [Google Scholar] [CrossRef]
- Yu, N.; Chen, L.; Leng, T.; Chen, Z.; Yi, X. An explainable deepfake of speech detection method with spectrograms and waveforms. J. Inf. Secur. Appl. 2024, 81, 103720. [Google Scholar] [CrossRef]
- Li, M.; Ahmadiadli, Y.; Zhang, X.P. A Survey on Speech Deepfake Detection. arXiv 2025, arXiv:2404.13914. [Google Scholar] [CrossRef]
- Pomponi, J.; Scardapane, S.; Uncini, A. A Probabilistic Re-Interpretation of Confidence Scores in Multi-Exit Models. Entropy 2021, 24, 1. [Google Scholar] [CrossRef] [PubMed]
- Heidemann, L.; Schwaiger, A.; Roscher, K. Measuring Ensemble Diversity and Its Effects on Model Robustness. In Proceedings of the 1st International Workshop on Artificial Intelligence Safety (SafeAI 2021) Co-Located with AAAI 2021, CEUR-WS, Virtually, 8 February 2021; Volume 2916, pp. 65–73. [Google Scholar]
- Lundberg, S.M.; Erion, G.; Chen, H.; DeGrave, A.; Prutkin, J.M.; Nair, B.; Katz, R.; Himmelfarb, J.; Bansal, N.; Lee, S.I. Explainable AI for Trees: From Local Explanations to Global Understanding. arXiv 2019, arXiv:1905.04610. [Google Scholar] [CrossRef] [PubMed]
- Hajjouz, A.; Avksentieva, E. Enhancing and extending CatBoost for accurate detection and classification of DoS and DDoS attack subtypes in network traffic. Sci. Tech. J. Inf. Technol. Mech. Opt. 2025, 25, 114–127. [Google Scholar] [CrossRef]
- Shen, J.; Pang, R.; Weiss, R.J.; Schuster, M.; Jaitly, N.; Yang, Z.; Chen, Z.; Zhang, Y.; Wang, Y.; Skerrv-Ryan, R.; et al. Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions. In Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada, 15–20 April 2018; pp. 4779–4783. [Google Scholar] [CrossRef]
- Siuzdak, H. Vocos: Closing the Gap Between Time-Domain and Fourier-Based Neural Vocoders for High-Quality Audio Synthesis. In Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Pons, J.; Pascual, S.; Cengarle, G.; Serrà, J. Upsampling Artifacts in Neural Audio Synthesis. In Proceedings of the ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing, Toronto, ON, Canada, 6–11 June 2021; pp. 3005–3009. [Google Scholar] [CrossRef]
- Tak, H.; Jung, J.-w.; Patino, J.; Kamble, M.; Todisco, M.; Evans, N. End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detection. In Proceedings of the 2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge, Online, 16 September 2021; pp. 1–8. [Google Scholar] [CrossRef]
- Zhang, Q.; Long, Y.; Cai, H.; Yu, S.; Shi, Y.; Tan, X. A multi-slice attention fusion and multi-view personalized fusion lightweight network for Alzheimer’s disease diagnosis. BMC Med. Imaging 2024, 24, 258. [Google Scholar] [CrossRef] [PubMed]
- Ben-Artzy, A.; Schwartz, R. Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers. arXiv 2024, arXiv:2409.03621. [Google Scholar] [CrossRef]
- Tak, H.; Patino, J.; Todisco, M.; Nautsch, A.; Evans, N.; Larcher, A. End-to-End anti-spoofing with RawNet2. In Proceedings of the ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing, Toronto, ON, Canada, 6–11 June 2021; pp. 6369–6373. [Google Scholar] [CrossRef]
- Rosello, V.; Evans, N. A Conformer-Based Classifier for Variable-Length Utterance Processing in Anti-Spoofing. In Proceedings of the INTERSPEECH, Dublin, Ireland, 20–24 August 2023; pp. 3632–3636. [Google Scholar] [CrossRef]
- D’Alterio, G.; Neghina, M.; Bestagini, P.; Tubaro, S. Attention-based Mixture of Experts for Robust Speech Deepfake Detection. In Proceedings of the IEEE International Workshop on Information Forensics and Security (WIFS), Perth, WA, Australia, 1–4 December 2025. [Google Scholar]
- Tran, H.M.; Amsaleg, L.; Ducq, E. Multi-level SSL Feature Gating for Audio Deepfake Detection. In Proceedings of the ACM International Conference on Multimedia Retrieval (ICMR), Chicago, IL, USA, 30 June–3 July 2025; ACM: New York, NY, USA, 2025. [Google Scholar]


























| Attack | EER (%) | Dominant Component(s) | Dominant Share (%) | Confidence Score | Identified Strategy |
|---|---|---|---|---|---|
| A09 | B2, B1 | , | , | Effective Specialization | |
| A14 | B2, B0 | , | , | Effective Specialization | |
| A07 | B2, B1 | , | , | Effective Specialization | |
| A11 | B0, B2 | , | , | Effective Consensus | |
| A16 | B2, B1 | , | , | Effective Consensus | |
| A19 | B1, B2 | , | , | Effective Specialization | |
| A13 | B1, B2 | , | , | Ineffective Specialization | |
| A15 | B2, B1 | , | , | Ineffective Consensus | |
| A08 | B1, B0 | , | , | Ineffective Specialization | |
| A12 | B1, B0 | , | , | Ineffective Specialization | |
| A17 | B1, B2 | , | , | Flawed Specialization (Vulnerability) | |
| A10 | B2, B1 | , | , | Ineffective Specialization | |
| A18 | B1, B2 | , | , | Flawed Specialization (Vulnerability) |
| Identified Strategy | Mean | Std | Var | Count |
|---|---|---|---|---|
| Effective Consensus | 18.76 | 0.63 | 0.39 | 3 |
| Effective Specialization | 23.18 | 2.00 | 4.02 | 3 |
| Flawed Specialization | 23.81 | 2.51 | 6.29 | 3 |
| Ineffective Consensus | 18.84 | 0.53 | 0.29 | 4 |
| Attack | ||||||
|---|---|---|---|---|---|---|
| A07 | ||||||
| A08 | ||||||
| A09 | ||||||
| A10 | ||||||
| A11 | ||||||
| A12 | ||||||
| A13 | ||||||
| A14 | ||||||
| A15 | ||||||
| A16 | ||||||
| A17 | ||||||
| A18 | ||||||
| A19 |
| Attack | B0 Share | B1 Share | B2 Share | B3 Share | GAT-S Share | GAT-T Share |
|---|---|---|---|---|---|---|
| A07 | ||||||
| A08 | ||||||
| A09 | ||||||
| A10 | ||||||
| A11 | ||||||
| A12 | ||||||
| A13 | ||||||
| A14 | ||||||
| A15 | ||||||
| A16 | ||||||
| A17 | ||||||
| A18 | ||||||
| A19 |
| Attack | Dominant Branch Kept | Baseline EER | EER with Only Dominant Branch |
|---|---|---|---|
| A18 | B0 | 28.63 | 63.34 |
| A17 | B1 | 14.26 | 66.81 |
| A10 | B2 | 17.31 | 67.55 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Viakhirev, I.; Borodin, K.; Gorodnichev, M.; Mkrtchian, G. Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance. Mathematics 2026, 14, 381. https://doi.org/10.3390/math14020381
Viakhirev I, Borodin K, Gorodnichev M, Mkrtchian G. Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance. Mathematics. 2026; 14(2):381. https://doi.org/10.3390/math14020381
Chicago/Turabian StyleViakhirev, Ivan, Kirill Borodin, Mikhail Gorodnichev, and Grach Mkrtchian. 2026. "Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance" Mathematics 14, no. 2: 381. https://doi.org/10.3390/math14020381
APA StyleViakhirev, I., Borodin, K., Gorodnichev, M., & Mkrtchian, G. (2026). Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance. Mathematics, 14(2), 381. https://doi.org/10.3390/math14020381

