Fusing Semantic and Structural Features for Code Error Detection
Abstract
1. Introduction
- We propose a novel framework that integrates LLMs and GNN for processing code and detecting multiple types of errors. This method fully utilizes the powerful semantic understanding ability of LLM and the structural reasoning ability of GNN, which can simultaneously capture the semantic context and structural patterns in the code, achieving accurate detection of diverse error types.
- We systematically studied different fusion strategies of GNN and LLMs, and deeply analyzed the impact of various fusion methods on the structural representation ability of the model. The research results indicate that designing an effective fusion mechanism plays a key role in leveraging the advantages of semantic and structural information and improving the effectiveness of code analysis.
- We introduced the PytraceBugs dataset for code error classification and verified through extensive experiments that our proposed model with cross-modal fusion outperforms other models, including model that relies solely on RoBERTa. The experimental results show that a reasonable fusion strategy can improve the understanding ability of code information.
2. Related Work
2.1. Code Error Detection
2.2. LLMs in Code Error Detection
2.3. GNNs in Code Error Detection
2.4. Integration of LLMs and GNNs
3. Methodology
3.1. Network Overview
3.1.1. Semantic Feature Extraction with RoBERTa
3.1.2. Structural Feature Extraction with GNN
3.1.3. Feature Fusion and Classification
3.2. Loss Function and Class Imbalance Handling
4. Experiment
4.1. Experimental Setup
4.1.1. Dataset
4.1.2. Implemetation Details
4.2. Evaluation Matrices
4.3. Experimental Results
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Devlin, J.; Chang, M.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv 2018, arXiv:1810.04805. [Google Scholar]
- OpenAI. GPT-4 Technical Report. arXiv 2023, arXiv:2303.08774. [Google Scholar] [CrossRef]
- Ray, B.; Posnett, D.; Filkov, V.; Devanbu, P. A Large Scale Study of Programming Languages and Code Quality in GitHub. In Proceedings of the 22nd ACM SIGSOFT International Symposium on the Foundations of Software Engineering (FSE), Seattle, WA, USA, 13–18 November 2016; Volume 22, pp. 1–12. [Google Scholar]
- White, M.; Vendome, C.; Linares-Vásquez, M.; Poshyvanyk, D. Toward Deep Learning Software Repositories. In Proceedings of the 12th Working Conference on Mining Software Repositories (MSR), Florence, Italy, 16–17 May 2015; Volume 12, pp. 1–10. [Google Scholar]
- Pradel, M.; Sen, K. DeepBugs: A Learning Approach to Name-Based Bug Detection. Proc. ACM Program. Lang. 2018, 2, 1–25. [Google Scholar] [CrossRef]
- Sun, C.; Huang, X.; Lo, D.; Liu, X. Bugbuster: Identifying and Fixing Software Bugs Using Semantic Learning. In Proceedings of the 43rd International Conference on Software Engineering (ICSE), Virtually, 23–29 May 2021; Volume 43, pp. 1–15. [Google Scholar]
- Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Zhang, X.; Qiu, X.; Fan, A.; Tolias, A.; Wenzek, G.; Cheng, H.; et al. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv 2019, arXiv:1907.11692. [Google Scholar]
- Feng, Z.; Guo, D.; Tang, D.; Duan, N.; Feng, X.; Gong, M.; Shou, L.; Qin, B.; Liu, T.; Jiang, D.; et al. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, 16–20 November 2020; pp. 1536–1547. [Google Scholar]
- Chen, H.; Zhang, Y.; Han, X.; Rong, H.; Zhang, Y.; Mao, T.; Zhang, H.; Wang, X.; Xing, L.; Chen, X. WitheredLeaf: Finding Entity-Inconsistency Bugs with LLMs. arXiv 2024, arXiv:2405.01668. [Google Scholar]
- Kim, Y. Convolutional Neural Networks for Sentence Classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar, 25–29 October 2014; Volume 1, pp. 1746–1751. [Google Scholar]
- Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C.D.; Ng, A.Y.; Potts, C. Recursive Deep Models for Semantic Compositionality over a Sentiment Treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP), Seattle, WA, USA, 18–21 October 2013; Volume 2, pp. 1631–1642. [Google Scholar]
- Bahdanau, D.; Cho, K.; Bengio, Y. Neural Machine Translation by Jointly Learning to Align and Translate. In Proceedings of the 3rd International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015; Volume 3, pp. 1–15. [Google Scholar]
- Chen, D.; Fisch, A.; Weston, J.; Bordes, A.; Buchwalter, W.; Weston, R.; Gardner, M. Reading Wikipedia to Answer Open-Domain Questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL), Vancouver, BC, Canada, 30 July–4 August 2017; Volume 55, pp. 1870–1879. [Google Scholar]
- Austin, T.; Smith, X.; Brown, Y. Static Analysis for Software Bug Detection: A Review. ACM Comput. Surv. 2018, 51, 1–45. [Google Scholar]
- Zhang, Y.; Zhao, J. Dynamic Debugging Techniques for Error Detection. Int. J. Softw. Eng. Appl. 2017, 8, 15–25. [Google Scholar]
- Jiang, L.; Zhou, Y. A Survey of Static Analysis Techniques for Software Bug Detection. ACM Comput. Surv. 2011, 43, 1–37. [Google Scholar]
- Hinton, G.E.; Osindero, S.; Teh, Y.W. A Fast Learning Algorithm for Deep Belief Nets. Neural Comput. 2006, 18, 1527–1554. [Google Scholar] [CrossRef]
- LeCun, Y.; Boser, B.; Denker, J.S.; Henderson, D.; Howard, R.E.; Hubbard, W.; Jackel, L.D. Backpropagation Applied to Handwritten Zip Code Recognition. Neural Comput. 1989, 1, 541–551. [Google Scholar] [CrossRef]
- LeCun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-Based Learning Applied to Document Recognition. Proc. IEEE 1998, 86, 2278–2324. [Google Scholar] [CrossRef]
- Elman, J.L. Finding Structure in Time. Cogn. Sci. 1990, 14, 179–211. [Google Scholar] [CrossRef]
- Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [PubMed]
- Anthropic. Claude: An AI Assistant for Thoughtful Conversations. In Anthropic AI Blog; Anthropic: San Francisco, CA, USA, 2023. [Google Scholar]
- Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H.W.; Sutton, C.; Gehrmann, S.; et al. PaLM: Scaling Language Modeling with Pathways. arXiv 2022, arXiv:2204.02311. [Google Scholar] [CrossRef]
- Touvron, H.; Martin, J.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, R.; Bhosale, S.; et al. LLaMA: Open and Efficient Foundation Language Models. arXiv 2023, arXiv:2302.13971. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.A.; Kaiser, Ł.; Polosukhin, I. Attention is All You Need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA, 4–9 December 2017; pp. 5998–6008. [Google Scholar]
- Alrashedy, K. Language Models are Better Bug Detectors Through Code-Pair Classification. arXiv 2023, arXiv:2311.07957. [Google Scholar]
- Scarselli, F.; Gori, M.; Tsoi, A.; Hagenbuchner, M.; Monfardini, G. The Graph Neural Network Model. IEEE Trans. Neural Netw. 2009, 20, 61–80. [Google Scholar] [CrossRef]
- Xu, K.; Hu, W.; Leskovec, J.; Jegelka, S. How Powerful are Graph Neural Networks? arXiv 2018, arXiv:1810.00826. [Google Scholar]
- Zhou, Y.; Liu, S.; Siow, J.; Du, X.; Liu, Y. Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural Networks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 8–14 December 2019; Volume 32. [Google Scholar]
- Kipf, T.N.; Welling, M. Semi-Supervised Classification with Graph Convolutional Networks. In Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France, 24–26 April 2017. [Google Scholar]
- Fraser, G.; Arcuri, A. EvoSuite: A Search-Based Unit Test Generation Tool for Java. In Proceedings of the 2012 ACM International Symposium on Software Testing and Analysis (ISSTA), Minneapolis, MN, USA, 15–20 July 2012; pp. 234–244. [Google Scholar]
- Aho, A.V.; Lam, M.S.; Sethi, R.; Ullman, J.D. Compilers: Principles, Techniques, and Tools, 2nd ed.; Addison-Wesley: Boston, MA, USA, 1986. [Google Scholar]
- Briem, J.A.; Smit, J.; Sellik, H.; Rapoport, P. Using Distributed Representation of Code for Bug Detection. arXiv 2019, arXiv:1911.12863. [Google Scholar] [CrossRef]
- Ferrante, J.; Ottenstein, K.J.; Warren, J.D. The Program Dependence Graph and Its Use in Optimization. ACM Trans. Program. Lang. Syst. 1987, 9, 319–349. [Google Scholar] [CrossRef]
- Lam, A.N.; Nguyen, A.T.; Nguyen, H.A.; Nguyen, T.N. Bug Localization with Combination of Deep Learning and Information Retrieval. In Proceedings of the 2017 IEEE/ACM 25th International Conference on Program Comprehension (ICPC), Buenos Aires, Argentina, 22–23 May 2017; pp. 218–229. [Google Scholar]
- Brockington, M.A. Control Flow Graphs for Flow of Control Analysis. ACM SIGPLAN Not. 1970, 5, 88–97. [Google Scholar]
- Salton, G.; Buckley, C. Term-Weighting Approaches in Automatic Text Retrieval. In Information Processing & Management; Elsevier: Amsterdam, The Netherlands, 1988; pp. 171–180. [Google Scholar]
- Zhang, C.; Yang, Q. A Generalized Cross-Entropy Loss for Imbalanced Classification. In Proceedings of the 2018 IEEE International Conference on Data Mining (ICDM), Singapore, 17–20 November 2018; pp. 800–809. [Google Scholar]
- Akimova, E.N.; Bersenev, A.Y.; Deikov, A.A.; Kobylkin, K.S.; Konygin, A.V.; Mezentsev, I.P.; Misilov, V.E. PyTraceBugs: A Large Python Code Dataset for Supervised Machine Learning in Software Defect Prediction. In Proceedings of the 2021 28th Asia-Pacific Software Engineering Conference (APSEC), Taipei, Taiwan, 6–9 December 2021; pp. 141–151. [Google Scholar]
- Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. arXiv 2017, arXiv:1711.05101. [Google Scholar]




| Model | Test Accuracy | Macro-Averaged F1 Score |
|---|---|---|
| RoBERTa | 69.35 | 57.34 |
| RoBERTa with Graph Neural Network (Direct Fusion) | 68.79 | 52.29 |
| RoBERTa with Graph Neural Network (Channel Attention) | 69.01 | 53.33 |
| RoBERTa with Graph Neural Network (Cross-Modal Fusion) | 71.06 | 57.27 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Zhang, Y.; Liu, W.; Jiang, F.; Ma, J.; Cao, J. Fusing Semantic and Structural Features for Code Error Detection. Entropy 2025, 27, 1229. https://doi.org/10.3390/e27121229
Zhang Y, Liu W, Jiang F, Ma J, Cao J. Fusing Semantic and Structural Features for Code Error Detection. Entropy. 2025; 27(12):1229. https://doi.org/10.3390/e27121229
Chicago/Turabian StyleZhang, Yiwen, Wei Liu, Fazhong Jiang, Jiquan Ma, and Jingtai Cao. 2025. "Fusing Semantic and Structural Features for Code Error Detection" Entropy 27, no. 12: 1229. https://doi.org/10.3390/e27121229
APA StyleZhang, Y., Liu, W., Jiang, F., Ma, J., & Cao, J. (2025). Fusing Semantic and Structural Features for Code Error Detection. Entropy, 27(12), 1229. https://doi.org/10.3390/e27121229

