Software Fault Localization Approach with Coverage Matrix Optimization Boosted by LLM-Based Code Naturalness
Abstract
1. Introduction
- Novel methodology. We propose CNFL, a fault localization approach that leverages code naturalness quantified by LLMs to optimize the coverage matrix. Unlike conventional SBFL techniques that treat all executed program elements uniformly, CNFL differentiates elements based on their naturalness scores, effectively suppressing noise and highlighting fault-relevant information. CNFL preserves the standard SBFL workflow, ensuring seamless integration with existing fault localization techniques while introducing a fundamentally new optimization dimension.
- Comprehensive empirical evaluation. We conduct extensive experiments on the Defects4J benchmark, comparing CNFL against five representative SBFL baselines and two state-of-the-art fault localization approaches. The experimental design explicitly isolates the impact of code-naturalness-based optimization from other confounding factors, providing a fair and rigorous assessment.
- New insights into the effectiveness and efficiency. We systematically analyze the experimental results to uncover when and why CNFL provides significant gains. Our findings reveal that CNFL not only outperforms traditional SBFL methods, but also exhibits consistent superiority over other fault localization approaches that focus on coverage matrix optimization. Notably, the findings also show the impacts of LLMs on the effectiveness and efficiency of CNFL.
2. Background
2.1. Spectrum-Based Fault Localization
- -
- : The number of failing test cases that execute the program element.
- -
- : The number of passing test cases that execute the program element.
- -
- : The number of failing test cases that do not execute program element.
- -
- : The number of passing test cases that do not execute program element.
- -
- : The total number of failing test cases in the test suite.
- -
- : The total number of passing test cases in the test suite.
- Jaccard [11] aims to quantify the suspiciousness of a program element by measuring the proportion of failing test cases in which the element is executed among all relevant test cases. The formula is defined as follows.
- Ochiai [7] adopts a non-linear formulation to balance the influence of passing and failing test cases, which helps mitigate the coincidental correctness problem, where faulty statements are executed but tests still pass. The formula is as follows.
- Dstar [8] exponentially increases the contribution of elements executed in failing tests (typically ), thereby amplifying fault signals and improving the distinguishability of faulty elements. The formula is defined as follows.
- Op2 [9] ranks program elements by computing the difference between execution frequencies in failing and passing test cases, thereby enhancing the distinguishability between faulty and non-faulty elements. The formula is as follows.
- Russell_Rao [10] computes the absolute probability of a program element appearing in failing tests. Its objective is to identify entities that exhibit high coverage consistency in failing tests. The formula is defined as follows.
2.2. Naturalness of Source Code
3. Related Work
4. Methodology
4.1. Overview
- Constructing the original coverage matrix. CNFL firstly constructs the original coverage matrix by executing the program with the given test suite. Each element in the coverage matrix is either 1 or 0, denoting whether a statement is covered by the relevant test case.
- Code naturalness calculation. CNFL collects statements that are covered by failing test cases. For each statement, it further performs code naturalness evaluation to obtain the naturalness score.
- Naturalness-aware statement weight optimization. CNFL integrates the naturalness scores of statements into the original coverage matrix. Specifically, the raw naturalness scores are first converted into standardized weights through a normalization process. These weights are then assigned to the corresponding statements, and the original binary coverage states (0 or 1) are multiplied by their respective weights to construct a weighted coverage matrix. In the weighted coverage matrix, executed statements are no longer treated as equivalent; instead, they are assigned distinct suspiciousness contributions according to their naturalness scores.
- Calculating suspicious scores. The statistical formulas are then applied to the weighted coverage matrix to deliver the list of statements and their suspicious scores.
4.2. Code Naturalness Calculation
4.3. Naturalness-Aware Statement Weight Optimization
4.4. A Motivating Example
5. Experimental Setup
5.1. Research Questions
5.2. Baselines
5.3. Dataset
5.4. LLMs for Code Naturalness Evaluation
- Decoder-only models. Three models, InCoder (6.7 B, Meta AI, Menlo Park, CA, USA) [12], DeepSeekCoder (6.7 B Base, DeepSeek-AI, Hangzhou, China) [13], and Qwen2.5-Coder (7 B Base, Alibaba Cloud, Hangzhou, China) [14] are selected. All of them adopt a Transformer decoder architecture and are trained using the CLM objective. Although they share the same training objective, each model differs in its approach: InCoder uses a causal infilling strategy, DeepSeekCoder adheres to the standard CLM framework, and Qwen2.5-Coder integrates instruction tuning to address complex programming and debugging tasks.
- Encoder-only model. CodeBERT (125 M Base, Microsoft, Redmond, WA, USA) [15] is selected to represent this architecture. It leverages MLM and replaced token detection to capture bidirectional semantic correspondences between source code and natural language.
- Encoder-decoder model. CodeT5Plus (6 B, Salesforce AI Research, Palo Alto, CA, USA) [16] is selected as the representative model of this category. It uses a span-denoising objective to reconstruct masked segments from a bidirectional context, providing strong capabilities in both code understanding and generation.
5.5. Evaluation Metrics
6. Results and Analysis
6.1. RQ1: Comparison of CNFL with SBFL Methods as Well as LLM-Based Methods
6.2. RQ2: Effect of Model Capacities
6.3. RQ3: Comparsion of CNFL with Other Weight-Based Fault Localization Methods
7. Threats to Validity
8. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| CLM | Causal Language Modeling |
| CNFL | Code Naturalness-Based Fault Localization |
| LLM | Large Language Model |
| MAR | Mean Average Rank |
| MFR | Mean First Rank |
| MLM | Masked Language Modeling |
| SBFL | Spectrum-Based Fault Localization |
| WSR | Wilcoxon Signed-Rank Test |
References
- Perez, A.; Abreu, R. A Qualitative Reasoning Approach to Spectrum-Based Fault Localization. In Proceedings of the 40th International Conference on Software Engineering: Companion (ICSE-Companion 2018), Gothenburg, Sweden, 27 May–3 June 2018; Association for Computing Machinery: New York, NY, USA, 2018; pp. 372–373. [Google Scholar]
- Tiwari, S.; Mishra, K.K.; Kumar, A.; Misra, A.K. Spectrum-based fault localization in regression testing. In Proceedings of the 2011 Eighth International Conference on Information Technology: New Generations (ITNG), Las Vegas, NV, USA, 11–13 April 2011; IEEE: New York, NY, USA, 2011; pp. 191–195. [Google Scholar]
- Chen, Z.Z.; Yan, M.; Xia, X.; Liu, Z.X.; Xu, Z.; Lei, Y. Research Progress of Code Naturalness and Its Application. J. Softw. 2021, 33, 3015–3034. [Google Scholar]
- Yang, A.Z.H.; Kolak, S.; Hellendoorn, V.; Martins, R.; Le Goues, C. Revisiting Unnaturalness for Automated Program Repair in the Era of Large Language Models. In Proceedings of the 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), Ottawa, ON, Canada, 27 April–3 May 2025; IEEE: New York, NY, USA, 2025; pp. 2561–2573. [Google Scholar]
- Kang, S.; Yoo, S. Language Models Can Prioritize Patches for Practical Program Patching. In Proceedings of the 3rd International Workshop on Automated Program Repair (APR 2022), Pittsburgh, PA, USA, 19 May 2022; Association for Computing Machinery: New York, NY, USA, 2022; pp. 8–15. [Google Scholar]
- Xia, C.S.; Wei, Y.; Zhang, L. Automated Program Repair in the Era of Large Pre-Trained Language Models. In Proceedings of the 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), Melbourne, Australia, 14–20 May 2023; IEEE: New York, NY, USA, 2023; pp. 1482–1494. [Google Scholar]
- Abreu, R.; Zoeteweij, P.; Golsteijn, R.; Van Gemund, A.J.C. A Practical Evaluation of Spectrum-Based Fault Localization. J. Syst. Softw. 2009, 82, 1780–1792. [Google Scholar] [CrossRef] [Scilit]
- Wong, W.E.; Debroy, V.; Gao, R.; Li, Y. The DStar Method for Effective Software Fault Localization. IEEE Trans. Reliab. 2013, 63, 290–308. [Google Scholar] [CrossRef] [Scilit]
- Naish, L.; Lee, H.J.; Ramamohanarao, K. A Model for Spectra-Based Software Diagnosis. ACM Trans. Softw. Eng. Methodol. 2011, 20, 1–32. [Google Scholar] [CrossRef] [Scilit]
- Pearson, S.; Campos, J.; Just, R.; Fraser, G.; Abreu, R.; Ernst, M.D.; Pang, D.; Keller, B. Evaluating and Improving Fault Localization. In Proceedings of the 2017 IEEE/ACM 39th International Conference on Software Engineering (ICSE), Buenos Aires, Argentina, 20–28 May 2017; IEEE: New York, NY, USA, 2017; pp. 609–620. [Google Scholar]
- Abreu, R.; Zoeteweij, P.; Van Gemund, A.J.C. On the Accuracy of Spectrum-Based Fault Localization. In Proceedings of the Testing: Academic and Industrial Conference Practice and Research Techniques-MUTATION (TAICPART-MUTATION 2007), Windsor, UK, 10–14 September 2007; IEEE Computer Society: Washington, DC, USA, 2007; pp. 89–98. [Google Scholar]
- Fried, D.; Aghajanyan, A.; Lin, J.; Wang, S.; Wallace, E.; Shi, F.; Zhong, R.; Yih, W.t.; Zettlemoyer, L.; Lewis, M. InCoder: A Generative Model for Code Infilling and Synthesis. arXiv 2022, arXiv:2204.05999. [Google Scholar]
- Guo, D.; Zhu, Q.; Yang, D.; Xie, Z.; Dong, K.; Zhang, W.; Chen, G.; Bi, X.; Wu, Y.; Li, Y.K.; et al. DeepSeek-Coder: When the Large Language Model Meets Programming—The Rise of Code Intelligence. arXiv 2024, arXiv:2401.14196. [Google Scholar]
- Hui, B.; Yang, J.; Cui, Z.; Yang, J.; Liu, D.; Zhang, L.; Liu, T.; Zhang, J.; Yu, B.; Lu, K.; et al. Qwen2.5-Coder Technical Report. arXiv 2024, arXiv:2409.12186. [Google Scholar]
- Feng, Z.; Guo, D.; Tang, D.; Duan, N.; Feng, X.; Gong, M.; Shou, L.; Qin, B.; Liu, T.; Jiang, D.; et al. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2020, Online, 16–20 November 2020; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 1536–1547. [Google Scholar]
- Wang, Y.; Le, H.; Gotmare, A.; Bui, N.; Li, J.; Hoi, S. CodeT5+: Open Code Large Language Models for Code Understanding and Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023), Singapore, 6–10 December 2023; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 1069–1088. [Google Scholar]
- Reps, T.; Ball, T.; Das, M.; Larus, J. The Use of Program Profiling for Software Maintenance with Applications to the Year 2000 Problem. In Proceedings of the 6th European Software Engineering Conference Held Jointly with the 5th ACM SIGSOFT International Symposium on Foundations of Software Engineering (ESEC/FSE), Zurich, Switzerland, 22–25 September 1997; Springer: Berlin/Heidelberg, Germany, 1997; pp. 432–449. [Google Scholar]
- Wong, W.E.; Debroy, V.; Choi, B. A Family of Code Coverage-Based Heuristics for Effective Fault Localization. J. Syst. Softw. 2010, 83, 188–208. [Google Scholar] [CrossRef] [Scilit]
- Debroy, V.; Wong, W.E.; Xu, X.; Choi, B. A Grouping-Based Strategy to Improve the Effectiveness of Fault Localization Techniques. In Proceedings of the 2010 10th International Conference on Quality Software (QSIC 2010), Zhangjiajie, China, 14–15 July 2010; IEEE: New York, NY, USA, 2010; pp. 13–22. [Google Scholar]
- Hindle, A.; Barr, E.T.; Su, Z.; Gabel, M.; Devanbu, P. On the Naturalness of Software. In Proceedings of the 34th International Conference on Software Engineering (ICSE 2012), Zurich, Switzerland, 2–9 June 2012; IEEE: New York, NY, USA, 2012; pp. 837–847. [Google Scholar]
- Jiang, Y.; Liu, H.; Zhang, Y.; Ji, W.; Zhong, H.; Zhang, L. Do Bugs Lead to Unnaturalness of Source Code? In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE 2022), Singapore, 14–18 November 2022; Association for Computing Machinery: New York, NY, USA, 2022; pp. 1085–1096. [Google Scholar]
- Li, Y.; Zhong, W.; Shen, Z.; Li, C.; Chen, X.; Ge, J.; Luo, B. An Empirical Study on the Code Naturalness Modeling Capability for LLMs in Automated Patch Correctness Assessment. Autom. Softw. Eng. 2025, 32, 35. [Google Scholar] [CrossRef] [Scilit]
- Ray, B.; Hellendoorn, V.; Godhane, S.; Tu, Z.; Bacchelli, A.; Devanbu, P. On the “Naturalness” of Buggy Code. In Proceedings of the 38th International Conference on Software Engineering (ICSE 2016), Austin, TX, USA, 14–22 May 2016; Association for Computing Machinery: New York, NY, USA, 2016; pp. 428–439. [Google Scholar]
- Raychev, V.; Vechev, M.; Yahav, E. Code Completion with Statistical Language Models. In Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI 2014), Edinburgh, UK, 9–11 June 2014; Association for Computing Machinery: New York, NY, USA, 2014; pp. 419–428. [Google Scholar]
- Nguyen, A.T.; Nguyen, T.D.; Phan, H.D.; Nguyen, T.N. A Deep Neural Network Language Model with Contexts for Source Code. In Proceedings of the 2018 IEEE 25th International Conference on Software Analysis, Evolution and Reengineering (SANER), Campobasso, Italy, 20–23 March 2018; IEEE: New York, NY, USA, 2018; pp. 323–334. [Google Scholar]
- Hellendoorn, V.J.; Devanbu, P. Are Deep Neural Networks the Best Choice for Modeling Source Code? In Proceedings of the 11th Joint Meeting on Foundations of Software Engineering (ESEC/FSE 2017), Paderborn, Germany, 4–8 September 2017; Association for Computing Machinery: New York, NY, USA, 2017; pp. 763–773. [Google Scholar]
- Yang, C.; Chen, J.; Jiang, J.; Huang, Y. Dependency-Aware Code Naturalness. Proc. ACM Program. Lang. 2024, 8, 2355–2377. [Google Scholar] [CrossRef] [Scilit]
- Baah, G.K.; Podgurski, A.; Harrold, M.J. Causal Inference for Statistical Fault Localization. In Proceedings of the 19th International Symposium on Software Testing and Analysis (ISSTA 2010), Trento, Italy, 12–16 July 2010; Association for Computing Machinery: New York, NY, USA, 2010; pp. 73–84. [Google Scholar]
- Mao, X.; Lei, Y.; Dai, Z.; Qi, Y.; Wang, C. Slice-Based Statistical Fault Localization. J. Syst. Softw. 2014, 89, 51–62. [Google Scholar] [CrossRef] [Scilit]
- Santelices, R.; Jones, J.A.; Yu, Y.; Harrold, M.J. Lightweight Fault-Localization Using Multiple Coverage Types. In Proceedings of the 2009 IEEE 31st International Conference on Software Engineering (ICSE 2009), Vancouver, BC, Canada, 16–24 May 2009; IEEE: New York, NY, USA, 2009; pp. 56–66. [Google Scholar]
- Le, T.D.B.; Oentaryo, R.J.; Lo, D. Information Retrieval and Spectrum Based Bug Localization: Better Together. In Proceedings of the 10th Joint Meeting on Foundations of Software Engineering (ESEC/FSE 2015), Bergamo, Italy, 30 August–4 September 2015; Association for Computing Machinery: New York, NY, USA, 2015; pp. 579–590. [Google Scholar]
- Li, Y.; Liu, C. Effective Fault Localization Using Weighted Test Cases. J. Softw. 2014, 9, 2112–2119. [Google Scholar] [CrossRef] [Scilit]
- Dutta, A. Enhancing Fault Localization by Incorporating Statement Frequency and Test Case Contribution. In Proceedings of the 2023 IEEE 23rd International Conference on Software Quality, Reliability, and Security (QRS), Chiang Mai, Thailand, 22–26 October 2023; IEEE: New York, NY, USA, 2023; pp. 128–137. [Google Scholar]
- Xuan, J.; Monperrus, M. Learning to Combine Multiple Ranking Metrics for Fault Localization. In Proceedings of the 2014 IEEE International Conference on Software Maintenance and Evolution (ICSME 2014), Victoria, BC, Canada, 29 September–3 October 2014; IEEE: New York, NY, USA, 2014; pp. 191–200. [Google Scholar]
- Sohn, J.; Yoo, S. Fluccs: Using Code and Change Metrics to Improve Fault Localization. In Proceedings of the 26th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA 2017), Santa Barbara, CA, USA, 10–14 July 2017; Association for Computing Machinery: New York, NY, USA, 2017; pp. 273–283. [Google Scholar]
- Campos, J.; Riboira, A.; Perez, A.; Abreu, R. GZoltar: An Eclipse Plug-In for Testing and Debugging. In Proceedings of the 27th IEEE/ACM International Conference on Automated Software Engineering (ASE 2012), Essen, Germany, 3–7 September 2012; IEEE: New York, NY, USA, 2012; pp. 378–381. [Google Scholar]
- Just, R.; Jalali, D.; Ernst, M.D. Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs. In Proceedings of the 2014 International Symposium on Software Testing and Analysis (ISSTA 2014), San Jose, CA, USA, 21–25 July 2014; Association for Computing Machinery: New York, NY, USA, 2014; pp. 437–440. [Google Scholar]
- Li, Y.; Wang, S.; Nguyen, T. Fault Localization with Code Coverage Representation Learning. In Proceedings of the 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE), Madrid, Spain, 22–30 May 2021; IEEE: New York, NY, USA, 2021; pp. 661–673. [Google Scholar]
- Rafi, M.N.; Chen, A.R.; Chen, T.H.P.; Wang, S. Revisiting Defects4J for fault localization in diverse development scenarios. In Proceedings of the 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR); IEEE: New York, NY, USA, 2025; pp. 63–75. [Google Scholar]
- Xie, H.; Lei, Y.; Yan, M.; Yu, Y.; Xia, X.; Mao, X. A Universal Data Augmentation Approach for Fault Localization. In Proceedings of the 44th International Conference on Software Engineering (ICSE 2022), Pittsburgh, PA, USA, 22–27 May 2022; Association for Computing Machinery: New York, NY, USA, 2022; pp. 48–60. [Google Scholar]
- Zhang, Z.; Lei, Y.; Mao, X.; Li, P. CNN-FL: An Effective Approach for Localizing Faults Using Convolutional Neural Networks. In Proceedings of the 2019 IEEE 26th International Conference on Software Analysis, Evolution and Reengineering (SANER), Hangzhou, China, 24–27 February 2019; IEEE: New York, NY, USA, 2019; pp. 445–455. [Google Scholar]
- Kochhar, P.S.; Xia, X.; Lo, D.; Li, S. Practitioners’ Expectations on Automated Fault Localization. In Proceedings of the 25th International Symposium on Software Testing and Analysis (ISSTA 2016), Saarbrücken, Germany, 18–20 July 2016; Association for Computing Machinery: New York, NY, USA, 2016; pp. 165–176. [Google Scholar]
- Richardson, A. Nonparametric Statistics for Non-Statisticians: A Step-by-Step Approach by Gregory W. Corder, Dale I. Foreman. Int. Stat. Rev. 2010, 78, 467–468. [Google Scholar] [CrossRef] [Scilit]
- Arcuri, A.; Briand, L. A Practical Guide for Using Statistical Tests to Assess Randomized Algorithms in Software Engineering. In Proceedings of the 33rd International Conference on Software Engineering (ICSE 2011), Waikiki, Honolulu, HI, USA, 21–28 May 2011; Association for Computing Machinery: New York, NY, USA, 2011; pp. 1–10. [Google Scholar]



| Code Snippet | Test Suite | Fault Information | |||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| L1: x1 = 0; | L8: else: {x1 = y + 1; | L15: {output (x1);} | The test suite contains two passing test cases and two failing test cases. | Fault statement is L10; its correct version is if (n > 6). | |||||||||||||||
| L2: x2 = 0; | L9: x2 = z + 1; | L16: else: {output (x2);} | |||||||||||||||||
| L3: x3 = 0; | L10: if (n > 0): | L17: output (x3);} | |||||||||||||||||
| L4: if (y < 0): | L11: {n = n + z;} | ||||||||||||||||||
| L5: {x1 = y; | L12: else: n = n + y} | ||||||||||||||||||
| L6: x2 = z; | L13: x3 = n + 1; | ||||||||||||||||||
| L7: x3 = n;} | L14: if (z > 0): | ||||||||||||||||||
| Detailed Score Annotation (L9–L11) | |||||||||||||||||||
| L9: x2 = z + 1; // Ochiai = 0.816, natural_score = 0.32 L10: if (n > 0): // Ochiai = 0.816, natural_score = 0.80 L11: {n = n + z;} // Ochiai = 0.816, natural_score = 0.34 | |||||||||||||||||||
| test | n, y, z | L1 | L2 | L3 | L4 | L5 | L6 | L7 | L8 | L9 | L10 | L11 | L12 | L13 | L14 | L15 | L16 | L17 | result |
| t1 | 9, 2, 6 | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 1 | 1 | 1 | 1 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
| t2 | 8, −2, 6 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
| t3 | 2, 2, −3 | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 1 | 1 | 1 | 1 | 0 | 1 | 1 | 0 | 1 | 1 | 1 |
| t4 | 1, 8, 5 | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 1 | 1 | 1 | 1 | 0 | 1 | 1 | 1 | 0 | 0 | 1 |
| Ochiai | susp | 0.707 | 0.707 | 0.707 | 0.707 | 0.00 | 0.00 | 0.00 | 0.816 | 0.816 | 0.816 | 0.816 | 0.00 | 0.707 | 0.707 | 0.408 | 0.707 | 0.707 | - |
| rank | 5 | 6 | 7 | 8 | 14 | 15 | 16 | 1 | 2 | 3 | 4 | 17 | 9 | 10 | 13 | 11 | 12 | - | |
| t1 | 9, 2, 6 | 0.12 | 0.15 | 0.13 | 0.20 | 0 | 0 | 0 | 0.30 | 0.32 | 0.80 | 0.34 | 0 | 0.28 | 0.36 | 0.32 | 0 | 0 | 0 |
| t2 | 8, −2, 6 | 0.12 | 0.15 | 0.13 | 0.20 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.28 | 0.36 | 0.32 | 0 | 0 | 0 |
| t3 | 2, 2, −3 | 0.12 | 0.15 | 0.13 | 0.20 | 0 | 0 | 0 | 0.30 | 0.32 | 0.80 | 0.34 | 0 | 0.28 | 0.36 | 0 | 0.40 | 0.22 | 1 |
| t4 | 1, 8, 5 | 0.12 | 0.15 | 0.13 | 0.20 | 0 | 0 | 0 | 0.30 | 0.32 | 0.80 | 0.34 | 0 | 0.28 | 0.36 | 0.32 | 0 | 0 | 1 |
| CNFL | selected | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | - |
| susp | 0.244 | 0.273 | 0.254 | 0.316 | 0 | 0 | 0 | 0.447 | 0.461 | 0.730 | 0.336 | 0 | 0.374 | 0.424 | 0.231 | 0.447 | 0.331 | - | |
| rank | 12 | 10 | 11 | 9 | 14 | 15 | 16 | 3 | 2 | 1 | 7 | 17 | 6 | 5 | 13 | 4 | 8 | - | |
| ID | Description | Faults | LoC (K) | Tests |
|---|---|---|---|---|
| Chart | JFreeChart | 26 | 96 | 2205 |
| Lang | Commons Lang | 65 | 22 | 2245 |
| Math | Commons Math | 106 | 85 | 3602 |
| Time | Joda-Time | 27 | 28 | 4130 |
| Mockito | Unit tests Framework | 38 | 67 | 1075 |
| Closure | Closure compiler | 133 | 90 | 7927 |
| Year | Model | Architecture | Size |
|---|---|---|---|
| 2020 | CodeBERT | Encoder-only | 125 M |
| 2022 | InCoder | Decoder-only | 6.7 B |
| 2024 | DeepSeekCoder | Decoder-only | 6.7 B |
| 2024 | Qwen2.5-Coder | Decoder-only | 7 B |
| 2023 | CodeT5Plus | Encoder–decoder | 6 B |
| SBFL Technique | Scenario/Model | Top-1 | Top-3 | Top-5 | MFR | MAR | p-Value | A-Test |
|---|---|---|---|---|---|---|---|---|
| Ochiai | Ochiai | 44 | 95 | 126 | 225.69 | 570.37 | - | - |
| OchiaiCNFLInCoder | 66 | 110 | 127 | 175.21 | 519.31 | 7.24 × 10−8 | 0.601 | |
| OchiaiCNFLCodeBERT | 60 | 109 | 134 | 213.58 | 551.16 | 5.09 × 10−4 | 0.574 | |
| OchiaiCNFLDeepSeekCoder | 60 | 112 | 133 | 193.83 | 533.08 | 7.44 × 10−8 | 0.582 | |
| OchiaiCNFLQwen2.5-Coder | 69 | 112 | 130 | 183.12 | 527.92 | 2.52 × 10−7 | 0.593 | |
| OchiaiCNFLCodeT5Plus | 43 | 76 | 98 | 250.24 | 654.50 | 0.99 | - | |
| Dstar | Dstar | 45 | 95 | 121 | 226.07 | 571.80 | - | - |
| DstarCNFLInCoder | 59 | 111 | 129 | 179.19 | 516.71 | 2.56 × 10−6 | 0.566 | |
| DstarCNFLCodeBERT | 50 | 98 | 122 | 222.93 | 561.51 | 0.051 | - | |
| 58 | 108 | 128 | 196.49 | 534.51 | 3.89 × 10−4 | 0.538 | ||
| DstarCNFLQwen2.5-Coder | 65 | 112 | 130 | 187.61 | 523.65 | 4.03 × 10−3 | 0.582 | |
| DstarCNFLCodeT5Plus | 36 | 72 | 94 | 251.06 | 651.83 | 0.99 | - | |
| Opt2 | Opt2 | 41 | 87 | 109 | 396.60 | 783.74 | - | - |
| Opt2CNFLInCoder | 43 | 76 | 98 | 303.17 | 736.13 | 2.37 × 10−3 | 0.517 | |
| Opt2CNFLCodeBERT | 38 | 73 | 91 | 383.76 | 773.43 | 0.057 | - | |
| Opt2CNFLDeepSeekCoder | 38 | 75 | 94 | 317.44 | 756.0 | 0.063 | - | |
| Opt2CNFLQwen2.5-Coder | 32 | 79 | 94 | 360.58 | 807.56 | 0.14 | - | |
| Opt2CNFLCodeT5Plus | 13 | 34 | 48 | 451.05 | 874.94 | 0.99 | - | |
| Jaccard | Jaccard | 46 | 92 | 124 | 212.21 | 557.10 | - | - |
| JaccardCNFLInCoder | 65 | 118 | 140 | 204.13 | 532.85 | 1.82 × 10−20 | 0.611 | |
| JaccardCNFLCodeBERT | 56 | 104 | 128 | 205.17 | 543.23 | 0.053 | - | |
| JaccardCNFLDeepSeekCoder | 66 | 122 | 143 | 199.54 | 525.08 | 6.59 × 10−22 | 0.622 | |
| JaccardCNFLQwen2.5-Coder | 74 | 117 | 145 | 204.46 | 538.34 | 1.14 × 10−11 | 0.591 | |
| JaccardCNFLCodeT5Plus | 46 | 88 | 115 | 228.05 | 611.85 | 0.694 | - | |
| Russell_rao | Russell_rao | 1 | 11 | 21 | 649.25 | 940.78 | - | - |
| Russell_raoCNFLInCoder | 22 | 53 | 71 | 402.03 | 820.70 | 2.06 × 10−26 | 0.664 | |
| Russell_raoCNFLCodeBERT | 14 | 29 | 37 | 620.19 | 920.30 | 4.02 × 10−3 | 0.539 | |
| Russell_raoCNFLDeepSeekCoder | 20 | 42 | 63 | 437.19 | 862.90 | 2.35 × 10−24 | 0.608 | |
| Russell_raoCNFLQwen2.5-Coder | 23 | 53 | 65 | 526.05 | 889.36 | 1.64 × 10−11 | 0.583 | |
| Russell_raoCNFLCodeT5Plus | 6 | 18 | 25 | 641.94 | 934.12 | 0.31 | - |
| LLMs for Code | Top-1 | Top-3 | Top-5 |
|---|---|---|---|
| CodeBERT | 15 | 35 | 59 |
| InCoder | 11 | 19 | 28 |
| DeepSeekCoder | 9 | 23 | 34 |
| CodeT5Plus | 10 | 23 | 32 |
| Qwen2.5-Coder | 8 | 24 | 36 |
| Wilcoxon Tests | Right-Tailed | Left-Tailed | Two-Tailed | Conclusion | |
|---|---|---|---|---|---|
| CNFL vs. WTCFL | Ochiai | 0.967 | 0.031 | 0.036 | better |
| Dstar | 0.985 | 0.012 | 0.008 | better | |
| Opt2 | 0.954 | 0.042 | 0.045 | better | |
| Jaccard | 0.992 | 6.41 × 10−3 | 1.20 × 10−4 | better | |
| Russell_rao | 0.998 | 1.12 × 10−4 | 2.24 × 10−5 | better | |
| Wilcoxon Tests | Right-Tailed | Left-Tailed | Two-Tailed | Conclusion | |
|---|---|---|---|---|---|
| CNFL vs. PFL | Ochiai | 0.991 | 0.038 | 0.042 | better |
| Dstar | 0.982 | 0.034 | 0.041 | better | |
| Opt2 | 0.951 | 0.049 | 0.048 | better | |
| Jaccard | 0.997 | 3.25 × 10−4 | 6.50 × 10−4 | better | |
| Russell_rao | 0.998 | 5.12 × 10−7 | 1.02 × 10−6 | better | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yao, W.; Jiang, M.; Zhou, Y. Software Fault Localization Approach with Coverage Matrix Optimization Boosted by LLM-Based Code Naturalness. Appl. Sci. 2026, 16, 4416. https://doi.org/10.3390/app16094416
Yao W, Jiang M, Zhou Y. Software Fault Localization Approach with Coverage Matrix Optimization Boosted by LLM-Based Code Naturalness. Applied Sciences. 2026; 16(9):4416. https://doi.org/10.3390/app16094416
Chicago/Turabian StyleYao, Wen, Mingyue Jiang, and Yuan Zhou. 2026. "Software Fault Localization Approach with Coverage Matrix Optimization Boosted by LLM-Based Code Naturalness" Applied Sciences 16, no. 9: 4416. https://doi.org/10.3390/app16094416
APA StyleYao, W., Jiang, M., & Zhou, Y. (2026). Software Fault Localization Approach with Coverage Matrix Optimization Boosted by LLM-Based Code Naturalness. Applied Sciences, 16(9), 4416. https://doi.org/10.3390/app16094416

