Next Article in Journal
Symmetric Dual-Domain Prototype Adaptation for Few-Shot Image-Based Malware Classification
Previous Article in Journal
A Nonlinear Hoek–Brown Criterion for Bedded Rock with Brittle–Ductile Transition
 
 
Correction published on 1 October 2026, see Symmetry 2026, 18(10), 1650.
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Bottleneck-Aware Heuristic and Metaheuristic Framework for Requirement-Based Test Case Prioritization in Sparse Traceability Matrices

by
Ahmed Enis Erkaya
1,2,*,
Sahin Emrah Amrahov
3 and
Fatih V. Çelebi
4
1
Department of Computer Engineering, Graduate School of Natural and Applied Sciences, Ankara University, 06110 Ankara, Türkiye
2
TUBITAK BILGEM Software Technologies Research Institute, 06530 Ankara, Türkiye
3
Department of Computer Engineering, Ankara University, 06830 Ankara, Türkiye
4
Department of Computer Engineering, Ankara Yildirim Beyazit University, 06010 Ankara, Türkiye
*
Author to whom correspondence should be addressed.
Symmetry 2026, 18(7), 1207; https://doi.org/10.3390/sym18071207
Submission received: 8 May 2026 / Revised: 7 July 2026 / Accepted: 15 July 2026 / Published: 17 July 2026 / Corrected: 1 October 2026
(This article belongs to the Section A: Computer Science)

Abstract

Regression testing is a critical activity for maintaining software quality by identifying faults and ensuring system reliability. However, in large-scale software systems, executing all available test cases within limited time and computational resources is often impractical. Therefore, test case prioritization aims to arrange test cases in an effective execution order to maximize testing effectiveness, particularly during the early stages of regression testing. Existing test case prioritization approaches consider various optimization objectives, including fault detection capability, code coverage, risk reduction, and execution cost. In this study, we focus on the requirement-based test case prioritization problem, where the main goal is to maximize the rate of requirement coverage as early as possible. Since the possible orderings of test cases form a factorial-sized search space, this problem exhibits NP-hard characteristics, leading to the widespread use of heuristic and metaheuristic optimization techniques. However, sparse requirement traceability matrices (RTMs) introduce additional challenges, particularly due to isolated requirements and delayed coverage of critical requirement elements. To address these challenges, this study proposes a bottleneck-aware heuristic and metaheuristic framework for requirement-based test case prioritization under sparse RTMs. The main empirical contribution of the study is the deterministic AG+BH strategy, which combines Additional Greedy with the proposed Bottleneck Hunter mechanism. This strategy uses the structure of the RTM to identify test cases associated with delayed coverage, especially singleton requirements. The MH-DBO-GA component is included as a secondary metaheuristic extension based on Dragon Boat Optimization, Genetic Algorithm operators, and memetic local search. Its role is to provide additional search diversity rather than to replace the deterministic AG+BH strategy in the current APRC setting. The proposed framework is evaluated using two sparse requirement traceability matrix datasets. The results show that AG+BH provides the strongest practical deterministic trade-off on the evaluated datasets. It obtains the highest APRC on Dataset 2 and a near-best APRC on Dataset 1, where 2-Optimal gives a slightly higher APRC but requires 23 h 23 min of execution time. MH-DBO-GA does not outperform AG+BH in this setting, but it performs better than the standard stochastic metaheuristic baselines and can be considered as an exploratory extension when additional search diversity is needed. Furthermore, an ablation analysis is performed to examine the individual contributions of informed initialization, Bottleneck Hunter, and hybrid optimization components. Overall, the findings indicate that sparse RTM-based test case prioritization benefits primarily from problem-specific bottleneck-aware heuristic reasoning, while metaheuristic refinement should be interpreted as a complementary layer for more complex or future multi-objective settings.

1. Introduction

The software development process is carried out within a holistic framework called the Software Development Life Cycle (SDLC), extending from requirements analysis to the delivery of the final product [1]. This process consists of planning, requirements analysis, design, development, testing, and maintenance phases, each of which has a direct impact on software quality. Among these phases, the testing process plays a critical role in ensuring the software meets reliability, performance, and functionality requirements. Thoroughly testing software systems before deploying them to the production environment enables early detection of potential errors and significantly reduces maintenance costs [2].
Regression testing aims to verify that new features added to the software or changes made to existing components do not negatively impact the system’s previous functionality [2]. In modern software projects that require continuous change and updates, regression testing is becoming increasingly critical. Neglecting these tests can lead to serious system-wide errors even with minor changes [3]. Furthermore, regression testing processes consume significant resources due to their high time and cost requirements. The literature reports that a single regression test cycle can take weeks [4]. Therefore, regression testing is increasingly automated through continuous integration (CI) and continuous delivery (CD) pipelines.
Three main approaches have been developed in the literature to reduce testing costs: test set reduction, test case selection, and test case prioritization [2,5]. Test set reduction permanently eliminates unnecessary or redundant test cases, while the test selection approach executes only tests related to current changes. Test case prioritization, on the other hand, aims to identify important requirements and errors at the earliest possible stage by reordering tests according to specific criteria. This study focuses on the test case prioritization problem among these approaches [4,6]. The main differences between these approaches are summarized in Table 1.
In the test case prioritization problem, various criteria such as code coverage, potential for error detection, historical test data, risk factors, and requirement priorities are used [7]. Different paradigms and techniques have been proposed in the literature, including Bayesian models, machine learning methods, and metaheuristic algorithms [3,8,9,10]. Requirement-based approaches aim to prioritize test cases according to characteristics such as customer priorities and implementation complexity, and require strong traceability mechanisms [5,11,12,13]. In this context, the Requirement Traceability Matrix (RTM) is used as an important tool in managing the relationship between requirements and test cases [14].
The effectiveness of test case prioritization methods is generally evaluated using the Average Percentage of Faults Detected (APFD) metric [3,7,8,15,16,17,18,19,20]. However, in requirements-based studies, requirements coverage performance takes precedence over error detection. Therefore, in this study, the APFD metric has been adapted as the Average Percentage of Requirements Covered (APRC) [21,22,23].
A substantial portion of existing studies on requirement-based test case prioritization has been evaluated using datasets with relatively dense requirement-test relationships. However, real-world software projects often contain sparse requirement traceability matrices (RTMs), where test cases are designed to cover specific functionalities with limited overlap. Such sparse structures present optimization problems, particularly due to the singleton requirements that are associated with only a single test case. These isolated requirements may become coverage bottlenecks by delaying the achievement of complete requirement coverage. This can result in slower coverage growth during the early stages of regression testing.
To address these challenges, this study proposes a bottleneck-aware heuristic and metaheuristic framework for requirement-based test case prioritization under sparse RTMs. The framework is centered on a deterministic problem-specific heuristic layer that combines Additional Greedy with the proposed Bottleneck Hunter mechanism to identify and prioritize test cases associated with delayed requirement coverage. A secondary metaheuristic refinement layer, based on Dragon Boat Optimization, Genetic Algorithm operators, and memetic local search, is incorporated to provide additional exploration capability in larger search spaces. By combining structural information from sparse RTMs with adaptive search mechanisms, the framework aims to improve early requirement coverage while keeping the main empirical emphasis on bottleneck-aware deterministic prioritization.
In light of the experimental results, we explicitly position the deterministic AG+BH component as the primary practical contribution of this study. The results indicate that AG+BH offers a favorable balance between APRC, saturation behavior, and runtime in the evaluated single-objective setting. Although 2-Optimal obtains a slightly higher APRC on Dataset 1, it requires 23 h 23 min to complete, whereas AG+BH reaches a near-best APRC with substantially lower runtime and obtains the highest APRC on Dataset 2. Therefore, MH-DBO-GA is not presented as a generally superior alternative to AG+BH or Additional Greedy. Instead, it is considered a secondary exploratory extension that may be useful when additional search diversity is required or when future versions of the problem include further constraints such as execution cost, requirement priority, or fault severity.
The main contributions of this study can be summarized as follows:
  • Identification of bottleneck characteristics in sparse requirement traceability matrices.
    This study analyzes the structural challenges of sparse RTMs, particularly the delayed coverage problem caused by isolated and singleton requirements. The study highlights that such bottlenecks may limit the effectiveness of conventional optimization approaches by reducing early requirement coverage capability.
  • Development of the deterministic AG+BH strategy as the primary bottleneck-aware heuristic.
    The main contribution of this work is a deterministic heuristic layer that combines Additional Greedy with the proposed Bottleneck Hunter mechanism. This component uses problem-specific information from the requirement–test relationship structure to identify critical test cases and improve early requirement coverage. In the evaluated APRC setting, AG+BH demonstrates that problem-specific bottleneck handling can produce highly competitive coverage results with low computational cost.
  • Positioning of MH-DBO-GA as a secondary exploratory metaheuristic extension.
    The framework also includes a metaheuristic refinement layer based on Dragon Boat Optimization, Genetic Algorithm operators, and memetic local search. This component is not claimed to be uniformly preferable to AG+BH under the current experimental setting. Instead, it is used as a complementary population-based extension that can provide additional search diversity and may be more relevant in future dynamic or multi-objective prioritization scenarios.
  • Comprehensive experimental evaluation and component-level analysis.
    The proposed framework is evaluated on sparse requirement traceability matrix datasets using the Average Percentage of Requirements Covered (APRC) metric and saturation analysis. The evaluation separates the deterministic AG+BH strategy from the stochastic MH-DBO-GA component and interprets their roles according to the observed results. In addition, an ablation study is conducted to investigate the individual effects of informed initialization, Bottleneck Hunter, and hybrid optimization components.
The rest of the article is structured as follows: The second section presents relevant studies, the third section presents problem definition, the fourth section provides preliminaries, the fifth section describes the proposed framework in detail, the sixth section reports the experimental setup and results, the seventh section analyzes the contribution of individual components, the eighth section discusses threats to validity and limitations, and the final section discusses the findings.

2. Related Work

Software testing processes are considered one of the most complex and critical stages of the Software Development Life Cycle (SDLC) [1]. In this context, test case prioritization (TCP) aims to reduce maintenance costs by increasing the failure detection rate at an early stage [4]. In the literature, a wide range of methods have been proposed for the TCP problem, from classical heuristic methods to nature-inspired metaheuristic algorithms.
In early studies, greedy algorithms were widely used [24]. Rothermel et al. [4] compared nine different prioritization techniques in their work on Siemens programs and showed that prioritization methods significantly increased the error detection rate. However, the inability of greedy approaches to reach the global optimum led to the emergence of improved methods such as Additional Greedy and 2-Optimal.
To overcome these limitations, evolutionary and metaheuristic algorithms have become widely used in the TCP problem. Li et al. [24] investigated the coverage performance in regression tests by comparing Genetic Algorithm (GA), Hill Climbing, and different greedy methods. Malhotra and Bharadwaj [25] proposed a GA model based on error detection capability. Mishra et al. [3] managed to reduce test costs with a GA-based method combining coverage rate and error potential. Wambua et al. [26] compared GA and Bat Algorithm in terms of APFD and memory usage, reporting that GA provided higher performance. Anuar et al. [27] examined the effectiveness of Ant Colony Optimization and GA-based models.
Some studies have developed specialized models that consider requirements traceability and software structural features. Demir et al. [28] optimized the test set according to requirements coverage by proposing a method based on the dominant cluster approach. Sabharwal et al. [29] developed a GA-based model using the information flow metric for scenarios derived from UML diagrams. Abubakar et al. [30] proposed the QAG-TCP method, which includes metrics such as connectivity, compliance, and fault severity.
In cases where coverage information is limited, Mukherjee and Patnaik [31] compared GA, Simulated Annealing, and Ant Colony methods by treating the problem as a 0/1 knapsack model. In scenarios involving multiple coverage measures, Di Nucci et al. [32] proposed a hybrid GA model driven by hypervolume indicators.
Similar motivations also appear in broader combinatorial optimization studies. Adasme et al. [33] combined Bender’s decomposition with a local-search metaheuristic for the quadratic p-median problem, showing that exact and heuristic strategies can play complementary roles in difficult optimization settings. Although this problem differs from requirement-based TCP, it provides additional motivation for using problem-specific heuristic search in large combinatorial spaces.
Recently, dragon boat inspired optimization has also been considered in TCP research. Dragon Boat Optimization was originally introduced by Li et al. as a human-based metaheuristic inspired by dragon boat racing [34]. Assiri [35] later adapted this optimization idea to test case prioritization and reported improvements in terms of APFD and execution time. In the present study, DBOA is not used as a standalone TCP method; instead, its movement principles are adapted to a random-key representation as part of the secondary MH-DBO-GA exploration layer.
However, a large portion of the studies in the literature have been evaluated on controlled and relatively balanced datasets. In real-world software projects, sparse requirement-test case matches and singleton requirement structures pose significant challenges for existing methods. In particular, studies incorporating specific mechanisms for addressing structural bottlenecks that delay full coverage are limited. Furthermore, the scalability of methods in large-scale and irregular data structures has not been adequately demonstrated. Vescan et al. show that Ant Colony Optimization (ACO) is successful in fault detection-based test case prioritization [36]. However, when applied alone to the ’sparse requirements matrices’ problem addressed in our study, it did not exhibit the expected performance.
This study addresses this gap by focusing on bottleneck-aware prioritization under sparse requirement traceability matrices. The proposed framework emphasizes a deterministic AG+BH strategy that directly uses sparse RTM structure to identify delayed-coverage test cases. A metaheuristic refinement layer is also included to provide additional population-based exploration, but the main empirical emphasis is placed on the bottleneck-aware heuristic behavior observed in the evaluated APRC setting.

3. Problem Definition

Software testing is often constrained by limited time and computational resources. Under such conditions, executing all test cases in an arbitrary order may delay achieving requirement coverage. The Test Case Prioritization (TCP) problem addresses this issue by determining an execution order that optimizes a predefined objective. In this study, the objective is to maximize requirement coverage as early as possible by rearranging the execution sequence of the available test cases.
As defined by Rothermel et al. [4], the TCP problem is formulated as follows:
  • Given: T is the set of test cases, and P T represents all possible permutations of the T set. f : P T → R is the objective function that calculates the fitness value corresponding to each permutation.
  • Objective: To find the best test order that satisfies the condition f ( T ′ ) ≥ f ( T ″ ) for all T ″ ∈ P T , where T ′ ∈ P T .
In this study, the Average Percentage of Requirements Covered (APRC) metric is used as the objective function. The APRC metric is defined as follows:
A P R C = 1 − ∑ i = 1 m T R i n m + 1 2 n
Here, n represents the number of test cases, m the number of requirements, and T R i represents the order of the test in which the i-th requirement is covered for the first time. A high APRC value indicates that critical requirements are covered in the early stages of the testing process, thus resulting in a more efficient test execution process.
Therefore, the objective of this study is to obtain an effective test case prioritization sequence that maximizes APRC. Since the denominator in the APRC formula is constant for a given dataset, maximizing APRC is equivalent to minimizing the sum of first-occurrence requirement coverage ranks. Accordingly, the prioritization strategy aims to place test cases that cover previously uncovered requirements as early as possible in the ordered sequence.
The problem is an NP-hard optimization problem due to the size and combinatorial nature of the solution space. Therefore, heuristic and metaheuristic approximation methods are commonly used to obtain effective prioritization sequences within practical time limits.

4. Preliminaries

This section briefly introduces the Genetic Algorithm (GA) and the Dragon Boat Optimization Algorithm (DBOA), which form the basis of the proposed hybrid approach.

4.1. Genetic Algorithm (GA)

Genetic Algorithms (GA) are population-based optimization methods inspired by natural selection and evolutionary processes. In these algorithms, each individual represents a solution in the problem space and is usually expressed by the chromosome structure.
In the context of the TCP problem, each chromosome represents a permutation of test cases. The initial population is iteratively evolved using selection, crossover, and mutation operators.
The selection operator ensures that individuals with high fitness values are passed on to subsequent generations, while the crossover operator combines the genetic information of different individuals to produce new solutions. The mutation operator increases diversity by adding random changes to the population and reduces the early convergence problem.
Overall, one of the main advantages of GA is its ability to search effectively in large solution spaces and reduce the risk of getting stuck in local minima.

4.2. Dragon Boat Optimization Algorithm (DBOA)

The Dragon Boat Optimization Algorithm (DBOA) is a metaheuristic optimization method proposed by Li et al. [34] and inspired by team dynamics in dragon boat racing.
DBOA performs an efficient search process in the solution space by modeling both social interactions among individuals and physical movement dynamics. At the heart of the algorithm are various control parameters that enable individuals to move synchronously towards a common goal.
The social behavior factor ( ψ ) represents the level of motivation and interaction between individuals, simulating the effects of social facilitation and social laziness. This mechanism helps prevent premature convergence of the population.
Furthermore, the acceleration coefficient ( λ ) determines the convergence speed of the search process, while the attenuation coefficient ( μ ) contributes to preserving solution diversity. Thanks to these parameters, DBOA exhibits a balanced search behavior between exploration and exploitation processes.

5. Proposed Framework

5.1. Overview of the Proposed Framework

In recent years, hybrid algorithms, which combine the strengths of different optimization methods, have been shown to produce successful results in various engineering problems [37,38,39,40,41,42,43,44,45].
This paper introduces a bottleneck-aware two-layer framework for requirement-based test case prioritization under sparse RTMs. The first layer is the deterministic AG+BH strategy, which combines Additional Greedy with the Bottleneck Hunter mechanism. This layer directly uses the structure of sparse RTMs to identify delayed-coverage test cases, especially those related to singleton requirements.
The second layer is the MH-DBO-GA metaheuristic refinement component. It combines Dragon Boat Optimization, Genetic Algorithm operators, and memetic local search to provide additional population-based exploration. In the current single-objective APRC evaluation, this layer is not presented as a replacement for AG+BH. Rather, it is treated as an exploratory extension that may be useful when deterministic bottleneck rules are not sufficient or when additional constraints are introduced.
As shown in Figure 1, the metaheuristic refinement component consists of four main components: (i) preprocessing, (ii) informed initialization, (iii) hybrid optimization process (DBOA-GA), and (iv) memetic local search mechanism.
The hybrid design combines DBOA and GA for complementary purposes. DBOA is responsible for rapidly identifying promising regions of the search space. On the other hand, GA operators maintain population diversity and generate alternative candidate solutions during the optimization process.
In addition, the memetic local search mechanism called Bottleneck Hunter, developed in this study, improves solution quality by identifying critical test cases that delay complete requirement coverage. This mechanism provides performance gains, especially in sparse requirement-test matrices.
Accordingly, the framework should be read as a heuristic-centered approach with an optional metaheuristic refinement layer. The experimental results indicate that the deterministic bottleneck-aware layer accounts for the main empirical improvement in the studied APRC setting, while MH-DBO-GA provides an additional search mechanism for cases where broader exploration is needed.

5.2. System Architecture

The proposed MH-DBO-GA framework has a multi-stage architecture that integrates global search and local optimization mechanisms. This architecture consists of four core modules: preprocessing, informed initialization, hybrid optimization, and memetic local search. They all complement each other to improve solution quality and make the search process more efficient.
In the preprocessing stage, the problem structure is analyzed to identify singleton requirements, which are covered by only a single test case.
In the second stage, an informed initialization strategy is used to improve the quality of the initial population. This approach starts from regions with high potential in the solution space instead of random initialization. It also contributes to obtaining better solutions in the early stages of the search process.
In the third stage, we carry out a hybrid optimization process using the Dragon Boat Optimization Algorithm (DBOA) and the Genetic Algorithm (GA) together. At this stage, DBOA guides the population globally. On the other hand, GA operators ensure the search process progresses in a balanced manner by maintaining solution diversity.
In the final stage, a memetic local search mechanism appears. In this stage, problem-specific improvements are made to the best available solution, enhancing solution quality. Specifically, the Bottleneck Hunter mechanism directly improves solution sequencing by identifying critical test cases that cause delayed requirement coverage.

5.3. Informed Initialization Strategy

In the proposed approach, the quality of the initial population is considered a critical factor directly affecting the success of the optimization process. Therefore, instead of purely random initialization, we propose an informed initialization strategy.
This strategy consists of two main steps. In the first step, requirements covered by only a single test case ( R s i n g l e t o n ) are identified. Such requirements are critical components that can negatively affect the total coverage time if they are delayed in the solution sequence. Therefore, test cases containing these requirements are prioritized in the initial phase.
In the second step, a large portion of the population (approximately 80%) is generated using the Additional Greedy approach. This method builds a stepwise solution by selecting the test case that contributes most to the uncovered requirements in each iteration. Thus, the generated initial individuals have a high coverage potential.
The remaining portion of the population (20%) is generated randomly to maintain diversity. This hybrid initialization approach allows for a rapid start in promising regions of the solution space while also contributing to maintaining sufficient diversity in the search process.
Consequently, the proposed knowledge-based initialization strategy reduces time loss that can be caused by low-quality initializations and makes it possible to obtain better solutions in the early stages of the optimization process.

5.4. Hybrid Optimization Mechanism (DBOA + GA)

The proposed MH-DBO-GA framework is based on a hybrid optimization mechanism established between the Dragon Boat Optimization Algorithm (DBOA) and the Genetic Algorithm (GA). This structure aims to improve solution quality by providing a balanced interaction between global exploration and local exploitation processes.
DBOA exhibits rapid convergence behavior by guiding the population around the global best individual (leader). While the movements of individuals in the solution space are controlled through acceleration ( λ ) and weakening ( μ ) coefficients, the social behavior factor ( ψ ) provides flexibility to the search process by modeling intra-population interactions. Thanks to this structure, DBOA can perform extensive exploration in the solution space.
However, early convergence and local minima problems, frequently observed in swarm intelligence-based methods, can limit performance when using DBOA alone. To overcome this disadvantage, GA operators are incorporated into the hybrid structure.
The GA component increases diversity by introducing new genetic variations into the population through crossover and mutation operators. Order crossover and random mutation mechanisms, particularly suitable for sequencing problems, allow for the recombination of existing solutions, enabling transitions to different regions of the search process.
Within the hybrid structure, low-performing individuals are eliminated at the end of each iteration, and new solutions are generated using elite individuals. This process increases exploration opportunities, further preventing population stabilization.
Within the proposed framework, DBOA and GA operators provide a population-based exploration layer. This layer can generate alternative candidate sequences and maintain diversity during the search process. However, in the current APRC-based experiments, this metaheuristic layer is considered complementary to the deterministic AG+BH strategy rather than the primary source of empirical improvement.

5.5. Memetic Local Search (Bottleneck Hunter)

One of the most distinctive components of the proposed MH-DBO-GA framework is the problem-specific memetic local search mechanism, Bottleneck Hunter. This mechanism directly targets delays in the requirements coverage process by improving upon the best available solution.
In the requirements-based test case prioritization problem, one of the most critical factors determining solution quality is the position at which each requirement is first covered. In this context, the T R i value represents the order of the test case that first covers the i-th requirement. Examining the APRC metric, it is seen that directly minimizing the sum of these values improves solution quality.
Accordingly, the saturation index ( S i n d e x ) indicates the last position at which all requirements are satisfied for the first time.
S i n d e x = max ( T R i )
Solutions with a high S i n d e x value show that at least one requirement is covered quite late, negatively affecting the APRC value. The test case causing this delay is considered a bottleneck in the solution sequence.
The proposed Bottleneck Hunter mechanism aims to identify this critical test case and increase its priority in the decoded sequence. In the implementation, this is done by assigning a high random-key value to the bottleneck test when the saturation index exceeds a predefined threshold. This intervention can reduce the T R i value of delayed requirements and, consequently, can reduce the sum of ∑ T R i .
The main advantage of this approach is that it directly targets the components of the objective function. In sparse requirements-test matrices, especially when requirements covered by only a single test case are located late in the sequence, the T R i values can increase disproportionately.
The memetic local search restricts one of the swap indices to the first min ( 300 , | T | ) positions of the test sequence, while the second index is sampled from the whole sequence. This design focuses the local refinement on the earlier part of the ranking, where improvements have a stronger effect on APRC. The value of 300 is used as a practical middle setting based on the sensitivity analysis. It improves over the 100-test-case setting on Dataset 1 and gives the highest mean APRC on Dataset 2, although the 500-test-case setting performs better on Dataset 1. Therefore, this value should be interpreted as a practical search-scope choice rather than as a universally optimal threshold.
In conclusion, the proposed memetic mechanism refines solutions obtained through the global search process at the local level, thereby increasing convergence speed and contributing to the achievement of higher-quality solutions.

5.6. Algorithm Workflow

The general operation of the proposed MH-DBO-GA algorithm is presented in Algorithm 1. The algorithm works iteratively in a multi-stage structure, gradually improving the quality of the solution.
In the first stage, a preprocessing step is performed to analyze the requirement-test relationships and specifically identify the requirements covered by only a single test case ( R s i n g l e t o n ). This information is stored for use in subsequent stages.
In the second stage, an initial population is created using an information-based initialization strategy. While a large portion of the population is generated using the Additional Greedy approach, the remaining individuals are randomly generated to ensure diversity.
In the third stage, the main optimization cycle of the algorithm is initiated. In each iteration, individuals are converted from their representations in the continuous space to discrete test sequences, and fitness values are calculated using the APRC metric. In this process, the individual with the best fitness value is updated as the global leader.
Algorithm 1 Memetic Hybrid Dragon Boat Optimization and Genetic Algorithm (MH-DBO-GA)
Require: T (Test Cases), R (Requirements), M a x I t e r (Maximum Iterations)
Ensure:  S b e s t (Optimized Test Case Sequence)
  •    Phase 1: Preprocessing
  1:
Identify R s i n g l e t o n (requirements covered by exactly one test case)
  2:
Initialize i d M a p to map test identifiers to vector indices
  •    Phase 2: Informed Initialization (Smart Seeding)
  3:
for each agent i ∈ P o p u l a t i o n S i z e  do
  4:
      if  i < 0.8× P o p u l a t i o n S i z e  then
  5:
            B o a t i ← Apply Additional Greedy Seeding prioritizing R s i n g l e t o n
  6:
      else
  7:
             B o a t i ← Initialize random position vector in [ 0 , 1 ]
  8:
      end if
  9:
end for
  •    Phase 3: Hybrid Optimization Loop
10:
for  i t e r = 1 to M a x I t e r  do
11:
      for each B o a t i  do
12:
          S e q u e n c e i ← Map continuous vector to discrete order (Random Key)
13:
          F i t n e s s i ← Compute APRC using S e q u e n c e i and R
14:
         if  F i t n e s s i > G l o b a l B e s t f i t n e s s  then
15:
             Update G l o b a l B e s t (Leader)
16:
         end if
17:
     end for
  •    Phase 4: Memetic Phase (Local Search)
18:
     if GlobalBest exists then
19:
          RefineLeader: Perform 100 candidate swaps with one index sampled from the first min ( 300 , | T | ) positions
20:
          BottleneckHunter: If the saturation index exceeds 500, assign random-key value 1.1 to the bottleneck test
21:
    end if
  •    Phase 5: DBOA Movement
22:
     Update acceleration ( λ ), attenuation ( μ ), and social factors ( ψ , H )
23:
     for each B o a t i  do
24:
           Update position using DBOA equations guided by G l o b a l B e s t
25:
           Clamp positions to [ 0 , 1.2]
26:
    end for
  •    Phase 6: Evolutionary Refill (Genetic Operators)
27:
     Sort population by fitness
28:
     Remove bottom 25 % of agents
29:
     while population size < P o p u l a t i o n S i z e  do
30:
           Select P a r e n t 1 , P a r e n t 2 from elite pool
31:
            C h i l d ← Apply Order Crossover and Mutation
32:
           Add C h i l d to population
33:
    end while
34:
end for
35:
return  S b e s t (sequence of the global best leader)
After each iteration, the best available solution is improved by applying a memetic local search phase. In this phase, Bottleneck Hunter and local optimization processes are performed on the leading individual to improve solution quality.
Following this, the DBOA mechanism is used to update the population positions and move individuals towards the global leader. Then, GA operators are activated to eliminate low-performing individuals and generate new solutions from elite individuals.
This iterative process continues until the maximum number of iterations is reached. The best solution found at the end of the algorithm is returned as the final test case ranking.
Algorithm 1 summarizes the workflow of the proposed method. The implementation-level details needed to follow the stochastic component, including the adapted DBOA update rules, local-search budget, genetic operators, replacement rule, stopping condition, and random-number handling, are given in Section 5.7.

5.7. Implementation Details for Reproducibility

This subsection gives the main implementation details of the stochastic metaheuristic component. Algorithm 1 presents the overall workflow, while this subsection describes the update rules, local-search budget, genetic operators, replacement strategy, stopping condition, and random-number handling used in the experiments.
Random-key decoding is performed by sorting the position values in descending order. If two test cases have the same key value, their test identifiers are used as a lexicographic tie-breaker. If a tie still remains, the original index order is used as the final tie-breaker.
The informed initialization step is applied to 80% of the population. First, test cases associated with singleton requirements are identified. These essential test cases are ordered according to their incremental contribution to uncovered requirements. After that, Additional Greedy selection continues until either 500 tests have been prioritized or no remaining candidate provides additional uncovered requirement coverage. At each greedy step, at most the first 1000 candidates in the remaining pool are evaluated. Non-prioritized positions are initialized with small random values in [ 0 , 0.2) . For prioritized tests, descending random-key values are assigned according to the greedy order.
The leader refinement step is applied once per iteration to the current global best solution. It performs 100 candidate swap trials. In each trial, the first index is sampled uniformly from the first min ( 300 , | T | ) positions, and the second index is sampled uniformly from the whole sequence. The two selected positions are swapped, APRC is recomputed, and the swap is accepted only if it improves the current best APRC value. Otherwise, the swap is reverted.
The Bottleneck Hunter step is also applied to the current global best solution. The current sequence is scanned until all requirements are covered for the first time. The test case located at this saturation position is treated as the bottleneck test. If the saturation index is greater than 500, the random-key value of this test case is set to 1.1 and APRC is recomputed. If the saturation index is 500 or less, no Bottleneck Hunter update is applied in that iteration.
The DBOA movement used in this study is based on the main factors of Dragon Boat Optimization, including the social behavior factor, acceleration factor, attenuation factor, imbalance term, and fastest-boat-based update mechanism [34,35]. Since the TCP problem is represented as a permutation problem, these rules were adapted to a random-key representation. In this representation, each boat is encoded as a continuous vector, and this vector is decoded into a test-case ordering by sorting the key values.
Let x i k ∈ R | T | denote the random-key position vector of the i-th boat at iteration k, where | T | is the number of test cases. The current global best position vector is denoted by x b e s t k . The iteration counter is defined as I = k + 1 , where k = 0 , … , K − 1 . In the experiments, K = 100 and the population size is N = 50 .
At each iteration, the social behavior factor ψ k is sampled using two integer random variables. Let a k be sampled uniformly from { 1 , … , 2 N } and b k be sampled uniformly from { 1 , … , N } . Then, ψ k is computed as follows:
ψ k = N b k , if a k < N , 1 , otherwise .
After ψ k is obtained, the acceleration factor λ k and attenuation factor μ k are computed as follows:
λ k = ψ k I − 1 ψ k I ,
μ k = 1 + ψ k I − K ψ k I .
The attenuation term follows the same general form as DBOA. In this implementation, the parameter l in the original formulation is set to the maximum iteration budget K. The imbalance term is computed by setting θ k = 2 π r k , where r k is sampled from U ( 0 , 1 ) , and H b = 0.01 :
H k = | cos ( θ k ) | I ψ k + H b , θ k = 2 π r k , H b = 0.01 .
The movement step is applied to the random-key position vectors. If the current boat has the same APRC fitness value as the current global best solution, the fastest-boat update term is computed as
R f , i , d k = x i , d k λ k ,
and the corresponding random-key position is updated by
x i , d k + 1 = clip x i , d k + R f , i , d k , 0 , 1.2 .
For the remaining boats, the update term is computed with reference to the current global best position:
R e , i , d k = x b e s t , d k + x i , d k 2 μ k λ k .
The position is then updated as
x i , d k + 1 = clip x i , d k + R e , i , d k H k , 0 , 1.2 .
Here, clip ( · , 0 , 1.2) restricts each random-key value to the interval [ 0 , 1.2] . This step is used because the DBOA movement is applied in a continuous random-key space, whereas the final TCP solution is a permutation.
The evolutionary refill stage uses an elitist replacement rule. At the end of each iteration, the population is sorted in descending order of the APRC fitness values available in that iteration. The nominal elimination rate is 25%. With N = 50 , the implementation keeps ⌊ 50 × ( 1 − 0.25) ⌋ = 37 boats and removes the 13 lowest-ranking boats. The top 10% of the original population size, corresponding to five boats, forms the elite pool. New offspring are generated until the population size returns to N = 50 . For each offspring, two parents are selected uniformly at random from the elite pool. Positions modified by the DBOA movement are evaluated in the next fitness evaluation step.
The crossover operator is an order-preserving one-point crossover with a fixed midpoint. The first half of the child sequence is copied from the first parent. The remaining positions are filled by scanning the second parent from the beginning and appending test cases that are not already present in the child. This produces a valid permutation without duplicate test cases.
Mutation is applied with probability 0.20. When mutation is triggered, five random swap operations are performed on the child sequence. In each swap operation, both positions are selected independently and uniformly from the whole child sequence. Therefore, a swap may leave the sequence unchanged if the same position is selected twice. After crossover and mutation, the child permutation is converted back into a random-key vector by assigning
x ( t j ) = 1 − j | T | ,
where j is the zero-based rank of test case t j in the child sequence. This conversion is used to represent the child ordering again in the random-key space.
The stopping condition is fixed: each stochastic run terminates after K = 100 iterations. No early stopping criterion is used. The implementation uses Java pseudo-random generators without manually fixed seed values. A local pseudo-random generator is used in the main optimization routine for DBOA parameter sampling and parent selection, while a static pseudo-random generator is used in the boat representation for random initialization, leader refinement, and mutation. For this reason, exact bit-level replay of the reported stochastic runs is not guaranteed. The reported stochastic results are therefore presented as statistical summaries over 50 independent executions.

5.8. Transformation from Continuous to Discrete Space and Position Bounding Strategies

Although the proposed framework is continuous, the Test Case Prioritization (TCP) problem is a discrete permutation problem. Therefore, the following strategies are used to align these two representations.
Random-Key Mapping: The position of each dragon boat within the search space is represented by a continuous vector (e.g., 0.25, 0.95, 0.60), which encodes the weights of the test cases. These steps transform the continuous vector into a valid test sequence.
  • Indexing and Sorting: The continuous weights belonging to the test cases are sorted in descending order. The test case with the highest weight is placed in the first position.
  • Tie-Breaking: If the position weights of two or more test cases are equal, a tie-breaking mechanism is activated to prevent non-deterministic behavior. In such instances, the related test cases are incorporated into the permutation by being sorted lexicographically according to their identifiers (Test ID).
Utilization of the 1.1 Value for Bottleneck Teleportation: When the Bottleneck Hunter mechanism identifies a test case associated with delayed coverage saturation, it assigns a random-key value of 1.1 to that test case.
  • In the normal random-key representation, most test weights are initialized or evolved within the range of 0.0 to 1.0.
  • Assigning a value slightly above 1.0 increases the priority of the detected bottleneck test and tends to move it toward the beginning of the decoded sequence. This intervention encourages earlier coverage of delayed requirements without directly rewriting the whole permutation.
Clamp Value: During random-key position updates, the maximum value of the position vector is restricted to 1.2. The upper bound is set above 1.0 to allow controlled priority increases while still keeping the random-key values within a bounded interval.
  • Providing headroom: The 1.2 upper bound gives a limited margin above the normalized [0.0, 1.0] range. This allows bottleneck-related tests assigned a value of 1.1 to remain distinguishable from normally evolved test keys.
  • Bounding random-key growth: The 1.2 threshold prevents unbounded growth of random-key values during position updates and keeps the decoding process stable.

6. Experimental Results

6.1. Datasets

The effectiveness and generalizability of the proposed bottleneck-aware hybrid heuristic and metaheuristic framework were evaluated using two different datasets derived from real-world software projects. These datasets contain varying scales of test cases and requirement complexity, allowing for an objective analysis of the method’s performance in both high-volume and enterprise-level software systems. The basic quantitative characteristics of the datasets are summarized in Table 2.
  • Dataset 1: Kaggle Car Rental Software Dataset
This dataset was compiled in 2020 from a web-based management system belonging to a car rental company (https://www.kaggle.com/datasets/zumarkhalid/a-test-case-data-set-with-requirements, accessed on 16 July 2026).
The dataset includes detailed attributes such as business requirements (B_Req), priority level, function points (FP), complexity level, duration, and cost. Comprising a total of 3399 test cases and 2000 requirements, this dataset contains 108 singleton requirements covered by only a single test. This highly sparse structure provides a suitable experimental environment for evaluating the capability of the proposed framework to identify and mitigate requirement coverage bottlenecks.
  • Dataset 2: Public Institution Traceability Matrix
The second dataset is an anonymized requirements-test traceability matrix obtained from a public institution operating in Türkiye (https://github.com/eniserkaya/RequirementBased-TCP-MH-DBO-GA/tree/main/data, accessed on 16 July 2026).
This enterprise-scale system consists of 14 modules, 1467 test cases, and 1237 requirements. As a result of the analyses, 72 singleton requirements were identified in this dataset. Due to its modular and hierarchical structure, this dataset was selected to evaluate the adaptability and scalability of the proposed framework in large-scale and complex enterprise software projects.
Table 2 presents the Requirement Traceability Matrix. Matrix density represents the ratio of traceability links(non-zero entries) to the total number of possible connections between test cases and requirements ( D e n s i t y = L / ( | T | × | R | ) ). A sparse RTM shows that most of the test cases cover a tiny subset of the requirements. This structural characteristic significantly restricts the search space.
In such situations, optimization approaches require problem-specific mechanisms capable of identifying structural bottlenecks. Therefore, the proposed framework incorporates the Bottleneck Hunter mechanism to exploit sparse RTM characteristics during prioritization. As detailed in Table 2, Dataset 1 consists of 3399 test cases and 2000 requirements, theoretically yielding a maximum of 6,798,000 matrix cells. However, it contains only 5835 traceability links, resulting in a very low matrix density of 0.0858% (sparsity: 99.9142%). Similarly, Dataset 2 contains 1467 test cases and 1237 requirements, resulting in a maximum of 1,814,679 possible entries. However, the dataset includes 5331 links, yielding a matrix density of 0.2938% (sparsity: 99.7062%). These density values (0.0858% and 0.2938%) quantitatively confirm that both RTMs are extremely sparse and support the study’s main motivation, namely the sparse RTM assumption.

6.2. Experimental Setup and Parameter Settings

We evaluated the proposed bottleneck-aware hybrid heuristic and metaheuristic framework on a MacBook Pro workstation equipped with an Apple M1 processor and 32 GB of RAM. The experiments included both deterministic heuristic components and stochastic metaheuristic optimization components. For stochastic algorithms, each configuration was executed for 50 independent runs, and each run consisted of 100 optimization iterations. The implementation uses Java pseudo-random generators without manually fixed seed values. Therefore, the reported stochastic results should be interpreted as statistical summaries of independent executions rather than as exact bit-level replayable runs. We adopted this evaluation protocol to reduce the influence of random variation and to provide a reliable statistical assessment of the obtained results.
The parameter settings of the metaheuristic refinement component are summarized in Table 3. We determined the population size to be 50 to provide a balance between population diversity and computational efficiency. Additionally, we fixed the maximum number of iterations at 100 for all stochastic optimization approaches to ensure a consistent comparison.
To incorporate problem-specific knowledge into the search process, 80% of the initial population was generated using a bottleneck-aware greedy initialization strategy, while 20% was randomly initialized to maintain exploration diversity. This strategy starts the search from promising regions derived from sparse requirement traceability matrices and avoids over-reliance on deterministic solutions.
In the memetic refinement phase, we focused on the first 300 test cases in the generated priority sequence. We selected this region because early requirement coverage is the primary optimization objective in regression testing scenarios. By concentrating local improvement efforts on the initial portion of the ranking, the method aims to enhance APRC performance without introducing excessive computational overhead.
The DBOA position update mechanism employs dynamic parameter schedules to control the balance between exploration and exploitation throughout the optimization process. The position clamping range was set to [ 0.0, 1.2] to prevent excessive movement outside the feasible search region while allowing controlled exploration beyond the normalized ranking interval. Similarly, the bottleneck teleportation value was set to 1.1 to provide a controlled perturbation mechanism for relocating bottleneck-related solutions without causing disruptive changes in the search trajectory.
In the genetic refinement stage, we employed an elite selection rate of 10%, an elimination rate of 25%, and a mutation rate of 20%. Furthermore, we selected these settings to maintain population diversity while preserving high-quality solutions during the evolutionary search. Lastly, we kept all parameter configurations identical across repeated experiments to ensure a fair comparison among the evaluated approaches.

6.3. Comparative Performance Analysis

We comparatively evaluated the contribution of the deterministic bottleneck-aware heuristic component and the metaheuristic refinement component within the proposed framework. We examined the effectiveness of the Additional Greedy and Bottleneck Hunter mechanisms, because sparse requirement traceability matrices introduce structural coverage challenges, particularly due to singleton requirements.
To avoid misleading visual comparisons between deterministic methods with zero variance and stochastic metaheuristic methods with repeated-run variability, we present them separately. The APRC scores of deterministic algorithms are shown as bar charts in Figure 2 and Figure 3, while the APRC distributions of stochastic algorithms over 50 independent runs are shown as boxplots in Figure 4 and Figure 5.
Subsequently, we evaluated the MH-DBO-GA component as a metaheuristic refinement strategy that enhances exploration capability through evolutionary and memetic operators. To evaluate the individual contributions of the proposed framework, we compared the bottleneck-aware heuristic configuration (AG+BH) and the metaheuristic refinement configuration (MH-DBO-GA) with representative baseline approaches, including Standard Greedy, Additional Greedy, 2-Optimal, GA, and DBOA.
Table 4 presents the APRC performance of all evaluated approaches. Deterministic methods are reported using their exact single-run APRC values. On the other hand, stochastic metaheuristic methods are evaluated over 50 independent runs. Statistical comparisons and practical effect sizes are reported only where stochastic distributions are available.
Table 4 shows that the deterministic methods provide the strongest APRC results in the current evaluation. On Dataset 1, 2-Optimal obtains the highest APRC with 89.96%, while AG+BH obtains a very close value of 89.92%. On Dataset 2, AG+BH obtains the highest APRC with 95.57%. The corresponding MH-DBO-GA results are 85.68% and 95.14%, respectively. These results indicate that, when the objective is early requirement coverage, problem-specific deterministic reasoning is effective under sparse RTMs. However, the practical preference among deterministic methods should also consider runtime, because 2-Optimal requires substantially higher computational cost on Dataset 1.
The MH-DBO-GA component has a different role in the framework. It does not improve over AG+BH in the evaluated setting, but it performs better than the standard stochastic metaheuristic baselines. For Dataset 1, MH-DBO-GA outperforms the strongest stochastic baseline, GA. For Dataset 2, it outperforms the strongest stochastic baseline, DBOA. These results suggest that MH-DBO-GA is useful as a population-based exploratory extension, but not as the main performance driver of the proposed framework.
These results show that bottleneck-aware deterministic prioritization is the main empirical finding of this study. AG+BH should be seen as a practical and strong deterministic strategy, not as the best APRC method on every dataset. It is very close to the best APRC on Dataset 1. It achieves the best APRC on Dataset 2. It also runs much faster than 2-Optimal. The metaheuristic layer is kept because it adds search diversity. It may also be useful in future settings with dynamic RTMs, changing test constraints, requirement weights, execution costs, or fault severity.

6.4. Saturation Point and Coverage Speed Analysis

The saturation point represents the position in the test case sequence at which all requirements are covered for the first time. This metric provides additional insight into early requirement coverage behavior beyond the final APRC value. Table 5 and Table 6 present the saturation points and execution times of the evaluated approaches for Dataset 1 and Dataset 2, respectively.
The 2-Optimal run on Dataset 1 completed in 84,180 s, corresponding to 23 h 23 min, and obtained 89.96% APRC. This value comes from a completed run rather than an extrapolated runtime. Although 2-Optimal reaches an early saturation point and gives the highest APRC on Dataset 1, its pairwise-swap evaluation leads to a very high computational cost on the larger sparse RTM. Therefore, this result is included mainly to illustrate the practical runtime limitation of pairwise local improvement on large test suites.
The saturation analysis provides further evidence that sparse RTM structures can benefit significantly from problem-specific bottleneck identification.
As Table 5 shows, the AG+BH configuration and Additional Greedy reach saturation at the 998th position on Dataset 1, outperforming the population-based metaheuristic approaches in terms of coverage speed. This result indicates that deterministic heuristic strategies can be effective when the sparse RTM contains bottleneck patterns.
Table 6 shows a similar pattern for Dataset 2. The AG+BH configuration reaches full coverage at the 282nd position, closely followed by Additional Greedy at the 283rd position. This result indicates that targeting singleton-induced coverage bottlenecks can accelerate early requirement coverage. The MH-DBO-GA configuration reaches full coverage at the 446th position while providing a population-based exploratory refinement mechanism.
These saturation results are consistent with the APRC findings. In both datasets, AG+BH or Additional Greedy reaches full requirement coverage earlier than MH-DBO-GA. Therefore, for sparse RTMs with clearly identifiable coverage bottlenecks, the deterministic heuristic layer is the more suitable option when the evaluation objective is early requirement coverage alone. The MH-DBO-GA component should be interpreted as a secondary exploratory mechanism. Its potential value is in providing population-based search diversity for more complex prioritization settings, rather than in replacing AG+BH in the current APRC experiments.

6.5. Computational Efficiency Analysis

Computational efficiency is an important consideration in regression testing, particularly for continuous integration and continuous delivery (CI/CD) environments. In those environments, prioritization decisions must be generated within limited time constraints. Therefore, we evaluate the computational cost of the proposed framework by considering both execution time and algorithmic complexity. We comparatively analyzed the deterministic heuristic components and the metaheuristic refinement component to reveal the trade-off between computational overhead and coverage effectiveness.

6.5.1. Execution Time Comparison

Table 7 presents the execution times of the evaluated approaches on both datasets. For stochastic methods, the reported values represent the mean execution time over repeated runs. For deterministic methods, the values correspond to the completed single-run execution time required to generate a complete test case prioritization sequence.
The execution-time results show that simple greedy-based deterministic strategies have low computational cost. Greedy and Additional Greedy are the fastest approaches, while AG+BH introduces only a limited additional cost because the Bottleneck Hunter mechanism performs a targeted refinement on the generated sequence. However, 2-Optimal is an exception among deterministic methods. Although it can reach an early saturation point, its pairwise-swap evaluation leads to a very high runtime, especially on Dataset 1.
The MH-DBO-GA component requires additional computational time due to population-based exploration, evolutionary operations, and memetic refinement. It is slower than DBOA, but it obtains substantially higher APRC values than the standard stochastic baselines. It is also faster than GA on both datasets. Therefore, the results indicate a trade-off: AG+BH is preferable when low runtime and early requirement coverage are the main objectives, while MH-DBO-GA provides a more exploratory population-based alternative with additional computational cost.

6.5.2. Computational Complexity Discussion

To avoid duplicate complexity analyses, we report a single sparse-link-based derivation in this subsection. Let N denote the population size, K the maximum number of iterations, | T | the number of test cases, | R | the number of requirements, and L the number of non-zero requirement–test links in the sparse RTM. Since the RTM is stored and traversed as sparse adjacency lists, the analysis is expressed in terms of L rather than the dense matrix size | T | × | R | .
For one candidate sequence, APRC evaluation is performed by scanning the test cases in the given order and visiting only the requirements linked to each visited test case. Therefore, the cost of one APRC computation is
O ( | T | + L ) .
This expression includes the traversal of the test sequence and the traversal of existing sparse links. It does not multiply | T | by L, because all non-zero links are visited through the adjacency lists, not re-scanned for each test case.
The stochastic metaheuristic component uses a random-key representation. Before APRC can be computed, the random-key vector of a candidate must be decoded into a test-case ordering. This decoding step requires sorting the | T | key values, which has the cost
O ( | T | log | T | ) .
Thus, evaluating one random-key candidate requires
O ( | T | log | T | + | T | + L ) .
For N candidates evaluated over K iterations, the main population-level evaluation cost is therefore
O N K | T | log | T | + | T | + L .
The informed initialization stage has a separate cost. The singleton-requirement preprocessing step scans the sparse RTM once and costs O ( L ) . For the greedy part of initialization, let G denote the maximum number of prioritized tests produced by the greedy construction and let C g denote the maximum number of candidate tests scanned at each greedy step. In our implementation, G = 500 and C g = 1000 . Let d max denote the maximum number of requirements linked to any test case. Under sparse traversal, the cost of constructing one greedy individual is bounded by
C AG = O | T | + G C g d max .
The | T | term accounts for assigning the remaining random-key values after the greedy prefix is constructed. The G C g d max term gives a conservative bound for checking the sparse requirement links of the scanned candidates. This is more precise for our implementation than the previous dense-style expression, because the method does not traverse a full | T | × | R | matrix.
Since 80% of the population is initialized using the informed greedy strategy and 20% is initialized randomly, the initialization cost can be written as
O L + 0.8N C AG + 0.2N | T | .
The memetic refinement step is applied only to the current leader. In each iteration, the leader refinement performs B L S = 100 candidate swap trials. Since APRC is recomputed after each candidate swap, the local-search cost over K iterations is bounded by
O K B L S ( | T | + L ) .
The Bottleneck Hunter step also scans the current leader sequence to identify the saturation position and the associated bottleneck test. Its cost is O ( | T | + L ) per application and is covered by the same sparse traversal argument.
The remaining population operations have lower-order costs in the current setting. DBOA position updates require O ( N | T | ) operations per iteration. The evolutionary refill stage sorts the population by fitness, which costs O ( N log N ) per iteration, and generates replacement offspring through crossover and mutation. Since the number of replaced individuals is proportional to N, the crossover cost is O ( N | T | ) per iteration. Mutation performs a fixed number of swaps and has a smaller contribution.
Combining these terms, the overall cost of the MH-DBO-GA component can be summarized as
O L + 0.8N C AG + 0.2N | T | + N K | T | log | T | + | T | + L + K B L S ( | T | + L ) .
In the experiments, N = 50 , K = 100 , G = 500 , C g = 1000 , and B L S = 100 are fixed parameter values. Therefore, the main dataset-dependent terms are the number of test cases | T | and the number of sparse links L. This explains why sparse traversal is important: the computation depends on the existing traceability links rather than on the theoretical dense matrix size | T | × | R | .
For the deterministic AG+BH configuration, there is no population-based iteration. Its cost is mainly the Additional Greedy construction and one Bottleneck Hunter refinement, which can be summarized as
O C AG + | T | + L .
This difference explains the lower runtime of AG+BH compared with MH-DBO-GA. The metaheuristic component introduces additional cost through repeated population evaluation, sorting, movement updates, and local refinement, whereas AG+BH directly applies the sparse greedy construction and bottleneck-aware correction.

6.6. Parameter Sensitivity Analysis

Control parameters can influence how MH-DBO-GA balances exploration, exploitation, and computational cost. We examine four parameters in this section: the greedy initialization ratio, local search region size, bottleneck teleportation value, and position clamping range.
For each parameter, we ran each alternative configuration 50 times on both datasets. The sensitivity tables report the mean APRC, standard deviation, 95% confidence interval of the mean, and runtime mean with standard deviation. Runtime values are reported in seconds. Since several parameter settings produce close APRC values, we use this analysis mainly to understand parameter sensitivity rather than to identify strictly dominant settings.
We compare runtime values mainly within the same parameter group, because each sensitivity experiment was executed as a separate set of runs. Therefore, repeated baseline settings across different parameter groups may show small runtime differences due to independent executions.

6.6.1. Greedy Initialization Ratio

Table 8 shows that the tested greedy initialization ratios produce close APRC values. The 80% ratio gives the highest mean APRC on both datasets, but the confidence intervals remain close to those of the other settings. The runtime values also do not show a consistent advantage for one ratio across both datasets. We therefore use the 80% setting as a practical configuration that combines a high proportion of informed greedy initialization with a limited amount of random initialization for population diversity.

6.6.2. Sensitivity Analysis of Local Search Region

The local search region determines which part of the priority sequence the memetic swap-based refinement can explore. We tested three region sizes: 100, 300, and 500 test cases. This parameter can affect both the search behavior and the computational cost, so Table 9 reports APRC variability and runtime values together.
Table 9 shows that the local search region affects the two datasets differently. On Dataset 1, the 500-test-case region gives the highest mean APRC. On Dataset 2, the 300-test-case region gives the highest mean APRC, while the 500-test-case region gives a slightly lower value. Increasing the region size therefore does not lead to a consistent APRC improvement across both datasets.
The runtime values do not increase monotonically with the local search region size. This is consistent with the implementation, because the algorithm keeps the number of local-search trials fixed and changes only the part of the sequence from which it samples swap positions. We use the 300-test-case region as a middle setting. It gives the best mean APRC on Dataset 2 and improves over the 100-test-case setting on Dataset 1, although it does not exceed the 500-test-case setting on Dataset 1.

6.6.3. Sensitivity Analysis of Bottleneck Teleportation Value

The bottleneck teleportation value controls the random-key value assigned to a bottleneck-related test case during the Bottleneck Hunter step. We tested three values: 1.0, 1.1, and 1.2. Since these settings produce very close APRC values, Table 10 reports standard deviations, confidence intervals, and runtime summaries in addition to the mean APRC values.
Table 10 shows very small APRC differences among the tested teleportation values. On Dataset 1, the mean APRC values of 1.0, 1.1, and 1.2 are almost indistinguishable. On Dataset 2, the 1.1 setting gives the highest mean APRC, but the difference from the other values remains small. The runtime values do not show a clear monotonic pattern with respect to the teleportation value.
We use 1.1 in the final configuration as a moderate teleportation value. It gives the highest mean APRC on Dataset 2 and remains close to the best value on Dataset 1, without using the largest tested perturbation.

6.6.4. Sensitivity Analysis of Position Clamping Range

The position clamping range defines the interval in which the algorithm keeps random-key values after DBOA movement. We tested three ranges: [0.0, 1.0], [0.0, 1.2], and [0.0, 1.5]. Since the APRC differences among these settings are small, Table 11 reports standard deviations, confidence intervals, and runtime summaries together with the mean APRC values.
Table 11 shows that the tested clamping ranges produce close APRC values on both datasets. The [0.0, 1.2] range gives the highest mean APRC on Dataset 1 and Dataset 2, but the differences from the other ranges remain small and the confidence intervals are close.
The runtime values do not show a fully monotonic trend as the upper clamping bound changes. However, the largest tested range, [0.0, 1.5], does not improve mean APRC on either dataset and gives the highest runtime on Dataset 1. We use the [0.0, 1.2] range in the final configuration as a moderate extension of the normalized [0.0, 1.0] interval. This range allows limited movement above 1.0 and remains consistent with the Bottleneck Hunter teleportation value of 1.1, while avoiding the largest tested upper bound.
Table 12 summarizes the parameter values used in the final MH-DBO-GA configuration. We select these values as practical choices based on the observed mean APRC values, variability, and runtime summaries. Since several tested settings produce close APRC values, this table documents the final configuration and the rationale behind each selected value.

7. Contribution of Individual Components

We performed an ablation study to examine how the main components of the metaheuristic refinement layer affect APRC performance. The analyzed components include informed initialization, Bottleneck Hunter, DBOA-based movement, and GA-based evolutionary operators. The purpose of this analysis is not only to confirm the usefulness of individual components, but also to identify whether the complete hybridization is necessary under different sparse RTM structures. The mean APRC values obtained from 50 independent executions for each variation are reported in Table 13.
The results in Table 13 show that the contribution of each component is not uniform across the two datasets. Informed initialization has a consistent positive role. Removing it decreases APRC from 0.856847 to 0.843575 on Dataset 1 and from 0.951380 to 0.943220 on Dataset 2. This suggests that starting the search from structurally informed solutions remains useful for sparse RTMs.
The effect of Bottleneck Hunter is more dataset-dependent. On Dataset 1, removing Bottleneck Hunter reduces APRC from 0.856847 to 0.842681, indicating that the bottleneck-oriented move helps the search process in this matrix structure. On Dataset 2, however, removing Bottleneck Hunter slightly increases APRC from 0.951380 to 0.951710. This does not mean that Bottleneck Hunter is generally harmful; rather, it suggests that when the sequence is already close to a high-coverage arrangement, the teleportation step can sometimes introduce a small disturbance.
The ablation results also show that the full combination of DBOA, GA operators, and Bottleneck Hunter is not always the best-performing configuration. On Dataset 1, the DBOA + Memetic Only variant obtains an APRC of 0.858211, slightly higher than the full MH-DBO-GA result of 0.856847. On Dataset 2, several reduced variants are very close to, or slightly better than, the full hybrid configuration. For example, GA + Memetic Only reaches 0.951539 and Without Bottleneck Hunter reaches 0.951710, while MH-DBO-GA obtains 0.951380. These differences are small, but they indicate that some components may be redundant under certain sparse RTM structures.
A possible explanation is that the GA crossover operator and the Bottleneck Hunter teleportation step may occasionally disrupt an already well-ordered sequence. This is especially possible when informed initialization and memetic refinement have already placed most bottleneck-related test cases near the beginning of the ranking. In such cases, additional exploration does not necessarily improve APRC and may slightly change a useful ordering. Therefore, the complete MH-DBO-GA configuration should not be interpreted as a necessary combination for all sparse RTMs.
Overall, the ablation study supports two main conclusions. First, informed initialization is an important component of the metaheuristic refinement process. Second, full hybridization is not uniformly beneficial. The reduced variants show that DBOA-based movement, GA operators, and Bottleneck Hunter may interact differently depending on the structure of the RTM. This finding is consistent with the positioning of MH-DBO-GA as a secondary exploratory extension rather than as the primary source of empirical improvement.

8. Threats to Validity and Limitations

8.1. Construct Validity

In this study, we evaluated test case prioritization primarily based on the Average Percentage of Requirements Covered and the saturation point. Although these metrics are widely used in the literature, they assume equal importance for all requirements. In practice, some requirements may be more critical than others. Not weighting requirements according to their importance or execution time is a limitation of our current metric evaluation.

8.2. Internal Validity

The main threat here is the stochastic nature of the metaheuristic component of the proposed framework. We ran the stochastic algorithms multiple times to mitigate this threat and evaluated the consistency of the obtained results using statistical analysis. This procedure reduces the possibility that the observed differences are caused by random variations in metaheuristic execution. Another threat is related to parameter settings (e.g., population size, mutation rate, and local search parameters). These parameters were initially selected based on preliminary experiments and further investigated through sensitivity analysis. The results indicate that moderate variations in parameter values do not lead to substantial performance degradation; however, different configurations may still influence the final performance. We also interpret the sensitivity analysis with caution, because several parameter settings produce close APRC values. Therefore, the analysis reports variability, confidence intervals, and runtime summaries, and does not treat small mean differences as strong evidence of parameter dominance.

8.3. External Validity

The proposed framework was specifically designed for sparse RTM. Although the proposed framework was designed to improve the balance between exploration and exploitation in sparse matrices, performance and runtime may vary in large-scale industrial environments. Additional validation on broader industrial datasets is required to generalize these findings to all software systems.

9. Conclusions

This study investigated the requirements-based test case prioritization problem under sparse requirement traceability matrices and proposed a bottleneck-aware heuristic and metaheuristic framework. The framework is centered on the deterministic AG+BH strategy, which combines Additional Greedy with the proposed Bottleneck Hunter mechanism. This component targets delayed requirement coverage caused by sparse requirement–test relationships, especially singleton requirements. The MH-DBO-GA component is included as a secondary population-based extension that combines Dragon Boat Optimization, Genetic Algorithm operators, and memetic local search.
Extensive experimental evaluations on two sparse RTM datasets show that When APRC, saturation behavior, and runtime are considered jointly, AG+BH provides the practical deterministic outcome in the evaluated single-objective setting. It achieves the highest APRC on Dataset 2 and a near-best APRC on Dataset 1, where 2-Optimal obtains higher APRC but requires 23 h 23 min to complete. AG+BH also reaches full requirement coverage earlier than MH-DBO-GA on the studied datasets. These results suggest that, when sparse RTM bottlenecks are directly identifiable, a deterministic bottleneck-aware heuristic can be more effective and simpler than applying the full metaheuristic framework.
The MH-DBO-GA component should therefore not be interpreted as a generally superior alternative to AG+BH or Additional Greedy. Its role is complementary. It provides a population-based search mechanism and performs better than the standard stochastic metaheuristic baselines, but it does not outperform AG+BH in the present APRC experiments.
The ablation results further show that full hybridization is not always necessary, since some reduced variants perform comparably to or slightly better than the complete MH-DBO-GA configuration on the evaluated datasets.
The computational evaluation also supports this interpretation. The deterministic AG+BH strategy achieves strong coverage behavior with low computational cost, while the metaheuristic refinement layer introduces additional overhead due to population-based search and memetic refinement. Thus, AG+BH is the preferable option for the current sparse RTM setting when the main objective is early requirement coverage. MH-DBO-GA may be considered when additional exploration is needed or when the prioritization problem is extended with further constraints.
In conclusion, this work shows that sparse RTM-based test case prioritization benefits primarily from domain-specific bottleneck identification. The main scientific message of the study is the effectiveness of bottleneck-aware deterministic prioritization, with metaheuristic refinement serving as a secondary extension rather than the primary source of empirical improvement.
Future work will focus on extending the framework toward multi-objective optimization scenarios by incorporating additional factors such as execution cost, fault severity, requirement priority, and code complexity. Such settings may provide a more suitable basis for evaluating when the MH-DBO-GA exploratory layer becomes preferable to deterministic heuristics. Furthermore, evaluating the approach on larger industrial datasets and adapting the framework to diverse software architectures, including distributed and microservice-based systems, will be investigated.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/sym18071207/s1, File S1: Source code and datasets.

Author Contributions

Conceptualization, A.E.E., S.E.A. and F.V.Ç.; methodology, A.E.E.; software, A.E.E.; formal analysis, S.E.A.; investigation, A.E.E.; data curation, S.E.A.; writing—original draft preparation, A.E.E.; writing—review and editing, S.E.A. and F.V.Ç.; visualization, A.E.E.; supervision, F.V.Ç. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The original contributions presented in this study are included in the article/Supplementary Materials. Further inquiries can be directed to the corresponding author.

Acknowledgments

The authors gratefully acknowledge support from the TUBITAK BILGEM Software Technologies Research Institute (YTE). The authors acknowledge the use of artificial intelligence-based tools for language editing and grammatical corrections. All scientific content, experimental design, analysis, and conclusions presented in this paper are solely the responsibility of the authors. Furthermore, the authors thank the Academic Editor and the anonymous reviewers for their constructive comments, which helped improve the clarity, methodological transparency, and presentation of the manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Ariffin, M.A.; Ibrahim, R.; Ibrahim, I.S.; Wahab, J.A. Test Cases Prioritization Using Ant Colony Optimization and Firefly Algorithm. Int. J. Eng. Trends Technol. 2022, 70, 22–28. [Google Scholar] [CrossRef] [Scilit]
  2. Khatibsyarbini, M.; Isa, M.A.; Jawawi, D.N.; Tumeng, R. Test case prioritization approaches in regression testing: A systematic literature review. Inf. Softw. Technol. 2018, 93, 74–93. [Google Scholar] [CrossRef] [Scilit]
  3. Mishra, D.B.; Panda, N.; Mishra, R.; Acharya, A.A. Total fault exposing potential based test case prioritization using genetic algorithm. Int. J. Inf. Technol. 2019, 11, 633–637. [Google Scholar] [CrossRef] [Scilit]
  4. Rothermel, G.; Untcn, R.H.; Chu, C.; Harrold, M.J. Prioritizing test cases for regression testing. IEEE Trans. Softw. Eng. 2001, 27, 929–948. [Google Scholar] [CrossRef] [Scilit]
  5. Yoo, S.; Harman, M. Regression testing minimization, selection and prioritization: A survey. Softw. Test. Verif. Reliab. 2012, 22, 67–120. [Google Scholar] [CrossRef] [Scilit]
  6. Do, H.; Rothermel, G.; Kinneer, A. Empirical studies of test case prioritization in a JUnit testing environment. In Proceedings of the 15th International Symposium on Software Reliability Engineering, Saint-Malo, France, 2–5 November 2004; pp. 113–124. [Google Scholar] [CrossRef] [Scilit]
  7. Mondal, S.; Nasre, R. Colosseum: Regression Test Prioritization by Delta Displacement in Test Coverage. IEEE Trans. Softw. Eng. 2022, 48, 4060–4073. [Google Scholar] [CrossRef] [Scilit]
  8. Samad, A.; Mahdin, H.; Kazmi, R.; Ibrahim, R. Regression Test Case Prioritization: A Systematic Literature Review. Int. J. Adv. Comput. Sci. Appl. 2021, 12, 2. [Google Scholar] [CrossRef] [Scilit]
  9. Rhmann, W.; Zaidi, T.; Saxena, V. Test Cases Minimization and Prioritization Based on Requirement, Coverage, Risk Factor and Execution Time. Br. J. Math. Comput. Sci. 2016, 14, 1–9. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Ahmed, B. Test Case Minimization Approach Using Fault Detection and Combinatorial Optimization Techniques for Configuration-Aware Structural Testing. J. Eng. Sci. Technol. 2016, 12, 737–753. [Google Scholar] [CrossRef] [Scilit]
  11. Naheed, M.; Zaman, Q.U.; Nadeem, A. A Requirement Based Approach to Test Case Prioritization for Regression Testing. In Proceedings of 2022 19th International Bhurban Conference on Applied Sciences and Technology, IBCAST 2022; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NJ, USA, 2022; pp. 358–363. [Google Scholar] [CrossRef] [Scilit]
  12. Tahvili, S.; Pimentel, R.; Afzal, W.; Ahlberg, M.; Fornander, E.; Bohlin, M. SOrTES: A Supportive Tool for Stochastic Scheduling of Manual Integration Test Cases. IEEE Access 2019, 7, 12928–12946. [Google Scholar] [CrossRef] [Scilit]
  13. Çetin, P.; Tanrıöver, Ö. Priority rule for resource constrained project planning problem with predetermined work package durations. J. Fac. Eng. Archit. Gazi Univ. 2020, 35, 1537–1549. [Google Scholar] [CrossRef] [Scilit]
  14. Krishnamoorthi, R.; Mary, S.A.S.A. Factor oriented requirement coverage based system test case prioritization of new and regression test cases. Inf. Softw. Technol. 2009, 51, 799–808. [Google Scholar] [CrossRef] [Scilit]
  15. Khatibsyarbini, M.; Isa, M.A.; Jawawi, D.N.; Hamed, H.N.A.; Suffian, M.D.M. Test Case Prioritization Using Firefly Algorithm for Software Testing. IEEE Access 2019, 7, 132360–132373. [Google Scholar] [CrossRef] [Scilit]
  16. Nazir, M.; Mehmood, A.; Aslam, W.; Park, Y.; Choi, G.S.; Ashraf, I. A Multi-Goal Particle Swarm Optimizer for Test Case Prioritization. IEEE Access 2023, 11, 90683–90697. [Google Scholar] [CrossRef] [Scilit]
  17. Ling, X.; Agrawal, R.; Menzies, T. How Different is Test Case Prioritization for Open and Closed Source Projects? IEEE Trans. Softw. Eng. 2022, 48, 2526–2540. [Google Scholar] [CrossRef] [Scilit]
  18. Mei, H.; Hao, D.; Zhang, L.; Zhang, L.; Zhou, J.; Rothermel, G. A static approach to prioritizing JUnit test cases. IEEE Trans. Softw. Eng. 2012, 38, 1258–1275. [Google Scholar] [CrossRef] [Scilit]
  19. Li, F.; Zhou, J.; Li, Y.; Hao, D.; Zhang, L. AGA: An Accelerated Greedy Additional Algorithm for Test Case Prioritization. IEEE Trans. Softw. Eng. 2022, 48, 5102–5119. [Google Scholar] [CrossRef] [Scilit]
  20. Biswas, S.; Bansal, A.; Mitra, P.; Mall, R. Fault-Based Regression Test Case Prioritization. IEEE Trans. Reliab. 2023, 72, 1176–1190. [Google Scholar] [CrossRef] [Scilit]
  21. Sugave, S.R.; Kulkarni, Y.R.; Jagdale, B.; Gutte, V. Fault-Aware Test Case Prioritization in Software Testing Using Jaya Archimedes Optimization Algorithm. J. Electron. Test. Theory Appl. (JETTA) 2025, 41, 41–61. [Google Scholar] [CrossRef] [Scilit]
  22. Iqbal, S.; Al-Azzoni, I. Test case prioritization for model transformations. J. King Saud Univ.-Comput. Inf. Sci. 2022, 34, 6324–6338. [Google Scholar] [CrossRef] [Scilit]
  23. J, S.; Sharma, S.; Tripathi, M.K. Hybrid optimization with constraints handling for combinatorial test case prioritization problems. Netw. Comput. Neural Syst. 2025, 1–31. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Li, Z.; Harman, M.; Hierons, R.M. Search algorithms for regression test case prioritization. IEEE Trans. Softw. Eng. 2007, 33, 225–237. [Google Scholar] [CrossRef] [Scilit]
  25. Malhotra, R.; Bharadwaj, A. Test Case Prioritization Using Genetic Algorithm. Int. J. Comput. Sci. Inform. 2014, 3, 151–154. [Google Scholar] [CrossRef] [Scilit]
  26. Wambua, A.W.; Wambugu, G.M. A Comparative Analysis of Bat and Genetic Algorithms for Test Case Prioritization in Regression Testing. Int. J. Intell. Syst. Appl. 2023, 15, 13–21. [Google Scholar] [CrossRef] [Scilit]
  27. Anuar, M.A.; Sahid, M.Z.; Zainal, N. Comparative Analysis of Test Case Prioritization Using Ant Colony Optimization Algorithm and Genetic Algorithm. J. Soft Comput. Data Min. 2023, 4, 52–58. [Google Scholar] [CrossRef] [Scilit]
  28. Demir, Z.C.; Amrahov, Ş.E. Dominating set-based test prioritization algorithms for regression testing. Soft Comput. 2022, 26, 8203–8220. [Google Scholar] [CrossRef] [Scilit]
  29. Sabharwal, S.; Sibal, R.; Sharma, C. Applying Genetic Algorithm for Prioritization of Test Case Scenarios Derived from UML Diagrams. IJCSI Int. J. Comput. Sci. Issues 2011, 8, 2. [Google Scholar]
  30. Abubakar, H.; Zambuk, F.U.; Ahmed, U.M.; Gital, A.Y. Quality-Aware Genetic Algorithm Based Cost Cognizant Test Case Prioritization for Object-Oriented Programs. J. Sci. Technol. Educ. 2023, 11, 81–92. [Google Scholar] [CrossRef]
  31. Mukherjee, R.; Patnaik, K.S. Prioritizing JUnit Test Cases Without Coverage Information: An Optimization Heuristics Based Approach. IEEE Access 2019, 7, 78092–78107. [Google Scholar] [CrossRef] [Scilit]
  32. Nucci, D.D.; Panichella, A.; Zaidman, A.; Lucia, A.D. A Test Case Prioritization Genetic Algorithm Guided by the Hypervolume Indicator. IEEE Trans. Softw. Eng. 2020, 46, 674–696. [Google Scholar] [CrossRef] [Scilit]
  33. Adasme, P.; Viveros, A.; Dehghan Firoozabadi, A. Quadratic p-Median Problem: A Bender’s Decomposition and a Meta-Heuristic Local-Based Approach. Symmetry 2024, 16, 1114. [Google Scholar] [CrossRef] [Scilit]
  34. Li, X.; Lan, L.; Lahza, H.; Yang, S.; Wang, S.; Yang, W.; Liu, H.; Zhang, Y. Dragon Boat Optimization: A Meta-Heuristic for Intelligent Systems. Expert Syst. 2025, 42, e13785. [Google Scholar] [CrossRef] [Scilit]
  35. Assiri, M. Test Case Prioritization Using Dragon Boat Optimization for Software Quality Testing. Electronics 2025, 14, 1524. [Google Scholar] [CrossRef] [Scilit]
  36. Vescan, A.; Pintea, C.M.; Pop, P. Test Case Prioritization—ANT Algorithm With Faults Severity. Log. J. IGPL 2020, 30, 277–288. [Google Scholar] [CrossRef] [Scilit]
  37. Azeem, S.; Javed, S.; Naseer, I.; Ali, O.; Ghazal, T. A New Hybrid PSO-HHO Wrapper Based Optimization for Feature Selection. IEEE Access 2025, 13, 87090–87099. [Google Scholar] [CrossRef] [Scilit]
  38. Cao, Y.; Li, S.; Shen, G.; Chen, H.; Liu, Y. Intelligent dynamic control of shield parameters using a hybrid algorithm and digital twin platform. Autom. Constr. 2025, 169, 105882. [Google Scholar] [CrossRef] [Scilit]
  39. Huang, W.; Zhang, J.; Li, X.; Zhou, X.; Qi, D.; Xi, J.; Liu, W. A Semantic and Intelligent Focused Crawler based on BERT Semantic Vector Space Model and Hybrid Algorithm. IEEE Access 2025, 14, 87713–87727. [Google Scholar] [CrossRef] [Scilit]
  40. Kartli, N. Hybrid algorithms for fixed charge transportation problem. Kybernetika 2025, 61, 141–167. [Google Scholar] [CrossRef] [Scilit]
  41. Kartli, N.; Cetin, P.; Ayhan, S. New gravitational algorithms for the detection of overlapping and disjoint communities in weighted complex networks. Kybernetika 2025, 61, 509–536. [Google Scholar] [CrossRef] [Scilit]
  42. Qi, R.; Jia, Y.H.; Chen, W.n.; Bi, Y.; Mei, Y. An evolutionary optimization-learning hybrid algorithm for energy resource management. Swarm Evol. Comput. 2025, 92, 101831. [Google Scholar] [CrossRef] [Scilit]
  43. Zhang, B.; Yin, Y.; Li, B.; He, S.; Song, J. A hybrid algorithm for predicting the remaining service life of hybrid bearings based on bidirectional feature extraction. Measurement 2025, 242, 116152. [Google Scholar] [CrossRef] [Scilit]
  44. Merze, A.; Çelebi, F.V. Advanced deep reinforcement learning for optimizing 3D printing toolpaths: A framework with enhanced agent architectures, Count-Prioritized Replay, and curriculum learning. Eng. Sci. Technol. Int. J. 2025, 72, 102205. [Google Scholar] [CrossRef] [Scilit]
  45. Fesli, U.; Ozdemir, M.B.; Akın, M. Grey Wolf Algorithm-Based source size and location determination method for capacity expansion planning in power systems. Eng. Sci. Technol. Int. J. 2025, 61, 101934. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Flowchart illustrating the main phases of the proposed MH-DBO-GA framework for test case prioritization.
Figure 1. Flowchart illustrating the main phases of the proposed MH-DBO-GA framework for test case prioritization.
Symmetry 18 01207 g001
Figure 2. APRC Scores of Deterministic Algorithms for Dataset 1.
Figure 2. APRC Scores of Deterministic Algorithms for Dataset 1.
Symmetry 18 01207 g002
Figure 3. APRC Scores of Deterministic Algorithms for Dataset 2.
Figure 3. APRC Scores of Deterministic Algorithms for Dataset 2.
Symmetry 18 01207 g003
Figure 4. Statistical distribution of APRC scores obtained in 50 independent runs for Dataset 1.
Figure 4. Statistical distribution of APRC scores obtained in 50 independent runs for Dataset 1.
Symmetry 18 01207 g004
Figure 5. Statistical distribution of APRC scores obtained in 50 independent runs for Dataset 2.
Figure 5. Statistical distribution of APRC scores obtained in 50 independent runs for Dataset 2.
Symmetry 18 01207 g005
Table 1. Comparison of test cost reduction techniques.
Table 1. Comparison of test cost reduction techniques.
TechniquePrimary ObjectiveStrategy
Test Case MinimizationTo permanently reduce the size of the test suite.Eliminates redundant test cases that verify the same requirements.
Test Case SelectionTo execute only relevant tests for specific modifications.Focuses on recent code changes and selects test cases associated with those changes.
Test Case PrioritizationTo optimize the execution order of test cases.Rearranges test cases based on specific criteria (e.g., coverage and fault-detection rate).
Table 2. Overview of Dataset Properties and RTM Sparsity Metrics.
Table 2. Overview of Dataset Properties and RTM Sparsity Metrics.
Feature/AttributeDataset 1 (Kaggle)Dataset 2 (Public)
DomainAutomotivePublic Sector
Total Test Cases ( | T | )33991467
Total Requirements ( | R | )20001237
Theoretical Max Entries6,798,0001,814,679
Total Non-Zero Links (L)58355331
Singleton Requirements10872
Matrix Density (%)0.0858%0.2938%
Matrix Sparsity (%)99.9142%99.7062%
Note: Singleton requirements refer to requirements covered by only one test case.
Table 3. Parameter Settings and Implementation Choices of the Metaheuristic Refinement Component.
Table 3. Parameter Settings and Implementation Choices of the Metaheuristic Refinement Component.
Parameter GroupParameterValue/Setting
General SetupPopulation Size (N)50
Maximum Iterations (K)100
Stopping ConditionFixed iteration budget; no early stopping
InitializationBottleneck-Aware Greedy Initialization Ratio80%
Random Initialization Ratio20%
Greedy Construction LimitUp to 500 prioritized tests; at most 1000 candidates scanned per greedy step
DBOA PhaseDynamic Parameters ( ψ , λ , μ , H )Adapted stochastic update rules in Section 5.7
Position Clamping Range [ 0.0, 1.2]
Memetic Refinement PhaseLocal Search RegionFirst 300 test cases
Local Search Trials per Iteration100 candidate swaps
Bottleneck Hunter ConditionApplied when the saturation index is greater than 500
Bottleneck Teleportation Value1.1
GA PhaseElite Selection Rate10%
Elimination Rate25%
Parent SelectionUniform random selection from the elite pool
Crossover OperatorOrder-preserving one-point crossover
Mutation OperatorFive random swaps
Mutation Rate20%
Randomness HandlingJava pseudo-random generators without manually fixed seeds
Table 4. APRC Performance Comparison of Deterministic and Metaheuristic Components.
Table 4. APRC Performance Comparison of Deterministic and Metaheuristic Components.
Panel A: Deterministic Methods (Single-Run Evaluations)
AlgorithmDataset 1 Exact APRC (%)Dataset 2 Exact APRC (%)
AG+BH89.9295.57
Additional Greedy89.9251.57
Standard Greedy84.6936.36
2-Optimal89.9668.84
Panel B: Metaheuristics (50 Runs)—Dataset 1
AlgorithmMean ± Std (%)95% CI (Mean)p-ValueSigned δ 95% CI ( δ )
MH-DBO-GA85.68 ± 0.15[85.64, 85.72]<0.001+1.00 (L)[1.00, 1.00]
GA74.61 ± 0.20 *[74.55, 74.67]–––
DBOA74.02 ± 0.22[73.96, 74.08]<0.001−0.94 (L)[−0.99, −0.86]
Panel C: Metaheuristics (50 Runs)—Dataset 2
AlgorithmMean ± Std (%)95% CI (Mean)p-ValueSigned δ 95% CI ( δ )
MH-DBO-GA95.14 ± 0.05[95.13, 95.15]<0.001+1.00 (L)[1.00, 1.00]
GA83.10 ± 0.47[82.97, 83.23]<0.001−0.08 (S)[−0.31, 0.15]
DBOA83.15 ± 0.61 *[82.98, 83.32]–––
Note: AG+BH denotes the bottleneck-aware heuristic component combining Additional Greedy and Bottleneck Hunter. MH-DBO-GA denotes the metaheuristic refinement component integrating Dragon Boat Optimization, Genetic Algorithm operators, and memetic local search. Since deterministic methods (Panel A) yield zero variance, we evaluate them using exact single-run APRC scores without statistical testing. On Dataset 1, 2-Optimal achieves the highest deterministic APRC, while AG+BH obtains a very close APRC with substantially lower runtime. On Dataset 2, AG+BH achieves the highest deterministic APRC observed in our experiments. Statistical analyses for the stochastic algorithms (Panels B and C) are conducted over 50 independent runs to summarize repeated-run variability. For pairwise comparisons (Mann–Whitney U test), we select the strongest standard metaheuristic baseline for each dataset (* GA for Dataset 1; DBOA for Dataset 2). p-values correspond to comparisons against these baselines, with Holm-Bonferroni corrections applied to control for multiple testing. δ denotes Cliff’s Delta (rank-biserial effect size), signed to indicate direction of superiority. Positive δ means the algorithm is better than the baseline, while a negative δ indicates the baseline is better. Confidence intervals for Cliff’s Delta (95% CI δ ) were computed via bootstrap resampling with 10,000 iterations to robustly estimate effect size uncertainty. We classify effect sizes as Large ( | δ | ≥ 0.474 ) and Small otherwise.
Table 5. Saturation Point and Execution Time—Dataset 1.
Table 5. Saturation Point and Execution Time—Dataset 1.
AlgorithmSaturation PointRunning Time (s)
MH-DBO-GA191310.09
AG+BH9981.292
DBOA33770.98
GA333729.15
2-Optimal97484,180
Additional Greedy9980.66
Greedy33000.08
Table 6. Saturation Point and Execution Time—Dataset 2.
Table 6. Saturation Point and Execution Time—Dataset 2.
AlgorithmSaturation PointRunning Time (s)
MH-DBO-GA4467.328
AG+BH2820.419
DBOA13970.721
GA139222.604
2-Optimal2843179.136
Additional Greedy2830.077
Greedy14640.051
Table 7. Execution Times for Dataset 1 and Dataset 2.
Table 7. Execution Times for Dataset 1 and Dataset 2.
AlgorithmDataset 1 (s)Dataset 2 (s)
MH-DBO-GA10.097.328
AG+BH1.2920.419
DBOA0.980.721
GA29.1522.604
2-Optimal84,1803179.136
Additional Greedy0.660.077
Greedy0.080.051
Table 8. Sensitivity Analysis of Greedy Initialization Ratio.
Table 8. Sensitivity Analysis of Greedy Initialization Ratio.
Panel A: Dataset 1
Greedy RatioMean ± Std (%)95% CI (Mean)Runtime Mean ± Std (s)
60%85.61 ± 0.12[85.58, 85.65]11.01 ± 0.78
80%85.68 ± 0.15[85.64, 85.72]10.09 ± 1.05
100%85.63 ± 0.16[85.59, 85.68]10.56 ± 0.06
Panel B: Dataset 2
Greedy RatioMean ± Std (%)95% CI (Mean)Runtime Mean ± Std (s)
60%95.07 ± 0.07[95.05, 95.09]6.21 ± 0.14
80%95.14 ± 0.05[95.13, 95.15]7.32 ± 0.54
100%95.11 ± 0.06[95.09, 95.12]7.18 ± 0.38
Note: Each setting includes 50 independent stochastic runs. The table reports the mean APRC, standard deviation, 95% confidence interval of the mean, and mean runtime with standard deviation.
Table 9. Sensitivity Analysis of Local Search Region Size.
Table 9. Sensitivity Analysis of Local Search Region Size.
Panel A: Dataset 1
Local Search RegionMean ± Std (%)95% CI (Mean)Runtime Mean ± Std (s)
100 Test Cases85.07 ± 0.15[85.03, 85.12]10.68 ± 0.22
300 Test Cases85.68 ± 0.15[85.64, 85.72]10.09 ± 0.15
500 Test Cases86.35 ± 0.13[86.32, 86.39]10.40 ± 0.06
Panel B: Dataset 2
Local Search RegionMean ± Std (%)95% CI (Mean)Runtime Mean ± Std (s)
100 Test Cases95.03 ± 0.04[95.02, 95.04]5.91 ± 0.05
300 Test Cases95.14 ± 0.05[95.13, 95.15]7.32 ± 0.14
500 Test Cases94.95 ± 0.06[94.94, 94.97]6.07 ± 0.12
Note: Each setting includes 50 independent stochastic runs. The table reports the mean APRC, standard deviation, 95% confidence interval of the mean, and mean runtime with standard deviation.
Table 10. Bottleneck Teleportation Values for Dataset 1 and Dataset 2.
Table 10. Bottleneck Teleportation Values for Dataset 1 and Dataset 2.
Panel A: Dataset 1
Teleportation ValueMean ± Std (%)95% CI (Mean)Runtime Mean ± Std (s)
1.085.676 ± 0.142[85.638, 85.717]10.58 ± 0.47
1.185.680 ± 0.150[85.640, 85.720]10.09 ± 0.28
1.285.684 ± 0.147[85.644, 85.724]10.29 ± 0.23
Panel B: Dataset 2
Teleportation ValueMean ± Std (%)95% CI (Mean)Runtime Mean ± Std (s)
1.095.124 ± 0.047[95.111, 95.137]6.08 ± 0.13
1.195.140 ± 0.050[95.130, 95.150]7.32 ± 0.24
1.295.133 ± 0.046[95.120, 95.145]6.28 ± 0.27
Note: Each setting includes 50 independent stochastic runs. The table reports the mean APRC, standard deviation, 95% confidence interval of the mean, and mean runtime with standard deviation. APRC values are shown with three decimal places because the differences among teleportation values are very small.
Table 11. Position Clamping Range Values for Dataset 1 and Dataset 2.
Table 11. Position Clamping Range Values for Dataset 1 and Dataset 2.
Panel A: Dataset 1
Clamping RangeMean ± Std (%)95% CI (Mean)Runtime Mean ± Std (s)
[0.0, 1.0]85.663 ± 0.135[85.625, 85.702]10.94 ± 0.48
[0.0, 1.2]85.680 ± 0.150[85.640, 85.720]10.09 ± 0.64
[0.0, 1.5]85.652 ± 0.130[85.615, 85.689]12.67 ± 1.19
Panel B: Dataset 2
Clamping RangeMean ± Std (%)95% CI (Mean)Runtime Mean ± Std (s)
[0.0, 1.0]95.125 ± 0.054[95.109, 95.140]6.29 ± 0.20
[0.0, 1.2]95.140 ± 0.050[95.130, 95.150]7.32 ± 0.27
[0.0, 1.5]95.116 ± 0.052[95.101, 95.130]7.29 ± 0.60
Note: Each setting includes 50 independent stochastic runs. The table reports the mean APRC, standard deviation, 95% confidence interval of the mean, and mean runtime with standard deviation. APRC values are shown with three decimal places because the differences among clamping ranges are very small.
Table 12. Summary of Selected Parameter Settings Based on Sensitivity Analysis.
Table 12. Summary of Selected Parameter Settings Based on Sensitivity Analysis.
ParameterTested ValuesSelected ValueInterpretation
Greedy Initialization Ratio60%, 80%, 100%80%Highest mean APRC among the tested ratios, with partial random initialization for diversity
Local Search Region100, 300, 500 test cases300 test casesMiddle search-scope setting; best mean APRC on Dataset 2 and higher APRC than 100 test cases on Dataset 1
Bottleneck Teleportation Value1.0, 1.1, 1.21.1Moderate perturbation value; APRC differences are small across tested values
Position Clamping Range[0.0, 1.0], [0.0, 1.2], [0.0, 1.5][0.0, 1.2]Moderate extension of the normalized range; the largest bound does not improve mean APRC
Table 13. Ablation Study Results Based on APRC.
Table 13. Ablation Study Results Based on APRC.
MethodDataset 1 APRCDataset 2 APRC
MH-DBO-GA0.8568470.951380
Without Bottleneck Hunter0.8426810.951710
Without Informed Initialization0.8435750.943220
DBOA + Memetic Only0.8582110.951424
GA + Memetic Only0.8564130.951539
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Erkaya, A.E.; Emrah Amrahov, S.; Çelebi, F.V. Bottleneck-Aware Heuristic and Metaheuristic Framework for Requirement-Based Test Case Prioritization in Sparse Traceability Matrices. Symmetry 2026, 18, 1207. https://doi.org/10.3390/sym18071207

AMA Style

Erkaya AE, Emrah Amrahov S, Çelebi FV. Bottleneck-Aware Heuristic and Metaheuristic Framework for Requirement-Based Test Case Prioritization in Sparse Traceability Matrices. Symmetry. 2026; 18(7):1207. https://doi.org/10.3390/sym18071207

Chicago/Turabian Style

Erkaya, Ahmed Enis, Sahin Emrah Amrahov, and Fatih V. Çelebi. 2026. "Bottleneck-Aware Heuristic and Metaheuristic Framework for Requirement-Based Test Case Prioritization in Sparse Traceability Matrices" Symmetry 18, no. 7: 1207. https://doi.org/10.3390/sym18071207

APA Style

Erkaya, A. E., Emrah Amrahov, S., & Çelebi, F. V. (2026). Bottleneck-Aware Heuristic and Metaheuristic Framework for Requirement-Based Test Case Prioritization in Sparse Traceability Matrices. Symmetry, 18(7), 1207. https://doi.org/10.3390/sym18071207

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop