1. Introduction
The software development process is carried out within a holistic framework called the Software Development Life Cycle (SDLC), extending from requirements analysis to the delivery of the final product [
1]. This process consists of planning, requirements analysis, design, development, testing, and maintenance phases, each of which has a direct impact on software quality. Among these phases, the testing process plays a critical role in ensuring the software meets reliability, performance, and functionality requirements. Thoroughly testing software systems before deploying them to the production environment enables early detection of potential errors and significantly reduces maintenance costs [
2].
Regression testing aims to verify that new features added to the software or changes made to existing components do not negatively impact the system’s previous functionality [
2]. In modern software projects that require continuous change and updates, regression testing is becoming increasingly critical. Neglecting these tests can lead to serious system-wide errors even with minor changes [
3]. Furthermore, regression testing processes consume significant resources due to their high time and cost requirements. The literature reports that a single regression test cycle can take weeks [
4]. Therefore, regression testing is increasingly automated through continuous integration (CI) and continuous delivery (CD) pipelines.
Three main approaches have been developed in the literature to reduce testing costs: test set reduction, test case selection, and test case prioritization [
2,
5]. Test set reduction permanently eliminates unnecessary or redundant test cases, while the test selection approach executes only tests related to current changes. Test case prioritization, on the other hand, aims to identify important requirements and errors at the earliest possible stage by reordering tests according to specific criteria. This study focuses on the test case prioritization problem among these approaches [
4,
6]. The main differences between these approaches are summarized in
Table 1.
In the test case prioritization problem, various criteria such as code coverage, potential for error detection, historical test data, risk factors, and requirement priorities are used [
7]. Different paradigms and techniques have been proposed in the literature, including Bayesian models, machine learning methods, and metaheuristic algorithms [
3,
8,
9,
10]. Requirement-based approaches aim to prioritize test cases according to characteristics such as customer priorities and implementation complexity, and require strong traceability mechanisms [
5,
11,
12,
13]. In this context, the Requirement Traceability Matrix (RTM) is used as an important tool in managing the relationship between requirements and test cases [
14].
The effectiveness of test case prioritization methods is generally evaluated using the Average Percentage of Faults Detected (APFD) metric [
3,
7,
8,
15,
16,
17,
18,
19,
20]. However, in requirements-based studies, requirements coverage performance takes precedence over error detection. Therefore, in this study, the APFD metric has been adapted as the Average Percentage of Requirements Covered (APRC) [
21,
22,
23].
A substantial portion of existing studies on requirement-based test case prioritization has been evaluated using datasets with relatively dense requirement-test relationships. However, real-world software projects often contain sparse requirement traceability matrices (RTMs), where test cases are designed to cover specific functionalities with limited overlap. Such sparse structures present optimization problems, particularly due to the singleton requirements that are associated with only a single test case. These isolated requirements may become coverage bottlenecks by delaying the achievement of complete requirement coverage. This can result in slower coverage growth during the early stages of regression testing.
To address these challenges, this study proposes a bottleneck-aware heuristic and metaheuristic framework for requirement-based test case prioritization under sparse RTMs. The framework is centered on a deterministic problem-specific heuristic layer that combines Additional Greedy with the proposed Bottleneck Hunter mechanism to identify and prioritize test cases associated with delayed requirement coverage. A secondary metaheuristic refinement layer, based on Dragon Boat Optimization, Genetic Algorithm operators, and memetic local search, is incorporated to provide additional exploration capability in larger search spaces. By combining structural information from sparse RTMs with adaptive search mechanisms, the framework aims to improve early requirement coverage while keeping the main empirical emphasis on bottleneck-aware deterministic prioritization.
In light of the experimental results, we explicitly position the deterministic AG+BH component as the primary practical contribution of this study. The results indicate that AG+BH offers a favorable balance between APRC, saturation behavior, and runtime in the evaluated single-objective setting. Although 2-Optimal obtains a slightly higher APRC on Dataset 1, it requires 23 h 23 min to complete, whereas AG+BH reaches a near-best APRC with substantially lower runtime and obtains the highest APRC on Dataset 2. Therefore, MH-DBO-GA is not presented as a generally superior alternative to AG+BH or Additional Greedy. Instead, it is considered a secondary exploratory extension that may be useful when additional search diversity is required or when future versions of the problem include further constraints such as execution cost, requirement priority, or fault severity.
The main contributions of this study can be summarized as follows:
Identification of bottleneck characteristics in sparse requirement traceability matrices.
This study analyzes the structural challenges of sparse RTMs, particularly the delayed coverage problem caused by isolated and singleton requirements. The study highlights that such bottlenecks may limit the effectiveness of conventional optimization approaches by reducing early requirement coverage capability.
Development of the deterministic AG+BH strategy as the primary bottleneck-aware heuristic.
The main contribution of this work is a deterministic heuristic layer that combines Additional Greedy with the proposed Bottleneck Hunter mechanism. This component uses problem-specific information from the requirement–test relationship structure to identify critical test cases and improve early requirement coverage. In the evaluated APRC setting, AG+BH demonstrates that problem-specific bottleneck handling can produce highly competitive coverage results with low computational cost.
Positioning of MH-DBO-GA as a secondary exploratory metaheuristic extension.
The framework also includes a metaheuristic refinement layer based on Dragon Boat Optimization, Genetic Algorithm operators, and memetic local search. This component is not claimed to be uniformly preferable to AG+BH under the current experimental setting. Instead, it is used as a complementary population-based extension that can provide additional search diversity and may be more relevant in future dynamic or multi-objective prioritization scenarios.
Comprehensive experimental evaluation and component-level analysis.
The proposed framework is evaluated on sparse requirement traceability matrix datasets using the Average Percentage of Requirements Covered (APRC) metric and saturation analysis. The evaluation separates the deterministic AG+BH strategy from the stochastic MH-DBO-GA component and interprets their roles according to the observed results. In addition, an ablation study is conducted to investigate the individual effects of informed initialization, Bottleneck Hunter, and hybrid optimization components.
The rest of the article is structured as follows: The second section presents relevant studies, the third section presents problem definition, the fourth section provides preliminaries, the fifth section describes the proposed framework in detail, the sixth section reports the experimental setup and results, the seventh section analyzes the contribution of individual components, the eighth section discusses threats to validity and limitations, and the final section discusses the findings.
2. Related Work
Software testing processes are considered one of the most complex and critical stages of the Software Development Life Cycle (SDLC) [
1]. In this context, test case prioritization (TCP) aims to reduce maintenance costs by increasing the failure detection rate at an early stage [
4]. In the literature, a wide range of methods have been proposed for the TCP problem, from classical heuristic methods to nature-inspired metaheuristic algorithms.
In early studies, greedy algorithms were widely used [
24]. Rothermel et al. [
4] compared nine different prioritization techniques in their work on Siemens programs and showed that prioritization methods significantly increased the error detection rate. However, the inability of greedy approaches to reach the global optimum led to the emergence of improved methods such as Additional Greedy and 2-Optimal.
To overcome these limitations, evolutionary and metaheuristic algorithms have become widely used in the TCP problem. Li et al. [
24] investigated the coverage performance in regression tests by comparing Genetic Algorithm (GA), Hill Climbing, and different greedy methods. Malhotra and Bharadwaj [
25] proposed a GA model based on error detection capability. Mishra et al. [
3] managed to reduce test costs with a GA-based method combining coverage rate and error potential. Wambua et al. [
26] compared GA and Bat Algorithm in terms of APFD and memory usage, reporting that GA provided higher performance. Anuar et al. [
27] examined the effectiveness of Ant Colony Optimization and GA-based models.
Some studies have developed specialized models that consider requirements traceability and software structural features. Demir et al. [
28] optimized the test set according to requirements coverage by proposing a method based on the dominant cluster approach. Sabharwal et al. [
29] developed a GA-based model using the information flow metric for scenarios derived from UML diagrams. Abubakar et al. [
30] proposed the QAG-TCP method, which includes metrics such as connectivity, compliance, and fault severity.
In cases where coverage information is limited, Mukherjee and Patnaik [
31] compared GA, Simulated Annealing, and Ant Colony methods by treating the problem as a 0/1 knapsack model. In scenarios involving multiple coverage measures, Di Nucci et al. [
32] proposed a hybrid GA model driven by hypervolume indicators.
Similar motivations also appear in broader combinatorial optimization studies. Adasme et al. [
33] combined Bender’s decomposition with a local-search metaheuristic for the quadratic p-median problem, showing that exact and heuristic strategies can play complementary roles in difficult optimization settings. Although this problem differs from requirement-based TCP, it provides additional motivation for using problem-specific heuristic search in large combinatorial spaces.
Recently, dragon boat inspired optimization has also been considered in TCP research. Dragon Boat Optimization was originally introduced by Li et al. as a human-based metaheuristic inspired by dragon boat racing [
34]. Assiri [
35] later adapted this optimization idea to test case prioritization and reported improvements in terms of APFD and execution time. In the present study, DBOA is not used as a standalone TCP method; instead, its movement principles are adapted to a random-key representation as part of the secondary MH-DBO-GA exploration layer.
However, a large portion of the studies in the literature have been evaluated on controlled and relatively balanced datasets. In real-world software projects, sparse requirement-test case matches and singleton requirement structures pose significant challenges for existing methods. In particular, studies incorporating specific mechanisms for addressing structural bottlenecks that delay full coverage are limited. Furthermore, the scalability of methods in large-scale and irregular data structures has not been adequately demonstrated. Vescan et al. show that Ant Colony Optimization (ACO) is successful in fault detection-based test case prioritization [
36]. However, when applied alone to the ’sparse requirements matrices’ problem addressed in our study, it did not exhibit the expected performance.
This study addresses this gap by focusing on bottleneck-aware prioritization under sparse requirement traceability matrices. The proposed framework emphasizes a deterministic AG+BH strategy that directly uses sparse RTM structure to identify delayed-coverage test cases. A metaheuristic refinement layer is also included to provide additional population-based exploration, but the main empirical emphasis is placed on the bottleneck-aware heuristic behavior observed in the evaluated APRC setting.
5. Proposed Framework
5.1. Overview of the Proposed Framework
In recent years, hybrid algorithms, which combine the strengths of different optimization methods, have been shown to produce successful results in various engineering problems [
37,
38,
39,
40,
41,
42,
43,
44,
45].
This paper introduces a bottleneck-aware two-layer framework for requirement-based test case prioritization under sparse RTMs. The first layer is the deterministic AG+BH strategy, which combines Additional Greedy with the Bottleneck Hunter mechanism. This layer directly uses the structure of sparse RTMs to identify delayed-coverage test cases, especially those related to singleton requirements.
The second layer is the MH-DBO-GA metaheuristic refinement component. It combines Dragon Boat Optimization, Genetic Algorithm operators, and memetic local search to provide additional population-based exploration. In the current single-objective APRC evaluation, this layer is not presented as a replacement for AG+BH. Rather, it is treated as an exploratory extension that may be useful when deterministic bottleneck rules are not sufficient or when additional constraints are introduced.
As shown in
Figure 1, the metaheuristic refinement component consists of four main components: (i) preprocessing, (ii) informed initialization, (iii) hybrid optimization process (DBOA-GA), and (iv) memetic local search mechanism.
The hybrid design combines DBOA and GA for complementary purposes. DBOA is responsible for rapidly identifying promising regions of the search space. On the other hand, GA operators maintain population diversity and generate alternative candidate solutions during the optimization process.
In addition, the memetic local search mechanism called Bottleneck Hunter, developed in this study, improves solution quality by identifying critical test cases that delay complete requirement coverage. This mechanism provides performance gains, especially in sparse requirement-test matrices.
Accordingly, the framework should be read as a heuristic-centered approach with an optional metaheuristic refinement layer. The experimental results indicate that the deterministic bottleneck-aware layer accounts for the main empirical improvement in the studied APRC setting, while MH-DBO-GA provides an additional search mechanism for cases where broader exploration is needed.
5.2. System Architecture
The proposed MH-DBO-GA framework has a multi-stage architecture that integrates global search and local optimization mechanisms. This architecture consists of four core modules: preprocessing, informed initialization, hybrid optimization, and memetic local search. They all complement each other to improve solution quality and make the search process more efficient.
In the preprocessing stage, the problem structure is analyzed to identify singleton requirements, which are covered by only a single test case.
In the second stage, an informed initialization strategy is used to improve the quality of the initial population. This approach starts from regions with high potential in the solution space instead of random initialization. It also contributes to obtaining better solutions in the early stages of the search process.
In the third stage, we carry out a hybrid optimization process using the Dragon Boat Optimization Algorithm (DBOA) and the Genetic Algorithm (GA) together. At this stage, DBOA guides the population globally. On the other hand, GA operators ensure the search process progresses in a balanced manner by maintaining solution diversity.
In the final stage, a memetic local search mechanism appears. In this stage, problem-specific improvements are made to the best available solution, enhancing solution quality. Specifically, the Bottleneck Hunter mechanism directly improves solution sequencing by identifying critical test cases that cause delayed requirement coverage.
5.3. Informed Initialization Strategy
In the proposed approach, the quality of the initial population is considered a critical factor directly affecting the success of the optimization process. Therefore, instead of purely random initialization, we propose an informed initialization strategy.
This strategy consists of two main steps. In the first step, requirements covered by only a single test case () are identified. Such requirements are critical components that can negatively affect the total coverage time if they are delayed in the solution sequence. Therefore, test cases containing these requirements are prioritized in the initial phase.
In the second step, a large portion of the population (approximately 80%) is generated using the Additional Greedy approach. This method builds a stepwise solution by selecting the test case that contributes most to the uncovered requirements in each iteration. Thus, the generated initial individuals have a high coverage potential.
The remaining portion of the population (20%) is generated randomly to maintain diversity. This hybrid initialization approach allows for a rapid start in promising regions of the solution space while also contributing to maintaining sufficient diversity in the search process.
Consequently, the proposed knowledge-based initialization strategy reduces time loss that can be caused by low-quality initializations and makes it possible to obtain better solutions in the early stages of the optimization process.
5.4. Hybrid Optimization Mechanism (DBOA + GA)
The proposed MH-DBO-GA framework is based on a hybrid optimization mechanism established between the Dragon Boat Optimization Algorithm (DBOA) and the Genetic Algorithm (GA). This structure aims to improve solution quality by providing a balanced interaction between global exploration and local exploitation processes.
DBOA exhibits rapid convergence behavior by guiding the population around the global best individual (leader). While the movements of individuals in the solution space are controlled through acceleration () and weakening () coefficients, the social behavior factor () provides flexibility to the search process by modeling intra-population interactions. Thanks to this structure, DBOA can perform extensive exploration in the solution space.
However, early convergence and local minima problems, frequently observed in swarm intelligence-based methods, can limit performance when using DBOA alone. To overcome this disadvantage, GA operators are incorporated into the hybrid structure.
The GA component increases diversity by introducing new genetic variations into the population through crossover and mutation operators. Order crossover and random mutation mechanisms, particularly suitable for sequencing problems, allow for the recombination of existing solutions, enabling transitions to different regions of the search process.
Within the hybrid structure, low-performing individuals are eliminated at the end of each iteration, and new solutions are generated using elite individuals. This process increases exploration opportunities, further preventing population stabilization.
Within the proposed framework, DBOA and GA operators provide a population-based exploration layer. This layer can generate alternative candidate sequences and maintain diversity during the search process. However, in the current APRC-based experiments, this metaheuristic layer is considered complementary to the deterministic AG+BH strategy rather than the primary source of empirical improvement.
5.5. Memetic Local Search (Bottleneck Hunter)
One of the most distinctive components of the proposed MH-DBO-GA framework is the problem-specific memetic local search mechanism, Bottleneck Hunter. This mechanism directly targets delays in the requirements coverage process by improving upon the best available solution.
In the requirements-based test case prioritization problem, one of the most critical factors determining solution quality is the position at which each requirement is first covered. In this context, the value represents the order of the test case that first covers the i-th requirement. Examining the APRC metric, it is seen that directly minimizing the sum of these values improves solution quality.
Accordingly, the saturation index (
) indicates the last position at which all requirements are satisfied for the first time.
Solutions with a high value show that at least one requirement is covered quite late, negatively affecting the APRC value. The test case causing this delay is considered a bottleneck in the solution sequence.
The proposed Bottleneck Hunter mechanism aims to identify this critical test case and increase its priority in the decoded sequence. In the implementation, this is done by assigning a high random-key value to the bottleneck test when the saturation index exceeds a predefined threshold. This intervention can reduce the value of delayed requirements and, consequently, can reduce the sum of .
The main advantage of this approach is that it directly targets the components of the objective function. In sparse requirements-test matrices, especially when requirements covered by only a single test case are located late in the sequence, the values can increase disproportionately.
The memetic local search restricts one of the swap indices to the first positions of the test sequence, while the second index is sampled from the whole sequence. This design focuses the local refinement on the earlier part of the ranking, where improvements have a stronger effect on APRC. The value of 300 is used as a practical middle setting based on the sensitivity analysis. It improves over the 100-test-case setting on Dataset 1 and gives the highest mean APRC on Dataset 2, although the 500-test-case setting performs better on Dataset 1. Therefore, this value should be interpreted as a practical search-scope choice rather than as a universally optimal threshold.
In conclusion, the proposed memetic mechanism refines solutions obtained through the global search process at the local level, thereby increasing convergence speed and contributing to the achievement of higher-quality solutions.
5.6. Algorithm Workflow
The general operation of the proposed MH-DBO-GA algorithm is presented in Algorithm 1. The algorithm works iteratively in a multi-stage structure, gradually improving the quality of the solution.
In the first stage, a preprocessing step is performed to analyze the requirement-test relationships and specifically identify the requirements covered by only a single test case (). This information is stored for use in subsequent stages.
In the second stage, an initial population is created using an information-based initialization strategy. While a large portion of the population is generated using the Additional Greedy approach, the remaining individuals are randomly generated to ensure diversity.
In the third stage, the main optimization cycle of the algorithm is initiated. In each iteration, individuals are converted from their representations in the continuous space to discrete test sequences, and fitness values are calculated using the APRC metric. In this process, the individual with the best fitness value is updated as the global leader.
| Algorithm 1 Memetic Hybrid Dragon Boat Optimization and Genetic Algorithm (MH-DBO-GA) |
| Require: T (Test Cases), R (Requirements), (Maximum Iterations) |
| Ensure: (Optimized Test Case Sequence) |
- 1:
Identify (requirements covered by exactly one test case) - 2:
Initialize to map test identifiers to vector indices
- 3:
for each agent do - 4:
if then - 5:
Apply Additional Greedy Seeding prioritizing - 6:
else - 7:
Initialize random position vector in - 8:
end if - 9:
end for
- 10:
for to do - 11:
for each do - 12:
Map continuous vector to discrete order (Random Key) - 13:
Compute APRC using and R - 14:
if then - 15:
Update (Leader) - 16:
end if - 17:
end for
- 18:
if GlobalBest exists then - 19:
RefineLeader: Perform 100 candidate swaps with one index sampled from the first positions - 20:
BottleneckHunter: If the saturation index exceeds 500, assign random-key value 1.1 to the bottleneck test - 21:
end if
- 22:
Update acceleration (), attenuation (), and social factors () - 23:
for each do - 24:
Update position using DBOA equations guided by - 25:
Clamp positions to - 26:
end for
- 27:
Sort population by fitness - 28:
Remove bottom of agents - 29:
while population size do - 30:
Select from elite pool - 31:
Apply Order Crossover and Mutation - 32:
Add to population - 33:
end while - 34:
end for - 35:
return (sequence of the global best leader)
|
After each iteration, the best available solution is improved by applying a memetic local search phase. In this phase, Bottleneck Hunter and local optimization processes are performed on the leading individual to improve solution quality.
Following this, the DBOA mechanism is used to update the population positions and move individuals towards the global leader. Then, GA operators are activated to eliminate low-performing individuals and generate new solutions from elite individuals.
This iterative process continues until the maximum number of iterations is reached. The best solution found at the end of the algorithm is returned as the final test case ranking.
Algorithm 1 summarizes the workflow of the proposed method. The implementation-level details needed to follow the stochastic component, including the adapted DBOA update rules, local-search budget, genetic operators, replacement rule, stopping condition, and random-number handling, are given in
Section 5.7.
5.7. Implementation Details for Reproducibility
This subsection gives the main implementation details of the stochastic metaheuristic component. Algorithm 1 presents the overall workflow, while this subsection describes the update rules, local-search budget, genetic operators, replacement strategy, stopping condition, and random-number handling used in the experiments.
Random-key decoding is performed by sorting the position values in descending order. If two test cases have the same key value, their test identifiers are used as a lexicographic tie-breaker. If a tie still remains, the original index order is used as the final tie-breaker.
The informed initialization step is applied to 80% of the population. First, test cases associated with singleton requirements are identified. These essential test cases are ordered according to their incremental contribution to uncovered requirements. After that, Additional Greedy selection continues until either 500 tests have been prioritized or no remaining candidate provides additional uncovered requirement coverage. At each greedy step, at most the first 1000 candidates in the remaining pool are evaluated. Non-prioritized positions are initialized with small random values in . For prioritized tests, descending random-key values are assigned according to the greedy order.
The leader refinement step is applied once per iteration to the current global best solution. It performs 100 candidate swap trials. In each trial, the first index is sampled uniformly from the first positions, and the second index is sampled uniformly from the whole sequence. The two selected positions are swapped, APRC is recomputed, and the swap is accepted only if it improves the current best APRC value. Otherwise, the swap is reverted.
The Bottleneck Hunter step is also applied to the current global best solution. The current sequence is scanned until all requirements are covered for the first time. The test case located at this saturation position is treated as the bottleneck test. If the saturation index is greater than 500, the random-key value of this test case is set to 1.1 and APRC is recomputed. If the saturation index is 500 or less, no Bottleneck Hunter update is applied in that iteration.
The DBOA movement used in this study is based on the main factors of Dragon Boat Optimization, including the social behavior factor, acceleration factor, attenuation factor, imbalance term, and fastest-boat-based update mechanism [
34,
35]. Since the TCP problem is represented as a permutation problem, these rules were adapted to a random-key representation. In this representation, each boat is encoded as a continuous vector, and this vector is decoded into a test-case ordering by sorting the key values.
Let denote the random-key position vector of the i-th boat at iteration k, where is the number of test cases. The current global best position vector is denoted by . The iteration counter is defined as , where . In the experiments, and the population size is .
At each iteration, the social behavior factor
is sampled using two integer random variables. Let
be sampled uniformly from
and
be sampled uniformly from
. Then,
is computed as follows:
After
is obtained, the acceleration factor
and attenuation factor
are computed as follows:
The attenuation term follows the same general form as DBOA. In this implementation, the parameter
l in the original formulation is set to the maximum iteration budget
K. The imbalance term is computed by setting
, where
is sampled from
, and
:
The movement step is applied to the random-key position vectors. If the current boat has the same APRC fitness value as the current global best solution, the fastest-boat update term is computed as
and the corresponding random-key position is updated by
For the remaining boats, the update term is computed with reference to the current global best position:
The position is then updated as
Here, restricts each random-key value to the interval . This step is used because the DBOA movement is applied in a continuous random-key space, whereas the final TCP solution is a permutation.
The evolutionary refill stage uses an elitist replacement rule. At the end of each iteration, the population is sorted in descending order of the APRC fitness values available in that iteration. The nominal elimination rate is 25%. With , the implementation keeps boats and removes the 13 lowest-ranking boats. The top 10% of the original population size, corresponding to five boats, forms the elite pool. New offspring are generated until the population size returns to . For each offspring, two parents are selected uniformly at random from the elite pool. Positions modified by the DBOA movement are evaluated in the next fitness evaluation step.
The crossover operator is an order-preserving one-point crossover with a fixed midpoint. The first half of the child sequence is copied from the first parent. The remaining positions are filled by scanning the second parent from the beginning and appending test cases that are not already present in the child. This produces a valid permutation without duplicate test cases.
Mutation is applied with probability 0.20. When mutation is triggered, five random swap operations are performed on the child sequence. In each swap operation, both positions are selected independently and uniformly from the whole child sequence. Therefore, a swap may leave the sequence unchanged if the same position is selected twice. After crossover and mutation, the child permutation is converted back into a random-key vector by assigning
where
j is the zero-based rank of test case
in the child sequence. This conversion is used to represent the child ordering again in the random-key space.
The stopping condition is fixed: each stochastic run terminates after iterations. No early stopping criterion is used. The implementation uses Java pseudo-random generators without manually fixed seed values. A local pseudo-random generator is used in the main optimization routine for DBOA parameter sampling and parent selection, while a static pseudo-random generator is used in the boat representation for random initialization, leader refinement, and mutation. For this reason, exact bit-level replay of the reported stochastic runs is not guaranteed. The reported stochastic results are therefore presented as statistical summaries over 50 independent executions.
5.8. Transformation from Continuous to Discrete Space and Position Bounding Strategies
Although the proposed framework is continuous, the Test Case Prioritization (TCP) problem is a discrete permutation problem. Therefore, the following strategies are used to align these two representations.
Random-Key Mapping: The position of each dragon boat within the search space is represented by a continuous vector (e.g., 0.25, 0.95, 0.60), which encodes the weights of the test cases. These steps transform the continuous vector into a valid test sequence.
Indexing and Sorting: The continuous weights belonging to the test cases are sorted in descending order. The test case with the highest weight is placed in the first position.
Tie-Breaking: If the position weights of two or more test cases are equal, a tie-breaking mechanism is activated to prevent non-deterministic behavior. In such instances, the related test cases are incorporated into the permutation by being sorted lexicographically according to their identifiers (Test ID).
Utilization of the 1.1 Value for Bottleneck Teleportation: When the Bottleneck Hunter mechanism identifies a test case associated with delayed coverage saturation, it assigns a random-key value of 1.1 to that test case.
In the normal random-key representation, most test weights are initialized or evolved within the range of 0.0 to 1.0.
Assigning a value slightly above 1.0 increases the priority of the detected bottleneck test and tends to move it toward the beginning of the decoded sequence. This intervention encourages earlier coverage of delayed requirements without directly rewriting the whole permutation.
Clamp Value: During random-key position updates, the maximum value of the position vector is restricted to 1.2. The upper bound is set above 1.0 to allow controlled priority increases while still keeping the random-key values within a bounded interval.
Providing headroom: The 1.2 upper bound gives a limited margin above the normalized [0.0, 1.0] range. This allows bottleneck-related tests assigned a value of 1.1 to remain distinguishable from normally evolved test keys.
Bounding random-key growth: The 1.2 threshold prevents unbounded growth of random-key values during position updates and keeps the decoding process stable.
6. Experimental Results
6.1. Datasets
The effectiveness and generalizability of the proposed bottleneck-aware hybrid heuristic and metaheuristic framework were evaluated using two different datasets derived from real-world software projects. These datasets contain varying scales of test cases and requirement complexity, allowing for an objective analysis of the method’s performance in both high-volume and enterprise-level software systems. The basic quantitative characteristics of the datasets are summarized in
Table 2.
The dataset includes detailed attributes such as business requirements (B_Req), priority level, function points (FP), complexity level, duration, and cost. Comprising a total of 3399 test cases and 2000 requirements, this dataset contains 108 singleton requirements covered by only a single test. This highly sparse structure provides a suitable experimental environment for evaluating the capability of the proposed framework to identify and mitigate requirement coverage bottlenecks.
This enterprise-scale system consists of 14 modules, 1467 test cases, and 1237 requirements. As a result of the analyses, 72 singleton requirements were identified in this dataset. Due to its modular and hierarchical structure, this dataset was selected to evaluate the adaptability and scalability of the proposed framework in large-scale and complex enterprise software projects.
Table 2 presents the Requirement Traceability Matrix. Matrix density represents the ratio of traceability links(non-zero entries) to the total number of possible connections between test cases and requirements (
). A sparse RTM shows that most of the test cases cover a tiny subset of the requirements. This structural characteristic significantly restricts the search space.
In such situations, optimization approaches require problem-specific mechanisms capable of identifying structural bottlenecks. Therefore, the proposed framework incorporates the Bottleneck Hunter mechanism to exploit sparse RTM characteristics during prioritization. As detailed in
Table 2, Dataset 1 consists of 3399 test cases and 2000 requirements, theoretically yielding a maximum of 6,798,000 matrix cells. However, it contains only 5835 traceability links, resulting in a very low matrix density of 0.0858% (sparsity: 99.9142%). Similarly, Dataset 2 contains 1467 test cases and 1237 requirements, resulting in a maximum of 1,814,679 possible entries. However, the dataset includes 5331 links, yielding a matrix density of 0.2938% (sparsity: 99.7062%). These density values (0.0858% and 0.2938%) quantitatively confirm that both RTMs are extremely sparse and support the study’s main motivation, namely the sparse RTM assumption.
6.2. Experimental Setup and Parameter Settings
We evaluated the proposed bottleneck-aware hybrid heuristic and metaheuristic framework on a MacBook Pro workstation equipped with an Apple M1 processor and 32 GB of RAM. The experiments included both deterministic heuristic components and stochastic metaheuristic optimization components. For stochastic algorithms, each configuration was executed for 50 independent runs, and each run consisted of 100 optimization iterations. The implementation uses Java pseudo-random generators without manually fixed seed values. Therefore, the reported stochastic results should be interpreted as statistical summaries of independent executions rather than as exact bit-level replayable runs. We adopted this evaluation protocol to reduce the influence of random variation and to provide a reliable statistical assessment of the obtained results.
The parameter settings of the metaheuristic refinement component are summarized in
Table 3. We determined the population size to be 50 to provide a balance between population diversity and computational efficiency. Additionally, we fixed the maximum number of iterations at 100 for all stochastic optimization approaches to ensure a consistent comparison.
To incorporate problem-specific knowledge into the search process, 80% of the initial population was generated using a bottleneck-aware greedy initialization strategy, while 20% was randomly initialized to maintain exploration diversity. This strategy starts the search from promising regions derived from sparse requirement traceability matrices and avoids over-reliance on deterministic solutions.
In the memetic refinement phase, we focused on the first 300 test cases in the generated priority sequence. We selected this region because early requirement coverage is the primary optimization objective in regression testing scenarios. By concentrating local improvement efforts on the initial portion of the ranking, the method aims to enhance APRC performance without introducing excessive computational overhead.
The DBOA position update mechanism employs dynamic parameter schedules to control the balance between exploration and exploitation throughout the optimization process. The position clamping range was set to to prevent excessive movement outside the feasible search region while allowing controlled exploration beyond the normalized ranking interval. Similarly, the bottleneck teleportation value was set to 1.1 to provide a controlled perturbation mechanism for relocating bottleneck-related solutions without causing disruptive changes in the search trajectory.
In the genetic refinement stage, we employed an elite selection rate of 10%, an elimination rate of 25%, and a mutation rate of 20%. Furthermore, we selected these settings to maintain population diversity while preserving high-quality solutions during the evolutionary search. Lastly, we kept all parameter configurations identical across repeated experiments to ensure a fair comparison among the evaluated approaches.
6.3. Comparative Performance Analysis
We comparatively evaluated the contribution of the deterministic bottleneck-aware heuristic component and the metaheuristic refinement component within the proposed framework. We examined the effectiveness of the Additional Greedy and Bottleneck Hunter mechanisms, because sparse requirement traceability matrices introduce structural coverage challenges, particularly due to singleton requirements.
To avoid misleading visual comparisons between deterministic methods with zero variance and stochastic metaheuristic methods with repeated-run variability, we present them separately. The APRC scores of deterministic algorithms are shown as bar charts in
Figure 2 and
Figure 3, while the APRC distributions of stochastic algorithms over 50 independent runs are shown as boxplots in
Figure 4 and
Figure 5.
Subsequently, we evaluated the MH-DBO-GA component as a metaheuristic refinement strategy that enhances exploration capability through evolutionary and memetic operators. To evaluate the individual contributions of the proposed framework, we compared the bottleneck-aware heuristic configuration (AG+BH) and the metaheuristic refinement configuration (MH-DBO-GA) with representative baseline approaches, including Standard Greedy, Additional Greedy, 2-Optimal, GA, and DBOA.
Table 4 presents the APRC performance of all evaluated approaches. Deterministic methods are reported using their exact single-run APRC values. On the other hand, stochastic metaheuristic methods are evaluated over 50 independent runs. Statistical comparisons and practical effect sizes are reported only where stochastic distributions are available.
Table 4 shows that the deterministic methods provide the strongest APRC results in the current evaluation. On Dataset 1, 2-Optimal obtains the highest APRC with 89.96%, while AG+BH obtains a very close value of 89.92%. On Dataset 2, AG+BH obtains the highest APRC with 95.57%. The corresponding MH-DBO-GA results are 85.68% and 95.14%, respectively. These results indicate that, when the objective is early requirement coverage, problem-specific deterministic reasoning is effective under sparse RTMs. However, the practical preference among deterministic methods should also consider runtime, because 2-Optimal requires substantially higher computational cost on Dataset 1.
The MH-DBO-GA component has a different role in the framework. It does not improve over AG+BH in the evaluated setting, but it performs better than the standard stochastic metaheuristic baselines. For Dataset 1, MH-DBO-GA outperforms the strongest stochastic baseline, GA. For Dataset 2, it outperforms the strongest stochastic baseline, DBOA. These results suggest that MH-DBO-GA is useful as a population-based exploratory extension, but not as the main performance driver of the proposed framework.
These results show that bottleneck-aware deterministic prioritization is the main empirical finding of this study. AG+BH should be seen as a practical and strong deterministic strategy, not as the best APRC method on every dataset. It is very close to the best APRC on Dataset 1. It achieves the best APRC on Dataset 2. It also runs much faster than 2-Optimal. The metaheuristic layer is kept because it adds search diversity. It may also be useful in future settings with dynamic RTMs, changing test constraints, requirement weights, execution costs, or fault severity.
6.4. Saturation Point and Coverage Speed Analysis
The saturation point represents the position in the test case sequence at which all requirements are covered for the first time. This metric provides additional insight into early requirement coverage behavior beyond the final APRC value.
Table 5 and
Table 6 present the saturation points and execution times of the evaluated approaches for Dataset 1 and Dataset 2, respectively.
The 2-Optimal run on Dataset 1 completed in 84,180 s, corresponding to 23 h 23 min, and obtained 89.96% APRC. This value comes from a completed run rather than an extrapolated runtime. Although 2-Optimal reaches an early saturation point and gives the highest APRC on Dataset 1, its pairwise-swap evaluation leads to a very high computational cost on the larger sparse RTM. Therefore, this result is included mainly to illustrate the practical runtime limitation of pairwise local improvement on large test suites.
The saturation analysis provides further evidence that sparse RTM structures can benefit significantly from problem-specific bottleneck identification.
As
Table 5 shows, the AG+BH configuration and Additional Greedy reach saturation at the 998th position on Dataset 1, outperforming the population-based metaheuristic approaches in terms of coverage speed. This result indicates that deterministic heuristic strategies can be effective when the sparse RTM contains bottleneck patterns.
Table 6 shows a similar pattern for Dataset 2. The AG+BH configuration reaches full coverage at the 282nd position, closely followed by Additional Greedy at the 283rd position. This result indicates that targeting singleton-induced coverage bottlenecks can accelerate early requirement coverage. The MH-DBO-GA configuration reaches full coverage at the 446th position while providing a population-based exploratory refinement mechanism.
These saturation results are consistent with the APRC findings. In both datasets, AG+BH or Additional Greedy reaches full requirement coverage earlier than MH-DBO-GA. Therefore, for sparse RTMs with clearly identifiable coverage bottlenecks, the deterministic heuristic layer is the more suitable option when the evaluation objective is early requirement coverage alone. The MH-DBO-GA component should be interpreted as a secondary exploratory mechanism. Its potential value is in providing population-based search diversity for more complex prioritization settings, rather than in replacing AG+BH in the current APRC experiments.
6.5. Computational Efficiency Analysis
Computational efficiency is an important consideration in regression testing, particularly for continuous integration and continuous delivery (CI/CD) environments. In those environments, prioritization decisions must be generated within limited time constraints. Therefore, we evaluate the computational cost of the proposed framework by considering both execution time and algorithmic complexity. We comparatively analyzed the deterministic heuristic components and the metaheuristic refinement component to reveal the trade-off between computational overhead and coverage effectiveness.
6.5.1. Execution Time Comparison
Table 7 presents the execution times of the evaluated approaches on both datasets. For stochastic methods, the reported values represent the mean execution time over repeated runs. For deterministic methods, the values correspond to the completed single-run execution time required to generate a complete test case prioritization sequence.
The execution-time results show that simple greedy-based deterministic strategies have low computational cost. Greedy and Additional Greedy are the fastest approaches, while AG+BH introduces only a limited additional cost because the Bottleneck Hunter mechanism performs a targeted refinement on the generated sequence. However, 2-Optimal is an exception among deterministic methods. Although it can reach an early saturation point, its pairwise-swap evaluation leads to a very high runtime, especially on Dataset 1.
The MH-DBO-GA component requires additional computational time due to population-based exploration, evolutionary operations, and memetic refinement. It is slower than DBOA, but it obtains substantially higher APRC values than the standard stochastic baselines. It is also faster than GA on both datasets. Therefore, the results indicate a trade-off: AG+BH is preferable when low runtime and early requirement coverage are the main objectives, while MH-DBO-GA provides a more exploratory population-based alternative with additional computational cost.
6.5.2. Computational Complexity Discussion
To avoid duplicate complexity analyses, we report a single sparse-link-based derivation in this subsection. Let N denote the population size, K the maximum number of iterations, the number of test cases, the number of requirements, and L the number of non-zero requirement–test links in the sparse RTM. Since the RTM is stored and traversed as sparse adjacency lists, the analysis is expressed in terms of L rather than the dense matrix size .
For one candidate sequence, APRC evaluation is performed by scanning the test cases in the given order and visiting only the requirements linked to each visited test case. Therefore, the cost of one APRC computation is
This expression includes the traversal of the test sequence and the traversal of existing sparse links. It does not multiply by L, because all non-zero links are visited through the adjacency lists, not re-scanned for each test case.
The stochastic metaheuristic component uses a random-key representation. Before APRC can be computed, the random-key vector of a candidate must be decoded into a test-case ordering. This decoding step requires sorting the
key values, which has the cost
Thus, evaluating one random-key candidate requires
For
N candidates evaluated over
K iterations, the main population-level evaluation cost is therefore
The informed initialization stage has a separate cost. The singleton-requirement preprocessing step scans the sparse RTM once and costs
. For the greedy part of initialization, let
G denote the maximum number of prioritized tests produced by the greedy construction and let
denote the maximum number of candidate tests scanned at each greedy step. In our implementation,
and
. Let
denote the maximum number of requirements linked to any test case. Under sparse traversal, the cost of constructing one greedy individual is bounded by
The term accounts for assigning the remaining random-key values after the greedy prefix is constructed. The term gives a conservative bound for checking the sparse requirement links of the scanned candidates. This is more precise for our implementation than the previous dense-style expression, because the method does not traverse a full matrix.
Since 80% of the population is initialized using the informed greedy strategy and 20% is initialized randomly, the initialization cost can be written as
The memetic refinement step is applied only to the current leader. In each iteration, the leader refinement performs
candidate swap trials. Since APRC is recomputed after each candidate swap, the local-search cost over
K iterations is bounded by
The Bottleneck Hunter step also scans the current leader sequence to identify the saturation position and the associated bottleneck test. Its cost is per application and is covered by the same sparse traversal argument.
The remaining population operations have lower-order costs in the current setting. DBOA position updates require operations per iteration. The evolutionary refill stage sorts the population by fitness, which costs per iteration, and generates replacement offspring through crossover and mutation. Since the number of replaced individuals is proportional to N, the crossover cost is per iteration. Mutation performs a fixed number of swaps and has a smaller contribution.
Combining these terms, the overall cost of the MH-DBO-GA component can be summarized as
In the experiments, , , , , and are fixed parameter values. Therefore, the main dataset-dependent terms are the number of test cases and the number of sparse links L. This explains why sparse traversal is important: the computation depends on the existing traceability links rather than on the theoretical dense matrix size .
For the deterministic AG+BH configuration, there is no population-based iteration. Its cost is mainly the Additional Greedy construction and one Bottleneck Hunter refinement, which can be summarized as
This difference explains the lower runtime of AG+BH compared with MH-DBO-GA. The metaheuristic component introduces additional cost through repeated population evaluation, sorting, movement updates, and local refinement, whereas AG+BH directly applies the sparse greedy construction and bottleneck-aware correction.
6.6. Parameter Sensitivity Analysis
Control parameters can influence how MH-DBO-GA balances exploration, exploitation, and computational cost. We examine four parameters in this section: the greedy initialization ratio, local search region size, bottleneck teleportation value, and position clamping range.
For each parameter, we ran each alternative configuration 50 times on both datasets. The sensitivity tables report the mean APRC, standard deviation, 95% confidence interval of the mean, and runtime mean with standard deviation. Runtime values are reported in seconds. Since several parameter settings produce close APRC values, we use this analysis mainly to understand parameter sensitivity rather than to identify strictly dominant settings.
We compare runtime values mainly within the same parameter group, because each sensitivity experiment was executed as a separate set of runs. Therefore, repeated baseline settings across different parameter groups may show small runtime differences due to independent executions.
6.6.1. Greedy Initialization Ratio
Table 8 shows that the tested greedy initialization ratios produce close APRC values. The 80% ratio gives the highest mean APRC on both datasets, but the confidence intervals remain close to those of the other settings. The runtime values also do not show a consistent advantage for one ratio across both datasets. We therefore use the 80% setting as a practical configuration that combines a high proportion of informed greedy initialization with a limited amount of random initialization for population diversity.
6.6.2. Sensitivity Analysis of Local Search Region
The local search region determines which part of the priority sequence the memetic swap-based refinement can explore. We tested three region sizes: 100, 300, and 500 test cases. This parameter can affect both the search behavior and the computational cost, so
Table 9 reports APRC variability and runtime values together.
Table 9 shows that the local search region affects the two datasets differently. On Dataset 1, the 500-test-case region gives the highest mean APRC. On Dataset 2, the 300-test-case region gives the highest mean APRC, while the 500-test-case region gives a slightly lower value. Increasing the region size therefore does not lead to a consistent APRC improvement across both datasets.
The runtime values do not increase monotonically with the local search region size. This is consistent with the implementation, because the algorithm keeps the number of local-search trials fixed and changes only the part of the sequence from which it samples swap positions. We use the 300-test-case region as a middle setting. It gives the best mean APRC on Dataset 2 and improves over the 100-test-case setting on Dataset 1, although it does not exceed the 500-test-case setting on Dataset 1.
6.6.3. Sensitivity Analysis of Bottleneck Teleportation Value
The bottleneck teleportation value controls the random-key value assigned to a bottleneck-related test case during the Bottleneck Hunter step. We tested three values: 1.0, 1.1, and 1.2. Since these settings produce very close APRC values,
Table 10 reports standard deviations, confidence intervals, and runtime summaries in addition to the mean APRC values.
Table 10 shows very small APRC differences among the tested teleportation values. On Dataset 1, the mean APRC values of 1.0, 1.1, and 1.2 are almost indistinguishable. On Dataset 2, the 1.1 setting gives the highest mean APRC, but the difference from the other values remains small. The runtime values do not show a clear monotonic pattern with respect to the teleportation value.
We use 1.1 in the final configuration as a moderate teleportation value. It gives the highest mean APRC on Dataset 2 and remains close to the best value on Dataset 1, without using the largest tested perturbation.
6.6.4. Sensitivity Analysis of Position Clamping Range
The position clamping range defines the interval in which the algorithm keeps random-key values after DBOA movement. We tested three ranges: [0.0, 1.0], [0.0, 1.2], and [0.0, 1.5]. Since the APRC differences among these settings are small,
Table 11 reports standard deviations, confidence intervals, and runtime summaries together with the mean APRC values.
Table 11 shows that the tested clamping ranges produce close APRC values on both datasets. The [0.0, 1.2] range gives the highest mean APRC on Dataset 1 and Dataset 2, but the differences from the other ranges remain small and the confidence intervals are close.
The runtime values do not show a fully monotonic trend as the upper clamping bound changes. However, the largest tested range, [0.0, 1.5], does not improve mean APRC on either dataset and gives the highest runtime on Dataset 1. We use the [0.0, 1.2] range in the final configuration as a moderate extension of the normalized [0.0, 1.0] interval. This range allows limited movement above 1.0 and remains consistent with the Bottleneck Hunter teleportation value of 1.1, while avoiding the largest tested upper bound.
Table 12 summarizes the parameter values used in the final MH-DBO-GA configuration. We select these values as practical choices based on the observed mean APRC values, variability, and runtime summaries. Since several tested settings produce close APRC values, this table documents the final configuration and the rationale behind each selected value.
7. Contribution of Individual Components
We performed an ablation study to examine how the main components of the metaheuristic refinement layer affect APRC performance. The analyzed components include informed initialization, Bottleneck Hunter, DBOA-based movement, and GA-based evolutionary operators. The purpose of this analysis is not only to confirm the usefulness of individual components, but also to identify whether the complete hybridization is necessary under different sparse RTM structures. The mean APRC values obtained from 50 independent executions for each variation are reported in
Table 13.
The results in
Table 13 show that the contribution of each component is not uniform across the two datasets. Informed initialization has a consistent positive role. Removing it decreases APRC from 0.856847 to 0.843575 on Dataset 1 and from 0.951380 to 0.943220 on Dataset 2. This suggests that starting the search from structurally informed solutions remains useful for sparse RTMs.
The effect of Bottleneck Hunter is more dataset-dependent. On Dataset 1, removing Bottleneck Hunter reduces APRC from 0.856847 to 0.842681, indicating that the bottleneck-oriented move helps the search process in this matrix structure. On Dataset 2, however, removing Bottleneck Hunter slightly increases APRC from 0.951380 to 0.951710. This does not mean that Bottleneck Hunter is generally harmful; rather, it suggests that when the sequence is already close to a high-coverage arrangement, the teleportation step can sometimes introduce a small disturbance.
The ablation results also show that the full combination of DBOA, GA operators, and Bottleneck Hunter is not always the best-performing configuration. On Dataset 1, the DBOA + Memetic Only variant obtains an APRC of 0.858211, slightly higher than the full MH-DBO-GA result of 0.856847. On Dataset 2, several reduced variants are very close to, or slightly better than, the full hybrid configuration. For example, GA + Memetic Only reaches 0.951539 and Without Bottleneck Hunter reaches 0.951710, while MH-DBO-GA obtains 0.951380. These differences are small, but they indicate that some components may be redundant under certain sparse RTM structures.
A possible explanation is that the GA crossover operator and the Bottleneck Hunter teleportation step may occasionally disrupt an already well-ordered sequence. This is especially possible when informed initialization and memetic refinement have already placed most bottleneck-related test cases near the beginning of the ranking. In such cases, additional exploration does not necessarily improve APRC and may slightly change a useful ordering. Therefore, the complete MH-DBO-GA configuration should not be interpreted as a necessary combination for all sparse RTMs.
Overall, the ablation study supports two main conclusions. First, informed initialization is an important component of the metaheuristic refinement process. Second, full hybridization is not uniformly beneficial. The reduced variants show that DBOA-based movement, GA operators, and Bottleneck Hunter may interact differently depending on the structure of the RTM. This finding is consistent with the positioning of MH-DBO-GA as a secondary exploratory extension rather than as the primary source of empirical improvement.
9. Conclusions
This study investigated the requirements-based test case prioritization problem under sparse requirement traceability matrices and proposed a bottleneck-aware heuristic and metaheuristic framework. The framework is centered on the deterministic AG+BH strategy, which combines Additional Greedy with the proposed Bottleneck Hunter mechanism. This component targets delayed requirement coverage caused by sparse requirement–test relationships, especially singleton requirements. The MH-DBO-GA component is included as a secondary population-based extension that combines Dragon Boat Optimization, Genetic Algorithm operators, and memetic local search.
Extensive experimental evaluations on two sparse RTM datasets show that When APRC, saturation behavior, and runtime are considered jointly, AG+BH provides the practical deterministic outcome in the evaluated single-objective setting. It achieves the highest APRC on Dataset 2 and a near-best APRC on Dataset 1, where 2-Optimal obtains higher APRC but requires 23 h 23 min to complete. AG+BH also reaches full requirement coverage earlier than MH-DBO-GA on the studied datasets. These results suggest that, when sparse RTM bottlenecks are directly identifiable, a deterministic bottleneck-aware heuristic can be more effective and simpler than applying the full metaheuristic framework.
The MH-DBO-GA component should therefore not be interpreted as a generally superior alternative to AG+BH or Additional Greedy. Its role is complementary. It provides a population-based search mechanism and performs better than the standard stochastic metaheuristic baselines, but it does not outperform AG+BH in the present APRC experiments.
The ablation results further show that full hybridization is not always necessary, since some reduced variants perform comparably to or slightly better than the complete MH-DBO-GA configuration on the evaluated datasets.
The computational evaluation also supports this interpretation. The deterministic AG+BH strategy achieves strong coverage behavior with low computational cost, while the metaheuristic refinement layer introduces additional overhead due to population-based search and memetic refinement. Thus, AG+BH is the preferable option for the current sparse RTM setting when the main objective is early requirement coverage. MH-DBO-GA may be considered when additional exploration is needed or when the prioritization problem is extended with further constraints.
In conclusion, this work shows that sparse RTM-based test case prioritization benefits primarily from domain-specific bottleneck identification. The main scientific message of the study is the effectiveness of bottleneck-aware deterministic prioritization, with metaheuristic refinement serving as a secondary extension rather than the primary source of empirical improvement.
Future work will focus on extending the framework toward multi-objective optimization scenarios by incorporating additional factors such as execution cost, fault severity, requirement priority, and code complexity. Such settings may provide a more suitable basis for evaluating when the MH-DBO-GA exploratory layer becomes preferable to deterministic heuristics. Furthermore, evaluating the approach on larger industrial datasets and adapting the framework to diverse software architectures, including distributed and microservice-based systems, will be investigated.