4.2.1. Overall Endpoint-Prediction and Recommended-Operation Results
For all 400 heats, the optimizer produced feasible recommendations satisfying the requirements for temperature, S, Mn, Als, O, operating bounds, and the historical support domain. Differences between the recommended and corresponding historical operations were compared for material and energy-input variables to determine whether the optimizer achieved endpoint-quality control by reallocating resource input.
Table 6 lists the mean values of the six principal operating quantities, and
Figure 3 shows their relative changes. Electric energy input and power-on duration primarily affect heat input; lime and slag agent jointly alter slag-forming conditions; and aluminum granules and high-carbon ferromanganese provide deoxidation and fine composition adjustment, respectively.
For resource item r in heat i, the absolute reduction is defined as the historical quantity minus the model-recommended quantity. The mean absolute reduction in
Table 6 is the average of this quantity over the 400 heats, and the mean reduction is the mean absolute reduction divided by the mean historical quantity.
The mean values of all six recommended operating quantities were lower than their corresponding historical values, although the magnitude of reduction differed because of their process functions and allowable adjustment ranges. The mean additions of lime, slag agent, aluminum granules, and high-carbon ferromanganese decreased from 398.49, 143.97, 192.62, and 77.28 kg to 370.66, 125.55, 174.62, and 70.44 kg, corresponding to mean relative reductions of 6.98%, 12.79%, 9.34%, and 8.85%, respectively. Electric energy input and power-on duration decreased from 4620.15 kWh and 11.32 min to 4255.55 kWh and 10.31 min, corresponding to mean relative reductions of 7.89% and 8.95%.
Figure 4 and
Table 7 further report the distributions and robust statistics of the per-heat relative reductions.
Figure 4 and
Table 7 reveal the between-heat distributions of the recommended operations. Positive values in the boxplots indicate that the recommended quantity is lower than the historical operation, whereas negative values indicate an increase. The boxes for all operating quantities lie predominantly above the zero baseline. In particular, the 25th percentiles of lime, aluminum granules, and power-on duration are clearly positive, indicating positive reductions in at least approximately 75% of the evaluation heats. The lower box edges for the slag agent and high-carbon ferromanganese lie at zero, but their median reductions remain 3.57% and 9.68%, respectively; the median reduction in electric energy input is 11.99%. Thus, the overall adjustment direction remains downward. When the endpoint-temperature, composition, and historical-support-domain constraints are satisfied, the resource objective consequently drives these variables toward their permitted lower levels.
The lower whiskers of some variables extend into the negative region, indicating that the recommended quantities exceed the corresponding historical values for a small number of heats. This result confirms that the method does not impose a uniform proportional reduction on every resource; instead, it allocates resources on a heat-specific basis to balance an overall reduction in resource input against endpoint feasibility. To determine whether the mean reductions were driven by a small number of extreme heats, bootstrap interval estimation, Wilcoxon signed-rank tests, and effect-size analysis were performed on the paired quantity differences for all 400 heats. The results are presented in
Table 8.
Table 8 shows that the 95% bias-corrected and accelerated (BCa) confidence intervals for the mean reductions in all six quantities lie entirely above zero, and all Holm-adjusted Wilcoxon signed-rank tests yield
p < 0.001. The overall reductions are therefore not driven by a few extreme heats. The rank-biserial effect sizes are 0.993, 0.994, 0.995, and 0.977 for lime, slag agent, aluminum granules, and high-carbon ferromanganese, respectively, and 0.940 for power-on duration, indicating a consistent downward direction across most heats. The effect size for electric energy input is 0.815; although lower than those of the other variables, it still indicates a strong directional change. These results agree with
Figure 4: individual inputs may increase in a small number of heats to maintain endpoint feasibility, but the overall distributions consistently indicate lower resource input.
To examine the influence of prediction error on hard-constraint screening, the independent-test MAE of each model in
Table 3 was used as the error scale for its corresponding output. The allowable prediction intervals for temperature, S, Mn, O, and Als were contracted inward by
times the relevant MAE, with
= 0, 0.25, 0.50, and 1.00. The feasible-recommendation acquisition rate is defined as the proportion of evaluation heats for which at least one recommendation satisfying all model constraints is obtained within the prescribed candidate-generation limit. The acquisition rate, temperature error, resource-input index reduction, and runtime under the different contraction levels are compared in
Table 9.
In the optimization experiments below, the mean absolute target-temperature deviation (MATD) is defined as the temperature metric, namely, the mean absolute difference between the endpoint temperature predicted for a recommended scheme and the externally specified target temperature. This metric differs from the prediction-model MAE in
Table 3, which is calculated relative to measured endpoint temperatures.
As
increases from 0 to 0.25, 0.50, and 1.00,
Table 9 shows that the feasible-recommendation acquisition rate decreases to 96.25%, 93.25%, and 88.00%, respectively, while the resource-input index reduction decreases from the baseline value of 8.67% to 8.45%, 8.15%, and 6.31%. Among the heats for which a recommendation is obtained, the MATD decreases progressively from 1.789 °C to 1.500 °C. Contracting the prediction intervals therefore requires candidate temperatures to lie closer to the target and also changes the composition of the feasible heats included in the statistics. The results indicate that moderate constraint contraction can absorb part of the prediction uncertainty while retaining a practical recommendation-acquisition rate.
4.2.2. Dynamic-Weight Evolution and Parameter Sensitivity
To further verify whether generational progress and feasible-population temperature-error feedback jointly coordinate the two objectives,
Figure 5 presents the temperature-term weight trajectories of 10 representative heats over 120 generations. Each trajectory is calculated using Equations (7)–(9). The generational annealing term prescribes a gradual transition from endpoint-feasibility priority to resource optimization, whereas the temperature-error feedback term adjusts the decay rate according to the current feasible population. The resource-input index term weight is determined simultaneously by λ2 = 1 − λ1.
As shown in
Figure 5, the temperature weight of every heat starts at its upper bound of 0.85, decreases smoothly as the search proceeds, and approaches the lower bound of 0.20 during the late stage. The trajectories differ visibly in the intermediate stage. Heats with larger feasible-population temperature errors exhibit slower decay to maintain emphasis on endpoint-temperature control, whereas those with smaller errors increase the resource-term weight more rapidly. The weighting mechanism therefore retains a clear process interpretation across the search stages while adaptively adjusting the objective-transition rate according to heat difficulty.
Fixed weights, different sigmoid parameters, and different characteristic index-change rates were further compared on 60 stratified heats; the results are presented in
Table 10.
Table 10 shows that the MATD values under fixed weights and the baseline dynamic weights are 2.167 °C and 1.960 °C, respectively, while the corresponding resource-input index reductions are 9.56% and 9.08%. Relative to fixed weights, the baseline dynamic weights reduce MATD by 0.207 °C while retaining a resource-input index reduction above 9%. The two alternative sigmoid settings produce MATD values of 1.911 °C and 1.989 °C and resource-input index reductions of 8.87% and 9.15%, indicating limited variation over the tested range. Because the prediction models and constraints are coupled nonlinearly, neither metric varies monotonically with an individual parameter. By contrast, increasing the characteristic index-change rate from 0.05 to 0.15 reduces MATD from 2.023 °C to 1.869 °C while decreasing the resource-input index reduction from 9.46% to 8.47%, revealing a clear temperature–resource trade-off. All six settings achieve a 100% feasible-recommendation acquisition rate, indicating that the baseline parameters lie within a stable feasible region.
Further, to test whether the conclusions obtained with equal typical contributions from the six resources depend on a single coefficient setting, 20 evenly spaced sampled heats and 10 random seeds were used. Taking the baseline coefficients as the reference, one of the six resource coefficients was increased or decreased by 20% at a time, producing 12 one-factor perturbation schemes. Each scheme comprised 200 identical heat–seed combinations, while all other algorithm parameters were held constant. The results are reported in
Table 11.
Under the baseline coefficients,
Table 11 gives a MATD of 1.819 °C and a resource-input index reduction of 8.01%. Across the 12 one-factor perturbations, the two metrics average 1.813 ± 0.010 °C and 7.97% ± 0.06%, respectively, and feasible recommendations are obtained in all 2400 perturbed runs. At the individual-scheme level, MATD ranges from 1.795 to 1.823 °C and the resource-input index reduction ranges from 7.89% to 8.06%. Although coefficient changes alter the relative contribution of each resource to the aggregate index, perturbations of ±20% do not change the principal conclusions of controlled temperature error, reduced overall resource input, and successful recommendation acquisition for every test task.
The number of similar cases,
, determines the range of local historical information used to establish the heat-specific reference level and initialize the population. To examine the influence of the retrieval range,
was varied from 5 to 60 on 60 evenly spaced sampled heats using five identical random seeds, producing 12 parameter settings and 3600 optimization tasks. All settings used the same search budget and were run in randomized interleaved order after warm-up to reduce batch effects in runtime comparisons. The results are presented in
Figure 6.
Figure 6 shows that the feasible-recommendation acquisition rate is 95.00% at
= 5, indicating that an excessively narrow case neighborhood cannot provide a stable reference level and initialization information for a small number of heats. All tasks obtain feasible recommendations when
≥ 10. MATD remains between 1.956 and 1.994 °C for
= 20–50, while the resource-input index reduction ranges from 8.84% to 9.01%, forming a stable performance region.
= 35 provides a favorable combination of MATD, resource-input index reduction, and runtime and is therefore selected as the baseline for the main experiment and algorithm comparison. Except at
= 5, P50 runtime ranges from 0.874 to 1.018 s and does not increase systematically with the number of retrieved cases, indicating that the additional computational cost of a wider retrieval range is small for the present case-base size and fixed search budget.
To distinguish the influence of historical case utilization intensity from that of local-search frequency, the case-initialization ratio
and local-search interval
were varied separately on the same 60 evenly spaced sampled heats and five random seeds, while all other parameters were held constant. The results are presented in
Table 12.
All parameter settings in
Table 12 achieve a 100% feasible-recommendation acquisition rate. When
increases from 0.25 to 0.50 and 0.75, the MATD values are 1.970, 1.956, and 1.963 °C; the resource-input index reductions are 8.84%, 9.01%, and 8.91%; and the P50 runtimes are 0.988, 0.930, and 1.050 s, respectively. Thus,
= 0.50 performs favorably on all three metrics. For
values of 5, 10, and 20, the MATD values are 1.986, 1.956, and 1.977 °C; the resource-input index reductions are 8.92%, 9.01%, and 8.99%; and the P50 runtimes are 1.238, 0.930, and 0.951 s, respectively. Among the three tested levels,
= 10 simultaneously yields the lowest MATD, the largest resource-input index reduction, and a short runtime. The baseline settings
= 0.50 and
= 10 therefore balance search quality, population diversity, and computational efficiency.
4.2.3. Convergence and Stability of the Hybrid Algorithm
To analyze the effects of case initialization and tabu search, standard GA, GA with case initialization, GA with tabu search, and the complete method were compared on 20 evenly spaced sampled heats using 10 identical random seeds. Differential evolution (DE) was included as a population-based reference algorithm using the best/1/bin strategy, a mutation factor of 0.5–1.0, a crossover probability of 0.70, 24 individuals, and 120 generations, without terminal local refinement. All algorithms used the same heats, random seeds, prediction-model interfaces, hard constraints, and maximum number of generations.
Using one candidate scheme evaluated through the complete endpoint-model interface as the common counting unit, standard GA and GA with case initialization require approximately 2.4 × 103 candidate evaluations for a complete 120-generation run. The methods incorporating tabu search require up to approximately 5.9 × 103 evaluations under the settings of 12 local-search triggers, two elite solutions per trigger, at most six local iterations per elite, and 24 neighborhood candidates per iteration. DE with a population size of 24 requires approximately 2.9 × 103 evaluations. If the no-improvement termination criterion is triggered early, the actual number of evaluations will be lower than these upper bounds. Candidate evaluation budget, solution quality, G95, and runtime are considered jointly in the following discussion.
Because the dynamic weights make search-objective values from different generations incomparable, all feasible candidates in each generation were re-evaluated using a common comparison index with equal 50% contributions from the temperature-error and resource-input terms, and the best value obtained up to each generation was recorded. This index is used only to compare convergence on a common scale and does not participate in individual selection or termination.
Figure 7 presents the median trajectories over 200 runs, and
Table 13 summarizes MATD, resource-input index reduction, feasible-recommendation acquisition rate, and runtime for the five algorithms. P50 and P95 are the 50th and 95th percentiles, respectively, of the 200-run runtime distribution. G95 is defined as the earliest generation at which 95% of the total decrease from the first-generation value to the final value of the common comparison index has been achieved.
Figure 7 shows that GA with case initialization and the complete method both begin at 0.358, which is 35.1% lower than the initial value of 0.551 for standard GA and GA with tabu search. Similar cases therefore guide the initial population toward a better search region. As iteration proceeds, standard GA, GA with case initialization, GA with tabu search, the complete method, and DE converge to final index values of 0.324, 0.320, 0.307, 0.300, and 0.317, respectively. Relative to standard GA, tabu search reduces the final index from 0.324 to 0.307; adding tabu search to case initialization further reduces it from 0.320 to 0.300. The complete method consequently achieves the lowest final index among the five algorithms, indicating that case guidance and local refinement are complementary. The corresponding G_95 values are 54, 28, 41, 25, and 27 generations. Case initialization shortens G_95 from 54 to 28 generations, tabu search shortens it to 41 generations, and their combination further reduces it to 25 generations and produces a stable plateau after approximately generation 25. Thus, case initialization primarily improves the starting point and early search efficiency, whereas tabu search improves solution quality during the middle and late stages; together, they produce the fastest principal improvement and the best final result.
Table 13 shows that all 200 repeated runs of each of the five algorithms yield recommendations satisfying the prescribed model constraints. The complete method achieves a MATD of 1.819 °C, a mean resource-input index reduction of 8.01%, and a standard deviation of 1.85 percentage points. Relative to standard GA, it lowers MATD by 0.039 °C and increases the resource-input index reduction by 0.15 percentage points. Relative to GA with case initialization, the improvements are 0.018 °C and 0.11 percentage points, while relative to GA with tabu search, they are 0.004 °C and 0.05 percentage points. The complete method achieves both the lowest MATD and the largest resource-input index reduction among the four genetic variants, together with the smallest standard deviation of resource-input index reduction, indicating that the combination of case initialization and tabu search improves both solution quality and repeated-run stability. DE yields a resource-input index reduction of 8.08%, 0.07 percentage points higher than that of the complete method, but its MATD increases to 2.069 °C, 0.250 °C above that of the complete method; thus, its additional resource reduction is accompanied by a larger target-temperature deviation. In terms of computational efficiency, the complete method has P50 and P95 runtimes of 0.814 and 1.873 s, respectively. Relative to standard GA, these values are shorter by 46.1% and 15.0%; relative to GA with tabu search, they are shorter by 44.6% and 21.3%; and relative to DE, they are shorter by 84.0% and 74.2%. GA with case initialization has the shortest P50 runtime of 0.618 s, but its final common comparison index and the temperature–resource results in
Table 13 are inferior to those of the complete method. Compared with that algorithm, the complete method increases P50 by 31.7% and P95 by only 0.3%. Together with
Figure 7, these results show that case initialization reduces the computation required to enter a high-quality region and offsets part of the cost of tabu local search, enabling the complete method to retain per-heat runtimes on the order of seconds while achieving the lowest common comparison index and a favorable temperature–resource balance.
To further examine the paired differences between the complete method and GA + case initialization, GA + tabu search, and DE, the mean paired differences in MATD and resource-input index reduction, their 95% bias-corrected and accelerated (BCa) confidence intervals obtained from 50,000 heat-level bootstrap resamples, the results of two-sided exact Wilcoxon signed-rank tests with Holm correction for the three comparisons, and the rank-biserial effect sizes are summarized in
Table 14. Specifically, the 10 random-seed results obtained by each algorithm for each heat were first averaged, and the subsequent analyses were conducted using the resulting 20 heat-level means.
Table 14 further clarifies the statistical and practical significance of the performance differences summarized in
Table 13. Compared with GA + case initialization, the complete method reduced the MATD by 0.018 °C. Its 95% BCa confidence interval lay entirely below zero, and the difference remained statistically significant after Holm correction (
p = 0.0306), with a rank-biserial effect size of −0.610. This result indicates that the incorporation of tabu search produced a modest improvement in temperature-control error, with a relatively consistent direction across the evaluated heats. By contrast, the differences between the complete method and GA + tabu search were only −0.004 °C for MATD and +0.050 percentage points for resource-input index reduction. Both confidence intervals crossed zero, and both Holm-adjusted
p-values were greater than 0.05. Therefore, within the scope of the present experiments, the primary contribution of case initialization was reflected in improved search efficiency, including a lower initial unified comparison metric, a smaller G95, and shorter P50 and P95 runtimes, while also providing feasible baseline candidate solutions for the subsequent search. Furthermore, compared with DE, the complete method reduced the MATD by 0.250 °C. The corresponding confidence interval lay entirely below zero, the Holm-adjusted
p-value was less than 0.001, and the rank-biserial effect size was −0.990. Although the resource-input index reduction achieved by DE was 0.070 percentage points greater than that achieved by the complete method, this difference was not statistically significant. These results further demonstrate that the complete method achieved a favorable combination of MATD and resource-input index reduction while maintaining a good overall balance among temperature-control performance, convergence behavior, and computational efficiency.