1. Introduction
Tin is an important non-ferrous metal that is widely used in solder, chemical products, tinplate, and advanced electronic materials. In industrial production, tin-bearing secondary materials and complex concentrates are commonly treated through pyrometallurgical processes because of their high processing capacity and adaptability to variable feedstocks. Among these processes, top-blowing furnaces provide intensive gas–liquid–solid interaction and are therefore suitable for the treatment of tin-bearing materials with complex compositions. However, the operation of a top-blowing tin-smelting furnace involves strongly coupled thermal, chemical, and flow phenomena. Small variations in operating parameters may simultaneously affect tin recovery, energy consumption, slag–metal separation, and the degradation of the magnesia–chrome refractory lining. Consequently, operating-parameter regulation should not be regarded as a single-objective control problem, but rather as a constrained multiobjective optimization problem.
In practical tin smelting, the coal-consumption rate, lance height, and belt-scale feed rate are the main adjustable operating parameters considered in this study. These parameters have different but strongly coupled effects on furnace operation. The coal-consumption rate directly affects the available heat input and the thermal balance of the furnace, thereby influencing the temperature required for the reduction and separation of tin-bearing phases. The lance height changes the position and intensity of gas injection and consequently affects gas dispersion, bath agitation, reaction-zone distribution, and local heat and mass transfer. The belt-scale feed rate determines the amount of material entering the furnace per unit time and therefore influences the material loading, residence time, slag generation, and thermal demand. The combined adjustment of these three variables may improve tin recovery or reduce energy consumption, but it may also increase thermal stress, chemical attack, and the mechanical erosion of the refractory lining [
1,
2,
3,
4,
5,
6]. Therefore, the relationship between the adjustable parameters and the process objectives is nonlinear, coupled, and subject to operational constraints.
The service life of the magnesia–chrome refractory lining is particularly important for the safe and economical operation of the furnace. Refractory degradation is associated with the combined effects of thermal shock, chemical interaction with molten slag and metal, and mechanical wear. In the present process, these effects are influenced by the three adjustable parameters through their effects on the furnace thermal field, gas–liquid–solid mixing, material loading, and local flow conditions. For example, an excessive coal-consumption rate may increase the local temperature and thermal gradients, while an inappropriate lance height may intensify local turbulence and particle erosion. An excessive belt-scale feed rate may increase the solid loading and slag generation, leading to stronger mechanical and chemical interaction with the lining. These coupled effects make it difficult to improve tin recovery and energy efficiency without accelerating refractory degradation. A practical optimization method should therefore consider the three adjustable parameters simultaneously and explicitly account for the trade-offs among tin recovery, energy consumption, and refractory-lining degradation.
Considerable efforts have been devoted to the modeling and optimization of metallurgical processes. In process modeling, first-principles kinetic and thermodynamic models have been developed to describe decarburization, oxygen consumption, heat balance, and impurity removal in high-temperature furnaces [
7]. Data-driven and hybrid intelligent models have also been applied to complex metallurgical processes. Examples include adaptive-network fuzzy inference systems combined with robust relevance vector machines, multivariate regression combined with Gaussian-process regression, and grey-wolf-optimization-based support-vector regression [
8,
9,
10]. In terms of optimization, evolutionary algorithms such as NSGA-II, NSGA-III, MOEA/D, and C-TAEA have been widely used for multiobjective search, whereas reinforcement learning and model-predictive strategies have been explored for sequential control and dynamic decision-making. Reviews of machine-learning applications in steelmaking further demonstrate the growing use of data-driven methods for process prediction, optimization, and operational guidance [
11].
Nevertheless, many existing approaches require either large training datasets or frequent objective evaluations. These requirements are difficult to satisfy in tin smelting. In particular, tin recovery and refractory degradation are not always available through continuous online measurements. Their evaluation may depend on periodic sampling, sample preparation, laboratory chemical analysis, and post-process assessment. Consequently, the interval between two reliable measurements may be substantially longer than the operating or adjustment cycle of the furnace. In addition, the furnace state may vary with changes in the coal-consumption rate, lance height, and belt-scale feed rate. These variations modify the thermal input, gas-injection environment, material loading, reaction-zone distribution, and slag–metal interaction. Although feed composition and other material properties may also fluctuate in industrial production, they are not treated as decision variables in this study; instead, their effects are reflected in the available industrial observations and process uncertainty. The combination of delayed measurements, costly laboratory analysis, variable process responses, and limited opportunities for plant trials makes it difficult to obtain a large and uniformly distributed dataset for repeated industrial optimization [
12].
Recent studies have also explored deep learning and generative modeling to address data scarcity in metallurgical applications. Semi-supervised variational autoencoders and conditional VAE–GAN architectures, for example, have been used to generate synthetic samples or improve data representation under limited-data conditions [
13]. Although these methods can enhance data representation, their complex network structures and training requirements may increase implementation difficulty in industrial environments [
14]. More importantly, many existing metallurgical optimization studies focus on single-objective control or static operating conditions, in which the process state is assumed to remain sufficiently stable during optimization. Such an assumption is restrictive for the present tin-smelting problem because changes in the three adjustable parameters can alter the thermal, chemical, and hydrodynamic conditions of the furnace. Variations in the coal-consumption rate affect the thermal balance; variations in lance height affect gas dispersion, bath agitation, and local reaction conditions; and variations in belt-scale feed rate affect material loading, residence time, and slag formation [
15,
16]. Therefore, the optimization problem should account for the dependence of process responses on the current operating state and should not rely on a single fixed input–output relationship.
Previous studies have also investigated data-driven prediction in tin-smelting processes. For example, ref. [
17] developed a multi-output regression model based on optimized multi-kernel support-vector regression to predict multiple energy-consumption media under small-sample and strongly coupled conditions; ref. [
18] further introduced a grey-wolf-optimized support-vector regression model combined with SHAP analysis to predict comprehensive energy consumption and interpret the contribution of individual features. These studies provide useful tools for process prediction and model interpretation. However, they are primarily concerned with static prediction and energy-related performance indicators. They do not simultaneously optimize tin recovery, energy consumption, and refractory-lining degradation by adjusting the coal-consumption rate, lance height, and belt-scale feed rate under limited data, delayed measurements, changing furnace states, and operational constraints. This limitation motivates the development of an integrated surrogate-assisted multiobjective optimization framework rather than a prediction model alone.
Surrogate-assisted optimization is particularly relevant when objective evaluations are expensive. Existing surrogate-assisted methods use approximate models, including Kriging, Gaussian-process regression, support-vector regression, and radial basis function networks, to reduce the number of direct evaluations required during optimization [
19,
20]. Combined with evolutionary multiobjective algorithms, these models can search for trade-off solutions while limiting costly process evaluations.
However, their application to industrial metallurgical processes still involves several challenges. First, surrogate prediction errors can become substantial in sparsely sampled regions, especially when the observations are limited and the responses change with the coal-consumption rate, lance height, and belt-scale feed rate. Second, a fixed surrogate model may become less representative when changes in these parameters lead to different thermal, reaction, mixing, and material-loading conditions. Third, the direct use of uncertain surrogate predictions may underestimate the risk of violating product-quality, process-safety, or refractory-related limits. Finally, a multiobjective algorithm generally produces a set of mathematically nondominated solutions, whereas plant operators require a practical procedure for selecting one executable operating scheme. Dynamic optimization methods based on transferable surrogates, reinforcement-learning-assisted evolution, and robust cost-sensitive operators have been investigated, but many studies remain focused on benchmark functions, simulated environments, or simplified process settings. Practical applications to constrained, dynamic, and expensive tin-smelting parameter optimization therefore remain limited [
21,
22,
23,
24].
Based on the above analysis, the main research gaps can be summarized as follows. First, the existing literature does not sufficiently formulate tin-smelting parameter regulation as a constrained, expensive, and time-varying multiobjective optimization problem involving tin recovery, energy consumption, and refractory degradation simultaneously [
25]. Second, the data available for this problem are limited not only in size but also in temporal coverage because key performance indicators are obtained through periodic and delayed measurements rather than continuous online sensing. Third, the responses of the process may change with the operating state induced by the three adjustable parameters, namely, the coal-consumption rate, lance height, and belt-scale feed rate. Fourth, surrogate uncertainty is not always incorporated into constraint handling, which may produce solutions that appear feasible in the surrogate space but are unsafe or infeasible in actual operation [
26,
27,
28,
29,
30]. Finally, existing frameworks do not sufficiently explain how a preferred executable operating scheme should be selected from a Pareto set when expert priorities and objective contributions provide different types of decision information.
To address these gaps, this study develops RAMOSTA, an RBFN-assisted dynamic multiobjective state transition algorithm for tin-smelting parameter optimization. The proposed method is designed specifically for the limited-data, expensive-evaluation, and time-varying characteristics of the process. Radial basis function neural-network surrogate models are constructed to approximate tin recovery, energy consumption, and refractory-lining degradation based on the three adjustable parameters. The surrogate models reduce the number of costly direct evaluations required during the optimization process. An uncertainty-aware adaptive offset penalty is introduced to provide conservative constraint handling. The penalty margin is adjusted according to the uncertainty of the surrogate prediction and the constraint status, so that solutions with uncertain feasibility are treated more cautiously. A Pareto-based dynamic state-transition search strategy is then used to explore time-varying trade-offs among the three conflicting objectives while adapting the search to changes in the operating state. Finally, a hybrid-weight TOPSIS decision layer is developed to select an executable operating scheme from the obtained Pareto solutions. In this layer, the fuzzy analytic hierarchy process (FAHP) is used to represent expert preferences while accounting for the uncertainty of expert judgments, whereas Shapley values are used to quantify the marginal contribution of each objective to the overall decision. An adaptive cosine-similarity gating mechanism regulates the relative influence of these two types of information before TOPSIS ranking.
The main contributions of this study are as follows:
A problem-driven constrained multiobjective optimization formulation is established for tin smelting by jointly considering tin recovery, energy consumption, and magnesia–chrome refractory degradation. The decision variables are explicitly limited to the coal-consumption rate, lance height, and belt-scale feed rate, thereby ensuring consistency between the optimization model and the actual adjustable operating parameters.
An RBFN-assisted data-efficient evaluation mechanism is developed for the proposed optimization framework. The mechanism approximates the three process objectives using limited industrial data and reduces the number of expensive direct evaluations required during multiobjective search.
An uncertainty-aware adaptive offset penalty mechanism is proposed for conservative constraint handling. The penalty margin is adjusted according to surrogate uncertainty and constraint violation information, reducing the risk of accepting solutions that appear feasible according to inaccurate surrogate predictions but may be unreliable in actual furnace operation.
A dynamic state-transition search strategy is developed to capture the time-varying effects of the three adjustable parameters. The method considers the effects of the coal-consumption rate on thermal input, the lance height on gas dispersion and bath mixing, and the belt-scale feed rate on material loading and slag formation, thereby searching for operating trade-offs under changing furnace states rather than assuming a single static optimum.
A hybrid decision-support procedure is developed to select an executable operating scheme from the Pareto set. FAHP is used to represent expert preferences under judgment uncertainty, while Shapley values quantify the marginal contribution of each objective. The adaptive cosine-similarity gating mechanism integrates these two types of decision information before TOPSIS ranking, providing an interpretable link between Pareto optimization and plant-oriented operating decisions.
The proposed framework is evaluated using an industrial tin-smelting dataset and compared with NSGA-II, NSGA-III, MOEA/D, and C-TAEA. The evaluation considers Pareto-front quality, hypervolume, inverted generational distance, and constraint satisfaction, as well as the practical characteristics of the selected operating scheme.
The remainder of this paper is organized as follows.
Section 2 describes the tin-smelting process and the three adjustable operating parameters, defines the constrained time-varying multiobjective optimization problem, and presents the RBFN surrogate models, uncertainty-aware adaptive offset penalty, dynamic state-transition search strategy, and hybrid-weight TOPSIS decision procedure.
Section 3 reports the surrogate-model validation, comparative optimization results, constraint-satisfaction analysis, ablation studies, sensitivity analysis, and evaluation of the selected operating scheme using the industrial dataset.
Section 4 discusses the optimization performance, the role of each methodological component, the industrial implications, and the limitations associated with offline surrogate-assisted validation. Finally,
Section 5 summarizes the main findings and outlines future work, including uncertainty calibration, online surrogate updating, and prospective industrial validation.
2. Dynamic Multiobjective Optimization and Decision-Making Framework for Tin Smelting
2.1. Tin-Smelting Process and Control Variables
The tin-smelting process considered in this study is a top-blowing reduction and separation process involving gas, liquid, and solid phases. The process includes the reduction of tin-bearing oxides, slag formation, phase separation, heat transfer, and impurity removal under high-temperature conditions. The overall reduction relationship can be represented by
where the equation represents the principal reduction relationship rather than a complete description of the multiphase reaction mechanism. Other reactions involving iron oxides, silica, carbon monoxide, slag components, and minor impurity-bearing phases may occur simultaneously. These reactions affect the reduction environment, slag properties, heat transfer, impurity removal, and degradation of the magnesia–chrome refractory bricks lining the molten pool.
The manipulated variables were selected through field investigation at the cooperating tin-smelting plant. They represent operating parameters that can be directly adjusted by the operator or the plant control system within predefined equipment, production, and safety limits. The selected variables include the coal feed rate, lance height, and the feed rates of nine belt-scale systems. The coal feed rate affects the reducing-agent supply and the thermal input associated with carbon oxidation. The lance height, defined as the vertical distance from the lance tip to the molten-bath surface, affects the gas-jet penetration condition, gas–liquid interaction distance, bath agitation, and heat and mass transfer near the molten bath. The nine belt-scale feed rates determine the total amount and blending proportions of the confidential furnace-feed streams.
Although the identities and detailed compositions of the nine feed streams are not disclosed because of industrial confidentiality, the streams have different material characteristics and therefore make different contributions to the overall furnace burden. Their feed rates influence the effective burden composition through the weighted blending of the individual streams. In general, for a material property or component
k, the effective burden value can be conceptually expressed as
where
is the feed rate of the
jth belt-scale stream and
denotes the corresponding, confidential material property or component fraction. The values of
are not reported in this paper. Equation (
2) is introduced only to describe the blending mechanism and does not disclose the identity or detailed composition of any individual feed stream.
The resulting effective burden composition and total feed rate influence the slag condition through several process-level mechanisms. First, they affect the loading of slag-forming components and consequently the amount of slag generated during smelting. Second, they modify the relative proportion of oxide components in the burden and may therefore change the effective acidity–basicity balance, liquidus-temperature tendency, liquid-phase fraction, and apparent slag viscosity. Third, changes in the amount of reducible material and other burden components influence the thermal demand and reduction environment required to maintain the desired smelting condition. These changes may further affect slag fluidity, gas–liquid–solid mixing, the settling and separation of metal and slag phases, and the transfer or retention of tin between the slag and metal phases. Therefore, the nine belt-scale feed rates do not influence the slag merely by changing the feed quantity; they also change the effective composition and physicochemical condition of the slag-forming system. In turn, the altered slag condition affects tin recovery, energy consumption, and the chemical and mechanical interaction between the molten slag and the magnesia–chrome refractory lining.
The control vector is defined as
where
denotes the plant-level coal feed rate,
denotes the vertical distance from the lance tip to the molten-bath surface, and
denotes the feed rate measured by the
jth belt scale,
. The nine belt scales correspond to confidential furnace-feed streams with different material characteristics. Their detailed identities, compositions, and numerical operating ranges are not disclosed because of industrial confidentiality.
When coal-bearing material is delivered through one or more of the belt-scale systems, represents the corresponding aggregate coal feed rate used for the thermal and reduction calculations. In this case, the aggregate coal rate is calculated consistently from the relevant belt-scale measurements and is not independently varied in a manner that would double-count the same material input. Thus, the optimization variables are limited to the plant-adjustable coal-feed condition, lance height, and belt-scale feed rates.
The feasible range of each manipulated variable is determined by equipment capacity, material-feeding capability, standard operating procedures, and safety requirements. The control vector includes only variables that are directly adjustable during plant operation. Measured furnace conditions, feed characteristics that are not directly adjustable, and external disturbances are treated as state variables or contextual variables rather than as manipulated variables.
The reduced process-state vector is defined as
where
represents the tin-related process state,
is the furnace-temperature state,
is the gas-phase oxygen-related state, and
represents the effective thermal state. These state variables are obtained from process measurements when available and are otherwise propagated through the reduced state-transition model.
For computational efficiency, the furnace is represented by a reduced material–energy balance model. This model is used to propagate the furnace-state trajectory and to evaluate state-related operating constraints. It does not replace direct plant measurements and does not attempt to reproduce every detailed chemical reaction, equilibrium relationship, or transport phenomenon occurring in the furnace. Its compact form is
Within the reduced state-transition model, the effects of the nine belt-scale feed rates are represented through their contributions to the total material input, effective burden composition, reducible-material loading, thermal demand, and slag-forming-material loading. The slag-related terms in therefore represent the process-level consequences of feed-rate changes, including changes in slag generation, effective slag condition, phase-separation behavior, and tin distribution between the slag and metal phases. This reduced representation captures the influence of the controllable feed streams without attempting to reconstruct the complete equilibrium composition or detailed multiphase transport field of the industrial furnace.
The effective combustion heat generation is represented by
where
denotes the effective combustion efficiency and
denotes the effective heat-release coefficient associated with carbon oxidation. The material-feeding variables
affect the state-transition model through the material-balance, effective-burden, thermal-load, reducible-material-loading, and slag-condition terms contained in
.
The mechanistic component is used for state propagation and state-related constraint checking. Three radial basis function network (RBFN) models are used separately to predict tin direct recovery, energy consumption, and magnesia–chrome refractory degradation. Therefore, the proposed method is a sequential mechanism-assisted and data-driven optimization framework, rather than a fully coupled hybrid model with jointly trained mechanistic and neural components. The effective coefficients in the reduced state model are calibrated using plant process information and are not disclosed because of industrial confidentiality.
2.2. Problem Formulation
The operating parameters are optimized over a finite production horizon to account for the time-varying furnace state and the effects of the operating schedule. In this study, the manipulated variables are limited to the coal feed rate, lance height, and the feed rates of the nine belt-scale systems. The furnace state evolves with these operating variables, and a control action applied at one time interval may influence the thermal condition, material loading, slag condition, and feasible operating region in subsequent intervals. Feed characteristics that cannot be directly adjusted and other external disturbances are treated as contextual variables or sources of process uncertainty rather than as manipulated variables. Because repeated high-fidelity simulation and real-plant trials are expensive, the objective functions are evaluated using surrogate models trained from historical plant data.
The RBFN surrogate input vector consists only of the directly adjustable operating variables. It includes one coal feed rate, one lance height, and nine belt-scale feed rates. Therefore, the surrogate input dimension is
and is defined as
Here, denotes the coal feed rate, denotes the vertical distance from the lance tip to the molten-bath surface, and denotes the feed rate of the jth belt-scale system, with . The reduced furnace-state vector is not included as an independent input of the RBFN models. Instead, it is propagated separately by the reduced state-transition model and is used for furnace-state constraint checking and for describing the dynamic process context.
Let and denote the start and end of the optimization horizon, respectively. Three conflicting objectives are considered.
The first objective is to maximize the average direct recovery of tin. Since the optimization algorithm is implemented in minimization form, the negative recovery is minimized:
The second objective is to minimize the average energy-consumption index:
The third objective is to minimize the average predicted refractory-degradation burden:
In this study, denotes a dimensionless refractory-degradation index for the magnesia–chrome refractory lining in the molten-pool region. It is used to quantify the predicted degradation burden imposed on the refractory lining under the evaluated operating condition. The index is constructed from refractory-related process information and operating-condition information available from the cooperating plant, including the effects associated with high-temperature exposure, thermal gradients, slag–metal interaction, and mechanical erosion. A higher value of indicates a greater predicted degradation burden and therefore a less favorable refractory-service condition. The index is normalized within the modeling and optimization procedure so that it can be compared across operating schedules without disclosing confidential plant-specific measurement units, coefficients, or calibration details.
The refractory-degradation index is not interpreted as a direct instantaneous measurement of refractory thickness loss. Instead, it is an empirical process-performance indicator that represents the relative refractory-degradation burden under the evaluated operating condition. Its dependence on the manipulated variables is learned by the RBFN model from historical plant data. At the process level, the coal feed rate may affect the refractory burden through changes in heat input and temperature gradients; the lance height may affect it through changes in gas-jet penetration, bath agitation, local turbulence, and slag–metal interaction; and the nine belt-scale feed rates may affect it through changes in material loading, effective burden composition, slag generation, slag fluidity, and contact between the molten phases and the refractory lining.
Although the furnace-state vector is not used as an independent RBFN input, it is propagated by the reduced state-transition model to evaluate the dynamic evolution and state-related constraints of the furnace. The state trajectory provides the process context for the operating schedule, whereas the RBFN models directly map the 11-dimensional manipulated-variable vector to the three process-response predictions.
Here, , , and denote the RBFN predictions of direct tin recovery, energy consumption, and the normalized magnesia–chrome refractory-degradation index, respectively. Therefore, represents the average predicted refractory-degradation burden over the optimization horizon, rather than an additional physical measurement or an independent optimization variable. Minimizing favors operating schedules that reduce the cumulative exposure of the refractory lining to unfavorable thermal, chemical, and mechanical conditions.
The operating variables are constrained by equipment capability, feeder capacity, lance-adjustment limits, and standard operating procedures:
The control vector is given by
where
denotes the coal feed rate,
denotes the vertical distance from the lance tip to the molten-bath surface, and
denotes the feed rate of the
jth belt-scale system, with
. The nine belt-scale systems correspond to confidential furnace-feed streams with different material characteristics. Their detailed material identities, compositions, and numerical operating ranges are not disclosed because of industrial confidentiality.
The main predicted process responses are restricted to their acceptable operating regions:
The bounds on R, E, and represent plant-specific acceptable regions. In particular, the upper bound on limits the predicted refractory-degradation burden, whereas the bounds on recovery and energy consumption prevent the optimization from obtaining apparently favorable refractory conditions by sacrificing essential production or energy-performance requirements.
The furnace-state constraints are expressed as
Here, denotes the furnace-temperature state, denotes the residual tin-related concentration or residual tin indicator used for process control, and is the desired terminal furnace-temperature state. These quantities are evaluated from available process information and the reduced state-transition model. Their detailed plant-specific definitions and numerical limits are not disclosed because of industrial confidentiality.
The safety margin is used only as a constraint-handling mechanism. It is not treated as an additional optimization objective and is not included in the objective vector. During optimization, an uncertainty-dependent offset is added to the nominal constraint function to reduce the risk that a candidate solution is accepted because of surrogate prediction error. In particular, when the prediction uncertainty of , , or is large, the corresponding feasible region is treated more conservatively. The original equipment, thermal, product-quality, energy-performance, and refractory-related limits remain the physical acceptance criteria. Thus, the tightened constraints are used to guide the search conservatively, whereas final feasibility is assessed against the original plant limits.
The complete constrained dynamic multiobjective optimization problem is formulated as
The objective vector therefore contains three process-level performance criteria: the negative average direct tin recovery, the average energy-consumption index, and the average normalized refractory-degradation burden. The optimization directly regulates only the 11 manipulated variables, namely, the coal feed rate, lance height, and nine belt-scale feed rates. The furnace-state trajectory is determined separately by the reduced state-transition model, and the effects of the manipulated variables on furnace-state evolution, effective burden condition, slag behavior, and refractory degradation are represented through the state model, the process constraints, and the RBFN surrogate models.
The numerical operating limits, normalization parameters, refractory-index construction details, and plant-specific calibration coefficients are maintained internally and are not disclosed because of industrial confidentiality. The reported optimization results are therefore interpreted in terms of normalized process indicators and relative performance comparisons among feasible operating schedules.
2.3. RBFN-Based Surrogate Modeling
The surrogate models are constructed to approximate the three process responses associated with the plant-adjustable operating variables. In accordance with the process description in
Section 2.1, the input vector contains only the coal feed rate, the lance height, and the feed rates of the nine belt-scale systems. Therefore, the input dimension of the RBFN surrogate models is
. The four reduced furnace-state variables are propagated separately by the reduced state-transition model and are used for furnace-state constraint checking; they are not included as independent inputs of the RBFN response models.
The 11-dimensional surrogate input vector is defined as
where
denotes the coal feed rate,
denotes the lance height, and
denotes the feed rate of the
jth belt-scale system, with
. Thus, the surrogate models use only the directly adjustable operating variables. The furnace-state trajectory
is calculated separately from the reduced state-transition model and is used to evaluate state-related constraints and to describe the dynamic process context.
Before model training, each of the 11 input variables and each output variable is transformed into a normalized representation. The normalization parameters are calculated only from the training subset and are then reused without modification for the validation subset, test subset, and optimization procedure. This treatment prevents information from the validation or test periods from entering the data-preprocessing stage. Missing observations are imputed using the median of the corresponding variable calculated from the training subset. Potential abnormal observations are screened using the interquartile-range rule and are removed only when they are confirmed to be measurement or recording errors. Observations that are statistically unusual but physically plausible are retained as valid process samples because they may represent actual operating conditions. The available observations are divided chronologically into training, validation, and test subsets. The chronological division is used to prevent information leakage from later production periods and to provide a more realistic evaluation of the ability of the surrogate models to generalize to subsequent operating conditions.
For a response variable
y, the normalized root mean squared error (RMSE) on a dataset
is calculated as
where
denotes the number of samples in the corresponding dataset,
is the normalized measured or assessed response, and
is the normalized RBFN prediction. The validation RMSE is calculated using observations that are not used to estimate the output-layer weights. It therefore provides an independent error measure for model development and hyperparameter selection. The test RMSE is calculated only after the model structure and all hyperparameters have been fixed.
An RBFN represents a nonlinear mapping from the 11-dimensional operating-variable vector to a process response by a weighted sum of radially symmetric basis functions:
where
M is the number of basis functions,
and
are the center and width of the
jth basis function, respectively, and
and
b are the output weight and bias. A Gaussian kernel is used:
Three independent RBFN models are trained for direct tin recovery, energy consumption, and the normalized magnesia–chrome refractory-degradation index. The models are trained separately because the three responses have different process meanings, scales, and error characteristics. The use of independent response models also allows the prediction error of each objective to be evaluated separately.
The RBF centers are selected by applying k-means++ clustering to the 11-dimensional training inputs. The clustering procedure is performed using the training subset only. Candidate values of the number of centers are evaluated using the validation subset. For each response model, the center number is selected by comparing the validation RMSE values obtained under the predefined candidate configurations. The model with the smallest validation RMSE is not selected because of its training fit; rather, it is selected because it provides the smallest prediction error on previously unused chronological validation observations among the candidate models. This criterion helps to balance approximation accuracy and generalization performance. In contrast, selecting a model according to the minimum training RMSE could favor an unnecessarily complex model and increase the risk of overfitting.
The initial RBF width is determined from the median nearest-center distance. This quantity provides a data-dependent reference scale for the local spacing of the RBF centers. Specifically, if
denotes the distance between center
and its nearest other center, the reference width is defined as
Three candidate width values are considered:
The factor produces narrower basis functions and more local approximation behavior. The factor 1 retains the width determined from the center distribution. The factor 2 produces wider basis functions and smoother approximation behavior. This validation grid examines the trade-off between local fitting and smooth generalization without introducing an excessively large hyperparameter-search cost.
Once a candidate number of centers and width have been specified, the RBF activation matrix for the training samples is constructed as
where
The bias term is included by augmenting the activation matrix with a column of ones. The output-layer weights are estimated using ridge-regularized linear least squares:
where
is the activation matrix augmented with the bias column,
contains the output weights and bias,
contains only the output weights, and
is the ridge coefficient. The bias is not penalized.
Ridge regularization is adopted because adjacent Gaussian basis functions may overlap substantially, which can produce correlated columns in the RBF activation matrix. This issue is particularly relevant when the industrial dataset is limited and contains measurement noise or correlated operating variables. The ridge term stabilizes weight estimation, reduces weight variance, and limits overfitting while preserving the computational efficiency of linear output-layer training.
For each response model, the candidate number of centers, width factor, and ridge coefficient are selected jointly using the chronological validation RMSE. Let
denote the predefined set of candidate hyperparameter combinations. The selected configuration is defined as
where
represents one candidate RBFN configuration. The minimum validation RMSE is used because it selects the candidate configuration with the smallest prediction error on the independent chronological validation period. This criterion is more appropriate than the training RMSE for model selection because the training RMSE may decrease as model complexity increases, even when the generalization performance deteriorates. The validation RMSE therefore provides a practical criterion for balancing fitting accuracy, model complexity, and prediction performance on subsequent operating periods.
After
is selected, the corresponding model is evaluated once on the untouched chronological test subset. The test RMSE, together with other error measures reported in
Section 3, is used to assess the out-of-sample predictive performance of the final surrogate model. The test subset is not used to select the center number, width, ridge coefficient, normalization parameters, or model structure. No model setting is adjusted according to the test results.
The initial algorithmic configuration is summarized in
Table 1. These values are surrogate-model settings rather than confidential plant operating data.
To estimate predictive uncertainty for constraint handling,
bootstrap RBFN members are constructed for each response. Each member is trained using a bootstrap resample of the training subset, while the chronological validation and test subsets remain untouched. For an input
, the ensemble prediction is calculated as
and the corresponding empirical predictive dispersion is estimated by
The ensemble mean is used as the nominal prediction, whereas is used as an empirical measure of model-disagreement uncertainty in the uncertainty-aware adaptive offset penalty. This uncertainty measure is intended to support conservative constraint handling; it is not interpreted as a complete probabilistic confidence interval for the industrial process.
For the three response models, the nominal predicted outputs are denoted by , , and , respectively. Here, denotes the predicted normalized magnesia–chrome refractory-degradation burden rather than a direct instantaneous measurement of refractory thickness loss. During candidate evaluation, the reduced state model propagates the furnace-state trajectory for a proposed schedule of the 11 operating variables. The RBFN models predict the corresponding process responses from the 11-dimensional operating-variable input. The predicted responses are aggregated over the optimization horizon to calculate the three objective values, while the state constraints are evaluated separately using the propagated state trajectory. The final surrogate performance is assessed using the chronological test subset before the models are used in the optimization procedure.
2.4. Dynamic Multiobjective State Transition Optimization
The control trajectory is represented by a piecewise-constant schedule:
Each control vector contains the eleven directly adjustable operating variables:
Here, denotes the coal feed rate, denotes the lance height, and denotes the feed rate of the jth belt-scale system. Thus, each candidate solution represents a complete schedule for one coal feed rate, one lance height, and nine belt-scale feed rates over the considered production horizon. The values of , the horizon length, and the control-interval duration are specified in the experimental protocol without disclosing confidential plant time scales.
For the search operators, the schedule is vectorized as
The RBFN surrogate input at each control interval is the eleven-dimensional operating-variable vector
The reduced furnace-state vector is propagated separately using the reduced state-transition model. It is not included as an independent RBFN input. Instead, it is used for dynamic state propagation, state-related constraint evaluation, and a description of the evolving furnace condition.
A Pareto archive is maintained during the optimization process. Candidate schedules are generated using rotation, translation, expansion, and axesion operators. The operators are mathematical search-space transformations; they do not represent physical rotation or movement of the furnace equipment. After each transformation, the candidate schedule is projected onto the admissible bounds of the eleven manipulated variables.
For a current candidate schedule
, the rotation operator is expressed as
where
is the dimension of the vectorized schedule,
is the rotation factor, and
is a random matrix whose entries are sampled from a bounded uniform distribution. The rotation operator changes the direction and magnitude of the current schedule vector in the normalized decision space. It creates a perturbed schedule that may differ simultaneously across multiple control intervals and control variables. Thus, the term “rotation” refers to a directional search transformation in the normalized decision space, rather than to a physical rotation of the lance.
The translation operator generates a candidate along the direction from the preceding schedule
to the current schedule
:
where
is the translation factor,
is a random scalar, and
is a small positive constant that prevents division by zero.
The expansion operator is defined as
where
is the expansion factor and
is a diagonal Gaussian random matrix. This operator enables a larger-scale exploration of the search space.
The axesion operator is defined as
where
is the axesion factor and
is a diagonal random matrix with only one nonzero diagonal element. The axesion operator performs a single-direction perturbation and is used to refine an individual decision dimension of the vectorized schedule.
After each search transformation, the generated schedule is projected onto the admissible bounds of the eleven manipulated variables:
The projected schedule is then propagated through the reduced state-transition model:
At each control interval, the three RBFN surrogate models are evaluated using the eleven-dimensional input
. The resulting predictions are aggregated over the optimization horizon to calculate the nominal objective vector:
where
,
, and
are the negative average direct tin recovery, the average energy-consumption index, and the average normalized magnesia–chrome refractory-degradation burden, respectively.
To make the constraint-handling procedure explicit, the violation of a scalar inequality constraint is defined using its positive-part operator:
For a predicted response
subject to
the normalized violation is defined as
where
and
are positive scaling constants based on the corresponding acceptable ranges. The scaling prevents variables with different units or numerical magnitudes from dominating the total violation measure.
For the process-response constraints, the instantaneous nominal violation is
where
The manipulated-variable violation is calculated as
In practice, the projection in Equation (
36) eliminates violations caused solely by the search operators. The term
is retained to provide a general definition of the constraint-handling framework and to identify any violation before projection or any violation caused by other implementation restrictions.
The furnace-state violation is defined as
The terminal-state violation is defined as
Here, the positive scaling constants in Equations (
40)–(
45) are calculated from the corresponding allowable operating ranges or internally defined normalization parameters. Their numerical values are not disclosed because of industrial confidentiality.
The surrogate models also provide empirical prediction-dispersion estimates. Let
,
, and
denote the ensemble prediction dispersions of direct tin recovery, energy consumption, and the refractory-degradation index, respectively. The uncertainty-dependent offsets are defined as
where
,
, and
are nonnegative safety factors. The uncertainty-tightened violation of the response constraints is then calculated as
The uncertainty offsets shrink the practically accepted response region during the search. For example, the lower recovery requirement is increased by and the upper energy and refractory-degradation limits are effectively reduced by their corresponding uncertainty offsets. The original plant limits remain the final physical acceptance criteria. The tightened violations are used only to guide the optimization conservatively.
The total nominal and uncertainty-tightened constraint violations over the optimization horizon are defined as
and
The nominal violation measures whether the candidate schedule satisfies the original plant constraints, whereas the tightened violation additionally accounts for surrogate-prediction uncertainty. The adaptive penalty coefficient is denoted by
at optimization iteration
q. It is updated according to the current constraint-handling stage and is nonnegative. The penalized objective values are explicitly defined as
where
Equivalently, the three penalized objective values are
Thus, a feasible schedule with retains its nominal objective values. For an infeasible schedule, a nonnegative penalty is added to all three minimization objectives. The penalty increases with the total normalized constraint violation and with the adaptive coefficient . Consequently, schedules with smaller tightened violations are favored when their nominal objective values are otherwise comparable.
Feasibility is considered before Pareto dominance. Specifically, a candidate with zero nominal violation is preferred to a candidate with positive nominal violation. Among candidates with the same feasibility status, the penalized objective vector in Equation (
50) is used for Pareto-dominance comparison. Archive diversity is maintained through a distance-based truncation rule. The nominal objective values and the nominal constraint violations are retained for final reporting, whereas the penalized objective values are used only during the search and archive-update procedures.
Each candidate schedule is therefore evaluated through the following sequence. First, the schedule is projected onto the bounds of the eleven manipulated variables. Second, the projected schedule is propagated by the reduced state-transition model. Third, the RBFN models are evaluated using the eleven-dimensional operating-variable input at each control interval. Fourth, the nominal objective values and prediction-dispersion estimates are calculated. Fifth, the nominal and uncertainty-tightened constraint violations are calculated. Finally, the adaptive penalty is added to the three nominal objectives to obtain the penalized objective vector used for constraint-aware Pareto comparison.
The optimization procedure is summarized as follows:
- Step 1.
Normalize the eleven manipulated variables and initialize a population of candidate schedules within the allowable bounds.
- Step 2.
Propagate the furnace-state trajectory of each candidate schedule using Equation (
37).
- Step 3.
Evaluate the three RBFN models at each control interval using the 11-dimensional input and aggregate the predicted responses into the nominal objective values.
- Step 4.
Calculate the nominal prediction-dispersion estimates and the uncertainty-dependent offsets for direct tin recovery, energy consumption, and the normalized refractory-degradation index.
- Step 5.
Calculate the nominal constraint violation
and the uncertainty-tightened constraint violation
using Equations (
48) and (
49).
- Step 6.
Update the adaptive penalty coefficient
and construct the penalized objective vector according to Equation (
50).
- Step 7.
Generate new schedules using the rotation, translation, expansion, and axesion operators.
- Step 8.
Project generated schedules onto the admissible bounds of the coal feed rate, lance height, and nine belt-scale feed rates.
- Step 9.
Update the nondominated archive according to nominal feasibility, penalized Pareto dominance, and archive-diversity criteria.
- Step 10.
Check newly available validation or replay samples. If the normalized prediction error exceeds the predefined threshold, append the samples to the training pool and retrain the affected RBFN model.
- Step 11.
Stop when the evaluation budget, maximum iteration number, or archive-stagnation criterion is reached.
All competing algorithms use the same initialization rule, 11-dimensional manipulated-variable representation, surrogate models, constraint definitions, penalty formulation, population size, evaluation budget, and stopping criteria. Surrogate evaluations are counted separately from real industrial evaluations. The present study is conducted offline and does not perform a real-plant query for every generated candidate solution.
2.5. Bootstrap-Based Constraint-Tightening Penalty Method
The prediction dispersion used for constraint handling is estimated from a bootstrap ensemble of RBFN surrogate models. The input of each RBFN model is the 11-dimensional manipulated-variable vector consisting of one coal feed rate, one lance height, and nine belt-scale feed rates:
The reduced furnace-state vector is propagated separately by the state-transition model and is used for state-related constraint evaluation. It is not included as an independent input of the RBFN models.
For each response model,
B bootstrap training datasets are generated by resampling the original training dataset with replacement. The same preprocessing, RBFN structure, and output-layer estimation procedure are applied to each bootstrap member. For an input
, the ensemble mean and empirical standard deviation are calculated as
The ensemble mean is used as the nominal prediction. The standard deviation represents the empirical disagreement among the bootstrap surrogate models and is used as a local measure of predictive dispersion. For the three response models, the corresponding quantities are denoted by , , and for direct tin recovery, energy consumption, and the normalized magnesia–chrome refractory-degradation index, respectively.
For a constraint written in the form
the uncertainty-dependent offset is defined as
where
is the baseline normalized engineering margin,
is the empirical standard deviation associated with the corresponding constraint quantity, and
is a nonnegative risk coefficient. A larger
produces a more conservative search region, whereas
removes the uncertainty-dependent component of the offset.
In the present implementation, is used for each normalized constraint. This value is adopted as a conservative scaling factor for converting the empirical bootstrap dispersion into an additional constraint-tightening term. It is not interpreted as a formal probabilistic guarantee or as evidence that the true constraint-violation probability is below a specified level.
The tightened constraint used during optimization is
The original plant constraint remains the physical acceptance criterion. The tightened constraint is introduced only to guide the search more conservatively in the presence of surrogate-model dispersion.
For numerical stability, the violation of the tightened constraint is calculated using a smooth approximation of the positive-part function:
where
controls the sharpness of the approximation. For a sufficiently large value of
,
closely approximates
Thus, the violation is approximately zero when the tightened constraint is satisfied and increases as the candidate schedule moves farther into the infeasible region.
The total penalty for a candidate schedule is calculated by
where
is the number of path constraints,
is the number of terminal constraints,
and
are the corresponding penalty coefficients, and
denotes the violation of the
rth terminal constraint. The integral term accounts for constraint violations occurring throughout the optimization horizon, whereas the second term accounts for terminal-condition violations.
The penalty coefficients are updated according to
where
is the population-average violation of the
mth constraint,
is the penalty-update learning rate, and
is a small positive constant used to avoid division by zero. Consequently, constraints that remain substantially violated across the population receive progressively larger penalty coefficients.
The initial implementation uses , , , , , , and in normalized constraint space. These values are algorithmic settings rather than confidential plant operating limits.
The penalized objective values are defined as
where
,
, and
denote the nominal objective values for negative average direct tin recovery, average energy consumption, and average normalized magnesia–chrome refractory-degradation burden, respectively.
A candidate satisfying all tightened constraints has and therefore retains its nominal objective values. For a candidate violating one or more tightened constraints, the nonnegative penalty is added to each minimization objective. The penalty increases with the magnitude and duration of the constraint violations. Therefore, candidates with lower uncertainty-tightened constraint violations are favored when their nominal objective values are otherwise comparable.
Feasibility is considered before Pareto dominance. A candidate satisfying the original plant constraints is preferred to a candidate violating those constraints. Among candidates with the same feasibility status, the penalized objective values are used for Pareto-dominance comparison and archive updating. The nominal objective values and the original, non-tightened constraint violations are retained for final reporting. The penalty values are used only for constraint handling during the optimization search.
The complete candidate-evaluation procedure is therefore as follows:
- Step 1.
Propagate the furnace-state trajectory using the reduced state-transition model for a proposed schedule of the eleven manipulated variables.
- Step 2.
Evaluate the three bootstrap RBFN ensembles using the 11-dimensional input .
- Step 3.
Calculate the ensemble-mean predictions and empirical prediction dispersions for direct tin recovery, energy consumption, and the normalized magnesia–chrome refractory-degradation index.
- Step 4.
Construct the uncertainty-dependent offsets for the corresponding path and terminal constraints.
- Step 5.
Evaluate the tightened constraints and calculate the smooth constraint violations.
- Step 6.
Integrate the path-constraint violations over the optimization horizon and add the terminal-constraint penalties to obtain .
- Step 7.
Add to each nominal minimization objective to obtain the penalized objective values , , and .
- Step 8.
Update the adaptive penalty coefficients according to the population-average constraint violations.
- Step 9.
Compare candidate schedules according to feasibility, penalized Pareto dominance, and archive diversity.
The bootstrap standard deviation is used only as an empirical model-disagreement measure for conservative constraint handling. It is not interpreted as a calibrated confidence interval or as a formal probabilistic guarantee of constraint satisfaction. The numerical operating limits, normalization parameters, penalty settings, and plant-specific calibration details are maintained internally and are not disclosed because of industrial confidentiality.
3. Performance Evaluation and Industrial Decision Support Results
This section validates the proposed method for tin smelting process parameter optimization. Considering enterprise confidentiality requirements, the batch wise numerical values of manipulated variables are not disclosed in tables or figures. Therefore, the experimental analysis focuses on objective side indicators, feasibility-related metrics, statistical performance, and industrial decision support effects. This setting allows the optimization effectiveness to be evaluated without revealing sensitive operating points.
3.1. Experimental Protocol and Implementation Details
To evaluate the proposed RAMOSTA framework [
31], confidential historical production data were collected from a real tin-smelting process operated under multiple production conditions. The supplied dataset is organized by production-cycle records rather than by isolated laboratory observations. Each record links a production period with the corresponding input materials, output materials, sampling information, process observations, and calculated performance indicators. The records cover consecutive production periods and include different combinations of tin-bearing feed materials, carbonaceous materials, intermediate products, and dust streams.
The data structure contains the following information groups: (i) production-cycle identification and process-period information; (ii) input-material names, material-batch identifiers, and input quantities; (iii) output-material names, material-batch identifiers, and output quantities; (iv) sample names, sampling information, sample identifiers, and tin-batch information; and (v) calculated fields associated with material balance, process operation, tin-related responses, coal-consumption-related responses, direct tin recovery, and coal-consumption loss. The material categories represented in the supplied records include roasting residue, calcined material, tin concentrate, granular anthracite, pulverized coal, purchased tin-bearing materials, electric-field dust, boiler dust, and related product or intermediate streams.
The production-cycle record is used as the effective modeling unit. Because the supplied data are organized by production cycle rather than by fixed-frequency sensor sampling, no uniform time-based sampling interval is assumed in this study. The exact number of production-cycle records, production batches, collection periods, and temporal recording intervals is treated as enterprise-sensitive information and is therefore not disclosed. The exact number of retained input fields and their field-level encoding is also not disclosed because it could reveal confidential process configuration. The retained model inputs consist of material-balance descriptors, process-operation descriptors, and process-context features that are available before the corresponding response is evaluated. Identifiers used only for record tracing or batch management are not treated as predictive inputs.
The response variables used in this study include direct tin recovery, a coal-consumption-related loss indicator, and refractory service life. Direct tin recovery is treated as a beneficial objective, whereas the coal-consumption-related response is transformed according to its optimization direction. Direct tin recovery is calculated as the ratio of the total tin contained in the output products to the total tin contained in all input materials during a production cycle. The output tin content is obtained from three Grade-A tin products and one Grade-B tin product, while the input tin content is calculated from all relevant input-material categories. Thus, direct tin recovery is defined as
where
and
denote the masses of the
jth Grade-A tin product and the Grade-B tin product, respectively;
and
denote their corresponding tin contents;
and
denote the mass and tin content of the
ith input-material category, respectively; and
n is the total number of input-material categories. The refractory-related objective represents the refractory life of the smelting furnace. Refractory degradation is primarily associated with the combined effects of physical erosion, caused by mechanical impacts, melt flow, particle abrasion, and bath agitation, and chemical corrosion, caused by reactions between the refractory lining and the molten bath, slag, gaseous species, or other chemically aggressive process constituents. The refractory-service-life objective is transformed according to its optimization direction and is evaluated using the same definition and constraint-handling procedure for RAMOSTA and all benchmark algorithms. The original mass, tin-content, and refractory-service-life records are not disclosed because of enterprise confidentiality.
Because the dataset is confidential, raw operating values, material quantities, batch identifiers, numerical ranges, physical units, time intervals, and batch-level trajectories are not disclosed. These restrictions also prevent the disclosure of information that could be used to infer plant capacity or production scale. The dataset is instead characterized by its production-cycle organization, variable roles, material categories, chronological structure, and data-processing procedure. This description provides information about the data structure and modeling scope without exposing sensitive enterprise information.
The preprocessing procedure was kept simple and reproducible. First, the data structure was checked for duplicated records, incomplete rows, inconsistent field types, and records that could not be linked to a valid production-cycle record. Records lacking the minimum identifiers required to associate the input, output, and response information with a production-cycle record were excluded before model construction. Duplicate records were removed only when the relevant identifying and response fields were identical. Potential abnormal numerical observations were screened using the interquartile-range rule. An observation was removed only when the abnormal value was confirmed to result from an invalid field format, a duplicated record, an impossible record link, or a measurement, transcription, or recording error. Process observations were not removed solely because they were statistically extreme, since such observations may represent valid operating conditions.
For missing numerical observations, the median calculated from the corresponding training subset was used for imputation. Missing categorical information was assigned to an internal unknown category and was not inferred from later production records. After data cleaning, the records were ordered chronologically according to their production-cycle information. The earliest partition was used for model training, the subsequent partition for validation, and the final partition for testing. No information from a later partition was used to calculate preprocessing parameters or fit model parameters for an earlier partition. Continuous variables were transformed into an internal normalized representation using parameters calculated from the training subset only. Categorical material fields were converted into a fixed internal encoding determined from the training subset, and the same encoding map was applied to the validation, test, replay, and optimization data. The raw normalization parameters and encoding map are retained internally and are not disclosed because of enterprise confidentiality.
The resulting confidential normalized representation was used for offline surrogate-assisted optimization. The RBFN models approximate the available response variables in the normalized feature space, while the optimization algorithm searches for feasible candidate operating schedules using the corresponding surrogate responses and constraint definitions. The same data representation, preprocessing procedure, objective definitions, constraint-handling protocol, and termination conditions were used for RAMOSTA and the benchmark algorithms.
The optimization problem was formulated as a constrained multiobjective problem. Direct tin recovery was treated as a beneficial objective, whereas the coal-consumption-related loss indicator and the refractory life were transformed according to their respective optimization directions.
To assess the competitiveness of RAMOSTA, four representative constrained or multiobjective evolutionary algorithms were selected as peer methods: NSGA-II, NSGA-III, MOEA/D, and C-TAEA. All methods used the same confidential normalized decision representation, objective definitions, constraint-handling protocol, initialization rule, population size, offline evaluation budget, and termination conditions. The parameter settings of the peer algorithms followed their commonly used configurations or the recommendations of their original publications. Each method was independently executed for the same number of runs under the same evaluation protocol. The offline evaluation budget and termination conditions were identical for all compared methods.
The uncertainty-aware penalty mechanism used in the experiments employed a distance-based surrogate uncertainty indicator. The indicator was calculated from the normalized distance between a candidate input and the RBFN center region together with the residual scale estimated during model fitting. A candidate located farther from the region covered by the training samples was therefore assigned a larger conservative uncertainty indicator. This quantity was used to adjust the penalty near constraint boundaries and was not interpreted as a statistically calibrated prediction interval or a guaranteed probability bound.
All experiments were implemented in the same Python-based computational environment and evaluated using the same data-processing pipeline. The results were assessed using Pareto-front quality, feasibility, constraint-violation measures, and the defined process-response criteria. Since the supplied dataset consists of historical production records and no prospective intervention trial was conducted, the experiments represent offline numerical replay and comparative decision-support evaluation. They do not constitute direct furnace trials, prospective production validation, or closed-loop industrial deployment.
3.3. Overall Optimization Performance
The Pareto fronts generated by the five compared algorithms, namely, NSGA-II, NSGA-III, MOEA/D, C-TAEA, and RAMOSTA, are presented in
Figure 3. RAMOSTA produces a competitive and relatively well-distributed set of nondominated solutions compared with the four benchmark algorithms. Its solution set shows improved convergence toward the desirable trade-off region while maintaining coverage across the feasible search domain. These results indicate that RAMOSTA can balance exploration and exploitation under the process constraints and the uncertainty associated with surrogate evaluation.
For consistent metric calculation, all objective values were transformed into the same normalized three-objective minimization space before the HV and IGD calculations. Let denote the feasible historical operating points that contain valid values for all three objectives after preprocessing and normalization. Let denote the feasible nondominated solutions obtained from an independent high-budget surrogate-assisted search conducted with the same objective definitions, normalization procedure, and constraint-handling rules. This independent search was not included in the 30 formal comparison runs.
The empirical reference Pareto set for IGD was constructed as
where
denotes nondominated filtering in the normalized three-objective minimization space. Before filtering, infeasible points and duplicate objective vectors were removed. The same
was used for all five algorithms and all independent runs.
For a nondominated solution set
generated by algorithm
a in run
r, the IGD is calculated as
The historical operating points anchor the reference set to the feasible process regime represented by the confidential production data, whereas the independent high-budget search improves the coverage of trade-off regions that may be sparsely represented in the historical records. Therefore, is interpreted as an empirical approximation of the data-supported Pareto front rather than as an exact global Pareto front.
For HV calculation, a single reference point was determined before comparing the algorithms. For each objective, the maximum value among all points in
was first calculated. The HV reference point was then defined as
where
and
are the component-wise maximum and minimum objective vectors in
,
is a fixed positive margin coefficient,
is a small numerical stabilizer, and
is a three-dimensional unit vector. The same normalized reference point was used for all five algorithms and all 30 independent runs. Because the problem was transformed into a minimization problem, the reference point was selected to be dominated by the objective vectors included in the HV calculation.
To quantify the optimization performance, Hypervolume (HV), Inverted Generational Distance (IGD), feasible ratio, and average runtime were used. HV measures the volume dominated by the obtained nondominated set with respect to the common reference point, and a larger value indicates better convergence and diversity. IGD measures the average distance from the empirical reference Pareto set to the obtained nondominated set, and a smaller value indicates better coverage of the reference set. The feasible ratio represents the proportion of evaluated solutions satisfying all specified operational and process constraints. Runtime represents the computational time required under the same implementation environment and evaluation protocol.
The statistical results over 30 independent runs are summarized in
Table 3. RAMOSTA obtains the largest mean HV and the smallest mean IGD among the five compared algorithms, indicating a more favorable balance between convergence and solution-set distribution relative to the empirical reference front. RAMOSTA also achieves a feasible ratio of
with zero reported standard deviation. This result should be interpreted in accordance with the feasibility definition and constraint-handling procedure adopted in this study: infeasible candidate solutions were identified through explicit process-constraint evaluation, and the implemented variable-bound projection, uncertainty-aware adaptive offset penalty, feasibility-prioritized selection, and archive-screening mechanisms prevented solutions retained for metric calculation from violating the specified constraints. Therefore, the reported
feasible ratio indicates that all solutions included in the evaluated final archives satisfied the prescribed feasibility criteria; this does not imply that every intermediate candidate generated during the search was feasible. The absence of variation across the 30 runs further indicates that this feasibility-preserving procedure operated consistently under the tested experimental settings.
RAMOSTA does not necessarily achieve the shortest runtime because its computational procedure involves additional operations beyond objective-function and surrogate-model evaluations. In particular, each iteration requires uncertainty-indicator calculation, uncertainty-dependent offset adjustment, explicit constraint checking, adaptive penalty updating, feasibility-prioritized Pareto comparison, and candidate-archive maintenance. These additional operations increase the per-iteration computational overhead, even though they improve constraint handling and may contribute to the superior HV, IGD, and feasibility results. Consequently, runtime reflects a trade-off between computational efficiency and the additional calculations required for robust feasibility management, rather than serving as a direct measure of optimization quality.
From the computational perspective, RAMOSTA remains competitive in runtime. Although the adaptive offset mechanism requires additional uncertainty-related calculations, the mean runtime of RAMOSTA is lower than those of NSGA-II, NSGA-III, and C-TAEA, and is higher than that of MOEA/D in the reported experiments. Thus, the additional constraint-handling calculations do not result in a disproportionate computational burden in the present surrogate-assisted setting.
3.4. Representative Compromise Solution Analysis
Although the Pareto front provides a set of candidate trade-off solutions, practical tin-smelting decision support requires the selection of one representative schedule. The decision maker therefore prefers a solution that provides a balanced combination of process performance, feasibility, and distance from active constraint boundaries rather than an extreme solution located near an uncertain boundary. Representative compromise schemes were selected from the nondominated sets generated by NSGA-II, NSGA-III, MOEA/D, C-TAEA, and RAMOSTA using the adaptive fusion decision mechanism.
The statistical comparison of the five algorithm-derived representative schemes is reported in
Table 4. The composite utility, stability margin, robustness score, and decision consistency are normalized decision-support indicators calculated according to the procedures defined in the experimental protocol, with larger values indicating preferable performance. RAMOSTA attains the largest mean value for all four indicators. This result is attributable to the combined effects of its search and solution-screening mechanisms rather than to the use of any algorithm-specific evaluation criterion. Specifically, the uncertainty-aware adaptive offset penalty and feasibility-prioritized selection reduce the likelihood that the final representative scheme is located near uncertain constraint boundaries, which is reflected in the stability-margin and robustness indicators. Meanwhile, the adaptive multiobjective search promotes solutions that maintain favorable trade-offs across the considered process objectives, contributing to higher composite utility. The archive-based retention and representative-scheme selection procedures further favor nondominated solutions with stable performance under the decision-support criteria, thereby improving decision consistency.
It should be noted that the four indicators are not fully independent, because they are calculated from the same normalized process responses, constraint information, and representative-scheme selection protocol. Therefore, RAMOSTA’s simultaneous superiority in these indicators should not be interpreted as four completely independent pieces of evidence. Rather, it indicates that, under the common evaluation framework used in this study, RAMOSTA more consistently identifies feasible and well-balanced representative schemes than the benchmark algorithms. All algorithms were evaluated using identical normalization rules, indicator definitions, and representative-scheme selection procedures; hence, the observed differences cannot be attributed to algorithm-specific post-processing or favorable metric definitions.
RAMOSTA obtains the largest mean values for all four indicators among the five compared algorithms. In particular, RAMOSTA achieves a composite utility of , a stability margin of , a robustness score of , and a decision consistency of in the reported comparison. These results indicate that, within the offline evaluation environment, the RAMOSTA-derived compromise scheme provides a favorable balance between objective performance and constraint-related conservativeness.
The observed advantage is consistent with the coordinated use of the uncertainty-aware adaptive offset penalty and the adaptive weight fusion mechanism. The former discourages risky boundary-seeking solutions in uncertain regions, whereas the latter selects a compromise solution by combining preference information with objective information from the candidate set.
3.7. Industrial Decision Support Analysis
To assess the decision-support value of the proposed framework, the nondominated solutions generated by NSGA-II, NSGA-III, MOEA/D, C-TAEA, and RAMOSTA were further processed using the defined compromise-solution selection procedure. Given the production conditions and operational requirements represented in the confidential offline data, the procedure first generates candidate nondominated solutions and then selects a final recommendation according to the adaptive weight fusion mechanism.
In the reported offline replay, the RAMOSTA-derived strategy shows improved decision consistency and lower fluctuation according to the defined normalized indicators. These results suggest that the selected compromise solutions have decision-support potential under the evaluated historical process conditions. They do not constitute prospective evidence of closed-loop production improvement.
Figure 6 compares the average normalized benefit-oriented performance of the five algorithms over the replay horizon. RAMOSTA achieves the highest normalized direct tin recovery and refractory-life scores, while obtaining the lowest normalized coal-consumption loss among the compared methods in the reported offline evaluation. The coal-consumption-related loss and refractory-life response are transformed according to their optimization directions so that larger normalized scores represent preferable performance.
Among the benchmark algorithms, NSGA-II, NSGA-III, MOEA/D, and C-TAEA also show optimization capability. The similar average normalized performance of NSGA-III, MOEA/D, and C-TAEA indicates that their representative schemes achieved comparable aggregated results under the common offline replay and normalization procedure. Because the plotted values are averaged normalized indicators, relatively small differences in individual objectives or constraint margins may not be visually distinguishable in the bar chart. Their relative performance should therefore be interpreted together with the Pareto-quality, feasibility, and statistical results. RAMOSTA provides the best reported balance among direct tin recovery, coal-consumption loss, and refractory life in the present offline replay.
For a more intuitive comparison,
Figure 7 presents a radar chart of the three normalized aggregated indicators. The coal-consumption-related loss is converted into a fuel-economy score for visualization, and the refractory-life response is transformed according to its optimization direction. The larger enclosed area of RAMOSTA indicates a more favorable balance among the three normalized indicators in the reported offline comparison.
Figure 7 presents the normalized benefit-oriented scores of direct tin recovery, fuel economy, and refractory-life performance. The larger enclosed area of RAMOSTA is associated with its relatively high direct-recovery and refractory-life scores, while its fuel-economy score remains competitive with those of the other methods. Because the area is jointly determined by the three normalized axes and their interactions, it cannot be attributed to any single objective. The radar chart therefore illustrates the overall balance of the three indicators in the offline replay rather than proving that one specific indicator is solely responsible for the observed difference.
To examine the stability of the final recommendation, a decision-weight sensitivity analysis was conducted. The baseline decision weights were perturbed within a prescribed relative range while preserving the normalization procedure and the candidate nondominated sets. For each perturbed weight setting, the adaptive fusion mechanism was reapplied to the same candidate solutions, and the resulting representative scheme was recorded. The analysis focused on the selected-scheme consistency, composite utility, and normalized refractory-life score.
The sensitivity-analysis results are summarized in
Table 7. All weight settings were evaluated using the same RAMOSTA candidate set and identical offline replay conditions. Thus, the observed changes in the selected representative scheme arise solely from changes in the decision-preference weights, rather than from stochastic variation in the optimization process or differences in the candidate solutions.
As the weight perturbation range increases, the variation in the decision-support scores generally becomes larger and the selected-scheme consistency correspondingly decreases. This behavior is expected because the final recommendation is determined by a weighted aggregation of multiple normalized criteria. Under small perturbations, the relative ranking of the leading candidate schemes is usually preserved, particularly when the baseline selected scheme has a sufficient score advantage over competing schemes. Consequently, the same scheme, or schemes with very similar performance, is repeatedly selected. When the perturbation magnitude increases, the contribution of individual criteria to the aggregated score changes more substantially. Candidate schemes with different strengths, such as higher utility but lower robustness, or greater stability but lower composite utility, may then exchange their ranking positions. This produces larger fluctuations in the selected-scheme statistics and reduces the frequency with which the baseline scheme remains the final recommendation.
The decrease in consistency should therefore not be interpreted as instability in the RAMOSTA optimization procedure. Instead, it reflects the expected sensitivity of a multi-criterion decision result when decision preferences are deliberately varied over a wider range. In particular, a more pronounced reduction in consistency may occur when several candidate schemes have similar baseline aggregate scores, because even moderate changes in criterion weights can alter their ranking. Conversely, the consistency values obtained under small and moderate perturbations quantify the robustness of the recommendation to plausible preference uncertainty. Overall,
Table 7 indicates the range over which the recommended scheme remains stable and identifies the point at which changes in decision preference become sufficiently large to favor alternative, yet still competitive, feasible schemes.
The sensitivity analysis should be interpreted as a robustness check for the decision layer rather than as a validation of the optimization layer. If moderate weight perturbations preserve the selected scheme or produce only small changes in the composite utility and refractory-life score, the final recommendation can be considered stable under the tested preference variations. If the selected scheme changes frequently, the decision layer should be reported as preference-sensitive, and the final recommendation should not be described as uniquely determined.
In summary, the offline replay results show differences in Pareto-front quality, feasibility, and compromise-solution performance among NSGA-II, NSGA-III, MOEA/D, C-TAEA, and RAMOSTA under the tested confidential evaluation protocol. The decision-weight sensitivity analysis further provides a check of the stability of the RAMOSTA-derived recommendation under moderate preference changes. These results should be interpreted within the stated dataset, surrogate-model assumptions, and offline evaluation conditions.