1. Introduction
Two-objective optimization arises when improving one criterion degrades another, so the solution is a set of non-dominated trade-offs rather than a single scalar optimum. Option-pricing model selection and calibration in financial mathematics and quantitative economic analysis provide one motivating context because pricing fit, parameter stability, hedging robustness, model complexity, and computational cost may conflict. The present article addresses the methodological problem of objective-space approximation and non-dominated extraction. Its experimental evidence is obtained from synthetic benchmark functions, while the financial context identifies the intended direction of subsequent application research [
1,
2].
Evolutionary multi-objective algorithms, represented by NSGA-II and SPEA2, remain widely used because they maintain candidate populations, rank them through Pareto dominance, and preserve diversity during search [
3,
4]. Modern implementations, component-wise configuration, decomposition, indicator-based selection, and surrogate-assisted search have improved reliability and sample efficiency [
5,
6,
7,
8,
9]. Neural approaches can additionally learn preference-conditioned trade-offs, approximate complex Pareto-front geometries, and guide policies through structured decision spaces. These developments are relevant to the present study, but they mainly improve candidate generation, surrogate prediction, or preference-conditioned search, whereas the proposed fixed kernel directly implements a two-objective dominance test after candidate objective vectors have been generated.
Modern toolkits and improved implementations make evolutionary algorithms easier to reproduce, but they do not remove the dependence on population size and repeated dominance checks. Studies on component-wise design and automatic algorithm configuration also show that performance gains increasingly arise from adaptive operators, alternative search structures, and problem-specific computational modules. These results motivate a more structural question: whether dominance itself can be converted into a representation that is easier for deterministic or neural operators to process. If the relation can be expressed in a regular objective-space form, part of the sorting workload can be transferred from nested vector comparison to matrix operations that are naturally parallel.
Surrogate-assisted and Bayesian methods reduce the cost of expensive objective evaluations through learned response models and acquisition functions. In particular, parallel Bayesian optimization of multiple noisy objectives can improve sampling efficiency by selecting candidate evaluations according to expected hypervolume improvement [
10]. Its principal mechanism nevertheless remains search and candidate selection rather than a direct reformulation of non-dominated sorting as a regular operation on a dense two-dimensional objective-space image.
Image representations have previously been used to visualize high-dimensional optimization states and to connect optimization trajectories with standard image-processing tools. Encoder-decoder and progressive feature-enhancement models further show how local boundaries and small structures can be preserved in high-resolution images. The present formulation differs in both the represented quantity and the operator: pixels encode occupied cells in a two-objective value plane, and the deterministic kernel is derived from Pareto dominance rather than learned as a generic segmentation filter. The supervised CNN remains an approximation module, while exact non-dominance is supplied only by the rule-based extractor together with the original-vector archive.
Recent application-oriented research illustrates the broader value of combining structured numerical models, multidimensional representations, information fusion, and data-adaptive computation. Interaction-aware robotic intervention, engineering-drawing information protection, multilevel particle-morphology simulation, geometry–intensity sensor calibration, agent-based spatial modelling, and interface-sensitive tomographic reconstruction all rely on the coordinated processing of explicit structural information and complex observational data [
11,
12,
13,
14,
15,
16]. Related developments include reduced-order projection models, multisource state-space prediction, multimodal system-state estimation, precise integration methods, collaborative routing optimization, and multivariate index construction [
17,
18,
19,
20,
21,
22]. Image restoration under rotational motion, discrete-element and finite-element analysis, machine-learning-based industrial detection, high-precision synchronization, and nonlinear structural shape-finding further demonstrate how physical constraints, geometric representations, and computational estimation can be organized within unified numerical frameworks [
23,
24,
25,
26,
27,
28].
A parallel methodological trend concerns the use of surrogate modelling, enhanced neural representations, interpretable learning, and secure intelligent computation. Hierarchical clustering weighted Gaussian-process regression provides a data-efficient approach to low-fidelity prediction, while neural-observer-based Mamba architectures, enhanced visual detection networks, interpretable ensemble models, AI-governance frameworks, and time–frequency deep learning demonstrate complementary strategies for representation learning, prediction, and model control [
29,
30,
31,
32,
33,
34]. Interpretable machine learning for material design, dynamic security computation, intelligent design automation, and learnable-parameter-driven Transformer models further show how structural constraints and adaptive models can be integrated in complex decision environments [
35,
36,
37,
38]. Neural-network-supported brain–computer-interface analysis, secure trajectory prediction and task offloading, data-science-assisted multi-criterion material optimization, and autonomous online-learning generation additionally highlight the importance of reliability, feature representation, and adaptive decision mechanisms in intelligent systems [
39,
40,
41,
42,
43]. Although these studies address different application domains, they collectively support the methodological premise that explicit structural knowledge and nonparametric learning can serve as complementary computational layers.
Symmetry is used here only as an interpretive device for the fixed two-objective operator. For a displacement δ = (δ1, δ2), the elementary central inversion J(δ) = −δ maps the local dominating quadrant to the opposite dominated quadrant, while repeated cross-correlation translates the same relation across the grid. Minimization is directional and asymmetric because only simultaneous objective improvement is tested for exclusion. This manuscript therefore does not claim a new symmetry theorem or global invariance of the optimization problem, neural architecture, or Pareto front.
These strands leave a specific computational gap. Existing optimizers primarily determine where candidate vectors should be evaluated, learned front models approximate trade-off sets, and image models provide spatial feature extraction. They do not, in general, convert the two-objective partial order itself into a fixed, explainable grid operator. This study therefore separates three roles: sampled front-image approximation, deterministic dominance extraction, and exploratory decision-space movement. The modules share an objective-space representation and may be combined, but they can also be invoked independently.
Reinforcement learning provides a complementary search mechanism. Value-based and policy-gradient methods can learn sequential decisions under sparse and delayed feedback [
44,
45]. Multi-objective reinforcement learning further considers vector rewards, coverage sets, preference-conditioned policies, and decomposition-based formulations [
46,
47,
48,
49,
50,
51,
52,
53,
54,
55]. For numerical multi-objective programming, the main difficulty is reward sparsity: only a small fraction of transitions improves the maintained non-dominated structure.
The contributions are fourfold. First, this study defines a modular two-objective representation in which sampled objective vectors are mapped to occupancy images without scalarizing the trade-off set; this representation is motivated by future model-selection and calibration research, including option-pricing applications. Second, a cell archive stores each sample index, decision vector, and original objective vector assigned to an occupied raster cell, thereby separating cell-level image processing from exact vector-level reporting. Third, a deterministic cross-correlation extractor is constructed from the two-objective dominance quadrant and equipped with an explicit coordinate convention, finite-range coverage condition, and archive refinement. Fourth, the supervised CNN and the three reinforcement-learning variants are positioned as optional approximation or exploration modules, and the evaluation distinguishes front-image approximation, deterministic extraction, and SCH feasibility evidence rather than combining them into a single competitive claim.
2. Two-Objective Foundations and Local-Symmetry Interpretation
2.1. Generic Two-Objective Formulation and Prospective Option-Pricing Interpretation
In the generic formulation, x may represent model parameters, hyperparameters, or other decision controls. In future option-pricing research, x could collect coefficients of a parametric pricing model together with controls of a nonparametric residual module, while paired objectives could represent pricing fit and stability or complexity, or calibration error and hedging error. In the present benchmark study, however, the objective mapping f is instantiated by synthetic two-objective functions so that the computational behavior of the proposed representation and extraction method can be examined independently of a particular market model.
Here X ⊆ ℝn, Y ⊆ ℝm, f : X → Y, and g : X → ℝp. Each feasible vector is mapped to y = f(x). A scalarization can return a single preferred solution, but a full Pareto front normally requires repeated scalar problems. The present framework keeps the vector form and processes the front directly in objective space.
The complete construction is restricted to two objectives and box-constrained decision variables. This scope is not merely an implementation choice: the binary image, the opposite-quadrant dominance relation, and the fixed two-dimensional kernel all depend on a planar objective space. Problems with three or more objectives would require a tensor grid, projection, or graph-based dominance representation and are outside the validated scope of this article.
The resulting numerical pipeline is therefore a general computational layer: feasible samples are drawn from X, mapped by f, encoded as occupied cells, and processed in objective space. A future financial implementation would replace the benchmark mapping with a specified pricing or calibration model while retaining the same distinction between candidate generation, raster representation, cell-level extraction, and vector-level archive refinement.
2.2. Pareto Dominance, Non-Dominated Set, and Pareto Front
For minimization, vector
xa dominates vector
xb only if it is no worse in every objective and strictly better in at least one objective:
The notation xa ≺ xb means that xa dominates xb. This relation is a partial order. Many feasible solutions are incomparable because one solution may improve one objective while worsening another, which is the central difference from scalar optimization.
Given a finite or sampled feasible set
D ⊆
X,
a vector
x* ∈
D is dominated if another vector
x ∈
D satisfies
x ≺
x*. Otherwise,
x* is non-dominated. The non-dominated set of the full feasible space is the Pareto set PS, and its image in objective space is the Pareto front PF:
This study focuses on PF because the convolutional modules operate in objective space. The associated decision vectors can be recovered through the sample-index mapping used during image construction. This separation clarifies two tasks that are often coupled in evolutionary algorithms: decision-space search and objective-space sorting.
2.3. Objective-Space Imaging and Dominance Symmetry
For two-objective minimization, the objective-space image is a finite occupancy representation of sampled objective vectors. A cell value of one means that at least one vector falls within the cell bounds. The image is not the continuous set f(X), and it becomes meaningful only together with the objective bounds, resolution, sampling density, and the archive described below.
Rasterization is many-to-one. Each occupied cell therefore stores not only a binary value, but also an archive of all original sample indices and objective vectors assigned to that cell. The binary extractor returns a cell-level non-dominated boundary. Before vector-level solutions are reported, the archived vectors in retained boundary cells are checked by the exact Pareto relation, and duplicate or mutually dominated vectors are removed. Without this archive refinement, two distinct samples colliding in one pixel would be indistinguishable and the binary result could only be interpreted as a quantized cell-level approximation.
Quantization error is controlled by the cell widths h1 and h2. If a cell center represents all vectors assigned to that cell, the coordinate-wise and Euclidean positional bounds are given in Equation (A2). Lower resolution increases this bound and increases the probability of collisions; higher resolution reduces positional error but increases grid memory, convolution cost, and the sampling density required to occupy the front continuously. The revised claims therefore distinguish the exact sampled vectors, the rasterized cell approximation, and the theoretical continuous Pareto front.
For a displacement
δ from an occupied cell, define
Q− as the simultaneous-improvement quadrant and
Q+ as the simultaneous-worsening quadrant. The central inversion
J in Equation (A3) maps
Q− to
Q+. This is the local dominance symmetry used to construct the kernel. The cross-correlation operator is translation equivariant because translating the occupancy image translates the response by the same amount. The asymmetric step is the minimization-directed decision rule: only occupancy in
Q− invalidates a candidate, so
J is not imposed as an invariance of the final output.
Equation (7) and the fixed kernel implement this convention. A retained cell has no occupied cell in the simultaneous-improvement quadrant. Other visible external boundaries are not Pareto-optimal under minimization, so the operator uses the local opposite-quadrant relation but keeps only the direction selected by the partial order.
Repeated cross-correlation expands the inspected simultaneous-improvement quadrant. The operation converts cell-level dominance testing into regular image filtering; exact vector-level output is obtained only after archive refinement.
2.4. Neural Computation Modules Used in This Study
The framework uses only the neural operations required by the proposed algorithms: affine layers, two-dimensional convolution, transposed convolution or up-sampling, back-propagation, and value or policy approximation for reinforcement learning. A fully connected layer is written as
where
W is the weight matrix,
b is the bias,
σ is the activation function, and
h is the output. The supervised solver mainly relies on two-dimensional convolution over binary images:
Deep-learning libraries implement two-dimensional cross-correlation, commonly named “convolution” in software interfaces. To avoid a notation conflict, fconv denotes this cross-correlation convention throughout Equations (6), (9) and (10): the kernel is translated without spatial reversal. Network I reconstructs a high-density sampled objective-space image, whereas Network II approximates its Pareto-front boundary.
The reinforcement-learning modules approximate either an action-value function or an Actor-Critic pair. Their states, actions, and rewards are defined from decision variables, dominance updates, and objective-space images, not from generic image labels.
2.5. Implementation Environment
The implementation uses Python 3.8.10 using TensorFlow 2.8.0, NumPy 1.21.6, SciPy 1.7.3, and Matplotlib 3.5.1. TensorFlow provides neural training and inference, while NumPy/SciPy support sampling and benchmark evaluation. The convolutional components are compatible with GPU execution.
Experiments were run on a Dell Precision 7530 mobile workstation with an Intel Core i7-8750H CPU, 16 GB RAM, and an Nvidia Quadro P3200M GPU. The hardware description contextualizes the archived runtime values. The revised evaluation reports method-specific evidence and does not use these results to rank independently tuned optimization algorithms.
3. Proposed Modular Two-Objective Computational Method
3.1. Nonparametric CNN Mapping from Objective-Space Images to Pareto-Front Approximations
The nonparametric component casts Pareto-front approximation as image-to-image prediction. For each benchmark instance, decision vectors are sampled uniformly from the decision box, evaluated by the explicit objective functions, and rasterized as a sparse binary occupancy image. In an option-pricing application, these explicit objectives would be supplied by a parametric or hybrid calibration model. A substantially larger finite sample of the same mapping provides a high-density sampled objective-space approximation, from which the sampled Pareto-front label is constructed. Neither target is an exact continuous image of f(X); both depend on the retained objective bounds, sample count, and raster resolution.
Two encoder-decoder CNNs are used. Network I maps a sparse objective-space image to a reconstructed full image. Network II maps the full or reconstructed image to a PF image. This design keeps both training and inference in objective space and avoids scalarizing the vector problem. The two-network design also makes the computational roles explicit: reconstruction and PF extraction are separated rather than hidden inside one opaque end-to-end network.
Network I reconstructs the objective-space image from sparse samples. Its stage-wise configuration is summarized in
Table 1. The encoder uses 3 × 3 convolutions and four max-pooling stages, while the decoder restores the resolution through transposed convolution and up-sampling.
Network II extracts the PF from the predicted or true objective-space image. Its compact stage-wise configuration is summarized in
Table 2. It is shallower than Network I because it isolates a boundary segment rather than reconstructing the full region.
The input representation is a binary occupancy grid in which a cell value of one denotes at least one sampled objective vector. After rasterization, decision coordinates are not fed to the CNN; the networks learn the geometry of the objective-space image. The complete phase-structured procedure for data construction, rasterization, target generation, CNN training, and subsequent inference is summarized in
Table 3.
During inference, Network I reconstructs a high-density sampled objective-space approximation from a sparse image, and Network II outputs a sampled PF image. This module is useful only when the target problem belongs to a structural family represented in training, and when rapid image approximation is more important than a formal dominance guarantee. Exact vector-level use requires the deterministic extractor and archive refinement. Sampling density, resolution, and collisions are therefore part of the approximation error rather than incidental plotting choices.
3.2. Convolutional Non-Dominated Sorting Based on PF Extraction Kernels
The deterministic component accepts a binary occupancy matrix A and its cell archive. Its direct output is a binary matrix B that retains an occupied cell only when no occupied cell in the dominance quadrant invalidates it at the selected resolution. The archive is then used for exact vector-level refinement. Unlike Network II, the kernel weights are not trained; they are constructed from the two-objective partial order, which makes the cell-level decision explainable and independent of the supervised training distribution.
A padding matrix  places A at the center of a larger zero matrix. Under the coordinate convention in Equation (A1), potentially dominating cells have larger row indices and smaller column indices. The fixed kernel places ones in that quadrant, and fconv translates the kernel without spatial reversal.
The integrated fixed-kernel example is summarized through representative raster coordinates in
Table 4. Unlike a trainable CNN, this is a feedforward convolutional computation with a fixed kernel. Repetition expands the inspected dominance region until the full image extent is covered.
Let
A = (
aij) be the binary image of a finite objective-vector set
S. If
α and
β are represented by occupied pixels
ap,q and
ar,s, respectively, the image-space dominance condition is
Equation (7) restates two-objective Pareto dominance under the selected matrix convention. A point that dominates α is no larger in both objective values and therefore falls in the corresponding dominance quadrant.
Let
K be a PF extraction kernel of size (2
u − 1) × (2
u − 1), and let
B =
fconv(
Â,
K). For an occupied pixel
aij representing
α, the kernel response satisfies
The kernel sums occupied pixels in the dominance quadrant. If only the central pixel contributes, the response equals one and α remains non-dominated. If another occupied pixel appears in that quadrant, the response exceeds one and α is removed after thresholding. The proof is therefore constructive: the value of the convolutional response has a direct dominance meaning, rather than being an empirical score produced by a trained classifier.
A binary cell with value one may contain several distinct objective vectors. Convolution alone cannot compare vectors that collide in the same cell because the occupancy value remains one. The implementation therefore associates every cell with a list of original objective vectors and sample indices. After the convolutional boundary cells are obtained, exact pairwise dominance is applied only to archived vectors associated with those cells and, when adjacent cells share a quantization boundary, to the relevant neighboring archives. The reported vector set is the refined archive result; the unrefined image is explicitly described as a cell-level approximation.
After each cross-correlation, thresholding suppresses cells whose response shows an additional occupied cell in the improvement quadrant, and intersection with A prevents empty cells from being introduced. The coverage condition in Proposition 1 yields the stopping rule in
Table 5. The proof is cell-level; the archive refinement is required for collisions.
The computational work and applicability conditions of this rule-based stage are analyzed in
Section 3.4 and
Appendix A.2; no asymptotic advantage is inferred from the use of convolution alone.
3.3. Exploratory Reinforcement-Learning Candidate Generation
The third component formulates candidate generation as a Markov decision process. States may be decision vectors, binary codes, or distribution parameters; actions update those states; rewards depend on the maintained non-dominated structure. The RL modules are independent candidate generators. Only MOP-AC-sample necessarily invokes the final convolutional extractor in the archived implementation.
MOP-DQN uses a
Q-network to approximate action values. The reward is one when the updated state is not dominated by the current set and zero otherwise. Transitions are stored in a replay queue, and mini-batches update the network through the discounted Bellman target. This design is intentionally simple, so the result should be read as a feasibility test for sparse dominance rewards rather than as a state-of-the-art MORL comparison. The complete phase-based MOP-DQN procedure, including initialization, ε-greedy action selection, state transition, replay-based value learning, and output of the maintained non-dominated set, is summarized in
Table 6.
The SCH configuration uses a three-layer
Q-network with two hidden layers of 20 ReLU units and two linear output units. The principal settings are given in Equation (11).
MOP-AC uses an Actor-Critic structure. The state is a binary code of the decision variable, the Actor outputs action probabilities, and the Critic estimates state value. Equation (12) assigns four ordinal reward levels to repeated, dominated, dominating, and newly non-dominated states. The complete phase-based MOP-AC procedure, including initialization, policy and value evaluation, state transition, temporal-difference updating, and maintenance of the non-dominated set, is summarized in
Table 7.
For SCH, both the Actor and the Critic contain two hidden layers with 128 ReLU units. The Actor output has 20 × 2 softmax units because state and action values are represented by 20-bit binary coding. The main settings are given in Equation (13).
MOP-AC-sample changes the state from a single decision vector to parameters of a normal distribution, including mean and standard deviation. One agent state can therefore generate multiple decision vectors in one step, increasing exploration relative to one-vector updates.
MOP-AC-sample evaluates multiple vectors from one distribution state, rasterizes them into the accumulated occupancy image, and uses newly occupied useful cells as a reward signal. At termination, the fixed extractor and cell archive produce the sampled non-dominated result. This is the only archived branch that directly couples RL and deterministic image extraction. The complete MOP-AC-sample procedure, including distribution-based sampling, occupancy-image updating, Actor–Critic optimization, archive maintenance, and final convolutional extraction, is summarized in
Table 8.
For the SCH trial, the decision space is [−1000, 1000] and the objective-space image resolution is 200 × 200. For one decision variable, the state shape is (
s,
w +
u), where
s = 1,
w = 20 encodes the mean, and
u = 8 encodes the standard deviation. Ten decision vectors are sampled from each decoded state. The main setting is given in Equation (14).
3.4. Computational Complexity and Local-Symmetry Interpretation
The framework separates two burdens: decision-space search and objective-space sorting. The supervised CNN handles search indirectly by learning image mappings from sampled data, while the DRL variants search through state transitions. The fixed convolutional extractor addresses sorting once objective vectors have been rasterized.
Let N be the number of sampled vectors. Direct general pairwise dominance checking is O(N2) for two objectives, while specialized two-objective sorting can reach O(N log N). Rasterization costs O(N), and the unoptimized grid operation has work O(Tk2(2u − 1)2), where T is defined above; practical GPU libraries reduce wall-clock time through parallel execution, but do not change the work bound. Storage is O(k2 + N) when the occupancy grid and cell archive are both retained. The method therefore has no universal asymptotic advantage over specialized sorting. Its potential benefit is regular parallel filtering when the grid is already available, objective vectors are dense, and sorting rather than objective evaluation is the bottleneck.
The map J is an elementary property of ℝ2 and is not claimed as a new mathematical symmetry. Its role is operational: it identifies the opposite local quadrants used to construct the fixed dominance kernel. The methodological contribution lies in encoding that relation as repeated cross-correlation under an explicit raster convention, proving finite-range coverage at the cell level, and recovering exact vector-level output through the cell archive.
The RL modules add reward asymmetry because improvements to the maintained non-dominated structure are intentionally weighted more strongly than repeated or dominated moves. This is a search heuristic, not a symmetry theorem, and its scale robustness is not established by the archived SCH experiment.
3.5. Full Modular Workflow and Cell-Archive Interaction
The architecture is modular rather than a mandatory serial solver. The supervised route maps sparse sampled objective vectors to a reconstructed occupancy image and then to a sampled boundary approximation. The rule-based route accepts candidate vectors from any sampler, evolutionary procedure, or reinforcement-learning agent and applies the fixed extractor directly. MOP-DQN and MOP-AC maintain non-dominated sets internally and do not require a CNN connection, whereas MOP-AC-sample accumulates an occupancy image and invokes the fixed extractor at termination.
During rasterization, each objective vector y = f(x) is assigned to a cell (i, j). The occupancy bit is set to one, while the tuple containing the sample index, x, and y is appended to the corresponding cell archive A
ij. The CNN branches read only the occupancy image. The deterministic branch uses the occupancy map for fixed cross-correlation and then retrieves A
ij from retained cells for exact vector-level dominance refinement.
Table 9 summarizes the complete data flow, archive construction, and admissible invocation relationships.
Throughout this manuscript, “objective-space occupancy image” denotes the rasterized sampled vectors, “sampled PF image” denotes the CNN boundary approximation, “cell-level boundary” denotes the fixed-kernel output, and “refined non-dominated vector set” denotes the archive-checked result. These terms are used consistently to prevent image approximation from being confused with exact vector-level dominance.
4. Experimental Setup and Evaluation Protocol
4.1. Benchmark Problems and Datasets
The supervised CNN and deterministic extractor are evaluated on two-objective synthetic benchmarks. These functions provide controlled reference geometries for examining front-image approximation, disconnected and nonlinear boundaries, and cell-level extraction. The benchmark role is methodological: the results characterize the proposed representation and operators rather than a market-specific pricing model.
The non-ZDT benchmark families are summarized in
Table 10. SCH has a continuous convex front, FON has a diagonal Pareto set, POL contains trigonometric components, and KUR produces a nonlinear and non-convex front. These cases test simple, symmetric, and irregular objective-space geometries.
The ZDT-series benchmark definitions and their corresponding geometric characteristics are summarized in
Table 11. ZDT1, ZDT2, and ZDT3 are generated through a unified parameterized form; ZDT3 provides a disconnected front. ZDT4 introduces multimodality, and ZDT6 has a nonlinear and unevenly distributed front.
The supervised labels are finite sampled approximations. For every problem instance, a sparse sample forms the input occupancy image and a substantially larger sample forms the high-density target occupancy image. The sampled reference front is extracted from the available objective vectors. We therefore refer to each target as a “high-density sampled objective-space approximation” rather than as a complete or exact image of the continuous objective space.
4.2. Training and Inference Protocol for the Supervised CNN Solver
The supervised solver was evaluated as an image-to-image mapping system. A pixel value of one indicates that at least one sampled objective vector falls into the target-space bin; no regression target is assigned to individual decision vectors.
Inference was performed on SCH, FON, POL, KUR, ZDT1, ZDT2, ZDT3, ZDT4, and ZDT6. The CNN output was evaluated against theoretical or high-density sampled reference fronts using Υ and, when applicable, Δ. These metrics characterize the archived front approximation and are not used to rank the method against independently tuned optimizers.
4.3. Evaluation Metrics for PF Approximation
Two indicators were used. The proximity indicator
Υ is the mean Euclidean distance from each obtained point to the nearest reference point; a smaller value indicates a closer front:
For connected fronts, the archived distribution indicator
Δ measures adjacent-point spacing and endpoint deviations as defined in Equation (16).
The proximity indicator Υ remains the mean nearest-reference distance. The endpoint-based Δ in Equation (16) is interpreted only for a connected front with two well-defined terminal points. ZDT3 contains five disconnected segments, so a single pair of global endpoint penalties does not represent its spacing.
4.4. Experimental Protocol for Convolutional Non-Dominated Sorting
The deterministic extractor was evaluated on 127 × 127, 512 × 512, 1024 × 1024, and 2048 × 2048 binary images, corresponding to 16,129, 262,144, 1,048,576, and 4,194,304 two-dimensional vectors when each pixel is occupied.
The archived implementation used u = 15 for all reported resolutions. This value was not selected by fitting benchmark outcomes. It is an engineering compromise: increasing u reduces T but enlarges the (2u − 1)2 kernel footprint and temporary convolution workload, whereas decreasing u produces a smaller kernel but more passes. Because no parameter-sensitivity runs were retained, the revision does not claim that u = 15 is optimal; it is reported as the fixed setting used for the archived runtime measurements. The automatic verification routine performs two exhaustive checks on the finite occupied set. First, for each retained cell, it searches the occupied grid for any cell satisfying Equation (7); a match is recorded as a false retention. Second, for each removed occupied cell, it confirms the existence of at least one dominating occupied cell; failure is recorded as a false removal. The extracted binary boundary is accepted only when both counts are zero and it equals the direct cell-level non-dominated set. The cell archive then applies the original continuous-valued dominance relation within retained boundary archives to remove collision-induced ambiguities.
4.5. Experimental Protocol for MOP-DQN, MOP-AC, and MOP-AC-Sample
The reinforcement-learning experiments are restricted to SCH, a one-variable benchmark with a known convex front. Their purpose is to examine whether the proposed state, action, and dominance-reward definitions can generate a recoverable candidate set. The protocol does not support conclusions for nonlinear, multimodal, disconnected, constrained, or high-dimensional problems.
MOP-DQN uses the network and settings in Equation (11). The decision range is [−30, 30]. Replay learning and an exploration schedule are used to stabilize value estimation.
MOP-AC uses the reward in Equation (12), the settings in Equation (13), and a decision range of [−1000, 1000]. State and action values are represented by 20-bit binary coding.
MOP-AC-sample uses distribution-parameter states and samples ten decision vectors at each step. Its training horizon is T = 20,000 and γ = 0.8, as shown in Equation (14). The final accumulated image is processed by the convolutional PF extractor.
The evaluation is organized by evidence type. The supervised CNN is assessed against theoretical or high-density sampled reference fronts. The fixed extractor is checked against the direct cell-level dominance relation and then refined through the original-vector archive. The reinforcement-learning variants are discussed only as SCH feasibility cases. This separation prevents modules with different computational roles from being treated as one competitive benchmark.