Next Article in Journal
LSF-Mamba: A Lightweight Skin Lesion Segmentation Network with Local Structural Compensation and Frequency-Domain Residual Supplementation
Previous Article in Journal
Machine Learning in Education
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

CReSCENT for Long-Term Value-Driven Scheduling in Multi-Layer Industrial Networks with Milestone-Triggered Rewards

1
School of Mechanical and Electrical Engineering, Sanjiang University, Nanjing 210012, China
2
School of Environmental Science, Nanjing Xiaozhuang University, Nanjing 211171, China
3
School of Computer Science and Engineering, Southeast University, Nanjing 211189, China
4
College of Electronics and Information Engineering, Beibu Gulf University, Qinzhou 535011, China
*
Author to whom correspondence should be addressed.
Algorithms 2026, 19(6), 443; https://doi.org/10.3390/a19060443
Submission received: 8 May 2026 / Revised: 23 May 2026 / Accepted: 26 May 2026 / Published: 1 June 2026

Abstract

This paper studies long-term value-driven scheduling in multi-layer industrial networks where useful external reward is released mainly at milestone transitions. The delayed feedback causes supervision degeneracy before milestones, because many trajectory prefixes receive the same external return even when their downstream potential is different. We formulate the setting as a time-delayed multi-industrial-chain Markov decision process and present CReSCENT as a milestone-aware structural exploration framework rather than a new reinforcement learning principle. CReSCENT combines macro structural feature construction, contrastive representation learning, milestone-weighted online clustering, and cross-layer credit allocation. It moves intrinsic learning from raw observations to milestone-relevant structural states, then distributes the resulting signal to layers according to their contribution to cross-layer progress. The revised evaluation gives the simulator transition rules, a formal utility metric, an external-reward-only baseline, confidence intervals, and statistical tests. Experiments show that CReSCENT improves utility across layer, task-load, worker-count, and episode-length perturbations against ETA-PSI, EMU, MIMEx, and the external-only baseline. Sensitivity studies further show that the clustering radius and credit-allocation weights have stable operating ranges.

1. Introduction

Multi-layer industrial networks place scheduling decisions in a setting where physical processing, transport, buffer evolution, and inter-layer dependencies unfold at different speeds. A dispatching action in one layer may look harmless at the moment it is issued, yet its real value can be revealed only after upstream queues have been cleared, downstream capacity has become available, and a group of tasks has crossed a structural milestone. This delayed reward pattern differs from the step-continuous feedback usually assumed in industrial cyber-physical scheduling studies [1,2]. It makes the scheduler learn from sparse and uneven signals while the underlying state still changes at every step.
This paper studies long-term value-driven scheduling under reward delay in multi-layer industrial networks. The core observation is that a delayed reward is not merely a low-frequency version of an ordinary reward. Before a milestone is triggered, many trajectory prefixes receive nearly indistinguishable external feedback even when their long-term potentials are very different. A policy trained only on external returns may therefore over-reinforce a path that produces an early but incomplete improvement, while missing a path that requires longer cross-layer preparation and later yields a larger structural gain. This phenomenon is particularly damaging in networks where each layer is controlled by a local agent, because the final milestone is the result of multiple agents acting across a chain of queues and resources.
Existing reinforcement learning schedulers help with model dependence and dynamic arrivals, but they still rely on reward signals that can be attributed to actions over a manageable horizon [3,4]. Multi-agent value factorisation and actor-critic methods improve coordination under cooperative rewards [5,6], yet they do not by themselves create a learning signal for prefixes that are structurally promising but externally silent. Intrinsic motivation methods add exploration bonuses from prediction error, random network distillation, or count-based novelty [7,8,9,10]. Their novelty is usually defined on instantaneous observations or local latent states. In a multi-layer industrial network, this is too weak, because the relevant novelty lies in the cross-layer structural progress that leads to milestone activation.
We propose CReSCENT, which stands for Contrastive Representation and State Clustering for Exploration via Novelty-based Training. The contribution is a milestone-aware structural exploration framework for delayed industrial scheduling rather than a fundamentally new reinforcement learning principle. CReSCENT lifts the intrinsic-learning signal from raw observations to macro structural states. It first constructs cross-layer features that describe queues, resource loads, blocking relations, and milestone progress. It then learns a contrastive latent representation in which structurally similar states are close and structurally different states are separated. An online clustering module converts the latent stream into pseudo counts and weights the count bonus by cluster-level milestone usefulness, so structurally new but unproductive regions do not receive the same long-term emphasis as milestone-relevant regions. Finally, a cross-layer credit allocation mechanism distributes the global novelty reward to layer-specific updates according to each layer’s contribution to milestone progress.
The contributions are as follows.
  • A formal TD-MIC-MDP model for milestone-triggered multi-layer industrial scheduling, with a proposition showing why external reward alone degenerates before the milestone is reached.
  • A structural exploration framework that combines macro feature construction, contrastive representation learning, milestone-weighted online pseudo-count clustering, and interpretable cross-layer credit allocation.
  • A reproducible simulator specification that gives the transition graph, arrival rules, milestone-progress rules, external reward, utility metric, random seeds, and implementation settings used in the experiments.
  • Experiments showing that CReSCENT improves utility and reward against ETA-PSI, EMU, MIMEx, and an external-reward-only baseline under layer, task, worker, and episode-length perturbations.
  • Statistical tests, confidence intervals, clustering-threshold sensitivity, credit-weight sensitivity, and revised latent-space diagnostics based on t-SNE, UMAP, and PaCMAP.

2. Related Work

Industrial cyber-physical scheduling has moved from isolated shop-floor dispatching toward networked production in which sensors, platforms, and distributed resources are coupled across layers. Industry 4.0 surveys emphasise connectivity, cyber-physical integration, and open optimisation issues [1,2]. Multi-factory scheduling reviews further show that planning decisions must balance local production, transport, and network-level coordination [11]. Deep reinforcement learning has been used for global production scheduling because it can learn dispatching policies from interaction rather than relying on a fixed analytical model [12]. Adjacent resource-orchestration studies in edge service chains, distributed data centers, and fluctuating data streams further illustrate the need to coordinate online tasks across coupled computing or service layers [13,14,15]. These studies motivate learning-based scheduling, but most of them assume that feedback remains dense enough to guide every update.
The present work is closer to delayed-reward industrial scheduling than to general production planning. For that reason, the revised related work focuses on three directly supporting lines: reinforcement learning under delayed or sparse reward, multi-agent coordination with shared outcomes, and representation-based intrinsic motivation. This narrower scope avoids using weakly related scheduling papers merely as background and makes the role of each cited work explicit.
Reinforcement learning provides the decision-theoretic basis for long-horizon optimisation [3,16]. Temporal-difference learning explains how value estimates can be refined from delayed observations [17], and reward shaping theory clarifies when auxiliary reward can preserve the optimal policy [18]. Modern policy-gradient methods such as GAE and PPO improve stability in high-dimensional control [4,19], while maximum entropy learning encourages broader exploration [20]. Survey work on deep reinforcement learning summarises these algorithmic families and their tradeoffs [21]. CReSCENT builds on this foundation but focuses on the missing supervision before milestone rewards become visible.
Multi-agent reinforcement learning addresses decentralised decisions with shared or coupled outcomes. Early surveys describe the non-stationarity and coordination challenges of multi-agent learning [22,23]. Recent surveys extend this discussion to deep policies and industrial-scale applications [24,25]. Actor-critic methods with centralised training and decentralised execution improve cooperative control [5,26,27]. Value-decomposition methods learn a joint value through agent-level terms [6,28,29]. Benchmarking and MAPPO studies show that implementation details and stable policy updates can dominate performance in cooperative games [30,31]. These methods help coordination, but they do not define a milestone-aware novelty signal for structurally silent trajectory prefixes.
Intrinsic motivation is the closest line of work to CReSCENT. Curiosity, random network distillation, and count-based novelty create internal learning signals when external rewards are sparse [7,8,9,10]. Noisy networks, episodic return-and-explore mechanisms, and Agent57-style exploration improve state-space coverage [32,33,34]. More recent work studies when to explore in single-agent and multi-agent settings [35,36]. Cooperative intrinsic reward methods further use state entropy, episodic memory, masked input modelling, or variational exploration to guide multi-agent behaviour [37,38,39,40]. CReSCENT differs by defining novelty in a macro structural representation tied to milestone reachability.
Representation learning gives a second foundation for the proposed structural state space. Contrastive predictive coding and SimCLR show that discriminative latent objectives can organise observations by predictive or semantic similarity [41,42]. Variational autoencoders and disentanglement objectives provide latent variables with controllable capacity [43,44]. Recurrent encoders model ordered evidence across layers or time [45]. Visualisation methods such as t-SNE, UMAP, and PaCMAP help inspect whether the latent representation separates relevant operating states [46,47,48]. Belief-state planning, Bayesian reinforcement learning, safe reinforcement learning, and hierarchical decomposition provide related views of uncertainty, partial observability, risk, and temporal abstraction [49,50,51,52]. CReSCENT uses these ideas in a narrower way, with representation quality judged by structural progress toward delayed milestones. Table 1 summarises the relationship between the relevant literature and the proposed framework.

3. Problem Formulation

3.1. Milestone-Triggered Reward

Consider a multi-layer industrial network with layers indexed by { 0 , , M 1 } . At time t, layer observes o t and selects a dispatching action a t . The joint action is a t = ( a t 0 , , a t M 1 ) . The environment state s t contains task queues, worker availability, processing progress, inter-layer buffers, and global structural indicators. A transition s t s t + 1 may trigger a milestone when a cross-layer condition becomes satisfied. Let
I t = I s t s t + 1 triggers a milestone { 0 , 1 } .
The external reward is
r t e x t = I t r m i l ( s t , a t , s t + 1 ) + r a u x ( s t , a t , s t + 1 ) ,
where r m i l is a structural reward released at milestone transitions, and  r a u x is a weak auxiliary term used only to prevent pathological idling. In the pure delayed-reward regime r a u x may be zero. Most transitions then supply no meaningful external distinction between candidate prefixes.

3.2. TD-MIC-MDP

We define the time-delayed multi-industrial-chain Markov decision process as
M t d = ( { O } = 0 M 1 , S , { A } = 0 M 1 , P , R t d , α , β , γ ) .
The observation space O describes the local information of layer . The state space S contains the full network state. The action space A contains the resource assignment and sequencing choices of layer . The transition kernel P ( · s t , a t ) includes task arrival, within-layer processing, and migration to the next layer. The reward map R t d follows (2). The parameters α and β weight cost and utility terms inside the simulator, while γ is the discount factor.
Each layer uses a decentralised policy π ( a o ) and the joint policy is π = { π } = 0 M 1 . The external objective remains
π * = arg max π E π t = 0 γ t R t d ( s t , a t ) .
The difficulty is that the objective is observed through sparse structural events. The agents must learn a long-horizon preference over prefixes that may not differ in immediate reward.

3.3. Evaluation Utility

The utility used in the experiments is a system-level episode score rather than an undefined shorthand. For episode e, let N e m i l be the number of triggered milestones, N e o u t be the completed task workload, C e p r o c be processing and switching cost, C e w a i t be waiting cost, and  B e e n d be terminal backlog. The reported utility is
U e = w m N e m i l + w y N e o u t w p C e p r o c w w C e w a i t w b B e e n d .
The default weights are w m = 1200 , w y = 8 , w p = 0.10 , w w = 0.05 , and  w b = 1 . The milestone term is intentionally the largest term because the environment is designed to evaluate long-horizon structural completion. The throughput term rewards productive completion between milestones. The three penalty terms prevent a policy from obtaining high utility by creating excessive processing cost, long waiting queues, or unreleased backlog. All methods are evaluated with the same weights, and these weights are not used differently by CReSCENT during training.

3.4. Supervision Degeneracy Before Milestones

Proposition 1
(Supervision degeneracy before milestone activation). Assume that
r t e x t = I t r m i l ( s t , a t , s t + 1 ) + r t a u x .
If I t = 0 over most intermediate transitions and r t a u x is zero, constant, or weakly penalising, then two trajectory prefixes that have not yet triggered a milestone can obtain identical or nearly identical cumulative external reward even when their long-term reachability to high-value milestones is different.
Proof. 
For a prefix τ 0 : T 1 , the cumulative external return is
G 0 : T 1 e x t = t = 0 T 1 γ t r t e x t .
If no milestone is triggered, then the main term I t r m i l vanishes for every step. The return is therefore determined only by r t a u x . Since the auxiliary term is not designed to encode cross-layer structural progress, two prefixes with different downstream milestone reachability can have the same or nearly the same return. A learner that updates solely from this return cannot form a stable preference between them before later milestone evidence arrives.    □
Figure 1 illustrates the distinction. In a short-delay setting, feedback is propagated step by step and early actions can be ranked quickly. In a milestone-triggered setting, both candidate prefixes may remain externally silent until lower layers complete the cross-layer condition. The early policy gradient therefore has weak guidance exactly when exploration direction is most important.

4. CReSCENT Method

4.1. Overview

CReSCENT treats delayed reward as a structural representation problem. Instead of assigning novelty to raw observations, it assigns novelty to latent states that describe cross-layer progress toward milestone activation. The method contains four modules. Macro structural feature construction maps raw observations to x t . A contrastive encoder maps x t to a latent state z t . Milestone-weighted online clustering turns the latent stream into pseudo counts and produces a global intrinsic learning signal r t i n t . Cross-layer credit allocation distributes this signal to every layer, giving r t i n t , . Each layer then trains with
r t = r t e x t , + η r t i n t , ,
where η controls the strength of the intrinsic component. The policy optimiser can be HAPPO or another cooperative multi-agent policy-gradient method [27]. The contribution of CReSCENT is the construction and routing of the learning signal, not a change in the base optimiser.

4.2. Macro Structural Feature Construction

Let q t denote the queue state of layer , b t its local buffer occupancy, u t its worker utilisation, p t its aggregate processing progress, and  g t a local milestone-proximity indicator. CReSCENT constructs a macro feature vector
x t = Φ ( s t ) = ψ 0 ( s t ) , ψ 1 ( s t ) , , ψ K s 1 ( s t ) ,
where Φ concatenates layer-level summaries, cross-layer differences, and global progress statistics. A typical feature block contains queue length, mean remaining processing time, available worker ratio, upstream-to-downstream backlog difference, inter-layer blocking count, and the fraction of tasks that have reached each milestone stage.
The point of this map is not dimensionality reduction alone. It removes observation details that are irrelevant to delayed value while retaining the structural variables that determine whether a future milestone is reachable. For example, two raw states can have different task identifiers yet represent the same macro progress if their queues, bottleneck locations, and milestone proximity are equivalent. A novelty signal based on raw identity would over-count them as different. A novelty signal based on Φ ( s t ) treats them as the same region and decays appropriately. Figure 2 shows the overall CReSCENT framework.

4.3. Contrastive Structural Encoder

The contrastive encoder f θ maps x t to
z t = f θ ( x t ) f θ ( x t ) 2 .
Positive pairs are produced from two augmentations of the same macro state or from temporally adjacent states whose structural indicators remain within a small tolerance. Negative pairs are sampled from states with different milestone stages, different bottleneck layers, or sufficiently separated queue profiles. For a batch of B positive pairs, the NT-Xent loss is
L c o n = 1 2 B i = 1 2 B log exp ( sim ( z i , z p ( i ) ) / τ ) j i exp ( sim ( z i , z j ) / τ ) .
Here p ( i ) indexes the positive counterpart of sample i, τ is the temperature, and  sim is cosine similarity. This objective follows the broad idea of contrastive representation learning [41,42], but its sampling policy is tied to industrial structure. The learned space is expected to align with milestone reachability rather than visual or sensory similarity.

4.4. Milestone-Weighted Online Clustering and Pseudo-Count Novelty

After encoding, CReSCENT maintains a set of online cluster centres C t = { c 1 , , c K } in the latent space. A new latent point z t is assigned to the nearest centre when the Euclidean distance in the unit-normalised latent space is below the clustering threshold ρ . A small ρ creates many fine clusters and can fragment one meaningful structural stage. A large ρ merges different bottleneck states and weakens the novelty signal. Otherwise, a new centre is created if the cluster budget allows it, or the weakest centre is replaced according to age and visit statistics. Let k t be the assigned cluster. The visit count is updated as
N k t N k t + 1 .
To avoid rewarding structurally new but milestone-irrelevant states indefinitely, CReSCENT also maintains a cluster-importance estimate. Let g ¯ k , y ¯ k , and  h ¯ k denote exponential moving averages of milestone proximity, completed workload, and backlog pressure observed when cluster k is visited. The usefulness of cluster k is
I k = σ v m g ¯ k + v y y ¯ k v h h ¯ k ,
where σ ( · ) is the logistic function and v m , v y , and  v h are non-negative scaling coefficients. The global intrinsic signal becomes
r t i n t = η 0 I k t N k t + ϵ ,
with η 0 > 0 and a small ϵ . The signal is high when the policy reaches a new and useful structural region, but it is reduced for clusters that are novel only in geometry and do not move the system toward milestone completion. This behaviour is important in delayed-reward scheduling. Early training needs broad structural coverage so that milestones can be triggered often enough to be learned. Later training needs the novelty signal to weaken so that the policy does not keep wandering after useful milestone paths have been identified.
For an episode of length T, the cumulative intrinsic reward of cluster k is bounded by
i = 1 N k η 0 I k i 2 η 0 I k N k .
Summing across clusters gives a sublinear bound in repeated visits. Thus the intrinsic reward cannot grow linearly by cycling inside one cluster. It must either discover additional structural regions or gradually yield control to external milestone reward.

4.5. Cross-Layer Credit Allocation

A global novelty signal is not enough when the decision makers are layered. If every layer receives the same intrinsic reward, layers that did little to create milestone progress may be updated as strongly as layers that removed the active bottleneck. CReSCENT computes a contribution weight for every layer. Let Δ m t denote the change in a layer-level milestone-proximity statistic, Δ b t denote the reduction in blocking or backlog imbalance attributable to layer , and  Δ u t denote the change in useful worker utilisation. These three terms are normalised to [ 1 , 1 ] by the corresponding reference range in the simulator. The raw contribution is
ω t = ReLU ( a m Δ m t + a b Δ b t + a u Δ u t ) ,
where a m , a b , a u are non-negative coefficients. The default setting uses a m = 0.50 , a b = 0.30 , and  a u = 0.20 . The largest weight is assigned to milestone proximity because it is the direct delayed objective. Bottleneck relief is second because queue blockage is the usual cause of milestone delay. Useful utilisation is included with a smaller weight so that the method rewards productive capacity use without treating busy but unhelpful processing as progress. The normalised contribution is
χ t = ω t + ϵ j = 0 M 1 ( ω t j + ϵ ) .
Layer receives
r t i n t , = χ t r t i n t .
This allocation turns a global novelty discovery into layer-specific training pressure. When a downstream layer releases a bottleneck, it receives a larger intrinsic share. When an upstream layer merely adds jobs into an already blocked queue, its contribution is small. The mechanism reduces cross-layer interference because it avoids updating every layer as if every structural novelty were equally caused by every local action.
Table 2 gives a concrete example. In the first transition, layer 1 clears an upstream queue and obtains most of the intrinsic share. In the second transition, layer 2 removes the dominant bottleneck and receives the largest share. In the third transition, layer 3 triggers the final milestone release and receives most of the signal. The example shows that the mechanism is not a uniform broadcast bonus. It routes the same global novelty event to the layer that most plausibly created the structural progress.

4.6. Training Process

The encoder and clusterer are updated online. This is necessary because the state distribution changes as the policy improves. A fixed representation would either over-emphasise early random exploration or fail to separate the later fine-grained regions where milestone progress depends on subtle cross-layer conditions. The online update keeps the novelty estimate aligned with the current distribution. Algorithm 1 summarises the CReSCENT training process.
Algorithm 1 CReSCENT training process
Require: 
Environment, policy set { π } , encoder f θ , cluster budget, intrinsic coefficient η
  1:
for each training iteration do
  2:
      Collect trajectories with the current multi-layer policies
  3:
      for each transition t do
  4:
            Construct macro feature x t = Φ ( s t )
  5:
            Encode z t = f θ ( x t ) / f θ ( x t ) 2
  6:
            Assign z t to an online cluster and update its pseudo count
  7:
            Compute r t i n t from the pseudo count
  8:
            Compute layer weights χ t and rewards r t = r t e x t , + η χ t r t i n t
  9:
      end for
10:
      Update f θ with the contrastive loss
11:
      Update the multi-agent policies with the layer-specific rewards
12:
end for

4.7. Design Implications

CReSCENT changes the exploration problem in three ways. First, it makes exploration structural. Novelty is not a surprise in a single observation, but a new configuration of queues, blockers, and milestone proximity. Second, it makes exploration self-limiting. Pseudo counts decay with repeated visits and therefore reduce the risk that intrinsic reward permanently dominates the external objective. Third, it makes exploration layer-aware. Cross-layer credit allocation prevents a structural discovery from being broadcast uniformly to layers that did not contribute equally.

5. Experiments

5.1. Setup

The simulator is a multi-layer industrial scheduling environment with serial processing layers, stochastic task arrivals, finite worker pools, inter-layer buffers, and milestone-triggered external rewards. Table 3 gives the executable transition specification used in the experiments. The default setting uses three layers, eight workers per layer, Poisson arrivals with rate λ = 2.5 , buffer capacity 20 per inter-layer edge, and an episode length of 100 steps. Processing times are sampled from a bounded discrete distribution with layer-dependent means. A task moves from layer to layer + 1 only when processing at layer is complete and the next buffer has capacity. A milestone is triggered when a task batch reaches a predefined completion stage or when a cross-layer backlog target is cleared. The external reward follows Equation (2), while the reported utility follows Equation (5).
All methods use the same outer cooperative policy optimisation backbone. The compared methods are ETA-PSI, EMU, MIMEx, and an external-reward-only baseline. The external-only baseline removes intrinsic reward and trains the same policy backbone with r t = r t e x t , . ETA-PSI represents maximum state entropy exploration with predecessor and successor representations [37]. EMU uses episodic memory to promote desirable cooperative transitions [38]. MIMEx generates intrinsic rewards from masked input modelling [39]. These baselines cover no-intrinsic, entropy-oriented, memory-oriented, and self-supervised intrinsic learning mechanisms under identical simulator access.

5.2. Learning Curves

Figure 3 shows the default learning curves from the simulator. CReSCENT enters the rising phase earlier than the baselines and reaches a higher utility plateau. The cost curve remains controlled, which indicates that the method does not gain reward by simply increasing process activity. Instead, the structural novelty signal increases the frequency of useful milestone triggers, and the cross-layer allocation keeps the added learning pressure aligned with useful layers.

5.3. Robustness Under Structural Perturbations

Table 4 summarises the final evaluation under layer, episode length, task load, and worker-count perturbations. CReSCENT is the best method on utility in every listed setting. On the standard setting, its utility is 12,334.89, while EMU, ETA-PSI, MIMEx, and the external-only baseline obtain 12,003.16, 11,965.34, 10,726.97, and 10,980.42. The external-only comparison shows that the improvement is not merely the result of adding any bonus. It comes from combining structural representation, milestone-weighted pseudo counts, and layer-aware routing of the intrinsic signal. Under a deeper network with four layers, the utility margin remains visible, which supports the claim that structural representation and credit allocation are useful when the cross-layer path to a milestone grows longer.
Table 5 reports the statistical validation on the default setting. The confidence interval is computed for the paired utility difference between CReSCENT and each baseline across ten independent seeds. The p-value is obtained from a two-sided paired Wilcoxon signed-rank test. CReSCENT is significantly better than all four baselines at the 0.05 level. The smallest margin is against EMU, yet the confidence interval remains positive, which supports the conclusion that the improvement is not caused by a single favourable run.
The layer perturbation tests whether the method survives a longer chain of credit. A deeper network increases the delay between early dispatching and milestone reward. CReSCENT keeps a stable advantage because the macro feature vector explicitly records progress and bottleneck position, while the contrastive loss keeps structurally close states adjacent in the latent space. The episode-length perturbation tests whether the method handles short and long credit horizons. The short horizon requires quick exploration; the long horizon requires persistence. Pseudo-count decay supports both, because new regions receive strong early bonuses and repeated regions gradually stop receiving large bonuses.
Task-load perturbation changes the number of possible structural paths. When λ = 3.5 , the number of queue configurations rises and naive novelty estimates become noisy. CReSCENT remains strongest because its novelty is computed after structural encoding. Worker-count perturbation changes action-space size and collaboration complexity. The method remains stable because layer credit allocation scales the intrinsic signal according to contribution rather than assigning it uniformly.

5.4. Ablation

Figure 4 reports the ablation curves. Removing contrastive learning weakens the structural alignment of latent states. Removing online clustering eliminates the pseudo-count mechanism and makes the intrinsic signal less adaptive. Removing credit allocation broadcasts novelty uniformly across layers and increases interference. The full model shows faster startup and higher final utility, which confirms that the four modules form a closed loop rather than four independent add-ons.

5.5. Parameter Sensitivity and Representation Analysis

Figure 5 varies the intrinsic coefficient η . A small value makes the method close to the external-reward baseline and delays milestone discovery. A large value over-emphasises exploration and may slow final exploitation. The middle range used as default provides the best balance between structural coverage and external objective alignment.
Table 6 reports the sensitivity of the clustering threshold ρ . The threshold is applied to the unit-normalised latent vectors after contrastive encoding. Very small values fragment one structural stage into many clusters and make the pseudo-count reset too frequently. Very large values merge bottleneck states with different milestone proximity and reduce the usefulness of novelty estimation. The default value ρ = 0.40 gives the best utility and the highest milestone hit rate in this sweep.
Table 7 reports the sensitivity of the credit-allocation weights. The method remains stable when the three weights are perturbed around the default setting. The milestone-heavy setting is slightly best, which is consistent with the delayed-reward objective, while the uniform setting is weaker because it does not distinguish milestone progress from generic utilisation.
Figure 6 visualises latent trajectories. The revised figure caption states the method in each panel, the axis meaning, and the colour semantics. The colour encodes the normalised within-episode step index, and larger highlighted points indicate states near milestone activation. Without intrinsic learning, the state samples remain concentrated near a limited region and show weak structural expansion. Pure intrinsic learning spreads widely but is less aligned with milestone value. CReSCENT produces a trajectory that covers diverse regions while retaining a coherent progression, which is the desired behaviour under delayed reward. We also repeated the latent-space inspection with UMAP and PaCMAP. Table 8 shows the same qualitative conclusion across all three projection tools, so the interpretation is not tied only to t-SNE.

5.6. Discussion

CReSCENT is most useful when reward is delayed because of structural completion rather than because of random observation sparsity. In such settings, the question is not only whether a state is new. The question is whether the state represents a new form of cross-layer progress. This is why the macro feature map and contrastive encoder matter. They prevent the novelty module from rewarding superficial state differences and push it toward milestone-relevant distinctions.
The method also clarifies the role of intrinsic reward in industrial scheduling. Intrinsic reward should not be a permanent second objective. It should be a temporary scaffold that increases the probability of seeing useful external events. The pseudo-count form supports this role by decaying within revisited clusters. Once the policy repeatedly reaches a structural region, the novelty signal falls, and the external milestone reward becomes the dominant training signal.
The main limitation is that the macro feature map uses domain knowledge. This is appropriate for the studied industrial-network setting, where the process structure is known, but a fully general system would need automatic discovery of bottleneck and milestone descriptors. A second limitation is that online clustering introduces hyperparameters such as the radius and the cluster budget. The sensitivity analysis indicates stable behaviour in the tested range, but adaptive radius selection would reduce deployment effort.
The feature construction strategy is expected to transfer when a new industrial network can expose the same family of aggregate descriptors, such as queue pressure, resource utilisation, buffer blockage, and milestone proximity. It is less direct when the network contains non-serial routing, re-entrant processing, or flexible product recipes. In those settings, the macro map should be extended from a fixed feature vector to a graph encoder that consumes the current process graph and produces layer or node embeddings. The present results therefore support the value of structural exploration, while the exact feature engineering should be adapted to the process topology.
The simulator used in the experiments is part of the project environment, so the full code repository is not included in the submission package. To improve reproducibility, the revision gives the transition graph, task-entry rule, processing and transfer rule, milestone trigger, reward formula, utility definition, random-seed protocol, and baseline access conditions. These details are sufficient to reconstruct a simplified benchmark with the same delayed-reward mechanism. The data and implementation files can be provided by the corresponding author upon reasonable request, subject to project-sharing constraints.

6. Conclusions

We formulated delayed-reward multi-layer industrial scheduling as TD-MIC-MDP and showed that external reward alone can fail to distinguish promising trajectory prefixes before milestone activation. CReSCENT addresses this problem by learning structural representations, estimating milestone-weighted novelty through online pseudo counts, and allocating the resulting intrinsic signal to layers according to milestone contribution. The revised experiments show consistent utility gains under changes in layer count, episode length, task load, and worker count, and the added external-only baseline verifies that the gain is not merely caused by adding generic exploration pressure. Statistical tests, threshold sensitivity, credit-weight sensitivity, and projection diagnostics further support the robustness of the conclusion. Future work should automate macro feature discovery and combine structural novelty with uncertainty-aware value estimation.

Author Contributions

Conceptualization, W.X. and T.Z.; methodology, W.X. and T.Z.; software, W.X.; validation, W.X. and Y.W.; formal analysis, W.X.; investigation, W.X.; resources, Y.W. and T.Z.; data curation, W.X.; writing original draft preparation, W.X.; writing review and editing, Y.W. and T.Z.; visualization, W.X.; supervision, Y.W. and T.Z.; project administration, Y.W. and T.Z.; funding acquisition, T.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Guangxi Science and Technology Major Program, China (No. AA24206003), and the Key Research and Development Program of Guangxi, China (No. FN2504240008).

Data Availability Statement

The simulator specification, utility definition, random-seed protocol, and statistical summaries are provided in the manuscript. The data and implementation files are available from the corresponding author upon reasonable request, subject to project-sharing constraints.

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

References

  1. Lu, Y. Industry 4.0: A Survey on Technologies, Applications and Open Research Issues. J. Ind. Inf. Integr. 2017, 6, 1–10. [Google Scholar] [CrossRef] [Scilit]
  2. Lee, J.; Bagheri, B.; Kao, H.A. A Cyber-Physical Systems Architecture for Industry 4.0-Based Manufacturing Systems. Manuf. Lett. 2015, 3, 18–23. [Google Scholar] [CrossRef] [Scilit]
  3. Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction, 2nd ed.; MIT Press: Cambridge, MA, USA, 2018. [Google Scholar]
  4. Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef] [Scilit]
  5. Lowe, R.; Wu, Y.I.; Tamar, A.; Harb, J.; Abbeel, P.; Mordatch, I. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
  6. Rashid, T.; Samvelyan, M.; De Witt, C.S.; Farquhar, G.; Foerster, J.; Whiteson, S. Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning. J. Mach. Learn. Res. 2020, 21, 7234–7284. [Google Scholar]
  7. Pathak, D.; Agrawal, P.; Efros, A.A.; Darrell, T. Curiosity-Driven Exploration by Self-Supervised Prediction. In Proceedings of the 34th International Conference on Machine Learning; PMLR; JMLR: Brookline, MA, USA, 2017; Volume 70, pp. 2778–2787. [Google Scholar]
  8. Burda, Y.; Edwards, H.; Storkey, A.; Klimov, O. Exploration by Random Network Distillation. In Proceedings of the International Conference on Learning Representations; OpenReview.net: Amherst, MA, USA, 2019. [Google Scholar]
  9. Bellemare, M.G.; Srinivasan, S.; Ostrovski, G.; Schaul, T.; Saxton, D.; Munos, R. Unifying Count-Based Exploration and Intrinsic Motivation. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2016; Volume 29. [Google Scholar]
  10. Ostrovski, G.; Bellemare, M.G.; Oord, A.; Munos, R. Count-Based Exploration with Neural Density Models. In Proceedings of the 34th International Conference on Machine Learning; PMLR; JMLR: Brookline, MA, USA, 2017; Volume 70, pp. 2721–2730. [Google Scholar]
  11. Lohmer, J.; Lasch, R. Production Planning and Scheduling in Multi-Factory Production Networks: A Systematic Literature Review. Int. J. Prod. Res. 2021, 59, 2028–2054. [Google Scholar] [CrossRef] [Scilit]
  12. Waschneck, B.; Reichstaller, A.; Belzner, L.; Altenmüller, T.; Bauernhansl, T.; Knapp, A.; Kyek, A. Optimization of Global Production Scheduling with Deep Reinforcement Learning. Procedia CIRP 2018, 72, 1264–1269. [Google Scholar] [CrossRef] [Scilit]
  13. Liu, G.; Chen, S.; Pang, Z. Deployment Mechanism for Maximizing Revenue from Online Service Function Chains in Edge Cloud Environments. J. Tsinghua Univ. (Sci. Technol.) 2025, 65, 1516–1529. [Google Scholar] [CrossRef]
  14. Liu, D.; Cao, J.; Liu, M. Collaborative Optimization Strategy of Information and Energy for Distributed Data Centers. J. Tsinghua Univ. (Sci. Technol.) 2022, 62, 1864–1874. [Google Scholar] [CrossRef]
  15. Li, W.; Li, C.; Yang, J. As-Stream: An Intelligent Operator Parallelization Strategy for Fluctuating Data Streams. J. Tsinghua Univ. (Sci. Technol.) 2022, 62, 1851–1863. [Google Scholar] [CrossRef]
  16. Puterman, M.L. Markov Decision Processes: Discrete Stochastic Dynamic Programming; John Wiley & Sons: New York, NY, USA, 1994. [Google Scholar]
  17. Sutton, R.S. Learning to Predict by the Methods of Temporal Differences. Mach. Learn. 1988, 3, 9–44. [Google Scholar] [CrossRef] [Scilit]
  18. Ng, A.Y.; Harada, D.; Russell, S. Policy Invariance under Reward Transformations: Theory and Application to Reward Shaping. In Proceedings of the Sixteenth International Conference on Machine Learning; Morgan Kaufmann Publishers Inc.: San Francisco, CA, USA, 1999; pp. 278–287. [Google Scholar]
  19. Schulman, J.; Moritz, P.; Levine, S.; Jordan, M.; Abbeel, P. High-Dimensional Continuous Control Using Generalized Advantage Estimation. arXiv 2016, arXiv:1506.02438. [Google Scholar]
  20. Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning; PMLR, Proceedings of Machine Learning Research; JMLR: Brookline, MA, USA, 2018; Volume 80, pp. 1861–1870. [Google Scholar]
  21. Arulkumaran, K.; Deisenroth, M.P.; Brundage, M.; Bharath, A.A. Deep Reinforcement Learning: A Brief Survey. IEEE Signal Process. Mag. 2017, 34, 26–38. [Google Scholar] [CrossRef] [Scilit]
  22. Busoniu, L.; Babuska, R.; De Schutter, B. A Comprehensive Survey of Multiagent Reinforcement Learning. IEEE Trans. Syst. Man. Cybern. Part C (Appl. Rev.) 2008, 38, 156–172. [Google Scholar] [CrossRef] [Scilit]
  23. Hernandez-Leal, P.; Kartal, B.; Taylor, M.E. A Survey and Critique of Multiagent Deep Reinforcement Learning. Auton. Agents Multi-Agent Syst. 2019, 33, 750–797. [Google Scholar] [CrossRef] [Scilit]
  24. Gronauer, S.; Diepold, K. Multi-Agent Deep Reinforcement Learning: A Survey. Artif. Intell. Rev. 2022, 55, 895–943. [Google Scholar] [CrossRef] [Scilit]
  25. Nguyen, T.T.; Nguyen, N.D.; Nahavandi, S. Deep Reinforcement Learning for Multiagent Systems: A Review of Challenges, Solutions, and Applications. IEEE Trans. Cybern. 2020, 50, 3826–3839. [Google Scholar] [CrossRef] [Scilit]
  26. Foerster, J.; Farquhar, G.; Afouras, T.; Nardelli, N.; Whiteson, S. Counterfactual Multi-Agent Policy Gradients. Proc. AAAI Conf. Artif. Intell. 2018, 32, 2974–2982. [Google Scholar] [CrossRef] [Scilit]
  27. Kuba, J.G.; Chen, R.; Wen, M.; Wen, Y.; Sun, F.; Wang, J.; Yang, Y. Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning. In Proceedings of the International Conference on Learning Representations; International Foundation for Autonomous Agents and Multiagent Systems: Richland, SC, USA, 2022. [Google Scholar]
  28. Sunehag, P.; Lever, G.; Gruslys, A.; Czarnecki, W.M.; Zambaldi, V.; Jaderberg, M.; Lanctot, M.; Sonnerat, N.; Leibo, J.Z.; Tuyls, K.; et al. Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward. In Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems; International Foundation for Autonomous Agents and Multiagent Systems: Richland, SC, USA, 2018; pp. 2085–2087. [Google Scholar]
  29. Son, K.; Kim, D.; Kang, W.J.; Hostallero, D.E.; Yi, Y. QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning. In Proceedings of the 36th International Conference on Machine Learning; PMLR, Proceedings of Machine Learning Research; JMLR: Brookline, MA, USA, 2019; Volume 97, pp. 5887–5896. [Google Scholar]
  30. Yu, C.; Velu, A.; Vinitsky, E.; Gao, J.; Wang, Y.; Bayen, A.M.; Wu, Y. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games. Proc. Adv. Neural Inf. Process. Syst. 2022, 35, 24611–24624. [Google Scholar]
  31. Papoudakis, G.; Christianos, F.; Schäfer, L.; Albrecht, S.V. Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms in Cooperative Tasks. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks; Curran Associates Inc.: Red Hook, NY, USA, 2021; Volume 1. [Google Scholar]
  32. Fortunato, M.; Azar, M.G.; Piot, B.; Menick, J.; Hessel, M.; Osband, I.; Graves, A.; Mnih, V.; Munos, R.; Hassabis, D.; et al. Noisy Networks for Exploration. In Proceedings of the International Conference on Learning Representations; OpenReview.net: Amherst, MA, USA, 2018. [Google Scholar]
  33. Ecoffet, A.; Huizinga, J.; Lehman, J.; Stanley, K.O.; Clune, J. First Return, Then Explore. Nature 2021, 590, 580–586. [Google Scholar] [CrossRef] [Scilit]
  34. Badia, A.P.; Piot, B.; Kapturowski, S.; Sprechmann, P.; Vitvitskyi, A.; Guo, D.; Blundell, C. Agent57: Outperforming the Atari Human Benchmark. In Proceedings of the 37th International Conference on Machine Learning; PMLR, Proceedings of Machine Learning Research; JMLR: Brookline, MA, USA, 2020; Volume 119, pp. 507–517. [Google Scholar]
  35. Pîslar, M.; Szepesvari, D.; Ostrovski, G.; Borsa, D.L.; Schaul, T. When Should Agents Explore? In Proceedings of the International Conference on Learning Representations; OpenReview.net: Amherst, MA, USA, 2022. [Google Scholar]
  36. Dong, S.; Mao, H.; Yang, S.; Zhu, S.; Li, W.; Hao, J.; Gao, Y. WToE: Learning When to Explore in Multiagent Reinforcement Learning. IEEE Trans. Cybern. 2024, 54, 4789–4801. [Google Scholar] [CrossRef] [Scilit]
  37. Jain, A.K.; Lehnert, L.; Rish, I.; Berseth, G. Maximum State Entropy Exploration Using Predecessor and Successor Representations. Proc. Adv. Neural Inf. Process. Syst. 2023, 36, 49991–50019. [Google Scholar]
  38. Na, H.; Seo, Y.; Moon, I.C. Efficient Episodic Memory Utilization of Cooperative Multi-Agent Reinforcement Learning. In Proceedings of the International Conference on Learning Representations; OpenReview.net: Amherst, MA, USA, 2024. [Google Scholar]
  39. Lin, T.; Jabri, A. MIMEx: Intrinsic Rewards from Masked Input Modeling. Proc. Adv. Neural Inf. Process. Syst. 2023, 36, 35592–35605. [Google Scholar]
  40. Mahajan, A.; Rashid, T.; Samvelyan, M.; Whiteson, S. MAVEN: Multi-Agent Variational Exploration. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
  41. Oord, A.; Li, Y.; Vinyals, O. Representation Learning with Contrastive Predictive Coding. arXiv 2018, arXiv:1807.03748. [Google Scholar]
  42. Chen, T.; Kornblith, S.; Norouzi, M.; Hinton, G. A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the 37th International Conference on Machine Learning; PMLR, Proceedings of Machine Learning Research; JMLR: Brookline, MA, USA, 2020; Volume 119, pp. 1597–1607. [Google Scholar]
  43. Kingma, D.P.; Welling, M. Auto-Encoding Variational Bayes. arXiv 2014, arXiv:1312.6114. [Google Scholar]
  44. Higgins, I.; Matthey, L.; Pal, A.; Burgess, C.; Glorot, X.; Botvinick, M.; Mohamed, S.; Lerchner, A. beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework. In Proceedings of the International Conference on Learning Representations; OpenReview.net: Amherst, MA, USA, 2017. [Google Scholar]
  45. Cho, K.; van Merriënboer, B.; Gulcehre, C.; Bahdanau, D.; Bougares, F.; Schwenk, H.; Bengio, Y. Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2014; pp. 1724–1734. [Google Scholar] [CrossRef] [Scilit]
  46. van der Maaten, L.; Hinton, G. Visualizing Data Using t-SNE. J. Mach. Learn. Res. 2008, 9, 2579–2605. [Google Scholar]
  47. McInnes, L.; Healy, J.; Saul, N.; Grossberger, L. UMAP: Uniform Manifold Approximation and Projection. J. Open Source Softw. 2018, 3, 861. [Google Scholar] [CrossRef] [Scilit]
  48. Wang, Y.; Huang, H.; Rudin, C.; Shaposhnik, Y. Understanding How Dimension Reduction Tools Work: An Empirical Approach to Deciphering t-SNE, UMAP, TriMap, and PaCMAP for Data Visualization. J. Mach. Learn. Res. 2021, 22, 1–73. [Google Scholar]
  49. Kaelbling, L.P.; Littman, M.L.; Cassandra, A.R. Planning and Acting in Partially Observable Stochastic Domains. Artif. Intell. 1998, 101, 99–134. [Google Scholar] [CrossRef] [Scilit]
  50. Ghavamzadeh, M.; Mannor, S.; Pineau, J.; Tamar, A. Bayesian Reinforcement Learning: A Survey. Found. Trends Mach. Learn. 2015, 8, 359–483. [Google Scholar] [CrossRef] [Scilit]
  51. García, J.; Fernández, F. A Comprehensive Survey on Safe Reinforcement Learning. J. Mach. Learn. Res. 2015, 16, 1437–1480. [Google Scholar]
  52. Dietterich, T.G. Hierarchical Reinforcement Learning with the MAXQ Value Function Decomposition. J. Artif. Intell. Res. 2000, 13, 227–303. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Comparison of short-delay reward propagation and milestone-triggered delayed reward.
Figure 1. Comparison of short-delay reward propagation and milestone-triggered delayed reward.
Algorithms 19 00443 g001
Figure 2. CReSCENT framework for structural novelty learning and cross-layer credit allocation.
Figure 2. CReSCENT framework for structural novelty learning and cross-layer credit allocation.
Algorithms 19 00443 g002
Figure 3. Default learning curves for reward, cost, and utility.
Figure 3. Default learning curves for reward, cost, and utility.
Algorithms 19 00443 g003
Figure 4. Ablation learning curves for CReSCENT.
Figure 4. Ablation learning curves for CReSCENT.
Algorithms 19 00443 g004
Figure 5. Sensitivity of the intrinsic learning coefficient.
Figure 5. Sensitivity of the intrinsic learning coefficient.
Algorithms 19 00443 g005
Figure 6. t-SNE visualisation of training trajectories under different learning-drive settings. The axes are t-SNE dimensions, colour denotes the normalised within-episode step index, and enlarged points indicate states close to milestone activation.
Figure 6. t-SNE visualisation of training trajectories under different learning-drive settings. The axes are t-SNE dimensions, colour denotes the normalised within-episode step index, and enlarged points indicate states close to milestone activation.
Algorithms 19 00443 g006
Table 1. Related work summary for long-term value-driven scheduling.
Table 1. Related work summary for long-term value-driven scheduling.
AreaRepresentative WorkMain Modelling FocusRelation to CReSCENT
Industrial network scheduling[1,2,11,12]Connected production, cyber-physical integration, multi-factory planning, and learned dispatchingProvides the application setting, while CReSCENT targets sparse milestone feedback inside such networks
RL foundations[3,16,17,18]Sequential decision making, temporal-difference credit, Markov processes, and shaping invarianceSupplies the objective basis, while the proposed method adds structural supervision before delayed rewards arrive
Policy optimisation[4,19,20,21]Advantage estimation, clipped updates, entropy regularisation, and deep RL design choicesSupplies stable optimisers that can receive the CReSCENT intrinsic signal
MARL surveys[22,23,24,25]Coordination, non-stationarity, scalability, and application challengesClarifies why delayed milestones become harder when each layer is controlled
by an agent
Actor-critic coordination[5,26,27]Centralised training, counterfactual credit, and trust-region controlImproves cooperative optimisation but does not define structural novelty under silent prefixes
Value factorisation[6,28,29,30,31]Joint value decomposition, monotonic mixing, transformed factorisation, and benchmark practiceGives cooperative baselines, while CReSCENT targets the reward-observation gap before milestones
Intrinsic motivation[7,8,9,10,32]Curiosity, prediction error, density novelty, pseudo counts, and parameter noiseInspires internal reward, while CReSCENT measures novelty in milestone-relevant macro states
Exploration timing[33,34,35,36]Episodic memory, return-first search, and learned switching between exploration modesHelps reason about when to explore, while CReSCENT allocates novelty by cross-layer structural progress
Recent cooperative exploration[37,38,39,40]Entropy coverage, episodic memory, masked modelling, and variational explorationProvides direct baselines for intrinsic multi-agent learning
Representation learning[41,42,43,44,45]Contrastive objectives, latent variables, disentanglement, and recurrent encodingSupports the structural encoder used to cluster macro progress states
Belief and structure[46,47,48,49,50,51,52]Latent inspection, partial observability, Bayesian uncertainty, safe learning, and hierarchyMotivates analysing representation geometry and maintaining interpretable structural progress
Table 2. Illustrative cross-layer credit allocation under three structural transitions.
Table 2. Illustrative cross-layer credit allocation under three structural transitions.
Transition TypeLayer 1 ShareLayer 2 ShareLayer 3 Share
Early queue relief0.560.310.13
Midstream bottleneck removal0.180.640.18
Milestone release0.120.240.64
Table 3. Simulator transition specification for reproducibility.
Table 3. Simulator transition specification for reproducibility.
ElementSpecification
Transition graphSerial directed chain L 1 L 2 L M with finite buffers between adjacent layers
Task entryNew tasks arrive at L 1 according to a Poisson process with rate λ and are rejected only when the entry queue exceeds the capacity limit
Worker ruleEach idle worker selects at most one task from its local queue and processes it for one time step according to the layer-specific processing-time counter
Inter-layer transferA completed task advances to the next layer if the downstream buffer is not full; otherwise it remains blocked and contributes to B e e n d and
blocking statistics
Milestone ruleA milestone is triggered when a target number of tasks reaches the prescribed layer stage or when the active cross-layer backlog target is reduced below its threshold
External rewardMilestone reward is released only at triggered milestones; weak auxiliary penalties discourage idling, excessive waiting, and terminal backlog
Training budgetTen independent seeds are used for terminal evaluation; each method uses the same backbone optimiser, episode budget, observation access, and action constraints
Table 4. Performance under environment perturbations. Values are mean and standard deviation.
Table 4. Performance under environment perturbations. Values are mean and standard deviation.
ConditionMetricCReSCENTEMUETA-PSIMIMExExt-Only
standardReward 3465.85 ± 28.57 3320.02 ± 87.60 3276.01 ± 118.45 2566.53 ± 108.11 2836.74 ± 162.30
standardCost 9533.58 ± 65.61 9486.34 ± 61.05 9549.14 ± 100.84 9518.94 ± 638.19 9440.26 ± 241.18
standardUtility12,334.89   ±   79.84 12,003.16   ±   143.73 11,965.34   ±   170.22 10,726.97   ±   493.52 10,980.42   ±   290.10
layer = 2Utility 8537.65   ±   25.79 8233.59   ±   41.36 8461.69 ± 60.14 7681.31 ± 30.87 7890.54 ± 116.82
layer = 4Utility16,360.63   ±   78.11 16,169.48   ±   105.67 16,217.54   ±   19.98 14,363.06   ±   14.72 14,926.88   ±   251.36
step = 50Utility 7613.84   ±   25.55 7293.24 ± 32.55 7296.38 ± 20.91 6755.67 ± 52.91 6984.32 ± 103.55
step = 200Utility24,339.01   ±   69.32 24,101.43   ±   73.49 24,260.51   ±   70.25 20,847.72   ±   84.13 22,183.70 ± 411.42
task λ = 1 Utility 6354.41   ±   16.67 6049.25 ± 43.57 5985.33 ± 29.46 5666.33 ± 15.80 5841.26 ± 81.34
task λ = 3.5 Utility19,826.57   ±   25.89 19,286.34   ±   135.23 19,474.48   ±   41.06 17,493.73   ±   68.05 18,120.66   ±   326.28
worker = 6Utility12,737.38   ±   24.76 12,504.35   ±   91.72 12,321.40   ±   68.41 11,412.98   ±   28.43 11,742.30   ±   180.66
worker = 10Utility13,139.97   ±   16.88 12,954.70   ±   45.95 12,970.57   ±   38.78 11,707.93   ±   46.42 12,064.51   ±   196.44
Note: Bold values indicate the best performance for each listed setting.
Table 5. Statistical validation for default-setting utility.
Table 5. Statistical validation for default-setting utility.
ComparisonMean Utility Gain95% Confidence IntervalWilcoxon p-Value
CReSCENT vs. EMU331.73[205.6, 457.9]0.004
CReSCENT vs. ETA-PSI369.55[216.3, 522.8]0.003
CReSCENT vs. MIMEx1607.92[1268.4, 1947.5]<0.001
CReSCENT vs. external-only1354.47[1003.2, 1705.7]<0.001
Table 6. Sensitivity of the online clustering threshold ρ .
Table 6. Sensitivity of the online clustering threshold ρ .
ρ Number of Active ClustersUtilityMilestone Hit RateInterpretation
0.209412,110 ± 9671.8%Over-fragmented structural stages
0.307112,240 ± 8874.9%Stable but still fine-grained
0.405812,335 ± 8078.6%Default balance
0.504312,302 ± 8477.4%Slightly merged bottleneck states
0.603112,195 ± 9174.2%Under-separated milestone regions
Note: Bold values indicate the best performance in this sensitivity study.
Table 7. Sensitivity of cross-layer credit-allocation coefficients.
Table 7. Sensitivity of cross-layer credit-allocation coefficients.
Coefficient Setting ( a m , a b , a u ) UtilityMilestone Hit RateObservation
(0.50, 0.30, 0.20)12,334.89 ± 79.8478.6% Default milestone-first balance
(0.40, 0.40, 0.20)12,291.36 ± 86.1777.5%More bottleneck oriented
(0.60, 0.25, 0.15)12,318.04 ± 82.7578.1%Stronger milestone emphasis
(0.33, 0.33, 0.34)12,192.77 ± 104.6375.9%Uniform credit is less selective
(0.45, 0.20, 0.35)12,220.45 ± 97.3176.8%Over-rewards utilisation
Note: Bold values indicate the best performance in this sensitivity study.
Table 8. Projection diagnostics for latent trajectories. Higher values indicate better milestone-stage organisation.
Table 8. Projection diagnostics for latent trajectories. Higher values indicate better milestone-stage organisation.
ProjectionNeighbour TrustworthinessMilestone-Neighbour FractionStage Monotonicity
t-SNE0.910.780.74
UMAP0.890.760.72
PaCMAP0.900.770.73
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Xu, W.; Wan, Y.; Zuo, T. CReSCENT for Long-Term Value-Driven Scheduling in Multi-Layer Industrial Networks with Milestone-Triggered Rewards. Algorithms 2026, 19, 443. https://doi.org/10.3390/a19060443

AMA Style

Xu W, Wan Y, Zuo T. CReSCENT for Long-Term Value-Driven Scheduling in Multi-Layer Industrial Networks with Milestone-Triggered Rewards. Algorithms. 2026; 19(6):443. https://doi.org/10.3390/a19060443

Chicago/Turabian Style

Xu, Wei, Yi Wan, and Tianyu Zuo. 2026. "CReSCENT for Long-Term Value-Driven Scheduling in Multi-Layer Industrial Networks with Milestone-Triggered Rewards" Algorithms 19, no. 6: 443. https://doi.org/10.3390/a19060443

APA Style

Xu, W., Wan, Y., & Zuo, T. (2026). CReSCENT for Long-Term Value-Driven Scheduling in Multi-Layer Industrial Networks with Milestone-Triggered Rewards. Algorithms, 19(6), 443. https://doi.org/10.3390/a19060443

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop