1. Introduction
In recent years, supply chains have become increasingly exposed to abrupt, large-scale disruptions that arise from diverse sources, including geopolitical instability, port congestion, and other logistical bottlenecks. For example, following the military escalation around the Strait of Hormuz in early 2026, transits through the strait fell sharply and some 2000 vessels were stranded [
1], immobilizing a large share of global crude oil shipments and forcing the remaining flows onto longer and less predictable routes [
2,
3]. Under such disruptions, the assumptions of stable lead times and reliable supply in conventional inventory models no longer hold. Classical replenishment policies such as the continuous-review (r, Q) policy are effective under stationary conditions, but their control parameters, in particular the reorder point, are determined for a stationary lead time distribution. During major supply-side shocks, a reorder point calibrated to cover demand over the normal lead time can no longer function as intended. The policy has no mechanism to incorporate information about an impending disruption, so it cannot respond until the shock has already taken effect. By then, severe stockouts are unavoidable and total system costs increase sharply.
These limitations have motivated growing interest in deep reinforcement learning (DRL) for inventory problems that are difficult to solve with classical methods. Most DRL applications, however, share the same reactive limitation. The state is typically confined to internal variables such as on-hand inventory and outstanding pipeline orders, with supply uncertainty represented as a stationary distribution rather than as something the agent can anticipate [
4]. Such an agent cannot foresee an impending shock and revises its policy only once lead times have already lengthened.
In practice, supply disruptions are often preceded by observable signals such as weather forecasts, port-congestion indices, or supplier risk alerts that become available before the disruption is realized. Augmenting the agent’s state with such exogenous information has been noted as a promising research direction [
4]. Recent studies have begun to bring exogenous information into the state, but almost exclusively on the demand side, in the form of demand histories, forecasts, and related covariates [
5,
6]. Advance information about the supply side has remained unexplored. We propose a risk sensing DRL framework that incorporates precursor risk signals into the state space of a PPO agent, thereby introducing a feedforward component into the replenishment decision. The agent learns to map these signals to adjustments in reorder timing and to position protective inventory ahead of a disruption. The value of this mechanism depends on the reliability of the signals and the length of the prediction window they provide. We therefore treat signal quality and predictive horizon as explicit design parameters rather than assuming perfect foresight. Preparing an organization for disruptions before they occur is also the purpose of scenario planning, the established managerial approach to this problem. The model developed here operates at a different level, and
Section 8 discusses how the two complement each other.
These considerations lead to three research questions.
- RQ1.
Does incorporating precursor risk signals into the state of a DRL agent reduce the long-run cost of inventory control, relative to a calibrated static policy and to an identically trained reactive policy?
- RQ2.
How does the value of the signal depend on its quality and on the predictive horizon? How reliable and how early must a warning be to be worth acting on?
- RQ3.
How much of the measured value comes from the information itself rather than from learning, and through what mechanism does the policy convert warnings into cost savings?
The contributions of this paper are as follows. First, we extend the conventional DRL state representation to incorporate observable precursor risk signals. Second, we develop a PPO-based learning scheme that maps these signals to proactive reorder decisions. Third, by treating signal quality and predictive horizon as design parameters, we measure the value of advance information directly, quantifying how cost and service-level performance improve as more and better information becomes available, relative to a conventional (s, Q) policy.
3. Problem Description
This study is abstracted from the component replenishment practice of a large Korean electronics manufacturer. Key components such as displays and secondary batteries are sourced from strategic single suppliers, and a shortage of any of them halts assembly. To manage the risk of such shortages, the manufacturer has operated a supply risk management process for a number of years, in which disruptions are classified by their source, whether exogenous, internal to the supplier, or logistical, graded by their severity, and recorded with their causes and their impact on each product. Within this process, leading indicators of supply conditions are monitored, and the detection accuracy and the horizon of the resulting warnings are tracked as measures of monitoring performance. The question this study addresses is how such warnings should be used in responding to disruptions.
Supply disruptions arise from many causes and differ in their effects. Ports close because of storms, congestion, or conflict, and shipments wait or take longer routes. Supplier plants stop because of labor disputes, fires, or earthquakes, and production resumes only after settlement or repair. A failure at a sub-tier supplier (the supplier’s own suppliers at any tier beyond the first) propagates the disruption downstream, and customs and transport delays occur routinely. Chopra and Sodhi [
25] classify supply chain risks into categories, each with its own drivers, and Sheffi and Rice [
26] describe a disruption as a profile that unfolds in phases from preparation through impact to recovery, whose shape differs by type.
Table 1 lists the main types. The responses available to a manufacturer are also diverse. Tomlin [
27] distinguishes mitigation tactics taken before a disruption, such as holding protective inventory and qualifying a second source, from contingency tactics taken after it, such as rerouting and expediting shipments. Tang [
28] surveys a broader set that includes postponement and flexible supply bases. This study considers protective inventory, which applies across disruption types and can be implemented at the operational level without changing the structure of the supply chain. The other responses are established risk management strategies but lie beyond the scope of the model, and
Section 8 discusses what this restriction means for the results and their implications.
The decision is then when to place a fixed-lot order for the component with its supplier, and the disruptions above take effect on this decision as a lengthening of the delivery lead time. Deliveries from the supplier are delayed while the disruption lasts and resume when it is over, so that a disruption appears to the planner as an interruption of replenishment followed by recovery. The model assumes that the supplier continues to accept orders during the interruption and delivers them after it ends, an assumption stated with its scope in
Section 4.1. A stoppage at the supplier’s plant, whether from a strike, a fire, or a failure further upstream, is represented in this way, since orders placed during the stoppage are queued and delivered once production resumes. Among the types in
Table 1, this holds except where a regulatory hold bars the product itself, in which case orders are rejected rather than delayed.
Within this response, what distinguishes one disruption type from another is how long deliveries are delayed, whether the disruption can be recognized before it begins, and how early it becomes recognizable. The length of the interruption determines how severe a disruption is for replenishment, since it sets how many periods of demand must be covered without deliveries, and we refer to this severity as the grade of the disruption. Whether and how early a disruption is recognized depends on the leading indicators that precede it, which come from the supplier itself, through advance notices, and from public and commercial sources, such as port authorities, weather services, and supplier risk monitoring services. The prediction that such indicators support is what
Section 4.2 defines as the precursor risk signal. We refer to the fraction of disruptions that such indicators reveal before the disruption onset as the detection rate, and to the number of periods between recognition and the disruption onset as the predictive horizon, since the earlier a disruption is recognized, the more time the planner has to build up stock.
Table 1 places the main disruption types on these dimensions. The model of
Section 4 formalizes these dimensions. The grade and its frequency define the disruption process (
Section 4.1), and the detection rate and the predictive horizon define the signal (
Section 4.2). Different disruption types are therefore represented as different positions on these dimensions rather than as a single event, and the learned policy responds to them differently, as
Section 7.3 shows. Which of these types apply to a given manufacturer and what grade and detection rate each warrants is a judgment about the firm’s own exposure. Scenario planning is designed to support such judgments, and
Section 8 discusses how it and the model studied here combine.
The model represents two parties, the manufacturer and its direct supplier. For a key component, the tiers above the direct supplier are monitored as well, because a disruption further upstream reaches the manufacturer only after it has propagated to a delay at the direct supplier, and watching the upstream tier reveals it that much earlier. The signal studied here does not restrict where the warning comes from. A warning obtained from an upstream tier is represented as a signal that detects the coming delay earlier, that is, with a longer predictive horizon, and the results of
Section 7.2 on the value of the horizon apply to it. What the model does not represent is the propagation of a disruption from an upstream tier to the direct supplier. Thus, decisions that depend on how a disruption propagates through the supply network, such as which tiers to monitor, how to integrate warnings that arrive from different tiers, and where along the chain to position protective stock, lie outside its scope and are identified as future research in
Section 9.
The disruptions in
Table 1 act on the inbound side of the manufacturer. The same causes, such as a port closure, a transport strike, or a regional conflict, often disrupt its outbound logistics as well, delaying or losing sales and stranding finished products. When both sides are disrupted, the effects compound. The model studied here accounts only for the holding and shortage costs of the component and holds its demand stationary to isolate the effect of supply-side information (
Section 6.1). It has no representation of outbound conditions, so outbound disruptions lie outside the model, as stated in
Section 8, and the results should be read as the value of advance supply information on the inbound side, taken in isolation from outbound conditions.
8. Discussion
The main finding of this study is that a precursor risk signal has substantial value in inventory control, but only under certain conditions. Incorporating the signal into the state of a PPO agent reduced the average cost by 9.3% against a calibrated (s, Q) policy and by 12.1% against an identically trained reactive policy at the reference configuration, and the advantage over the recalibrated (s, Q) policy persisted across every environment variation. The decomposition of
Section 7.3 identifies where the gains arise. Against the static (s, Q) policy, the saving arises in the calm periods, where the policy holds substantially less of the insurance buffer. Against the reactive policy, the saving arises in the exposed periods, where stock positioned upon a warning prevents shortages. The signal thus lets the policy substitute information for inventory, holding less stock in calm periods and more stock ahead of predicted disruptions. This extends the classical demand-side result of Hariharan and Zipkin [
20] to supply-side advance information in a learning setting.
Three aspects of the results have practical implications. First, the horizon response shows a plateau rather than an interior optimum. Once the effective horizon is covered, that is, the few ordering periods needed to accumulate protective stock (about four here, under the one-lot-per-period action space), additional look-ahead brings neither measurable cost nor benefit. The plateau persists even when the training run length is doubled and the neural networks are widened. For practice, this makes horizon selection forgiving. A forecast need not reach far into the future to be useful and providing a longer horizon than necessary does no harm within the tested range.
Second, the critical requirement is signal quality. Advance information is valuable only above a reliability threshold. A signal that detects half of upcoming disruptions offers only a marginal gain over a well-calibrated static rule, and a weaker one falls behind it. Investments in sensing infrastructure should therefore be judged by the detection rate they can deliver, not by the length of the forecast horizon.
Third, the value of the signal depends on the disruption frequency, while no dependence on the shortage penalty was detected within the tested range. The value is largest where severe disruptions are rare, because a static buffer held against them is idle most of the time. The value decreases considerably where disruptions are frequent, because the buffer is then in regular use and warnings arrive faster than stock can be accumulated. Precursor signals are therefore most useful in environments with low-probability, high-impact disruptions.
The leading indicators of
Table 1 are observable but reading them as a warning that a disruption is coming is a judgment, and that judgment is a matter of perception. Managers’ perception of supply risk is subjective, and their actions follow their perception [
35], and indicators that look clear in hindsight are often ambiguous beforehand. Instead of representing this judgment, the model relies on what the indicators and the judgment together leave behind. Looking back over a period, one can count how many disruptions were recognized before they began and how many periods in advance, and these two quantities are the detection rate and the predictive horizon. They are properties of a firm’s monitoring that are established in hindsight, not quantities known in advance. As noted in
Section 3, the practice from which the problem is abstracted tracks both quantities. The model takes them as given, and the policy is learned for a given detection rate and horizon, so that it prescribes how to act on a signal of that quality, whose value is then measured in the experiments (
Section 7.2).
Acting on advance information before a disruption arrives is also the purpose of scenario planning [
36], a managerial practice with a long history, and the two approaches can be used in a complementary way. Developed at Royal Dutch/Shell in the 1970s [
37], scenario planning constructs multiple internally consistent scenarios of how the business environment could unfold, uses them to help managers reperceive their exposure, and selects the leading indicators that would show which scenario is unfolding [
38]. Its application to supply chains and logistics has been developed in later studies [
39,
40,
41]. Scenario planning identifies the strategic options each scenario calls for, and the model studied here can be applied based on that selection at the operational level. In the terms of
Section 3, each scenario corresponds to a disruption type in
Table 1, with its grade and the indicators that precede it. For each such scenario, the model derives the ordering response, as a learned policy or as the transparent rule of
Section 6.2, and quantifies its expected cost and service level under the detection rate and horizon achievable with the selected indicators. The results then return to the scenario stage. They show which scenarios are worth monitoring and how reliable their indicators must be. In this study, that threshold is found to be a detection rate above one half. The two tools thus form a loop rather than a sequence.
A drawback of a DRL policy is its lack of transparency, which scenario planning provides through scenarios and indicators that a manager can follow. A rule can be considered in place of the DRL policy to address this. In this study, the signal-triggered heuristic, which reads the same signal through five interpretable parameters, retains about 83% of the measured benefit, indicating that most of the value comes from the information itself rather than from the learning method. A manager who requires a transparent instrument can adopt the rule, and the additional value of the less transparent DRL policy becomes a quantified increment rather than an assumption.
For a manager considering adoption, the results suggest the following steps. First, classify the firm’s exposure by the disruption types of
Table 1 and identify the indicators available for each, which is the task that scenario planning supports. Second, estimate the detection rate and the horizon of the current monitoring from past disruptions, as the practice studied does (
Section 3). Third, if the detection rate exceeds about one half, act on the signal, initially through the transparent rule of
Section 6.2, which requires only a base reorder point and grade-specific increments and captures most of the gain, and through a learned policy where the remaining gain justifies it. Fourth, evaluate investments in monitoring by the detection rate they add rather than by the horizon they extend, since the value of the signal saturates beyond a few periods of warning but rises steeply with the detection rate. The saving to expect relative to a well-calibrated static rule depends on the disruption frequency, from about 3% where severe disruptions are frequent to about 12% where they are rare (
Section 7.4).
Several limitations should be noted. The model contains a single response lever, the timing of replenishment orders. In practice, disruption types differ in the responses they call for, such as rerouting or expediting for a transport disruption, a second source for a supplier outage, or product substitution for a discontinued component. Earlier ordering was chosen because it is available for every disruption type and requires no contractual or design change, and fixing the lever isolates the value of the information from the value of the means to act on it. Additional levers would give a signal-aware policy more ways to convert a warning into protection, but a static policy would gain the same levers, so their net effect on the value of the signal cannot be determined without modeling them. The results should therefore be read as the value of advance information when the only available response is to reorder earlier, which is the response common to all disruption types.
The action space permits at most one lot per period, which limits the rate at which a warning can be converted into protective inventory. An action space allowing multiples of the lot size (an (s, nQ)-type policy) would speed up this conversion, shortening the effective horizon of
Section 7.2, and could change the measured value of the signal in either direction, since faster accumulation benefits the signal-based policies and the benchmark alike.
The model also assumes that the supplier keeps accepting orders during a disruption. Where a disruption removes the supplier’s capacity so that orders are rejected, the manufacturer cannot replenish at all until recovery, and protective inventory built before the onset becomes the only source of supply for the duration. For disruptions short enough to be covered by pre-positioned stock, the value of a warning would then be larger than measured here, since pre-positioning is the sole response available, whereas for longer ones both the signal-based and the static policies run out of stock and the gap narrows. The value of a warning under rejected orders thus depends on the length of the disruption relative to the stock that can be pre-positioned, and establishing this dependence requires modeling the rejection explicitly, which is a direction for extension.
The signal model is stylized. It models missed detections, with probability , but not false alarms, and a detected disruption carries its true severity and timing, whereas real leading indicators also produce false positives and imprecise estimates. The warning lead is also constant, fixed at H for every detected disruption, whereas real disruptions differ in the pace at which they develop and become recognizable. A false alarm would trigger a protective buildup for a disruption that never arrives, and its cost is the temporary holding of unneeded stock, whereas a missed detection exposes the system to the full shortage cascade of an unanticipated disruption. Under the cost structure studied here, false alarms are therefore expected to erode the value of the signal gradually rather than eliminate it, but quantifying this erosion, and the false-alarm rate at which a signal stops being worth acting on, requires extending the signal model and is left for future work.
Demand is stationary, the system is a single-item, two-echelon chain, and only the inbound supply process is modeled, so disruptions to the manufacturer’s outbound logistics lie outside the scope of the model.
Section 3 relates this restriction to the disruptions that affect both sides. When both sides are disrupted, the manufacturer must decide whether to reduce production while shipments are blocked and components are themselves in short supply, and whether to adjust its component inventory accordingly. An outbound warning may thus call for the opposite of what an inbound warning calls for, and a planner who receives both must weigh them together. Extensions to outbound disruptions, interactions between demand-side and supply-side information, and the propagation of warnings through deeper supplier networks are left for future research.
Finally, although disruptions produce structural breaks in realized lead times, the underlying risk profile is fixed over time, so the environment is stationary in the stochastic process sense and the risk profile can be learned. Genuine non-stationarity, that is, structural change in the risk profile itself, would require online adaptation beyond the scope of this paper.
9. Conclusions
This paper developed a risk sensing deep reinforcement learning framework for inventory control under supply disruptions. The state of a PPO agent is augmented with precursor risk signals, which predict the severity and timing of upcoming disruptions, so that the replenishment decision acquires a feedforward component. The learned policy operates a state-dependent reorder point that stays low in calm conditions and rises with the severity and proximity of a predicted threat, positioning protective inventory before lead times lengthen rather than after.
Treating signal quality and predictive horizon as explicit design parameters allowed the value of advance information to be measured directly. The risk sensing policy outperforms both a per-environment calibrated (s, Q) policy and an identically trained reactive policy whenever the signal is sufficiently reliable, and the comparison with the reactive policy isolates the value of the signal itself (RQ1). That value requires only a short horizon, enough to complete the protective buildup, but it does require a sufficiently high detection rate (RQ2). The value is greatest where severe disruptions are infrequent and static buffers are therefore most wasteful. In this environment, neither deep reinforcement learning by itself nor an unreliable signal produced a gain. A transparent signal-triggered rule captured about 83% of the gain, with learning contributing the rest (RQ3). The gains reported here come from reliable information, not from the learning method alone. The practical message for supply chain managers is that the transition from reactive response to proactive mitigation depends on sensing infrastructure. Even an imperfect early-warning capability can replace a substantial share of safety stock, provided that it detects a majority of upcoming disruptions a few periods in advance. These findings rest on a stylized simulation abstracted from a single industrial setting, and the quantitative magnitudes are specific to the environment studied. More general real-world settings remain to be addressed.
There are several directions for future research. On the modeling side, an action space that allows more than one lot per period and a signal model with false alarms, graded confidence, and warning leads that vary with the pace of the disruption would bring the setting closer to practice. Applying the combination of scenario planning and the model outlined in
Section 8 to an actual firm would show how the disruption types and indicators selected at the scenario stage translate into the value of monitoring them. In multi-echelon systems, warnings could propagate across stages, and the questions become which tiers to monitor, how to integrate warnings that arrive from different tiers, and where in the chain protective stock should be positioned when a disruption is predicted upstream. Extending the action space with the type-specific responses of
Table 1, such as expediting or a second source, would show how the value of a warning changes when more than one response is available. Finally, validation on empirical leading indicators, such as port-congestion indices or supplier risk alerts, would establish the detection rates and horizons that real signals deliver, the two parameters on which the value of the signal was shown to depend.