Skip to Content
GamesGames
  • Article
  • Open Access

29 May 2026

A Cooperative Pollution Control Differential Game with Randomly Switching Payoffs

and
Faculty of Applied Mathematics and Control Processes, St. Petersburg State University, 199034 St. Petersburg, Russia
*
Author to whom correspondence should be addressed.
This article belongs to the Section Cooperative Game Theory and Bargaining

Abstract

We study a continuous-time cooperative differential game of pollution control in which the pollution stock accumulates emissions and affects long-run welfare. The key feature is a one-time random increase in the public damage weight, interpreted as a regime shift in environmental policy, social damage assessment, or regulatory pressure. Using dynamic programming, we characterize the grand-coalition feedback solution from the Hamilton–Jacobi–Bellman equations and derive closed-form expressions for cooperative emissions, pollution dynamics, regime-specific steady states, and transition paths. Under emission caps, we construct the coalition characteristic function using a conservative worst-case benchmark for outsider behavior rather than an unlimited-pollution assumption. For payoff allocation, we derive a dynamic payment schedule that implements the Shapley allocation along the stochastic pollution path and keeps the remaining payoff consistent with the corresponding continuation game. Finally, we extend the framework to a threshold-triggered shifted-exponential switching mechanism. This extension gives a computable objective for the optimal threshold-hitting time and clarifies how the pollution threshold and switching hazard can be interpreted as policy-relevant indicators of regulatory or ecological regime change.

1. Introduction

In pollution-control problems, current emissions affect not only present payoffs but also the future state of the environment. Since emissions accumulate into a pollution stock that persists over time, the choice of current emissions changes the future damage faced by all players. At the same time, the existing pollution stock influences optimal emission decisions and the value of cooperation. Continuous-time differential games provide a natural framework for studying this interaction among strategic players and for analyzing optimal cooperative behavior and the allocation of cooperative gains over time (Dockner & Long, 1993; Kossioris et al., 2008; Rubio & Casino, 2002).
In environmental cooperation, such as international environmental agreements or joint regulatory arrangements, participants face a tension between commitment and flexibility: cooperation requires a stable allocation rule, while the perceived severity of environmental damages, regulatory pressure, and political priorities may change over time.
This paper studies a continuous-time cooperative differential game of pollution control with a single state variable P ( t ) , the pollution stock, and a one-time stochastic switch in the public damage weight. The regime switch is interpreted as a change in environmental policy, social damage assessment, or regulatory pressure. In the baseline model, the damage-weight increase occurs at an exponentially distributed random time. This specification captures a sudden change in environmental policy or damage assessment while preserving a constant hazard rate, which makes the dynamic programming problem analytically tractable. This baseline case is later extended to a threshold-triggered switching mechanism, where switching risk becomes active only after pollution reaches a specified level. Under this regime-switching structure, we derive feedback policies, regime-specific steady states, and transition paths from the Hamilton–Jacobi–Bellman equations (Fleming & Soner, 2006; Yin & Zhu, 2010).
The model includes emission caps, interpreted as short-run technological or regulatory constraints on feasible emissions. When evaluating the worth of a coalition, we use a conservative worst-case benchmark for outsiders: nonmembers choose feasible emissions that are least favorable to the coalition, in the spirit of a characteristic-function construction for cooperative games (von Neumann & Morgenstern, 1953). This is a pessimistic benchmark, not an assumption of unlimited or unconstrained pollution by nonmembers.
The main contributions are as follows. First, we characterize the grand-coalition feedback solution from the Hamilton–Jacobi–Bellman equations and obtain closed-form expressions for cooperative emissions, pollution dynamics, steady states, and transition paths. Second, we construct the characteristic function under emission caps and a conservative outsider benchmark. Third, after deriving the coalition values, we compute the corresponding Shapley allocation and construct a payment schedule for its implementation over time. This schedule keeps the remaining unpaid payoff equal to the Shapley value of the continuation game at each date (Petrosyan & Danilov, 1979; Shapley, 1953). Fourth, we extend the model to a threshold-triggered shifted-exponential switching mechanism. In this extension, switching risk becomes active only after the pollution stock reaches a prescribed threshold, which can be interpreted as an ecological warning level, a regulatory trigger, or a point at which public pressure for stricter policy increases. The post-threshold hazard rate captures the speed of institutional or regulatory reaction after this level has been reached. This extension gives a tractable objective for the optimal threshold-hitting time, which summarizes the trade-off between current emission benefits and the risk of entering a high-damage regime (Seierstad & Sydsæter, 1987). This interpretation is relevant for regulators and cooperative agreements in which emission paths and participation incentives must be managed jointly.
The remainder of this paper is organized as follows. Section 2 introduces the baseline exponential-switch model and derives the cooperative feedback solution. Section 3 develops the characteristic function under emission caps. Section 4 constructs the Shapley-value-based dynamic payment schedule and proves its time-consistency property. Section 5 analyzes the threshold-triggered switching extension and provides a numerical illustration. Section 6 concludes this paper.

2. Pollution-Control Game with an Exponential Damage Switch

We study a cooperative differential game of pollution control with a single state variable, the pollution stock P ( t ) , following the standard stock-pollution modeling tradition (e.g., Dockner & Long, 1993; Rubio & Casino, 2002; Tur & Gromova, 2018/2020). The stock evolves as the accumulation of emissions net of natural decay:
P ˙ ( t ) = i = 1 n e i ( t ) α P ( t ) , α > 0 , P ( 0 ) = P 0 0 .
where e i ( t ) denotes player i’s emission rate at time t, and α is the natural decay rate. We assume that emissions are bounded by
0 e i ( t ) k i , i = 1 , , n ,
where k i > 0 is the maximum feasible emission rate of player i during the game. These bounds are interpreted as short-run technological, regulatory, or production-capacity constraints. Thus, the model focuses on optimal emission choices and payoff allocation under given feasible-emission ranges. The law of motion captures the dynamic externality of stock pollutants: current emissions affect future payoffs through the state P ( t ) , thereby generating strategic interaction in feedback form (Kossioris et al., 2008).
Each player’s private benefit, excluding public environmental damage, is given by
π i ( e i ) = b i e i 1 2 e i 2 , i = 1 , , n .
The parameter b i > 0 measures the marginal private gain from emissions. The quadratic term ensures strict concavity at the individual level, which supports closed-form first-order conditions and facilitates the subsequent characterization of stationary feedback policies and steady-state properties.
In the cooperative setting, a grand-coalition planner chooses emission paths { e i ( · ) } i = 1 n to maximize discounted welfare. The planner discounts at rate r > 0 and counts the public damage once for each of the n players:
max { e i } E 0 e r t i = 1 n ( b i e i 1 2 e i 2 ) n γ ( t ) P ( t ) d t .
The key uncertainty is a one-time regime shift in the public damage weight γ ( t ) : at a random time τ Exp ( λ ) with λ > 0 , the damage parameter jumps from γ 1 to γ 2 ,
γ ( t ) = γ 1 , t < τ , γ 2 , t τ .
Differential games with random switching of the payoff functional were considered in Zaremba (2022); however, the analysis was carried out in open-loop strategies, and a transformation of the expected payoff was demonstrated to construct solutions via the maximum principle, analogous to games with random duration (Kostyunin & Shevkoplyas, 2011). In contrast, we are interested in constructing solutions in feedback strategies using the dynamic programming method. Because the exponential distribution is memoryless, the switching mechanism can be represented by a constant hazard rate, so that the switching risk enters the continuous-time HJB equation in a tractable linear form within a regime-switching (hybrid) control framework (see Fleming & Soner, 2006; Lv, 2020; Mao & Yuan, 2006; Yin & Zhu, 2010). For compact notation, we write S b = i = 1 n b i and S b 2 = i = 1 n b i 2 .
We denote the instantaneous payoff of player i in the first regime by
h i 1 ( e i ( t ) , P ( t ) ) = e r t ( b i e i 1 2 e i 2 ) γ 1 P ( t )
and in the second regime by
h i 2 ( e i ( t ) , P ( t ) ) = e r t ( b i e i 1 2 e i 2 ) γ 2 P ( t ) .

2.1. Feedback (HJB) Strategies

This subsection derives the positional cooperative optimum. Given the one-dimensional stock dynamics and the LQ structure, the solution can be obtained from the Hamilton–Jacobi–Bellman (HJB) equations using an affine value-function ansatz, a standard approach in differential games and continuous-time dynamic programming (see, e.g., Basar & Olsder, 1999; Dockner et al., 2000).

2.1.1. Post-Switch Problem ( t τ , γ = γ 2 )

After the switch, the damage weight is constant at γ 2 . Let V 2 ( P ) denote the regime-2 cooperative value function. The HJB equation is
r V 2 ( P ) = max { e i } i = 1 n ( b i e i 1 2 e i 2 ) n γ 2 P + V 2 P ( P ) i = 1 n e i α P .
The first-order condition for each i is
b i e i + V 2 P ( P ) = 0 e i ( 2 ) ( P ) = b i + V 2 P ( P ) .
With the affine ansatz
V 2 ( P ) = A 2 + B 2 P V 2 P ( P ) B 2 ,
the optimal emissions are constant, e i = b i + B 2 , and the projection onto [ 0 , k i ] yields
e i ( 2 ) = Π [ 0 , k i ] ( b i + B 2 ) , i N ,
where Π [ 0 , k i ] ( x ) : = min { k i , max { 0 , x } } . Matching the P-coefficients gives
r B 2 = n γ 2 α B 2 , B 2 = n γ 2 r + α ,
Define
E 2 : = i N e i ( 2 ) , C 2 : = i N b i e i ( 2 ) 1 2 ( e i ( 2 ) ) 2 .
Then, the constant term satisfies
r A 2 = C 2 + B 2 E 2 ,
and the regime-2 steady state is
P 2 * = E 2 α .
For any initial value P ( τ ) , the transition path is
P 2 ( t ) = P 2 * + P ( τ ) P 2 * e α ( t τ ) ,
so P 2 * is globally exponentially stable since α > 0 .
Note that if 0 b i n γ 2 r + α k i for all i = 1 , , n , then
e i ( 2 ) = b i n γ 2 r + α i = 1 , , n ,
A 2 = 1 2 S b 2 + B 2 S b + 1 2 n B 2 2 r ,
P 2 * = S b n 2 γ 2 r + α α ,
P 2 ( t ) = P 2 * + P ( τ ) P 2 * e α ( t τ ) , t τ .

2.1.2. Pre-Switch Problem ( t < τ , γ = γ 1 , Hazard λ )

Before the switch, the damage weight is γ 1 and the system faces a constant switching hazard λ toward regime 2. Let V 1 ( P ) denote the regime-1 cooperative value. With exponential switching, the HJB equation is
( r + λ ) V 1 ( P ) = max { e i } i = 1 n ( b i e i 1 2 e i 2 ) n γ 1 P + V 1 P ( P ) i = 1 n e i α P + λ V 2 ( P ) .
Assume V 1 ( P ) = A 1 + B 1 P . The FOC has the same form as (6):
b i e i + V 1 P ( P ) = 0 e i ( P ) = b i + V 1 P ( P ) .
The projected optimal emissions are
e i ( 1 ) = Π [ 0 , k i ] ( b i + B 1 ) , i N ,
and matching coefficients gives
( r + λ ) B 1 = n γ 1 α B 1 + λ B 2 ,
hence
B 1 = n r + λ + α γ 1 + λ γ 2 r + α .
Define
E 1 : = i N e i ( 1 ) , C 1 : = i N b i e i ( 1 ) 1 2 ( e i ( 1 ) ) 2 .
Then,
( r + λ ) A 1 = C 1 + B 1 ( E 1 ) + λ A 2 ,
and the regime-1 steady state (if it were to persist) is
P 1 * = E 1 α .
If the system were to remain in regime 1 indefinitely, then
P 1 ( t ) = P 1 * + P 0 P 1 * e α t .
Note that if 0 b i + B 1 k i , 0 b i + B 2 k i for all i = 1 , , n , then
A 1 = 1 2 S b 2 + B 1 S b + 1 2 n B 1 2 + λ A 2 r + λ ,
e i ( 1 ) ( t ) = b i n r + λ + α γ 1 + λ γ 2 r + α ,
P 1 * = S b + n B 1 α = S b n 2 γ 1 + λ γ 2 / ( r + α ) r + λ + α α ,
P 1 ( t ) = P 1 * + P 0 P 1 * e α t , t < τ ,
so P 1 * is exponentially stable. In the full model, regime 1 lasts a random time: while t < τ , P ( t ) moves toward P 1 * , remains continuous at t = τ , and then converges toward P 2 * under regime 2.

3. Characteristic Function

To analyze the distribution of the cooperative payoff, we construct a characteristic function for each coalition S N . The worth of a coalition is defined as the payoff that the coalition can guarantee when its members coordinate their admissible emissions, while players outside the coalition are evaluated through a conservative benchmark. This construction follows the classical characteristic-function approach for cooperative games (von Neumann & Morgenstern, 1953) and its use in cooperative differential games (Petrosjan, 2005; Yeung & Petrosyan, 2006; Zaccour, 2003).
For a nonempty coalition S N , let
s : = | S | , K S : = j N S k j .
The conservative benchmark assumes that nonmembers choose feasible emissions that are least favorable to coalition S, subject to the same emission caps 0 e j ( t ) k j . This gives a lower-bound assessment of the coalition’s value and should not be interpreted as an assumption that outsiders can pollute without limit.
Since pollution damages enter negatively in the payoff of coalition S, the least favorable feasible choice of each outsider is the upper emission bound. Hence, under this benchmark,
e j ( t ) k j , j N S .
Coalition members then re-optimize their own admissible controls, taking the aggregate outsider emission K S as given. We now characterize the value of coalition S under this benchmark.

3.1. Post-Switch Regime ( t τ , γ = γ 2 )

Let V 2 ( S ; P ) denote the coalition value in regime 2. The HJB is
r V 2 ( S ; P ) = max { e i [ 0 , k i ] } i S i S b i e i 1 2 e i 2 s γ 2 P + V 2 , P ( S ; P ) i S e i + K S α P .
With the affine ansatz
V 2 ( S ; P ) = A 2 , S + B 2 , S P ,
the unconstrained FOC gives e i = b i + B 2 , S , and projecting to [ 0 , k i ] yields
e i ( 2 , S ) = Π [ 0 , k i ] ( b i + B 2 , S ) , i S .
Matching the P-coefficients gives
r B 2 , S = s γ 2 α B 2 , S B 2 , S = s γ 2 r + α .
Define
E 2 , S : = i S e i ( 2 , S ) , C 2 , S : = i S b i e i ( 2 , S ) 1 2 ( e i ( 2 , S ) ) 2 .
Then, the constant term satisfies
r A 2 , S = C 2 , S + B 2 , S ( E 2 , S + K S ) .

3.2. Pre-Switch Regime ( t < τ , γ = γ 1 , Hazard λ )

Let V 1 ( S ; P ) denote the coalition value in regime 1. The HJB is
( r + λ ) V 1 ( S ; P ) =                   max { e i [ 0 , k i ] } i S i S b i e i 1 2 e i 2 s γ 1 P + V 1 , P ( S ; P ) i S e i + K S α P + λ V 2 ( S ; P ) .
With
V 1 ( S ; P ) = A 1 , S + B 1 , S P ,
the projected optimal emissions are
e i ( 1 , S ) = Π [ 0 , k i ] ( b i + B 1 , S ) , i S ,
and matching coefficients gives
( r + λ ) B 1 , S = s γ 1 α B 1 , S + λ B 2 , S ,
hence
B 1 , S = s r + λ + α γ 1 + λ γ 2 r + α .
Define
E 1 , S : = i S e i ( 1 , S ) , C 1 , S : = i S b i e i ( 1 , S ) 1 2 ( e i ( 1 , S ) ) 2 .
Then,
( r + λ ) A 1 , S = C 1 , S + B 1 , S ( E 1 , S + K S ) + λ A 2 , S .
Now, write the coalition value:
V 1 ( S ; P 0 ) = E 0 τ e r t C 1 , S s γ 1 P ( t ) d t + τ e r t C 2 , S s γ 2 P ( t ) d t = A 1 , S + B 1 , S P 0 .
In the subgame starting at time t from P 1 ( t ) , the value of the characteristic function is
e r t V 1 ( S ; P 1 ( t ) ) = e r t A 1 , S + e r t B 1 , S P 1 ( t ) .

4. Time-Consistent Imputation Distribution Procedure

After constructing the characteristic function of a cooperative game, one of the established cooperative solution concepts—such as the core, the Shapley value, or nucleolus—can be applied to obtain a fair imputation of the total payoff among players. Equally important is the dynamic payoff distribution over time, which ensures the time-consistency property of the chosen solution concept. This process involves making continuous payments to players such that the vector of remaining unpaid balances at any time t belongs to the same solution concept in the subgame starting at t. Such a distribution guarantees that players adhere to the chosen imputation at every moment. The notion of imputation distribution procedure and time-consistent solutions was introduced in Petrosyan and Danilov (1979). This approach was adapted for games with random switching in payoff functions in Zaremba (2022), but no explicit rule for constructing the IDP was considered. Here, we introduce such a rule. A distinctive feature of the game is that, prior to the switching time, players in the current subgames base their decisions on expected payoffs, whereas afterward, they rely on realized payoffs; this must be accounted for in designing the payoff distribution procedure.

4.1. Shapley-Based IDP After the Switching Time

To illustrate this feature, consider the Shapley value (Shapley, 1953) as the optimality principle. We begin with subgames starting after the switching time.
Let τ ^ be a realization of the random variable τ . Consider the game starting at t τ ^ from P 2 ( t ) , where the players already know that they are in the second regime. The Shapley value is defined by the following formula:
S h i ( 2 ) ( t , P 2 ( t ) ) = S N : i S ( s 1 ) ! ( n s ) ! n ! e r t V 2 ( S ; P 2 ( t ) ) e r t V 2 ( S i ; P 2 ( t ) ) .
By the time-consistent imputation distribution procedure (IDP) in such a subgame, we mean a vector-function β ( 2 ) ( t ) = ( β 1 ( 2 ) ( t ) , , β n ( 2 ) ( t ) ) that, for all t τ ^ and all i N , satisfies the following condition:
S h i ( 2 ) ( t , P 2 ( t ) ) = t β i ( 2 ) ( υ ) d υ .
Here, β i ( 2 ) ( t ) can be interpreted as the instantaneous payment to player i at time t, with t τ ^ . Then, the fulfillment of (13) for every t τ ^ guarantees that, under this imputation distribution, the remaining part of player i’s payoff that has not yet been paid by time t τ ^ coincides with their payoff according to the Shapley value in the subgame starting at that moment t from the corresponding point on the cooperative trajectory. This ensures the time-consistency of the Shapley value in the sense that, at any intermediate stage of the game, players have no incentive to switch to another cooperative optimality principle.
It is well-known (see, for example, Petrosyan & Danilov, 1979) that if S h i ( 2 ) ( t , P 2 ( t ) ) is continuously differentiable, then the IDP
β i ( 2 * ) ( t ) = d S h i ( 2 ) ( t , P 2 ( t ) ) d t
is time-consistent in such a subgame.

4.2. Shapley-Based IDP Before the Switching Time

Now, consider a subgame starting at time t < τ ^ at state P 1 ( t ) , where players do not yet know when the switch will occur.
To distribute the expected joint payoff e r t V 1 ( N ; P 1 ( t ) ) among the players in such subgames, we again use the Shapley value
S h i ( 1 ) ( t , P 1 ( t ) ) = S N : i S ( s 1 ) ! ( n s ) ! n ! e r t V 1 ( S ; P 1 ( t ) ) e r t V 1 ( S i ; P 1 ( t ) ) .
However, building the IDP for the subgames under consideration is more complicated than in the subgames after switching. It should be taken into account that, in the optimization problem, the payoff is represented by the expected value, whereas the actually realized payoffs must be allocated. Therefore, it is necessary to take into account the structure of the expected payoff. For this purpose, we refer to the work (Zaremba, 2022). The expected payoff in such a subgame is computed using the probability that switching occurs after time t, conditional on the switch not having occurred before that time.
The probability that switching occurs before time υ > t , given that switching has not occurred before time t, is given by
F t ( υ ) = F ( υ ) F ( t ) 1 F ( t ) .
Then, the expected payoff of the players in this case is given by
e r t V 1 ( N ; P 1 ( t ) ) = i N t t υ h i 1 ( e i ( 1 ) ( ε ) , P 1 ( ε ) ) d ε + υ h i 2 ( e i ( 2 ) ( ε ) , P 2 ( ε ) ) d ε d F t ( υ ) ] .
Let f ( t ) denote the density function of the switching time. Taking into account that (see Kostyunin & Shevkoplyas, 2011; Zaremba, 2022)
t t υ h i 1 ( e i ( 1 ) ( ε ) , P 1 ( ε ) ) d ε d F t ( υ ) = 1 1 F ( t ) t ( 1 F ( υ ) ) h i 1 ( e i ( 1 ) ( υ ) , P 1 ( υ ) ) d υ
and
t υ h i 2 ( e i ( 2 ) ( ε ) , P 2 ( ε ) ) d ε d F t ( υ ) = 1 1 F ( t ) t f ( υ ) υ h i 2 ( e i ( 2 ) ( ε ) , P 2 ( ε ) ) d ε d υ ,
we derive the formula
e r t V 1 ( N ; P 1 ( t ) ) =           i N 1 1 F ( t ) t ( 1 F ( υ ) ) h i 1 ( e i ( 1 ) ( υ ) , P 1 ( υ ) ) + f ( υ ) υ h i 2 ( e i ( 2 ) ( ε ) , P 2 ( ε ) ) d ε d υ .
Now that the structure of the expected payoff of the players is clear, we can proceed with constructing the imputation distribution procedure.
By the time-consistent imputation distribution procedure in this subgame, we mean a vector-function β ( 1 ) ( t ) = ( β 1 ( 1 ) ( t ) , , β n ( 1 ) ( t ) ) that, for all t < τ ^ and all i N , satisfies the following condition:
S h i ( 1 ) ( t , P 1 ( t ) ) = 1 1 F ( t ) t ( 1 F ( υ ) ) β i ( 1 ) ( υ ) + f ( υ ) υ β i ( 2 ) ( ε ) d ε d υ .
where β ( 2 ) ( t ) is an IDP satisfying condition (13).
Note that β i ( 1 ) ( t ) is interpreted as the instantaneous payment to the player before the switching time, while β i ( 2 ) ( t ) is interpreted as the instantaneous payment after the switch. Then, the expression on the right-hand side of equality (15) gives us the expected payoff of the player in the subgame starting at time t < τ , conditional on no switch having occurred before time t. Moreover, the fulfillment of equality (15) implies that this expected payoff coincides with the payoff of player i in the considered subgame according to the Shapley value.
Applying (13), we get
S h i ( 1 ) ( t , P 1 ( t ) ) = 1 1 F ( t ) t ( 1 F ( υ ) ) β i ( 1 ) ( υ ) + f ( υ ) S h i ( 2 ) ( υ , P 1 ( υ ) ) d υ .
We now show that if S h i ( 1 ) ( t , P 1 ( t ) ) is continuously differentiable, then the vector function β ( 1 * ) ( t ) constructed by the rule
β i ( 1 * ) ( t ) = f ( t ) 1 F ( t ) S h i ( 1 ) ( t , P 1 ( t ) ) S h i ( 2 ) ( t , P 1 ( t ) ) d S h i ( 1 ) ( t , P 1 ( t ) ) d t
is a time-consistent IDP. Note that according to (17), we have
1 1 F ( t ) t ( 1 F ( υ ) ) β i ( 1 * ) ( υ ) + f ( υ ) S h i ( 2 ) ( υ , P 1 ( υ ) ) d υ = 1 1 F ( t ) t f ( υ ) S h i ( 1 ) ( υ , P 1 ( υ ) ) ( 1 F ( υ ) ) d S h i ( 1 ) ( υ , P 1 ( υ ) ) d υ d υ = 1 1 F ( t ) t d ( 1 F ( υ ) ) S h i ( 1 ) ( υ , P 1 ( υ ) ) d υ d υ = S h i ( 1 ) ( t , P 1 ( t ) ) .
Thus, condition (16) is satisfied, and β ( 1 * ) ( t ) is the time-consistent IDP.
Note that
i N h i 1 ( e i ( 1 ) ( t ) , P 1 ( t ) ) = i N β i ( 1 * ) ( t ) ,
i N h i 2 ( e i ( 2 ) ( t ) , P 2 ( t ) ) = i N β i ( 2 * ) ( t ) .
Thus, to maintain time-consistency of the Shapley value, players must redistribute instantaneous payoffs according to the imputation procedures β i ( 1 * ) ( t ) and β i ( 2 * ) ( t ) . Specifically, the procedure β i ( 1 * ) ( t ) applies before the switching time, while β i ( 2 * ) ( t ) is used thereafter.
It should be noted that the Shapley value here can be replaced by any other imputation chosen as the cooperative optimality principle.

4.3. Numerical Illustration of the Characteristic Function and IDP

To illustrate the computation of the characteristic function, the Shapley value, and the corresponding imputation distribution procedure, we consider a three-player example. The parameters are
N = 3 , ( b 1 , b 2 , b 3 ) = ( 12.0 , 11.2 , 10.7 ) ,
r = 0.05 , α = 0.10 , λ = 0.30 , γ 1 = 0.05 , γ 2 = 0.15 ,
with
P 0 = 0.20 , ( k 1 , k 2 , k 3 ) = ( 12.0 , 4.0 , 2.7 ) .
Table 1 reports the values of the characteristic function V 1 ( S ; P 0 ) for all nonempty coalitions. The column K S denotes the aggregate emission of outsiders under the conservative feasible benchmark, and E 1 ( S ) denotes the aggregate emission of coalition members in regime 1.
Table 1. Characteristic-function values V 1 ( S ; P 0 ) in the three-player example.
The grand-coalition value is
V 1 ( N ; P 0 ) = 1678.9730 .
The Shapley value is calculated according to formula (14):
S h ( 1 ) ( 0 , P 0 ) = ( 1121.7504 , 394.1613 , 163.0613 ) .
Player 1 receives the largest Shapley component, which is consistent with their larger marginal contribution to coalition values. Players 2 and 3 receive smaller components, reflecting their lower emission benefits and smaller marginal contributions under the chosen parameters.
We now illustrate the dynamic implementation of this allocation by the IDP. For one realization of the switching time, the simulation gives
τ ^ = 0.510522 .
For the grand coalition, the computed regime-specific emission vectors are
e ( 1 ) = ( 9.6667 , 4.0000 , 2.7000 ) , e ( 2 ) = ( 9.0000 , 4.0000 , 2.7000 ) .
Panels (a)–(c) of Figure 1 plot the IDP payment rates and the corresponding discounted instantaneous payoff flows for players 1, 2, and 3, respectively. For player i, the solid blue curve represents the pre-switch IDP payment rate β i ( 1 * ) ( t ) , while the dashed green curve represents the pre-switch discounted instantaneous payoff flow h i 1 ( e i ( 1 ) ( t ) , P 1 ( t ) ) . After the realized switching time τ ^ , the solid orange curve represents the post-switch payment rate β i ( 2 * ) ( t ) , and the dashed red curve represents the post-switch discounted instantaneous payoff flow h i 2 ( e i ( 2 ) ( t ) , P 2 ( t ) ) .
Figure 1. IDP payment rates and discounted instantaneous payoff flows for the three players. The vertical dashed line marks the realized switching time τ ^ = 0.510522 . Before the switch, the curves correspond to β i ( 1 * ) ( t ) and h i 1 ( e i ( 1 ) ( t ) , P 1 ( t ) ) ; after the switch, they correspond to β i ( 2 * ) ( t ) and h i 2 ( e i ( 2 ) ( t ) , P 2 ( t ) ) , i = 1 , 2 , 3 .
The figures show that the IDP payment rates change at the realized switching time. For all three players, both the payment rates and the instantaneous payoff flows decrease over time, reflecting discounting and the evolution of the pollution stock along the cooperative trajectory.
The comparison between β i ( j * ) ( t ) and h i j ( t ) illustrates that the IDP is not simply a pointwise assignment of each player’s own instantaneous payoff flow. Instead, the payment rate is obtained from the change in the player’s continuation Shapley value. Hence, the IDP implements the Shapley allocation dynamically: at each time, the remaining unpaid payoff is consistent with the Shapley value of the corresponding continuation game. This is precisely the time-consistency requirement imposed by the imputation distribution procedure.

5. Threshold-Triggered Exponential Switch

We extend the baseline model by allowing the switching risk to become active only after the pollution stock reaches a threshold. This captures cases where policy or damages escalate only when pollution becomes sufficiently high, while the implementation still has a random delay.
Fix a pollution threshold P 1 > 0 and define the first hitting time
t 1 : = inf { t 0 : P ( t ) = P 1 } .
Switching cannot occur before the threshold is reached. After reaching P 1 , switching occurs with a constant hazard λ :
τ = t 1 + X , X Exp ( λ ) , λ > 0 .
The damage weight is
γ ( t ) = γ 1 , t < τ , γ 2 , t τ .
where k i > 0 is the emission constraint of player i:
0 e i ( t ) k i , i = 1 , , n , t 0 .

5.1. Post-Threshold Continuation Problem ( t t 1 , Hazard λ Active)

Once the threshold is reached, the model becomes the same constant-hazard switching environment as in Section 2 (now interpreted as a post-threshold subproblem). Therefore, the post-threshold continuation value at the threshold state is
e r t 1 V 1 th ( P 1 ) = e r t 1 ( A 1 + B 1 P 1 ) .

5.2. Pre-Threshold Subproblem on [ 0 , t 1 ] (Hazard 0, Fixed Endpoint P ( t 1 ) = P 1 )

Fix a candidate hitting time t 1 > 0 . Since switching cannot occur before t 1 , the planner solves a finite-horizon control problem with fixed endpoints:
max { e i } t [ 0 , t 1 ] 0 t 1 e r t i = 1 n b i e i ( t ) 1 2 e i ( t ) 2 n γ 1 P ( t ) d t , P ( 0 ) = P 0 , P ( t 1 ) = P 1 ,
subject to P ˙ = i e i α P and the constraint (20). This is a fixed-terminal-state problem with a free terminal time parameter t 1 .
Let the current-value Hamiltonian be
H pre = i = 1 n b i e i 1 2 e i 2 n γ 1 P + μ i = 1 n e i α P .
The unconstrained FOC would give e i = b i + μ ( t ) and the costate equation
μ ˙ ( t ) = ( r + α ) μ ( t ) + n γ 1 .
With the constraint, the implemented control is the projection
e i ( 0 ) ( t ) = Π [ 0 , k i ] ( b i + μ ( t ) ) , i = 1 , , n .
Writing κ : = r + α and d : = n γ 1 κ , the costate has the explicit form
μ ( t ) = c ( t 1 ) e κ t d ,
where the constant c ( t 1 ) is chosen so that the state trajectory under (24) satisfies the fixed endpoint P ( t 1 ) = P 1 .

5.3. Optimal Hitting Time t 1 *

For a given candidate t 1 > 0 , let the pre-threshold value be
J 1 ( t 1 ) : = 0 t 1 e r t i = 1 n b i e i ( 0 ) ( t ) 1 2 ( e i ( 0 ) ( t ) ) 2 n γ 1 P 0 ( t ) d t ,
where P 0 ( t ) is the solution of (1) under controls (24) with P ( 0 ) = P 0 , P ( t 1 ) = P 1 .
The total value associated with t 1 is
W ( t 1 ) = J 1 ( t 1 ) + e r t 1 V 1 th ( P 1 ) .
Hence, the optimal threshold-hitting time is
t 1 * arg max t 1 > 0 W ( t 1 ) .
In computations, we solve for t 1 * by one-dimensional numerical maximization of W ( t 1 ) . This direct approach is robust because the constraints can create kinks (at both 0 and k i ), so W ( t 1 ) is generally only piecewise smooth.
The pre-threshold current-value Hamiltonian at t = t 1 is
H pre ( t 1 ) = i = 1 n b i e i ( 0 ) ( t 1 ) 1 2 ( e i ( 0 ) ( t 1 ) ) 2 n γ 1 P 1 + μ 1 i = 1 n e i ( 0 ) ( t 1 ) α P 1 ,
where μ 1 : = μ ( t 1 ) .
The standard free-terminal-time envelope identity with fixed terminal state P ( t 1 ) = P 1 yields (see, e.g., Seierstad & Sydsæter, 1987)
d d t 1 J 1 ( t 1 ) = e r t 1 H pre ( t 1 ) .
Differentiating (26) and using (27), we obtain
W ( t 1 ) = e r t 1 H pre ( t 1 ) r V 1 th ( P 1 ) .

Diagnostic Equation for t 1 *

On any smooth branch, a necessary condition for an interior maximizer is W ( t 1 ) = 0 . By (28), this is equivalent to
g ( t 1 ) : = H pre ( t 1 ) r V 1 th ( P 1 ) = 0 .
If, in addition, the terminal control is interior for all players,
0 < b i + μ 1 < k i , i = 1 , , n ,
then e i ( 0 ) ( t 1 ) = b i + μ 1 and the Hamiltonian simplifies to
H pre ( t 1 ) = 1 2 S b 2 + S b μ 1 + 1 2 n μ 1 2 n γ 1 P 1 α μ 1 P 1 ,
which gives the explicit interior-branch diagnostic equation
1 2 S b 2 + S b μ 1 + 1 2 n μ 1 2 n γ 1 P 1 α μ 1 P 1 r V 1 th ( P 1 ) = 0 .

5.4. Numerical Illustration

This subsection illustrates how the endogenous hitting time t 1 * and the realized switching time τ = t 1 * + X shape the pollution path in the threshold-triggered model under the emission constraints (20).
For this threshold-triggered path illustration, we consider a five-player example with
N = 5 , ( b i ) i = 1 5 = ( 12.0 , 11.2 , 10.7 , 10.1 , 9.6 ) , r = 0.05 , α = 0.10 , λ = 0.30 , γ 1 = 0.10 , γ 2 = 0.30 ,
and
P 0 = 0.20 , P 1 = 2.00 .
The emission caps are set uniformly as
( k 1 , k 2 , k 3 , k 4 , k 5 ) = ( 3.0 , 3.0 , 3.0 , 3.0 , 3.0 ) .
These caps make the projection operator active in the numerical example. In the post-threshold pre-switch regime, the unconstrained emission vector
( 4.222 , 3.422 , 2.922 , 2.322 , 1.822 )
is projected to
( 3.000 , 3.000 , 2.922 , 2.322 , 1.822 ) .
We maximize W ( t 1 ) numerically to obtain the optimal threshold-hitting time. For the above parameters, the maximization yields
t 1 * 0.327792 , W ( t 1 * ) 117.093 .
Figure 2 illustrates the one-dimensional maximization problem. The objective W ( t 1 ) reaches its maximum near t 1 * , after which the value gradually decreases as the threshold-hitting time is delayed.
Figure 2. The total value W ( t 1 ) under the uniform emission caps k i = 3 . The marked point corresponds to the computed optimum t 1 * 0.327792 , with W ( t 1 * ) 117.093 .
To generate the simulated pollution path in Figure 3, conditional on t 1 * , we draw one realization of X Exp ( λ ) . With the fixed random seed used in the numerical simulation, this gives
X 0.510522 , τ = t 1 * + X 0.838314 .
Figure 3. Simulated pollution path P ( t ) under the uniform emission caps 0 e i ( t ) 3 . The solid blue curve shows the simulated pollution path. The black dashed horizontal line denotes the threshold P 1 = 2.00 , the red dashed vertical line marks the optimal hitting time t 1 * 0.327792 , and the blue dashed vertical line marks the realized switching time τ 0.838314 . The red dot indicates the threshold-hitting point, and the blue dot indicates the pollution stock at the realized switching time.
On [ 0 , t 1 * ) , the hazard is zero and the planner drives the stock to the fixed endpoint P 1 . On [ t 1 * , τ ) , the hazard is active but the regime has not switched yet. At t = τ , the damage weight jumps from γ 1 to γ 2 , while the state P ( t ) remains continuous.

6. Conclusions

This paper studies a pollution control differential game featuring a one-time increase in public damage weight at a random, exponentially distributed switching time. The model is solved using standard continuous-time dynamic programming, resulting in closed-form cooperative feedback policies for each regime in the baseline scenario. The characteristic function is constructed under pointwise emission caps and a conservative outsider benchmark. For the cooperative solution, a time-consistent imputation distribution procedure (IDP) is proposed to implement the Shapley allocation along the stochastic pollution path.
The numerical illustration shows how the characteristic function, the Shapley vector, and the corresponding IDP payment rule can be computed in practice. It also shows how active emission caps affect the constrained cooperative path and how the threshold-triggered mechanism links the switching risk to the pollution stock.
Finally, we extend the model to a threshold-triggered shifted-exponential specification, where switching risk activates only after the pollution stock reaches a threshold. The problem is separated into (i) a post-threshold continuation value solved via the HJB equation and (ii) a pre-threshold fixed-endpoint problem solved using the truncated maximum principle. The optimal hitting time t 1 * is obtained by maximizing the one-dimensional objective W ( t 1 ) , subject to the diagnostic condition g ( t 1 ) = 0 on smooth branches. The threshold and hazard parameters provide a tractable interpretation of ecological or regulatory regime changes.

Author Contributions

Methodology, F.X.; Formal analysis, F.X.; Investigation, F.X.; Writing—original draft preparation, F.X. and A.T.; Writing—review & editing, A.T.; Supervision, A.T. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding. No external funding was received for the APC.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Basar, T., & Olsder, G. J. (1999). Dynamic noncooperative game theory (2nd ed.). SIAM. [Google Scholar] [CrossRef] [Scilit]
  2. Dockner, E. J., Jørgensen, S., Van Long, N., & Sorger, G. (2000). Differential games in economics and management science. Cambridge University Press. [Google Scholar] [CrossRef] [Scilit]
  3. Dockner, E. J., & Long, N. V. (1993). International pollution control: Cooperative versus noncooperative strategies. Journal of Environmental Economics and Management, 25, 13–29. [Google Scholar] [CrossRef] [Scilit]
  4. Fleming, W. H., & Soner, H. M. (2006). Controlled Markov processes and viscosity solutions (2nd ed.). Springer. [Google Scholar] [CrossRef] [Scilit]
  5. Kossioris, G., Plexousakis, M., Xepapadeas, A., de Zeeuw, A., & Mäler, K.-G. (2008). Feedback Nash equilibria for non-linear differential games in pollution control. Journal of Economic Dynamics and Control, 32, 1312–1331. [Google Scholar] [CrossRef] [Scilit]
  6. Kostyunin, S. Y., & Shevkoplyas, E. V. (2011). On simplification of integral payoff in differential games with random duration. Vestnik of St. Petersburg University. Series 10. Applied Mathematics. Informatics. Control Processes, 4, 47–56. Available online: https://www.mathnet.ru/eng/vspui57 (accessed on 24 May 2026).
  7. Lv, S. (2020). Two-player zero-sum stochastic differential games with regime switching. Automatica, 114, 108819. [Google Scholar] [CrossRef] [Scilit]
  8. Mao, X., & Yuan, C. (2006). Stochastic differential equations with Markovian switching. Imperial College Press. [Google Scholar] [CrossRef] [Scilit]
  9. Petrosjan, L. A. (2005). Cooperative differential games. In A. Haurie, S. Muto, L. A. Petrosjan, & T. E. S. Raghavan (Eds.), Advances in dynamic games: Applications to economics, finance, optimization, and stochastic control (pp. 183–200). Birkhäuser. [Google Scholar] [CrossRef] [Scilit]
  10. Petrosyan, L. A., & Danilov, N. N. (1979). Stability of solutions in non-zero sum differential games with transferable payoffs. Vestnik Leningradskogo Universiteta, 1, 52–79. (In Russian) [Google Scholar]
  11. Rubio, S. J., & Casino, B. (2002). A note on cooperative versus non-cooperative strategies in international pollution control. Resource and Energy Economics, 24, 251–261. [Google Scholar] [CrossRef] [Scilit]
  12. Seierstad, A., & Sydsæter, K. (1987). Optimal control theory with economic applications. North-Holland. [Google Scholar]
  13. Shapley, L. S. (1953). A value for n-person games. In H. W. Kuhn, & A. W. Tucker (Eds.), Contributions to the theory of games II (pp. 307–317). Princeton University Press. [Google Scholar] [CrossRef] [Scilit]
  14. Tur, A. V., & Gromova, E. V. (2020). On optimal control of pollution emissions: An example of the largest industrial enterprises of Irkutsk Oblast. Automation and Remote Control, 81, 548–565, (Original work published 2018). [Google Scholar] [CrossRef] [Scilit]
  15. von Neumann, J., & Morgenstern, O. (1953). Theory of games and economic behavior (3rd ed.). Princeton University Press. [Google Scholar]
  16. Yeung, D. W. K., & Petrosyan, L. A. (2006). Cooperative stochastic differential games. Springer. [Google Scholar] [CrossRef] [Scilit]
  17. Yin, G. G., & Zhu, C. (2010). Hybrid switching diffusions: Properties and applications. Springer. [Google Scholar] [CrossRef] [Scilit]
  18. Zaccour, G. (2003). Computation of characteristic function values for linear-state differential games. Journal of Optimization Theory and Applications, 117, 183–194. [Google Scholar] [CrossRef] [Scilit]
  19. Zaremba, A. P. (2022). Cooperative differential games with the utility function switched at a random time moment. Matematicheskaya Teoriya Igr i Prilozheniya, 14, 31–50. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.