Next Article in Journal
Quantum-Verified Environmental Sensing: Integrating Atmospheric Data into Sustainable Finance
Previous Article in Journal
Deep Learning-Based Mapping of Check Dams and Sediment Volume Estimation in Ningxia Province, China
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Transition-Sensitive Congestion Dynamics in Heterogeneous Urban Traffic Networks Under Coordinated Reinforcement Learning

1
Transportation College, Jilin University, Changchun 130022, China
2
School of Electronic Engineering and Computer Science, Peking University, Beijing 100871, China
*
Author to whom correspondence should be addressed.
Sustainability 2026, 18(11), 5561; https://doi.org/10.3390/su18115561
Submission received: 3 May 2026 / Revised: 29 May 2026 / Accepted: 30 May 2026 / Published: 1 June 2026
(This article belongs to the Section Sustainable Transportation)

Abstract

Urban traffic networks under high-demand and incident-like perturbations can evolve from stable operation to cascading congestion, increasing delay, stop-and-go traffic, fuel or energy consumption, and traffic-related emissions. These effects make congestion regulation an important component of sustainable urban traffic management. Existing signal control methods still focus mainly on local delay reduction or short-horizon response, limiting their ability to regulate congestion propagation and stress-induced network degradation. This paper proposes Mamba-PTC, a coordinated reinforcement learning framework for urban signal control in heterogeneous traffic networks. The framework combines centralized multi-intersection control with a simplified Mamba-style sequence encoder and a transition-aware objective optimized by PPO. To connect control with network-level traffic dynamics, we introduce a transition risk indicator for online regulation and macroscopic observables for evaluation, including a composite congestion measure and an instability-amplification proxy. Experiments on stressed heterogeneous urban networks show that Mamba-PTC improves the throughput–duration profile while reducing congestion degradation indicators under heavy load and perturbation. Matched control comparisons, ablation analysis, and cross-network validation further show that these gains arise from the joint effect of temporal representation, transition-aware objective design, and coordinated control. The results suggest that coordinated reinforcement learning can support sustainable network operation by regulating congestion growth in stressed urban traffic networks. The findings provide a basis for designing congestion-aware signal control strategies, robustness evaluation protocols, and future intelligent traffic management systems for stressed urban networks.

1. Introduction

Urban traffic congestion is a persistent obstacle to implementing sustainable urban mobility. When demand exceeds the effective capacity of signalized networks, vehicles experience longer travel times, more frequent stops, idling, and unreliable arrival times. These operating conditions increase fuel or energy consumption and traffic-related emissions while also reducing the service quality and resilience of urban transport systems [1,2,3]. For this reason, traffic signal control is not only an operational efficiency problem, but also an important component of sustainable traffic management. Recent studies in Sustainability have accordingly evaluated intelligent traffic control with respect to delay, fuel or energy use, emissions, and network-level coordination, rather than only local throughput [4,5]. At the same time, urban traffic demand is not purely random: peak hours, weekday–weekend differences, and public holiday patterns create recurrent demand structures that shape the baseline operating conditions of a road network. Empirical evidence has shown systematic differences between public holiday and workday traffic distributions on selected road network elements [6], indicating that signal control strategies should be examined under both recurrent demand pressure and non-recurrent disruptions.
However, congestion in urban networks is not simply a result of the accumulation of local delays at isolated intersections. In heterogeneous networks with irregular junctions, roundabouts, different road capacities, and strong spillback coupling, a local queue can propagate beyond its original bottleneck and gradually affect a much larger part of the network. This network-level degradation is especially important under high-demand or incident-like perturbations, where small disruptions may trigger persistent queue growth and instability amplification. Therefore, sustainable signal control requires not only reducing local delay, but also describing and regulating how congestion develops and spreads under stressed operating conditions.
This motivates a transition-sensitive view of urban congestion. In this study, “transition-sensitive” means that signal control design and evaluation explicitly account for the tendency of an urban traffic network to shift from relatively stable operation toward severe congestion under increasing demand or incident-like perturbations. Operationally, we use the term as a control-oriented, transition-aware perspective that targets congestion growth risk, spillback tendency, and network-level degradation during this shift. Traffic flow and complex network studies show that local vehicle interactions, network topology, and external loading jointly shape macroscopic transport states [7,8,9]. As demand increases, the system may evolve nonlinearly from relatively stable flow to congestion-dominated operation, with strong sensitivity to perturbations and cascade-like degradation under high loading conditions [10,11,12,13]. In heterogeneous urban networks, localized disruptions may therefore trigger congestion propagation well beyond the original bottleneck and produce transition-like deterioration at the network scale [14,15,16]. This perspective provides a useful bridge between sustainable traffic operation objectives and the macroscopic dynamics of congestion growth.
Many existing traffic signal control studies remain organized primarily around engineering performance metrics and controller architectures. Classical responsive strategies such as Max-Pressure [17] and more recent reinforcement learning approaches [18,19] have improved operational efficiency in many settings, but they still focus mainly on local delay, queue balance, or short-horizon responses. In most cases, congestion propagation, instability amplification, and transition-sensitive degradation are not incorporated into the control design through explicit macroscopic observables. As a result, even when modern learning-based controllers improve reported traffic performance, their connection to the network-level evolution of congestion under stressed heterogeneous traffic conditions remains underdeveloped.
Motivated by this gap, this paper studies how coordinated reinforcement learning can be used as a regulation mechanism for transition-sensitive congestion dynamics in stressed heterogeneous urban networks. We develop Mamba-PTC as a spatiotemporal deep reinforcement learning controller that combines a transition-aware objective, macroscopic congestion observables, a task-specific simplified Mamba-style sequence backbone [20], and a centralized coordinated signal control interface. Built on this design, the method aims to improve aggregate traffic performance while also regulating congestion growth and instability under sustained stress. The main contributions of this study are summarized as follows:
  • Transition-aware reward and macroscopic congestion observables: We define a congestion-sensitive reward together with control-oriented macroscopic observables for evaluation, so that the controller can be trained with transition-aware signals and assessed in terms of network-level degradation under stressed traffic conditions.
  • History-aware spatiotemporal representation of congestion evolution: We employ the Mamba-PTC backbone as the main spatiotemporal feature extractor. Implemented as a task-specific simplified Mamba-style encoder, it provides a history-aware representation of congestion evolution under strong perturbations and is better aligned with the temporally accumulated dynamics of heterogeneous traffic flow.
  • Coordinated multi-intersection regulation of congestion propagation: To translate the Mamba-PTC spatiotemporal representation into effective intervention, we formulate urban traffic signal control as a centralized joint decision problem over multiple signalized intersections. The controller receives a network-level observation and outputs coordinated signal–actions across the controlled intersections, enabling the policy to respond to congestion propagation through inter-intersection coordination rather than isolated local reaction.
The remainder of this paper is organized as follows: Section 2 reviews prior studies on signal control objectives, congestion transition views, spatiotemporal modeling, and realistic topology validation. Section 3 presents the observation design, reward formulation, macroscopic observables, and Mamba-PTC policy architecture. Section 4 describes the simulation environment, perturbation design, comparison groups, training protocol, and evaluation metrics. Section 5 reports the training behavior, stress-response comparisons, ablation studies, extended seed analysis, and cross-network validation. Section 6 interprets the findings and discusses practical feasibility. Section 7 summarizes the main conclusions and future research directions.

2. Related Work

To position the present study more clearly, this section reviews prior work from four aspects that match the motivation given in the Introduction: traffic signal control objectives and coordination structure, macroscopic views of congestion transition, spatiotemporal sequence modeling, and validation on realistic heterogeneous topologies.

2.1. Traffic Signal Control Objectives and Coordination Structure

Traffic signal control provides the main operational lever through which recurrent queues, spillback, and intersection-level delay can be regulated in urban networks. Early methods focused on stable and interpretable timing rules. Fixed-time methods such as Webster’s method [21] perform reasonably well under relatively stable flow conditions, while adaptive systems such as SCOOT [22] and SCATS [23] improve responsiveness by adjusting signal settings according to detector information. These approaches established the operational foundation of urban signal control, but their objectives are usually defined in terms of local delay, queue discharge, or corridor-level responsiveness.
As network interactions became more important, coordination-oriented objectives received increasing attention. Max-Pressure control [17] became a strong decentralized baseline because it uses queue differentials to promote network throughput, and PressLight later incorporated max-pressure intuition into a reinforcement learning framework [18]. With the development of deep reinforcement learning, high-dimensional traffic states and larger action spaces became more tractable. The DQN framework introduced by Mnih et al. [24] provided a basis for value-based control, and Li et al. [25] demonstrated the effectiveness of DRL for signalized intersection control. Subsequent studies further explored movement-centric signal modeling [26,27], network-level cooperation and communication [28,29,30], and spatial extensions of pressure-based control [31]. These works improved adaptivity and coordination, but their evaluation still tends to emphasize aggregate performance indicators such as throughput, delay, queue length, or waiting time.
Recent studies have further broadened the objective of signal control beyond conventional delay and queue reduction. Related work has examined signal optimization for delay, fuel consumption, and emissions [1], reinforcement learning control for reducing vehicle emissions [2], CO2-oriented DRL control in mixed traffic flow [3], and priority-metric signal optimization based on real-time-measured traffic information [32]. Other studies have extended this line to hierarchical multi-agent DRL and coordinated multi-intersection control [4,5]. These works show that signal control objectives are becoming broader and more system-oriented. However, even in these settings, congestion propagation and transition-sensitive network degradation are still rarely represented through explicit macroscopic observables. This gap motivates the present study: the controller should not only improve conventional traffic metrics, but should also be evaluated by how it affects the growth and amplification of congestion under network stress.

2.2. Macroscopic Views of Congestion Transition

Macroscopic traffic flow and complex network studies provide a useful way to describe how local vehicle interactions, bottlenecks, and network topology produce network-level congestion states [7,8]. Earlier work examined traffic instability, metastability, and congestion onset in traffic networks [10,11]. Subsequent studies further showed that urban traffic degradation can exhibit percolation-like transition behavior, macroscopic collapse, and switching between different congested regimes under increasing stress [12,13,14,33]. These findings suggest that traffic degradation should not be understood only as accumulated local delay, but also as a collective network process with transition-sensitive behavior.
Complex network studies have further examined urban congestion through topological complexity, congestion-warning communities, cascading degradation, topology-dependent traffic performance, and large-scale spatiotemporal congestion patterns. Prior work has examined the topological characteristics and flow complexity of urban traffic congestion [34], identified congestion-warning communities in heterogeneous traffic networks [35], analyzed cascading dynamics in urban road systems [36], and studied how road network structure shapes the performance of urban traffic systems [37]. Other studies also investigated spatial heterogeneity and migration characteristics of traffic congestion from large-scale trajectory data [38], while recent work uncovered interpretable spatiotemporal congestion patterns through a complex network approach [39]. Together, these works support the interpretation of congestion as a network-level dynamical process rather than merely as isolated local delay.
These studies provide important insight into congestion propagation, topology-dependent vulnerability, and transition-like degradation, but their integration into closed-loop traffic signal control remains limited. Most DRL-based traffic controllers still rely mainly on local observables such as queue length, waiting time, or phase pressure. By contrast, macroscopic congestion observables can provide a more informative description of system-level congestion risk and degradation. In this study, our goal is not to construct a strict phase-transition model. Instead, we adopt a control-oriented framing in which congestion observables are used to describe and regulate network-level degradation under stress.

2.3. Spatiotemporal Sequence Modeling for Traffic Control

Capturing congestion propagation across a road network requires models that can represent network-scale spatiotemporal dependence. Earlier traffic forecasting and control studies often used recurrent architectures such as LSTMs to process sequential traffic states [40]. These models are effective for short- to medium-range sequence modeling, but their long-range memory is limited and their parallel efficiency is relatively low. Graph-based and attention-based approaches later strengthened the modeling of inter-intersection interaction and nonlocal dependency, while coordination-aware methods such as Wang et al. [29] further emphasized scalable large-network cooperation. Transformer-style architectures [41,42] made global dependency modeling more explicit. Recent studies have further broadened this direction through deep spatiotemporal graph experts, multi-view attention, and federated graph learning for traffic state prediction [43,44,45]. Related traffic management studies have also coupled travel time prediction with route optimization and distributed multi-agent signal coordination [46,47]. Recent multi-agent signal control studies also suggest that coordinated learning architectures are needed when the control objective shifts from isolated intersection response toward network-level spillback mitigation [4,5].
However, congestion evolution in heterogeneous urban networks often combines slowly accumulated spillback structure with rapidly fluctuating local disturbances. In such settings, the main challenge is not only long-sequence representation, but also the causal accumulation of traffic history relevant to transition-sensitive degradation. These recent studies show the value of richer temporal, topological, and cooperative traffic representations, but most of them focus on prediction, route guidance, or intersection-level coordination rather than on embedding history-aware representation inside a closed-loop signal control policy and evaluating macroscopic degradation under matched stressed traffic realizations. More recently, Gu and Dao [20] proposed Mamba, which is based on selective state-space models and offers a promising alternative for long-sequence modeling through linear-time recursive processing. For online traffic signal control, this direction is attractive because congestion evolution often shows long memory, sensitivity to perturbations, and non-Markovian propagation. In this context, we use a simplified Mamba-style module as a history-aware state encoder within the reinforcement learning controller, linking recent congestion evolution to coordinated signal decisions for regulating spillback and network-level degradation.

2.4. Validation on Realistic Heterogeneous Topologies

A substantial portion of traffic RL research is still evaluated on synthetic grids or Manhattan-style layouts, where the topology is relatively regular and interactions are easier to isolate [19]. In contrast, real urban networks contain irregular junctions, roundabouts, heterogeneous road capacities, and stronger spillback coupling. These structural heterogeneities are not merely engineering complications; they also influence congestion propagation pathways, vulnerability to non-recurrent disruption, and the growth of delay-related externalities [35,37,39].
To examine algorithmic robustness under realistic topological stress, this study adopts the “Joined Network” scenario from the DLR-TS Bologna dataset released by Bieker et al. [48] as the primary comparison setting. This network combines signalized urban grids with non-Euclidean structures such as roundabouts and irregular multi-branch intersections. It therefore provides a more demanding test bed for studying network-level congestion regulation than standard synthetic layouts. In addition, experiments are conducted on additional network settings to examine whether the observed performance trends remain qualitatively stable beyond a single topology.
Taken together, the literature shows that current traffic signal control research has made substantial progress in coordination structure, movement-level inductive bias, and sequence representation, while traffic flow and complex network studies have clarified the importance of congestion propagation, topology-dependent degradation, and transition-sensitive dynamics. However, the joint consideration of macroscopic congestion observables, simplified state-space sequence modeling, and coordinated signal control under stressed heterogeneous urban topologies remains limited. In this context, the present study develops a coordinated signal control framework that combines transition-sensitive observables with the Mamba-PTC backbone under a unified stress-testing setting.

3. Methodology

This section presents the methodological framework of Mamba-PTC. We first formulate the coordinated traffic signal control problem and define the observation, action, and reward components. We then introduce control-oriented macroscopic observables for characterizing transition-sensitive congestion evolution and network-level degradation under stress, and finally describe the Mamba-PTC spatiotemporal control framework optimized with PPO [49,50].

3.1. Problem Formulation, Observation, Action, and Reward Design

We formulate the task as a partially observable Markov decision process (POMDP) [51] for coordinated urban traffic signal control. The policy takes a centralized network-level observation, including a whole-network microscopic trajectory sequence and auxiliary traffic descriptors, and outputs coordinated signal decisions for the controlled intersections. In this paper, transition-sensitive congestion regulation means using signal control to slow congestion growth and reduce the likelihood that the network moves from relatively stable operation toward severe congestion under sustained load and perturbation. To keep the formulation physically interpretable, the main state variables, reference quantities, and control thresholds are initialized from traffic flow and complex network intuition and then kept fixed in all reported experiments.

3.1.1. Microscopic Observation and Transition Risk Descriptor

The controller uses a centralized microscopic kinematic observation X t as its main input. Instead of relying only on aggregated quantities such as queue length or waiting time, this observation keeps vehicle-level motion information at the network level and allows the Mamba-PTC backbone to capture the temporal evolution of congestion. This design preserves richer traffic dynamics while remaining tractable for coordinated online signal control.
  • Microscopic spatiotemporal trajectory tensor ( X t ).
At each decision step, the environment builds a microscopic spatiotemporal trajectory tensor from the active vehicles in the controlled network. For each selected vehicle, the recent kinematic states within a fixed time window are collected and arranged into a tensor of size
X t R T w i n × N k e y × F ,
where T w i n is the temporal window length, N k e y is the maximum number of tracked active vehicles, and F = 5 is the feature dimension, including normalized position, velocity, acceleration, heading angle, and lane index.
To keep the observation consistent across time steps, the buffered sequence is aligned by vehicle identity whenever a vehicle remains active in consecutive frames. If the number of active vehicles exceeds the observation budget, the excess vehicles are truncated. If some historical entries are missing, they are filled by current snapshot carryover or zero padding. In this way, X t provides a fixed-size network-level microscopic observation that can be processed directly by the policy backbone.
To limit the observation size, we use a fixed vehicle budget N k e y instead of the full active vehicle set. This keeps the network-level observation compact and makes large-scale roll-out training computationally feasible.
We set N k e y using a simple capacity-based estimate. For a representative topology with L = 8 lanes, a perception radius of R o b s = 150   m , and an effective vehicle length of l e f 7.5   m , the corresponding jam-scale capacity is approximated by
N j a m = L · R o b s l e f 160 .
Since congestion transitions usually emerge before full jam saturation [7,10], we set
N k e y 0.5 × N j a m = 80 .
This choice keeps the observation size manageable while retaining the microscopic traffic states that are most relevant to coordinated control.
  • Transition-risk indicator ( ϕ t ).
To provide the controller with a compact online descriptor of congestion risk, we define a transition risk indicator ϕ t [ 0 , 1 ] from two factors: traffic accumulation and speed reduction,
ϕ t = w ρ min 1 , N t N s a t + w v max 0 , 1 v t v r e f ,
where w ρ and w v are the weights for the accumulation term and the speed term, respectively. Here, N t is the number of active vehicles in the monitored junction-centered neighborhood at the current decision step, and v t is the mean speed of the same vehicle set. The first term increases when accumulation approaches saturation, and the second term increases when the average speed falls below a reference level. Accordingly, a larger ϕ t indicates a higher risk that the traffic system is moving away from relatively stable operation toward severe congestion.
This type of descriptor is consistent with recent studies that use parsimonious macroscopic quantities to characterize congestion propagation, warning patterns, and network-level traffic states in urban traffic systems [4,15,35,52].
The thresholds are chosen as follows:
  • N s a t = 40 is a saturation-level accumulation threshold motivated by MFD-style network accumulation analysis [53].
  • v r e f = 5.0   m / s is a reference speed for the onset of strongly slowed urban traffic [10].
In this paper, ϕ t is used as a control-oriented online descriptor of transition-sensitive degradation rather than as a universal traffic state variable.
At each decision step, the policy receives a centralized observation consisting of the trajectory tensor X t and the scalar risk indicator ϕ t [ 0 , 1 ] . In implementation, the environment operates on the same capped active vehicle set used to build X t , which keeps the observation interface fixed and reproducible while providing both microscopic trajectory information and a compact summary of congestion risk.

3.1.2. Joint Signal–Action Space and Switching Regularization

To model coordinated control over multiple intersections, we define a discrete joint signal–action space
A = A 1 × × A D ,
where D is the number of controlled signalized intersections in the network and A d is the admissible signal–action set for intersection d. At each decision step, the controller outputs a joint action
a t = ( a t , 1 , , a t , D ) A ,
where a t , d denotes the signal decision applied to intersection d. The scenario-specific value of D is given later in the experimental setup.
To discourage excessive retiming, we regularize the number of intersections whose signal decision changes between two consecutive control steps. Specifically, we define
N s w t = d = 1 D I ( a t , d a t 1 , d ) ,
where N s w t is the number of controlled intersections whose signal–action is switched at the current decision step.
This regularization encourages parsimonious coordinated control: the controller is still allowed to adapt signals when traffic conditions change, but unnecessary switching across the network is penalized.

3.1.3. Reward Design

The reward is designed to balance three coupled objectives: maintaining transport efficiency, avoiding unnecessarily aggressive coordinated switching, and penalizing traffic states with elevated transition-sensitive congestion risk. At decision step t, the reward is defined as
R t = R e f f i c i e n c y E c t r l ( a t ) P ( ϕ t ) .
The first term encourages effective traffic discharge, the second regularizes coordinated control effort, and the third penalizes states that indicate increasing risk of network-level degradation.
  • Efficiency term ( R e f f i c i e n c y ).
R e f f i c i e n c y = w 1 · N o u t , t w 2 · W ¯ s y s w 3 · Q s y s ,
where N o u t , t is the number of vehicles that complete their trips at the current control step, W ¯ s y s is the mean waiting time over the monitored controlled lanes, and Q s y s is the number of active vehicles currently remaining in the network. This term encourages the policy to improve throughput while reducing delay and network accumulation.
  • Switching penalty ( E c t r l ( a t ) ).
E c t r l ( a t ) = w 4 · N s w t ,
where N s w t is the number of controlled intersections whose signal decision changes at the current step. This penalty discourages excessive coordinated retiming and helps keep the control policy stable.
  • Risk penalty ( P ( ϕ t ) ).
P ( ϕ t ) = w 5 · ϕ t ,
where ϕ t is the transition risk indicator defined above. This term penalizes traffic states with high congestion transition risk and encourages the controller to act before severe congestion fully develops.
The reward terms are used only for online policy optimization. They are distinct from the trip-based post hoc evaluation metrics reported later in the experiments.

3.2. Macroscopic Observables for Transition-Sensitive Congestion Evaluation

Macroscopic descriptions of traffic state and congestion diffusion have increasingly been used to evaluate network degradation, identify stressed traffic regimes, and compare traffic control outcomes under different operating conditions [4,52,53]. This subsection defines the macroscopic observables used to characterize congestion evolution in the reported experiments. We first introduce a calibrated low-stress reference state, and then define the composite congestion observable Φ together with an instability-amplifying proxy χ . These quantities are not used as direct online control inputs; instead, they are used to evaluate how different controllers affect network-level degradation and transition-sensitive congestion growth under stressed traffic conditions.

3.2.1. Calibration of the Free-Flow Reference State

The evaluation metrics defined below are expressed relative to a baseline traffic state. We therefore first calibrate a free-flow reference state, which is taken as a low-stress operating regime used to normalize the congestion-related quantities. We estimate free flow empirically from a low-load, low-incident reference setting and record two baseline quantities:
  • Reference trip duration ( T r e f ): the mean travel time under the free-flow reference state.
  • Reference waiting time ( W r e f ): the irreducible signal-related waiting component under the same reference state.
These two reference quantities are kept fixed during evaluation and are used to measure how far the system deviates from stable transport conditions. Because W r e f is estimated from an empirical nonzero operating regime rather than from an ideal zero-delay limit, the normalization remains bounded away from zero in the reported evaluation setting. In this study, the calibrated reference values are T r e f = 5.916 min and W r e f = 35.819 s, and the same values are used for all compared methods, flow scales, perturbation levels, and seed settings.
Using a low-stress empirical reference serves three purposes. First, it makes the duration-based and waiting-based contributions dimensionless and therefore comparable across stress slices with different absolute scales. Second, it anchors the observables to a reproducible operating regime of a signalized urban network rather than to an unattainable zero-delay idealization. Third, it allows the resulting quantities to be interpreted as coarse-grained departures from relatively stable operation rather than as raw performance scores. In this sense, the reference-state normalization is introduced to support comparative analysis of degradation under stress, not to define a universal physical baseline.

3.2.2. Composite Congestion Observable and Instability-Amplifying Proxy

To evaluate congestion severity under different control settings, we introduce two related macroscopic quantities based on the calibrated low-stress reference state. The composite congestion observable Φ summarizes the overall departure of the traffic system from stable operation by combining travel efficiency loss and waiting-related blockage, while the instability-amplifying proxy χ further emphasizes degradation accompanied by large signal-induced waiting, which is closely associated with queue accumulation and spillback pressure. In defining Φ , the duration-based and waiting-based terms are normalized differently because they capture different traffic mechanisms. The duration-based term is normalized by the observed trip duration, so it represents the congestion-related share of actual travel time and remains stable when absolute trip duration varies across stress regimes. By contrast, the waiting-based term is normalized by the low-stress reference waiting time W r e f , because signal-related waiting is itself a blockage indicator and its relative increase over the reference state directly reflects queueing and spillback formation. These two normalizations allow Φ to combine travel efficiency degradation and waiting-related blockage in a dimensionless form while preserving their distinct physical meanings.
  • Duration-based congestion rate ( I d u r ).
I d u r = max 0 , T ¯ o b s T r e f T ¯ o b s
where T ¯ o b s is the observed mean trip duration over the evaluation horizon and T r e f is the reference trip duration defined above. This term measures the fraction of realized travel time that can be attributed to congestion-induced delay.
  • Waiting-based congestion rate ( I w a i t ).
I w a i t = max 0 , W ¯ o b s W r e f W r e f
where W ¯ o b s is the observed mean waiting time and W r e f is the reference waiting time. This term captures the growth of signal-related waiting relative to the low-stress reference regime.
  • Composite congestion observable ( Φ ).
Φ = 1 2 ( I d u r + I w a i t )
Under this definition, Φ 0 corresponds to a near-reference operating state, whereas larger Φ indicates a stronger departure from stable transport conditions. In the reported experiments, Φ is used to compare how different controllers affect the combined growth of travel efficiency loss and waiting-related blockage under the same stress-testing protocol. Φ is introduced here as a compact macroscopic observable of degradation, in the same spirit as recent traffic management studies that summarize system state through aggregate performance indicators or spillback-related measures [1,4,52].
  • Instability-amplification proxy ( χ ).
We define another proxy that reflects how strongly network-level degradation is amplified by waiting-related stress. The instability-amplifying proxy χ is defined as
χ = Φ · 1 + W ¯ o b s 100
where W ¯ o b s is the observed mean waiting time in seconds. The multiplicative factor increases χ when the composite congestion level is accompanied by a substantial waiting burden. The constant 100 s is used as a fixed waiting time scale that converts the observed waiting burden into a moderate dimensionless amplification factor, so that χ remains comparable across stress slices while still responding to severe queueing conditions. In signalized networks, large waiting time is closely related to queue accumulation, blocked discharge, and spillback pressure. Therefore, χ highlights degradation states in which the departure from stable operation is amplified by waiting-related blockage, rather than reflected only in longer trip duration. This interpretation is consistent with studies that emphasize congestion warning, cascading degradation, spillback mitigation, and macroscopic signatures of network-wide instability [5,35,36,52].

3.3. Mamba-PTC Framework

An overview of the Mamba-PTC framework is shown in Figure 1. At each decision step, the framework takes a centralized microscopic observation and a transition risk descriptor as input, encodes them with the Mamba-PTC backbone, and uses a PPO-based actor–critic policy to generate coordinated signal decisions under the physics-guided reward. In this sense, the framework can be viewed as a coordinated regulation mechanism acting on transition-sensitive congestion evolution in a heterogeneous urban traffic network.

3.3.1. Framework Overview

The framework can be described as a three-stage process. First, the traffic environment constructs the microscopic trajectory tensor X t and the transition risk descriptor ϕ t from the current network state. Second, the Mamba-PTC backbone transforms these inputs into a compact feature representation that summarizes recent congestion evolution. Third, the PPO-based actor–critic policy maps the encoded feature to coordinated signal decisions for the controlled intersections and is optimized with the physics-guided reward defined in Section 3.1.3.

3.3.2. Sequence Processing Motivation

Effective traffic signal control requires more than the current traffic snapshot. In heterogeneous urban networks, congestion propagation is temporally persistent, sensitive to perturbations, and only partially observable at each decision step. Under such conditions, the relevant traffic state is not purely instantaneous; it depends on the recent accumulation of spillback, delayed discharge, and local disturbances across the network. For this reason, the policy backbone should be able to accumulate traffic evidence over time rather than respond only to instantaneous local states.
A state-space sequence model provides a natural way to represent this type of causal temporal accumulation. For a single feature channel u ( t ) , the hidden state h ( t ) evolves as
h ˙ ( t ) = A h ( t ) + B u ( t ) .
Under a sampling interval Δ , a standard zero-order-hold discretization yields the recursive form
h k = A ¯ h k 1 + B ¯ u k ,
where A ¯ and B ¯ are the discretized transition matrices. This recurrence motivates the use of state-space-style sequence processing in the present framework, because it preserves causal temporal accumulation while avoiding the quadratic sequence cost associated with full-attention-based processing.
Based on this motivation, the proposed framework uses the Mamba-PTC backbone to model the temporal evolution of network traffic before the encoded feature is passed to the control policy.

3.3.3. Mamba-PTC Backbone

To encode spatiotemporal features under PPO-scale roll-outs, we design the Mamba-PTC backbone as a task-specific sequence module tailored to coordinated traffic signal control and inspired by the sequence processing idea of Mamba [20].
  • Spatial flattening and projection.
At each time step, the whole-network microscopic observation tensor is flattened across the tracked vehicles and projected into a hidden feature space:
u t = Linear p r o j vec ( X t ) .
This operation produces a compact network-level representation that can be processed efficiently during large-scale reinforcement learning.
  • Causal local temporal filtering.
Following the input projection, the feature stream is divided into a content branch and a gating branch. On the content branch, we apply a causal one-dimensional convolution followed by a SiLU activation:
x ˜ t = SiLU Conv 1 d ( x 1 : t ) t .
This step acts as a learnable local temporal filter that suppresses short-term fluctuations while preserving trend information relevant to congestion propagation.
  • Gated cumulative state update.
To maintain linear-time sequence processing, we use a simplified input-dependent gated accumulation mechanism:
d t = σ Linear Δ Linear x ( x ˜ t ) ,
h t = h t 1 + x ˜ t d t ,
with output
y t = h t SiLU ( z t ) ,
where z t denotes the projected gating branch. In implementation, this update is equivalent to an input-gated cumulative scan along the time axis and is then mapped to the policy feature vector through an output projection.
This simplified operator preserves the properties most relevant to the present task, including causal temporal accumulation, input-dependent gating, and linear-time sequence processing. These properties make it suitable for history-aware traffic representation under reinforcement learning roll-outs in heterogeneous urban networks. Accordingly, the Mamba-PTC backbone should be understood as a task-specific simplified Mamba-style encoder designed for coordinated traffic signal control, rather than as a literal module-by-module reproduction of the original selective-scan architecture.

3.3.4. PPO-Based Policy Learning

The encoded sequence feature is fed into an actor–critic policy and optimized end-to-end with Proximal Policy Optimization (PPO) [49]. In the present framework, PPO is used to learn coordinated signal decisions from high-dimensional and partially observable traffic observations under the physics-guided reward defined in Section 3.1.3.
This learning setup allows the spatiotemporal encoder and the control policy to be trained jointly within a single reinforcement learning framework. PPO is adopted here because it supports stable on-policy optimization while remaining practical for large-scale traffic control roll-outs under high-variance stressed dynamics.

3.3.5. Network Architecture Instantiation

To instantiate the proposed framework in the reported setting, each temporal frame of the microscopic observation is flattened to dimension 400 and projected to a hidden width of d = 512 . The Mamba-PTC backbone is implemented as a single task-specific simplified Mamba-style block with a depthwise Conv1d local temporal filter of kernel size 3, an input-dependent gated recursive update, and an output projection back to dimension 512. No LayerNorm is used in the present implementation.
The backbone output is fused with the transition risk descriptor and mapped to a 512-dimensional extractor representation. Thus, the “512-d” label in Figure 1 refers to the final extractor output rather than only to the hidden width inside the sequence block. On top of this representation, PPO uses separate actor and critic multilayer perceptrons with hidden widths [ 512 , 512 ] . In the reported signal control setting, the final policy head is MultiDiscrete, with one categorical branch for each controlled intersection.

4. Experimental Setup

We evaluate Mamba-PTC using Eclipse SUMO version 1.24.0 [54] with the standardized SUMO-RL interface version 1.4.5 [55]. This section describes the simulation environment, the perturbation modeling, the shared signal control interface, the common training protocol, and the stress-testing program used to examine transition-sensitive congestion regulation under heterogeneous urban traffic conditions. The experiments focus on mobility and congestion degradation indicators; explicit fuel, energy, or emission models are left for future work.

4.1. Simulation Environment and Controlled Network

To test the proposed framework under realistic topological stress, we use the Joined Network from the DLR-TS Bologna scenario family [48] as the primary evaluation setting. Unlike regular synthetic grids, this network combines heterogeneous urban structures, including standard signalized intersections, roundabouts, and irregular multi-branch junctions, which substantially increase spillback coupling and the difficulty of congestion propagation control. This setting therefore provides a demanding test bed for studying network-level congestion regulation and transition-sensitive degradation in non-Euclidean urban traffic systems.
In the reported setting, the directly controlled elements are the signalized intersections in the network. For the Joined Network setting used in the main experiments, the number of controlled signalized intersections is D = 29 . The coordinated signal control action is therefore instantiated over these 29 controlled nodes. Roundabouts and unsignalized structures remain part of the traffic-propagation dynamics, but they are not directly actuated by the controller. The signal control interface is updated at a fixed decision interval of 15 s, and standard transition constraints are enforced through a 2 s yellow phase and a 5 s minimum green duration. These constraints keep the reported setting within the operational scope of coordinated urban signal control, rather than broader simulator-side intervention.
Vehicle routes follow the dataset-provided SUMO-RL scenario configurations for the Joined, Acosta, and Pasubio Networks. This preserves the released demand profiles and movement patterns of these real-network scenarios. Within each stress slice, all compared controllers use the same route realization and departures; demand scaling λ is handled uniformly in the stress-testing matrix below.
Figure 2 shows the topology of the Bologna Joined Network used as the primary evaluation setting. In addition to this main scenario, we also evaluate the framework on the Acosta and Pasubio [48] Networks later in the paper as a cross-network validation under the same stress-testing protocol.

4.2. Perturbation Modeling

To evaluate whether the controller can still respond effectively under non-recurrent disruptions, we introduce explicit perturbations into the SUMO environment. The purpose is to create controlled local blockages that can trigger queue growth, spillback, and congestion propagation. The perturbation levels follow a monotonic stress-gradient design based on traffic operational mechanisms. The design links perturbation severity to three factors that directly affect spillback risk: local disruption occupancy, queue-accumulation time, and the recovery opportunity before the next disruption. This stress-testing choice is consistent with recent work that evaluates intelligent signal control under accident-like or incident-driven scenarios [56]. In this way, the experiments probe how different controllers regulate stressed traffic dynamics rather than only how they perform under near-stationary demand conditions.
In the simulator, each perturbation is implemented as a scheduled temporary blockage at the vehicle level. At a trigger time, a prescribed number of vehicles near junction centers are selected, their speeds are set to zero, and they remain blocked for a specified duration. This design creates localized flow interruption and network-scale stress without changing the road topology itself. Operationally, the perturbations are implemented in SUMO as deterministic vehicle-level blockage events within the active time window of 200–5800 s. At each trigger slot, a predefined number of vehicles are temporarily halted for the configured hold duration and then released. The three levels are chosen as a mechanism-based incident stress gradient: the low level represents short local obstruction with sufficient recovery opportunity, the middle level increases both blockage occupancy and queue-retention time, and the high level combines repeated triggers, longer holding duration, and shorter recovery intervals to create sustained spillback pressure. This design provides a controlled representation of urban incident-like disruptions while keeping disruption timing and severity matched across controllers within each stress slice.
The perturbation schedule is defined by the parameter tuple
( N e v t , Δ e v t , H e v t , [ t s t a r t , t e n d ] ) ,
where N e v t is the number of blocked vehicles triggered in each event, Δ e v t is the interval between two consecutive events, H e v t is the blockage duration, and [ t s t a r t , t e n d ] is the active time window of the perturbation schedule. Increasing N e v t raises the local blockage intensity, increasing H e v t allows queues to accumulate for a longer time, and decreasing Δ e v t reduces the time available for the network to dissipate the previous disruption.
Different perturbation levels therefore form an ordered stress gradient while keeping the topology, demand pattern, and signal control scope unchanged.

4.3. Comparison Groups and Control Scope

To clarify the comparison scope, the reported experiments are organized into three comparison tracks: practical adaptive baselines, shared-interface learned baselines, and movement-centric reference controllers. Across all tracks, the perturbation schedule, demand scaling settings, and signal control layer are kept consistent so that the reported differences can be interpreted under the same stressed traffic realization.
Within this structure, practical adaptive baselines include Max-Pressure and PressLight [18], with fixed-time control reported separately in the framework-level comparison. Shared-interface learned baselines include Wang et al. [29], LSTM [40], and Transformer [41,42]; this group is used for backbone-attribution analysis. Movement-centric reference controllers include FRAP [26] and MoveLight [27], which are treated as controller-level references because their internal policy structures are method-specific.

4.4. Training Protocol and Baseline Tuning

Given the large state–action space of the Joined Network, we adopt an end-to-end distributed PPO training protocol with parallel simulation for the learning-based controllers. All learned controllers were trained on a server equipped with dual Intel Xeon Gold 6148 CPUs (2.40 GHz, 40 physical cores/80 threads), 128 GB RAM, and one NVIDIA GeForce RTX 4090 GPU. This hardware setup was sufficient to support the parallel PPO training and large-scale SUMO roll-out collection reported in this paper.
For the shared-interface learned baselines, baseline tuning follows a common training protocol: the environment configuration, observation schema, signal–action interface, reward definition, PPO optimizer, training budget, extractor output width, and actor–critic head widths are held fixed, while only the representation backbone is changed. Mamba-PTC and the shared-interface learned baselines are trained from scratch and selected under the same validation rule. Practical adaptive baselines keep their established control formulations, including queue-pressure or pressure-based decision rules, while movement-centric baselines keep their method-specific policy structures and movement-level inductive biases. These baselines are therefore tuned and reported according to their intended controller families, and their evaluation conditions are matched later through the common route files, demand scaling, perturbation schedule, signal control layer, and seed protocol.
  • Parallel sampling.
We launch 72 independent SUMO environments to reduce temporal correlation among samples and improve gradient estimation stability. For the main Joined Network training runs, the total interaction horizon reaches approximately 5.5 × 10 6 decision timesteps, where one timestep corresponds to one policy decision interval of 15 s in simulation time.
  • Optimization.
Policy learning is performed with PPO [49]. The main hyperparameters are summarized in Table 1. Only the settings most relevant to reproducibility and control behavior are retained here.
  • Checkpoint selection.
For the main learned controllers, model selection is based on the shared validation protocol under matched departures and perturbations, rather than on manually choosing a visually favorable run. This rule is kept consistent across the shared-interface comparisons and, where applicable, the movement-centric controller-level comparison so that the reported differences are not driven by inconsistent checkpoint selection.
These hyperparameters serve three roles in the reported experiments. The PPO optimization settings, including the learning rate, roll-out length, batch size, and update epochs, determine how data are collected and how policy updates are performed during training. The risk calibration and reward weight parameters specify how congestion risk, waiting cost, queue-related penalty, switching cost, and completion reward are balanced in the training objective. The remaining policy-stabilization settings, namely γ , λ GAE , the clip range, and the entropy coefficient, are used to keep policy learning stable under high-variance traffic dynamics.

4.5. Stress-Testing Matrix and Evaluation Slices

Based on the perturbation mechanism defined above, we construct a 5 × 3 factorial stress matrix to examine how different controllers respond as the traffic system is driven from moderate load toward increasingly stressed regimes.
In the stress tests, traffic load is controlled by the demand scaling factor λ . This factor is implemented through SUMO’s built-in demand scaling applied to the fixed evaluation route files. The demand scaling dimension represents recurrent mobility pressure, while the perturbation dimension represents non-recurrent disruptions. For paired comparisons, controllers in the same flow–perturbation slice use the same route files, flow scale, perturbation schedule, and evaluation seed. The primary matrix uses five matched seeds ( M = 5 , seeds 42–46), while the extended seed protocol uses ten matched seeds on selected high-severity slices ( M = 10 , seeds 47–56). The resulting settings are summarized in Table 2.
Accordingly, the experiments are organized into training-stability analysis, stress-response comparisons on the Joined Network, ablation studies, and generalization tests under different evaluation settings.

4.6. Evaluation Metrics and Reporting Protocol

The primary operational metrics are throughput N out and mean trip duration T ¯ . Mean waiting time W ¯ is reported as an additional operational indicator to characterize waiting-related trade-offs among controllers. These metrics are relevant to sustainable traffic operation because prolonged travel, waiting, and stopping are associated with inefficient vehicle operation and have been used in emissions-aware signal control analyses [57]. Alongside these operational metrics, the macroscopic quantities Φ and χ are reported as transition-sensitive observables for characterizing network-level degradation and instability amplification under increasing load and perturbation.
Accordingly, the main empirical claims are based on paired comparisons under matched departures and matched perturbations, so that the reported differences can be attributed to controller behavior rather than unrelated stochastic variation. The results tables therefore emphasize pooled means together with matched-run robustness indicators. Where paired sign tests are reported, they are two-sided non-parametric checks of whether the direction of improvement is consistent across matched traffic realizations, while the extended seed analysis characterizes run-to-run variability on selected high-stress slices within the same evaluation protocol. In the tables below, ↑ and ↓ indicate whether higher or lower values are preferred, respectively, and bold values identify the best result for the corresponding metric within each comparison group.

5. Experimental Results

This section reports the empirical evidence in four steps. Section 5.1 first examines whether Mamba-PTC shows stable and progressively improving training behavior on the heterogeneous Joined Network. Section 5.2 then presents the main results on the Joined Network from the perspective of transition-sensitive congestion dynamics under coordinated regulation. Section 5.3 separates the effect of the framework itself from that of the reward design. Section 5.4 finally checks whether the main qualitative trends remain visible under extended seeds and on other road networks.

5.1. Training Stability of Mamba-PTC

Before comparing traffic control performance across baselines, we first examine whether Mamba-PTC can be trained stably on the heterogeneous Joined Network. This step is necessary because the training environment is high-dimensional, strongly non-stationary, and subject to stressed traffic dynamics.
Figure 3 reports the evolution of the episode reward during training. Since the reward is the quantity directly optimized by the policy, it provides the clearest training-side evidence of whether the controller is progressively learning a better operating strategy.
The curve shows a clear three-stage pattern. In the early stage, the reward rises rapidly from about 4000 to around 2500 , indicating that the policy quickly escapes the poorest exploration regime. In the middle stage, the reward continues to improve overall while exhibiting noticeable oscillations, especially between approximately 0.5 × 10 6 and 1.5 × 10 6 timesteps, where it moves upward from the 2500 range toward roughly 1500 to 1300 . In the later stage, after roughly 1.5 × 10 6 timesteps, the trajectory enters a relatively stable high-reward regime and fluctuates mainly between about 1400 and 1100 .
Overall, Figure 3 shows that Mamba-PTC can be trained without evident optimization collapse and that the learned policy progressively reaches a stable and substantially improved operating regime. This provides the training basis for the stressed regime congestion analyses reported in the following subsections.

5.2. Stress-Response of Network Degradation Observables and Supporting Comparisons

This section examines how the heterogeneous Joined Network responds as demand and perturbation push traffic from stable operation toward congestion growth. The main question is whether a controller improves throughput and trip duration while also limiting propagation, degradation, and instability amplification under the same stressed realization. Matched control comparisons provide the primary evidence because the observation, action, reward, and training interfaces are aligned; practical controller and movement-centric comparisons are retained as supporting operational results.

5.2.1. Practical Controller Baselines: Max-Pressure and PressLight

This subsection evaluates whether Mamba-PTC can regulate stressed network dynamics more effectively than practical baselines with clear operational meaning. Table 3 reports the direct pooled comparison, and Table 4 reports the robustness check separately.
Table 3 gives the direct practical controller comparison. Relative to Max-Pressure, Mamba-PTC increases throughput from 592.44 to 664.80 vehicles, an increase of 72.36 vehicles or about 12.21 % . At the same time, mean trip duration decreases from 6.10 to 5.44 min, mean waiting time decreases from 210.53 to 174.53 s, Φ decreases from 2.460 to 1.937 , and χ decreases from 8.829 to 3.368 . The comparison with PressLight shows larger gains in throughput and trip duration: throughput rises from 548.60 to 664.80 vehicles ( + 21.18 % ), mean trip duration decreases from 6.58 to 5.44 min ( 17.33 % ), and mean waiting time decreases from 230.24 to 174.53 s. In the macroscopic observables, Mamba-PTC also attains the lowest Φ in this practical controller track, while its χ remains close to the best baseline value ( 3.368 versus 3.046 for PressLight). Taken together, these results show that the advantage of Mamba-PTC is not limited to a single operational metric, but remains visible in both transport performance and the reported macroscopic observables in this comparison setting.
Table 4 shows further checks of whether the practical controller gains are stable at the paired-run level. The mean effects remain positive after accounting for run-to-run variation: against Max-Pressure, Mamba-PTC improves throughput by 72.36 vehicles and reduces duration by 0.67 min; against PressLight, the corresponding gains are 116.20 vehicles and 1.14 min. This indicates that the practical controller advantage is not only visible in pooled averages, but also remains consistent under matched traffic realizations.

5.2.2. Matched Control Evidence for Suppression of Congestion Amplification

This subsection uses a stricter matched control comparison: the environment, observation format, signal control interface, PPO optimizer, and reward definition are fixed, and only the learned backbone is changed. The resulting comparison is therefore used to examine whether different sequence representations lead to different regulation of transition-sensitive congestion growth under the same stressed traffic dynamics.
Table 5 provides the main matched-RL evidence in compact form. Across the full set of 15 slices, Mamba-PTC attains the highest mean throughput and the shortest mean trip duration. At the slice level, it achieves the best throughput in 14 of the 15 matched scenarios and the shortest trip duration in 14 of the 15 scenarios; the only exception in both metrics is the ( λ = 1.4 , high ) slice, where LSTM is marginally better.
The regime-wise summary also clarifies where the advantage becomes most visible. When only the high-perturbation branch is retained, the throughput gap of Mamba-PTC relative to LSTM widens to about 39.9 vehicles on average, while the mean trip duration decreases by about 0.42 min. On the two most-stressed slices, ( λ = 1.8 , high ) and ( λ = 2.0 , high ) , the separation grows further: Mamba-PTC reaches an average throughput of 631.5 vehicles, compared with 559.7 for LSTM and 550.6 for Wang et al., while mean trip duration falls to 5.71 min, compared with 6.46 and 6.56 min. This pattern indicates that the backbone advantage is not expressed as a uniform shift at all operating points, but becomes clearer as the system is pushed deeper into the stressed regime.
Table 6 reports the matched-RL comparison with the Transformer baseline. Relative to the Transformer baseline, Mamba-PTC increases throughput by 28.74 vehicles, reduces trip duration by 0.24 min, reduces waiting time by 38.55 s, and lowers both Φ and χ . This result is consistent with the regime-level summary above and suggests not only better operational performance, but also a weaker departure from the low-stress reference regime and reduced instability amplification under the same control interface.
To make the role of the macroscopic observables explicit, Table 7 reports Φ and χ on the two most-stressed slices of the matched-RL comparison, where the throughput–duration separation is already strongest in Table 5. At ( λ = 1.8 , high ) , Mamba-PTC reduces Φ from 2.686 and 2.756 to 1.782 , and reduces χ from 9.565 and 10.056 to 4.834 , relative to LSTM and Wang et al., respectively. At ( λ = 2.0 , high ) , Φ decreases from 2.364 and 3.390 to 1.991 , while χ decreases from 7.998 and 14.111 to 5.647 . These observable-level differences show that, once the network enters the severe stressed regime, the throughput–duration advantage of Mamba-PTC is accompanied by a weaker departure from the low-stress reference regime and lower instability amplification under the same control interface.
The robustness test in Table 8 confirms the same pattern at the matched-run level. Mamba-PTC wins 88.0 % of the runs against LSTM on both throughput and duration, and 93.3 % against Wang et al. on the same two metrics. Against the Transformer baseline, the win rates are 78.67 % for throughput, 82.67 % for duration, and 88.0 % for waiting time. The paired statistics show the same direction: the mean gains are 53.43 , 64.35 , and 32.11 vehicles in throughput, and 0.48 , 0.60 , and 0.34 min in duration against LSTM, Wang et al., and Transformer, respectively. Taken together with Table 5, Table 6 and Table 7, these results indicate that, once the control interface is matched, Mamba-PTC provides the strongest regulation effect in the stressed regime, with the clearest separation appearing near the transition from moderate to severe congestion.

5.2.3. Supporting Comparison with Movement-Centric Controllers

This subsection provides a controller-level comparison with FRAP and MoveLight, which are stronger movement-centric signal control methods with their own inductive biases. Because these methods are not matched backbone replacements, the comparison is interpreted mainly at the operational level.
Table 9 shows that Mamba-PTC attains the strongest throughput–duration profile in this controller-level comparison, while FRAP and MoveLight remain more competitive on waiting time. This comparison should not be interpreted as uniform superiority on all metrics. Instead, it indicates that the main strength of Mamba-PTC lies in network-level discharge efficiency and trip completion under stress, whereas movement-centric controllers may retain an advantage on some waiting-related slices.
The robustness test in Table 10 confirms this interpretation. Mamba-PTC shows a stable throughput–duration advantage. Compared with FRAP, the paired gains are 159.00 vehicles and 1.73 min; compared with MoveLight, they are 85.47 vehicles and 0.79 min. Accordingly, this controller-level comparison is used to show the strongest throughput–duration profile among the tested controller families, rather than to claim that Mamba-PTC dominates every metric against FRAP and MoveLight.

5.3. Ablation Study

This section explains where the gains in Section 5.2 come from. The first ablation asks whether the learning-based Mamba-PTC framework itself is better than a fixed signal plan on the same network. The second keeps the framework fixed and asks whether the reward design is responsible for the observed improvement.

5.3.1. Framework-Level Comparison with Fixed-Time Control

This ablation addresses a different question from the comparison analyses above. The issue here is whether the overall learning-based Mamba-PTC framework provides a stronger control regime than a non-learning fixed signal plan on the same road network.
Figure 4, together with Table 11 and Table 12, clarifies the first ablation result at the framework level. The figure shows that, across the tested flow range and under all three perturbation levels, Mamba-PTC generally remains above fixed-time in throughput and below fixed-time in mean trip duration, indicating that the advantage of the learning-based framework is not confined to a single operating point. This visual pattern is consistent with the pooled summary in Table 11, where Mamba-PTC raises throughput from 601.23 to 664.80 vehicles, reduces mean trip duration from 6.02 to 5.44 min, and reduces mean waiting time from 246.56 to 174.53 s. Importantly, the same pooled comparison also shows a clear reduction in the macroscopic degradation observables, with Φ decreasing from 2.903 to 1.913 and χ decreasing from 10.061 to 5.251 . This means that the framework-level gain is not only an operational improvement, but is also accompanied by lower values of the macroscopic degradation observables under the same network stress.

5.3.2. Reward Design Ablation on Matched High-Stress Slices

The second ablation keeps the network architecture, training protocol, and evaluation setting fixed and changes only the reward design. The purpose is to test whether the transition risk term and the switching regularization term contribute useful improvements beyond pure efficiency shaping, especially in terms of the resulting degradation profile under stressed traffic conditions.
Table 13 shows that the reward design changes the control profile in a systematic way. The efficiency-only variant attains almost the same throughput as the full objective ( 608.00 versus 607.67 ), so the difference in discharge is negligible. However, the full objective reduces waiting time from 117.55 to 77.44 s, a drop of 40.11 s or about 34.12 % . It also reduces Φ from 1.306 to 0.747 and χ from 3.163 to 1.329 , corresponding to reductions of about 42.80 % and 57.98 % . Once the risk term is removed, throughput falls to 501.67 and waiting time rises to 197.40 s. Once the switching regularization is removed, the collapse is even stronger: throughput drops to 304.33 , which is about 49.92 % lower than the full objective, while the mean waiting time increases to 311.18 s. Therefore, the reward ablation does not merely show that the full objective is better overall; it shows that the additional transition risk and switching-related terms help prevent the controller from exchanging short-term efficiency for more unstable stressed regime behavior, as reflected in both the operational metrics and the transition-sensitive observables.

5.4. Generalization Under Different Evaluation Settings

After establishing the main Joined Network analysis and the two ablation layers, this section examines whether the same qualitative pattern of transition-sensitive congestion regulation remains visible under different evaluation settings. LSTM is used here as the reference learned baseline because it is the strongest non-Mamba baseline in the matched-RL comparison and is also closest to Mamba-PTC in model scale. The first validation increases the seed budget on the high-stress slices. The second validation transfers the protocol from the Joined Network to the heterogeneous Acosta and Pasubio Networks.
Table 14 and Table 15 show a consistent stressed regime trend across the extended seed high-stress slices. The per-flow means in Table 14 show that Mamba-PTC maintains higher throughput and shorter trip duration from λ = 1.6 to λ = 2.0 . When these high-stress slices are pooled in Table 15, the paired mean gains remain positive across all reported metrics: throughput increases by 27.10 vehicles, mean trip duration decreases by 0.258 min, mean waiting time decreases by 54.84 s, Φ decreases by 0.774 , and χ decreases by 4.391 . The positive-pair ratios range from 21 / 30 to 23 / 30 , and the sign-test results remain below 0.05 , indicating that the high-stress advantage is visible across seeds rather than being driven by a single flow level.

Cross-Network Validation on Acosta and Pasubio

The cross-network validation transfers the same 5 × 3 stress matrix from the Joined Network analysis to two heterogeneous road networks, Acosta ( D = 16 ) and Pasubio ( D = 15 ). In this validation layer, Mamba-PTC is compared with LSTM, which is selected as the strongest overall baseline in the Joined Network experiments. Each network contains 15 stress scenarios and 75 paired runs.
Table 16 and Table 17 show a consistent cross-network trend. On Acosta, Mamba-PTC improves from 393.41 to 455.28 in N o u t (+15.73%), while reducing T ¯ from 9.206 to 7.953 min (−13.61%) and W ¯ from 208.89 to 161.73 s (−22.43%). On Pasubio, gains are smaller but remain positive: N o u t increases from 513.75 to 530.15 (+3.19%), with T ¯ and W ¯ reduced by 3.12% and 13.79%. Overall, the pooled cross-network results show that Mamba-PTC maintains a favorable transfer pattern across heterogeneous road networks: mean N o u t increases from 453.58 to 492.72 (+8.63%), mean T ¯ decreases from 8.117 to 7.381 min (−9.07%), and mean W ¯ decreases from 214.89 to 176.07 s (−18.06%). These results indicate that the observed advantage is not confined to the primary Joined Network setting, although its magnitude remains topology-dependent.
Taken together, the evidence indicates that Mamba-PTC is most useful when the network moves from moderate load into sustained high-stress operation. Its main advantage appears in the throughput–duration profile, while the matched control, ablation, extended seed, and cross-network results suggest weaker degradation under stress and point to the joint contribution of temporal representation, transition-aware reward design, and coordinated control.

6. Discussion

The results suggest that the advantage of Mamba-PTC becomes clearer when the traffic system enters a sustained high-stress regime. Under these conditions, congestion is not determined only by the current local state, but also by the recent evolution of spillback, delayed discharge, and inter-intersection interactions. High-stress traffic states contain delayed effects: blocked vehicles form queues, queues reduce discharge at neighboring intersections, and the resulting spillback may appear several decision steps after the original disruption. A controller based mainly on instantaneous local response may therefore react after the network has already entered a degraded state. This provides a plausible explanation for why the proposed framework shows more stable gains in the stressed slices: compared with methods that rely mainly on instantaneous response, it is better able to use temporally accumulated traffic information for the coordinated regulation of congestion growth. Mamba-PTC combines this history-aware representation with a transition risk penalty and coordinated multi-intersection actions, making the policy more likely to preserve network discharge before local disruption develops into broader congestion amplification.
The matched comparisons sharpen this interpretation. Because the environment, observation format, action interface, PPO setup, and reward definition are fixed, the remaining difference is mainly the sequence representation. The better throughput–duration profile and lower transition-sensitive observables therefore suggest that Mamba-PTC benefits from representing recent traffic evolution more effectively under stressed dynamics.
At the same time, the results do not support a claim of uniform superiority on all metrics or across all controller families. The movement-centric comparisons show that the main strength of Mamba-PTC lies in network-level discharge efficiency and trip duration, while waiting time performance remains more method-dependent. In connection to this, the transition-sensitive observables are most informative when they are used to interpret degradation under matched or closely comparable control settings, rather than as the sole ranking criterion across all heterogeneous controller families. The ablation and robustness results point in the same direction. The reported gain appears to come from the combined effect of temporal representation, transition-aware objective design, and coordinated control, rather than from any single component alone.
Taken together, the evidence suggests that Mamba-PTC is most useful in operating regimes where congestion growth has a clear temporal and network-coupled structure. In the tested heterogeneous networks, this is associated with better throughput retention, shorter trip duration, and weaker transition-sensitive degradation under stress. From this perspective, the main value of the framework lies not only in improved control performance, but also in linking coordinated signal control with the macroscopic regulation of congestion growth in a stressed urban traffic network.
Regarding practical feasibility, the centralized microscopic observation used in Mamba-PTC should be interpreted as an information-rich benchmark rather than a required field-deployment condition. In real traffic networks, such a state description could only be approximated in well-instrumented areas through connected-vehicle trajectories, roadside detectors, camera- or radar-based tracking, and data-fusion or state-estimation modules. Therefore, the present experiments mainly test the control value of rich network-state information under stress; practical deployment would require replacing the full microscopic observation with partial and noisy estimates and evaluating performance under different sensing-penetration levels.
From the perspective of sustainable transport networks, the practical implication is that a controller that weakens congestion growth can also reduce the operating conditions that normally increase fuel or energy consumption and emissions, such as excessive waiting, repeated stopping, and network-wide spillback [1,2,3,57]. The present experiments do not directly estimate pollutant emissions; therefore, the sustainability interpretation should be understood as an operational pathway rather than a direct emissions inventory. A natural next step is to couple the proposed transition-sensitive control framework with SUMO-based emission models so that congestion regulation, energy consumption, and CO2 or pollutant emissions can be evaluated within the same stress-testing matrix.

7. Conclusions

This paper investigated transition-sensitive congestion dynamics in heterogeneous urban traffic networks under coordinated reinforcement learning. Mamba-PTC combines a transition-aware control objective, recursively accumulated spatiotemporal representation, and coordinated signal control within a unified reinforcement learning framework. Across the Joined Network stress tests and the Acosta–Pasubio validation, the method attains a favorable throughput–duration profile under stress and, under matched control conditions, produces lower macroscopic degradation observables. The robustness and ablation analyses further suggest that these gains arise from the joint effect of representation, objective design, and coordinated control. Accordingly, the main contribution of the present study is not only a coordinated reinforcement learning framework for urban signal control, but also a control formulation that links network-level signal coordination with the macroscopic regulation of congestion growth in stressed heterogeneous urban networks.
Several limitations define the scope of the present evidence and indicate the next steps. The current observation model still assumes relatively reliable microscopic trajectory access; in real-world deployment, however, sparse sensor coverage, limited connected-vehicle penetration, and partial connected-vehicle information may reduce the quality and availability of such observations [58,59,60]. Thus, field use would not require observing every vehicle in exactly the same form as the simulator; instead, an implementable system would need to approximate the centralized state through connected-vehicle samples, roadside sensing, and state estimation. The signal control setting remains a simulator-level abstraction of field deployment, and the observables Φ and χ are used here as control-oriented macroscopic indicators of degradation and instability amplification rather than as universal measures of congestion transition. In addition, although the current validation spans two heterogeneous urban networks, it is not yet a broad multi-topology study. The Acosta and Pasubio experiments should therefore be viewed as preliminary cross-network validation rather than as evidence of full multi-city generalization. Accordingly, the most relevant follow-up directions are stricter operational comparisons under more tightly matched signal–transition rules, improved realism under partial observability, and broader evaluation across additional irregular urban networks. Future work should further examine online adaptation, transfer across cities, and integration with connected-vehicle systems so that the framework can operate under changing demand, incomplete observations, and heterogeneous sensing infrastructure. This study also evaluates sustainability through congestion-related operational proxies; future work should incorporate explicit fuel, energy, and emissions models to quantify the environmental benefits of transition-sensitive coordinated control more directly.

Author Contributions

Conceptualization, Z.O.; methodology, Z.O. and C.L.; software, Z.O.; formal analysis, Z.O.; writing—original draft preparation, Z.O.; writing—review and editing, Z.O., C.L., Y.T., Y.S., Z.W., Y.M. and T.D.; supervision, C.L. and T.D.; funding acquisition, T.D. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Pilot Research Project for Building China into a Transport Power, grant number 2025ZDGC-06.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are available on request from the corresponding author.

Conflicts of Interest

The authors declare no conflict of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

References

  1. Al-Turki, M.; Jamal, A.; Al-Ahmadi, H.M.; Al-Sughaiyer, M.A.; Zahid, M. On the Potential Impacts of Smart Traffic Control for Delay, Fuel Energy Consumption, and Emissions: An NSGA-II-Based Optimization Case Study from Dhahran, Saudi Arabia. Sustainability 2020, 12, 7394. [Google Scholar] [CrossRef]
  2. Kovári, B.; Szoke, L.; Bécsi, T.; Aradi, S.; Gáspár, P. Traffic Signal Control via Reinforcement Learning for Reducing Global Vehicle Emission. Sustainability 2021, 13, 11254. [Google Scholar] [CrossRef]
  3. Wang, Z.; Xu, L.; Ma, J. Carbon Dioxide Emission Reduction-Oriented Optimal Control of Traffic Signals in Mixed Traffic Flow Based on Deep Reinforcement Learning. Sustainability 2023, 15, 16564. [Google Scholar] [CrossRef]
  4. Cao, Q.; Li, J.; Trucco, P. Sustainability-Oriented Urban Traffic System Optimization Through a Hierarchical Multi-Agent Deep Reinforcement Learning Framework. Sustainability 2026, 18, 1606. [Google Scholar] [CrossRef]
  5. Ma, L.; Liu, Y.; Liu, Y.; Ma, C.; Wang, S. Coordinated Multi-Intersection Traffic Signal Control Using a Policy-Regulated Deep Q-Network. Sustainability 2026, 18, 1510. [Google Scholar] [CrossRef]
  6. Macioszek, E.; Kurek, A. Road Traffic Distribution on Public Holidays and Workdays on Selected Road Transport Networks Elements. Transp. Probl. 2021, 16, 127–138. [Google Scholar] [CrossRef]
  7. Helbing, D. Traffic and Related Self-Driven Many-Particle Systems. Rev. Mod. Phys. 2001, 73, 1067–1141. [Google Scholar] [CrossRef]
  8. Chowdhury, D.; Santen, L.; Schadschneider, A. Statistical Physics of Vehicular Traffic and Some Related Systems. Phys. Rep. 2000, 329, 199–329. [Google Scholar] [CrossRef]
  9. Schadschneider, A. Traffic Flow: A Statistical Physics Point of View. Phys. A 2002, 313, 153–187. [Google Scholar] [CrossRef]
  10. Kerner, B.S. The Physics of Traffic: Empirical Freeway Pattern Features, Engineering Applications, and Theory; Springer: Berlin/Heidelberg, Germany, 2004. [Google Scholar] [CrossRef]
  11. Zhao, L.; Lai, Y.-C.; Park, K.; Ye, N. Onset of Traffic Congestion in Complex Networks. Phys. Rev. E 2005, 71, 026125. [Google Scholar] [CrossRef] [PubMed]
  12. Li, D.; Fu, B.; Wang, Y.; Lu, G.; Berezin, Y.; Stanley, H.E.; Havlin, S. Percolation Transition in Dynamical Traffic Network with Evolving Critical Bottlenecks. Proc. Natl. Acad. Sci. USA 2015, 112, 669–672. [Google Scholar] [CrossRef]
  13. Olmos, L.E.; Colak, S.; Shafiei, S.; Saberi, M.; Gonzalez, M.C. Macroscopic Dynamics and the Collapse of Urban Traffic. Proc. Natl. Acad. Sci. USA 2018, 115, 12654–12661. [Google Scholar] [CrossRef] [PubMed]
  14. Zeng, G.; Li, D.; Guo, S.; Gao, L.; Gao, Z.; Stanley, H.E.; Havlin, S. Switch between Critical Percolation Modes in City Traffic Dynamics. Proc. Natl. Acad. Sci. USA 2019, 116, 23–28. [Google Scholar] [CrossRef]
  15. Saberi, M.; Hamedmoghadam, H.; Ashfaq, M.; Hosseini, S.A.; Gu, Z.; Shafiei, S.; Nair, D.J.; Dixit, V.; Gardner, L.; Waller, S.T.; et al. A Simple Contagion Process Describes Spreading of Traffic Jams in Urban Networks. Nat. Commun. 2020, 11, 1616. [Google Scholar] [CrossRef]
  16. Mendes, G.A.; da Silva, L.R.; Herrmann, H.J. Traffic Gridlock on Complex Networks. Phys. A 2012, 391, 362–370. [Google Scholar] [CrossRef]
  17. Varaiya, P. Max Pressure Control of a Network of Signalized Intersections. Transp. Res. Part C Emerg. Technol. 2013, 36, 177–195. [Google Scholar] [CrossRef]
  18. Wei, H.; Chen, C.; Zheng, G.; Wu, K.; Gayah, V.; Xu, K.; Li, Z. PressLight: Learning Max Pressure Control to Coordinate Traffic Signals in Arterial Network. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Anchorage, AK, USA, 4–8 August 2019; pp. 1290–1298. [Google Scholar] [CrossRef]
  19. Noaeen, M.; Naik, A.; Goodman, L.; Crebo, J.; Abrar, T.; Abad, Z.S.H.; Bazzan, A.L.C.; Far, B. Reinforcement Learning in Urban Network Traffic Signal Control: A Systematic Literature Review. Expert Syst. Appl. 2022, 199, 116830. [Google Scholar] [CrossRef]
  20. Gu, A.; Dao, T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar] [CrossRef]
  21. Webster, F.V. Traffic Signal Settings; Road Research Technical Paper No. 39; Road Research Laboratory, H.M.S.O.: London, UK, 1958. [Google Scholar]
  22. Robertson, D.I.; Bretherton, R.D. Optimizing Networks of Traffic Signals in Real Time: The SCOOT Method. IEEE Trans. Veh. Technol. 1991, 40, 11–15. [Google Scholar] [CrossRef]
  23. Lowrie, P.R. The Sydney Co-ordinated Adaptive Traffic System: Principles, Methodology, Algorithms. IEE Conf. Publ. 1982, 207, 67–70. [Google Scholar]
  24. Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-Level Control through Deep Reinforcement Learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef]
  25. Li, L.; Lv, Y.; Wang, F.-Y. Traffic Signal Timing via Deep Reinforcement Learning. IEEE/CAA J. Autom. Sin. 2016, 3, 247–254. [Google Scholar] [CrossRef]
  26. Zheng, G.; Xiong, Y.; Zang, X.; Feng, J.; Wei, H.; Zhang, H.; Li, Y.; Xu, K.; Li, Z. Learning Phase Competition for Traffic Signal Control. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management, Beijing, China, 3–7 November 2019; pp. 1963–1972. [Google Scholar] [CrossRef]
  27. Shao, J.; Zheng, C.; Chen, Y.; Huang, Y.; Zhang, R. MoveLight: Enhancing Traffic Signal Control through Movement-Centric Deep Reinforcement Learning. arXiv 2024, arXiv:2407.17303. [Google Scholar] [CrossRef]
  28. Wei, H.; Xu, N.; Zhang, H.; Zheng, G.; Zang, X.; Chen, C.; Zhang, W.; Zhu, Y.; Xu, K.; Li, Z. CoLight: Learning Network-Level Cooperation for Traffic Signal Control. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management, Beijing, China, 3–7 November 2019; pp. 1913–1922. [Google Scholar] [CrossRef]
  29. Wang, T.; Cao, J.; Hussain, A. Adaptive Traffic Signal Control for Large-Scale Scenario with Cooperative Group-Based Multi-Agent Reinforcement Learning. Transp. Res. Part C Emerg. Technol. 2021, 125, 103046. [Google Scholar] [CrossRef]
  30. Wang, W.; Zhang, H.; Qiao, T.; Ma, J.; Jin, J.; Li, Z.; Wu, W.; Jiang, Y. Real-Time Network-Level Traffic Signal Control: An Explicit Multiagent Coordination Method. IEEE Trans. Intell. Transp. Syst. 2024, 25, 19688–19698. [Google Scholar] [CrossRef]
  31. Li, X.; Wang, X.; Smirnov, I.; Sanner, S.; Abdulhai, B. Multi-Hop Upstream Anticipatory Traffic Signal Control with Deep Reinforcement Learning. IEEE Open J. Intell. Transp. Syst. 2025, 6, 554–567. [Google Scholar] [CrossRef]
  32. Kim, M.; Schrader, M.; Yoon, H.-S.; Bittle, J.A. Optimal Traffic Signal Control Using Priority Metric Based on Real-Time Measured Traffic Information. Sustainability 2023, 15, 7637. [Google Scholar] [CrossRef]
  33. Lampo, A.; Borge-Holthoefer, J.; Gomez, S.; Sole-Ribalta, A. Multiple Abrupt Phase Transitions in Urban Transport Congestion. Phys. Rev. Res. 2021, 3, 013267. [Google Scholar] [CrossRef]
  34. Wen, T.-H.; Chin, W.-C.-B.; Lai, P.-C. Understanding the Topological Characteristics and Flow Complexity of Urban Traffic Congestion. Phys. A 2017, 473, 166–177. [Google Scholar] [CrossRef]
  35. Guo, Y.; Yang, L.; Hao, S.; Gao, J. Dynamic Identification of Urban Traffic Congestion Warning Communities in Heterogeneous Networks. Phys. A 2019, 522, 98–111. [Google Scholar] [CrossRef]
  36. Yin, R.-R.; Yuan, H.; Wang, J.; Zhao, N.; Liu, L. Modeling and Analyzing Cascading Dynamics of the Urban Road Traffic Network. Phys. A 2021, 566, 125600. [Google Scholar] [CrossRef]
  37. Wu, C.-Y.; Hu, M.-B.; Jiang, R.; Hao, Q.-Y. Effects of Road Network Structure on the Performance of Urban Traffic Systems. Phys. A 2021, 563, 125361. [Google Scholar] [CrossRef]
  38. Fu, X.; Xu, C.; Liu, Y.; Chen, C.-H.; Hwang, F.J.; Wang, J. Spatial Heterogeneity and Migration Characteristics of Traffic Congestion—A Quantitative Identification Method Based on Taxi Trajectory Data. Phys. A 2022, 588, 126482. [Google Scholar] [CrossRef]
  39. Zeng, J.; Xiong, Y.; Liu, F.; Ye, J.; Tang, J. Uncovering the Spatiotemporal Patterns of Traffic Congestion from Large-Scale Trajectory Data: A Complex Network Approach. Phys. A 2022, 604, 127871. [Google Scholar] [CrossRef]
  40. Ma, X.; Tao, Z.; Wang, Y.; Yu, H.; Wang, Y. Long Short-Term Memory Neural Network for Traffic Speed Prediction Using Remote Microwave Sensor Data. Transp. Res. Part C Emerg. Technol. 2015, 54, 187–197. [Google Scholar] [CrossRef]
  41. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. In Advances in Neural Information Processing Systems 30; Curran Associates, Inc.: Red Hook, NY, USA, 2017; pp. 5998–6008. [Google Scholar]
  42. Xu, M.; Dai, W.; Liu, C.; Gao, X.; Lin, W.; Qi, G.-J.; Xiong, H. Spatial-Temporal Transformer Networks for Traffic Flow Forecasting. arXiv 2020, arXiv:2001.02908. [Google Scholar] [CrossRef]
  43. Edalatpanah, S.A.; Pourqasem, J. DSTGN-ExpertNet: A Deep Spatio-Temporal Graph Neural Network for High-Precision Traffic Forecasting. Mechatron. Intell. Transp. Syst. 2025, 4, 28–40. [Google Scholar] [CrossRef]
  44. Wu, F.; Zheng, C.; Zhang, C.; Ma, J.; Sun, K. Multi-View Multi-Attention Graph Neural Network for Traffic Flow Forecasting. Appl. Sci. 2023, 13, 711. [Google Scholar] [CrossRef]
  45. Feng, J.; Du, C.; Mu, Q. Traffic Flow Prediction Based on Federated Learning and Spatio-Temporal Graph Neural Networks. ISPRS Int. J. Geo-Inf. 2024, 13, 210. [Google Scholar] [CrossRef]
  46. Aljanabi, M.R.; Kadhim, A.R.; Hammoodi, M.R.; Borna, K. Hybrid Computational-Intelligence Framework for Dynamic Travel-Time Prediction and Route Optimization in Traqi Urban Transportation Networks. Int. J. Comput. Methods Exp. Meas. 2025, 13, 597–611. [Google Scholar] [CrossRef]
  47. Yan, L.; Wang, J. Deep Reinforcement Learning for Ecological and Distributed Urban Traffic Signal Control with Multi-Agent Equilibrium Decision Making. Electronics 2024, 13, 1910. [Google Scholar] [CrossRef]
  48. Bieker, L.; Krajzewicz, D.; Morra, A.P.; Michelacci, C.; Cartolano, F. Traffic Simulation for All: A Real World Traffic Scenario from the City of Bologna. In Modeling Mobility with Open Data; Behrisch, M., Weber, M., Eds.; Springer: Cham, Switzerland, 2015; pp. 47–60. [Google Scholar] [CrossRef]
  49. Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef]
  50. Engstrom, L.; Ilyas, A.; Santurkar, S.; Tsipras, D.; Janoos, F.; Rudolph, L.; Madry, A. Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO. arXiv 2020, arXiv:2005.12729. [Google Scholar] [CrossRef]
  51. Kaelbling, L.P.; Littman, M.L.; Cassandra, A.R. Planning and Acting in Partially Observable Stochastic Domains. Artif. Intell. 1998, 101, 99–134. [Google Scholar] [CrossRef]
  52. Ambühl, L.; Menendez, M.; González, M.C. Understanding Congestion Propagation by Combining Percolation Theory with the Macroscopic Fundamental Diagram. Commun. Phys. 2023, 6, 26. [Google Scholar] [CrossRef]
  53. Geroliminis, N.; Daganzo, C.F. Existence of Urban-Scale Macroscopic Fundamental Diagrams: Some Experimental Findings. Transp. Res. Part B Methodol. 2008, 42, 759–770. [Google Scholar] [CrossRef]
  54. Lopez, P.A.; Wiessner, E.; Behrisch, M.; Bieker-Walz, L.; Erdmann, J.; Flotterod, Y.-P.; Hilbrich, R.; Lucken, L.; Rummel, J.; Wagner, P. Microscopic Traffic Simulation Using SUMO. In Proceedings of the 2018 21st International Conference on Intelligent Transportation Systems (ITSC), Maui, HI, USA, 4–7 November 2018; pp. 2575–2582. [Google Scholar] [CrossRef]
  55. Alegre, L.; Terry, J.; Kalus, F.; Schumacher, M.; Kwiatkowski, A.; Kelso, K.; Gregor, M. LucasAlegre/sumo-rl: SUMO-RL 1.4.5 Release; Zenodo: Genève, Switzerland, 2024. [Google Scholar] [CrossRef]
  56. Zeinaly, Z.; Sojoodi, M.; Bolouki, S. A Resilient Intelligent Traffic Signal Control Scheme for Accident Scenario at Intersections via Deep Reinforcement Learning. Sustainability 2023, 15, 1329. [Google Scholar] [CrossRef]
  57. Alshayeb, S.; Stevanovic, A.; Dobrota, N. Impact of Various Operating Conditions on Simulated Emissions-Based Stop Penalty at Signalized Intersections. Sustainability 2021, 13, 10037. [Google Scholar] [CrossRef]
  58. Dulac-Arnold, G.; Levine, N.; Mankowitz, D.J.; Li, J.; Paduraru, C.; Gowal, S.; Hester, T. Challenges of Real-World Reinforcement Learning: Definitions, Benchmarks and Analysis. Mach. Learn. 2021, 110, 2419–2468. [Google Scholar] [CrossRef]
  59. Guo, Q.; Li, L.; Ban, X. Urban Traffic Signal Control with Connected and Automated Vehicles: A Survey. Transp. Res. Part C Emerg. Technol. 2019, 101, 313–334. [Google Scholar] [CrossRef]
  60. Bin Al Islam, S.M.A.; Hajbabaie, A.; Aziz, H.M.A. A Real-Time Network-Level Traffic Signal Control Methodology with Partial Connected Vehicle Information. Transp. Res. Part C Emerg. Technol. 2020, 121, 102830. [Google Scholar] [CrossRef]
Figure 1. Overall architecture of Mamba-PTC.
Figure 1. Overall architecture of Mamba-PTC.
Sustainability 18 05561 g001
Figure 2. Topology of the Bologna Joined Network used as the primary evaluation setting.
Figure 2. Topology of the Bologna Joined Network used as the primary evaluation setting.
Sustainability 18 05561 g002
Figure 3. Episode reward during training.
Figure 3. Episode reward during training.
Sustainability 18 05561 g003
Figure 4. Framework-level ablation against fixed-time control.
Figure 4. Framework-level ablation against fixed-time control.
Sustainability 18 05561 g004
Table 1. Main PPO hyperparameters.
Table 1. Main PPO hyperparameters.
CategoryParameterValue
Learning RateInitial LR 10 4
Sampling N steps 1024
SamplingBatch Size8192
Sampling N epochs 4
Risk Calibration ( w ρ , w v ) 0.6, 0.4
Reward Weights ( w 1 , w 2 , w 3 , w 4 , w 5 ) ( 60.0 ,   0.035 ,   0.08 ,   0.05 ,   15.0 )
Advantage Estimation γ 0.99
Advantage Estimation λ GAE 0.95
Policy ConstraintClip Range0.03
Policy ConstraintEntropy Coefficient0.0003
Table 2. A 5 × 3 stress-testing matrix for transition-sensitive congestion evaluation.
Table 2. A 5 × 3 stress-testing matrix for transition-sensitive congestion evaluation.
DimensionParameter/LevelSetting and Purpose
Flow Intensity λ { 1.2 , 1.4 , 1.6 , 1.8 , 2.0 } Covers traffic evolution from moderate load to heavily congested regimes.
Perturbation ScheduleLowLight perturbation.
N e v t = 4 , Δ e v t = 180   s , H e v t = 240   s
MidMedium-strength perturbation.
N e v t = 8 , Δ e v t = 120   s , H e v t = 360   s
HighStrong and persistent perturbation.
N e v t = 12 , Δ e v t = 90   s , H e v t = 480   s
Statistical ProtocolPrimary matrix: M = 5 Evaluation with matched sets.
Extended seed protocol: M = 10 Selected high-stress slices with matched seeds.
Table 3. Direct comparison with practical controller baselines on the Joined Network.
Table 3. Direct comparison with practical controller baselines on the Joined Network.
Method N out  ↑ T ¯ (min) ↓ W ¯ (s) ↓ Φ  ↓ χ  ↓
Max-Pressure592.446.10210.532.4608.829
PressLight548.606.58230.242.3113.046
Mamba-PTC664.805.44174.531.9373.368
Table 4. Compact paired statistical summary in the practical controller track ( n = 75 paired runs per baseline).
Table 4. Compact paired statistical summary in the practical controller track ( n = 75 paired runs per baseline).
Baseline Δ N out  ↑ Δ T ¯ (min) ↓Sign-Test p ( N out / T ¯ )
Max-Pressure72.36 ± 39.98
95% CI [63.16, 81.56]
0.67 ± 0.37
95% CI [0.58, 0.75]
7.55 × 10 20
7.55 × 10 20
PressLight116.20 ± 34.07
95% CI [108.36, 124.04]
1.14 ± 0.30
95% CI [1.07, 1.21]
2.65 × 10 23
2.65 × 10 23
Table 5. Matched-RL summary across overall, high-perturbation, and most-stressed regimes.
Table 5. Matched-RL summary across overall, high-perturbation, and most-stressed regimes.
RegimeMetricWang et al.LSTMMamba-PTC
All 15 slices N o u t  ↑600.45611.37664.80
T ¯ (min) ↓6.035.925.44
High perturbation only N o u t  ↑575.20586.96626.84
T ¯ (min) ↓6.296.175.75
Most-stressed average N o u t  ↑550.60559.70631.50
T ¯ (min) ↓6.566.465.71
Table 6. Direct comparison with the Transformer baseline.
Table 6. Direct comparison with the Transformer baseline.
Method N out  ↑ T ¯ (min) ↓ W ¯ (s) ↓ Φ  ↓ χ  ↓
Transformer593.466.28301.284.00317.30
Mamba-PTC622.206.04262.733.46514.21
Table 7. Transition-sensitive observables on the most-stressed matched-RL slices.
Table 7. Transition-sensitive observables on the most-stressed matched-RL slices.
SliceMethod Φ  ↓ χ  ↓
( λ = 1.8 , high ) Wang et al.2.75610.056
LSTM2.6869.565
Mamba-PTC1.7824.834
( λ = 2.0 , high ) Wang et al.3.39014.111
LSTM2.3647.998
Mamba-PTC1.9915.647
Table 8. Compact paired statistical summary against matched-RL learned baselines.
Table 8. Compact paired statistical summary against matched-RL learned baselines.
Baseline Δ N out  ↑ Δ T ¯ (min) ↓Sign-Test p ( N out / T ¯ )
LSTM53.43 ± 45.07
95% CI [43.06, 63.80]
0.48 ± 0.42
95% CI [0.39, 0.58]
3.83 × 10 12
3.83 × 10 12
Wang et al.64.35 ± 49.38
95% CI [52.99, 75.71]
0.60 ± 0.47
95% CI [0.49, 0.71]
4.91 × 10 16
4.91 × 10 16
Transformer32.11 ± 44.32
95% CI [21.91, 42.30]
0.34 ± 0.37
95% CI [0.25, 0.42]
1.09 × 10 6
4.20 × 10 9
Table 9. Overall comparison with movement-centric controllers on the Joined Network.
Table 9. Overall comparison with movement-centric controllers on the Joined Network.
Method N out  ↑ T ¯ (min) ↓ W ¯ (s) ↓
FRAP505.807.17144.23
MoveLight579.336.23144.28
Mamba-PTC664.805.44174.53
Table 10. Compact paired statistical summary in the movement-centric track.
Table 10. Compact paired statistical summary in the movement-centric track.
Baseline Δ N out  ↑ Δ T ¯ (min) ↓Sign-Test p ( N out / T ¯ )
FRAP159.00 ± 49.22
95% CI [147.68, 170.32]
1.73 ± 0.59
95% CI [1.60, 1.87]
2.65 × 10 23
2.65 × 10 23
MoveLight85.47 ± 40.19
95% CI [76.22, 94.71]
0.79 ± 0.36
95% CI [0.71, 0.87]
7.55 × 10 20
7.55 × 10 20
Table 11. Pooled summary of the fixed-time ablation.
Table 11. Pooled summary of the fixed-time ablation.
Method N out  ↑ T ¯ (min) ↓ W ¯ (s) ↓ Φ  ↓ χ  ↓
Fixed-time601.236.02246.562.90310.061
Mamba-PTC664.805.44174.531.9135.251
Table 12. Robustness check for waiting time W ¯ (s) gains against fixed-time control.
Table 12. Robustness check for waiting time W ¯ (s) gains against fixed-time control.
FlowMean Gain (s)Positive Pair RatioSign-Test p
1.247.8786.67% 1.01 × 10 3
1.423.8993.33% 9.16 × 10 5
1.630.8893.33% 6.10 × 10 5
1.815.4580.00% 7.54 × 10 3
2.074.6386.67% 4.18 × 10 3
Table 13. Performance comparison of reward design variants on matched high-stress slices.
Table 13. Performance comparison of reward design variants on matched high-stress slices.
Variant N out  ↑ T ¯ (min) ↓ W ¯ (s) ↓ Φ  ↓ χ  ↓
Full objective607.678.8877.440.7471.329
Efficiency only608.008.91117.551.3063.163
No risk term501.6712.58197.402.5018.094
No switching regularization304.3317.68311.184.17517.403
Table 14. Extended seed summary with uncertainty on the high-stress slices.
Table 14. Extended seed summary with uncertainty on the high-stress slices.
FlowMethod N out T ¯ (min) Φ χ
1.6LSTM584.9 ± 28.8 [564.3, 605.5]6.169 ± 0.310 [5.947, 6.390]3.543 ± 1.474 [2.489, 4.597]15.134 ± 10.172 [7.858, 22.411]
1.6Mamba-PTC606.6 ± 41.7 [576.8, 636.4]5.961 ± 0.421 [5.660, 6.262]2.497 ± 1.055 [1.742, 3.252]8.538 ± 5.970 [4.267, 12.808]
1.8LSTM610.5 ± 18.1 [597.5, 623.5]5.902 ± 0.179 [5.774, 6.029]2.321 ± 0.816 [1.738, 2.905]7.431 ± 4.819 [3.983, 10.878]
1.8Mamba-PTC635.9 ± 18.7 [622.5, 649.3]5.666 ± 0.169 [5.545, 5.786]2.029 ± 0.748 [1.494, 2.564]6.065 ± 3.831 [3.324, 8.805]
2.0LSTM592.4 ± 22.8 [576.1, 608.7]6.085 ± 0.235 [5.917, 6.253]3.097 ± 0.931 [2.431, 3.763]11.594 ± 5.683 [7.529, 15.660]
2.0Mamba-PTC626.6 ± 27.8 [606.7, 646.5]5.755 ± 0.247 [5.579, 5.932]2.113 ± 0.707 [1.607, 2.618]6.385 ± 3.696 [3.741, 9.029]
Table 15. Pooled, paired, and extended seed statistics on the high-stress slices (Mamba-PTC vs. LSTM).
Table 15. Pooled, paired, and extended seed statistics on the high-stress slices (Mamba-PTC vs. LSTM).
Metric Δ Mean ± SD95% CI of Δ Positive PairsSign-Test p
N o u t 27.10 ± 36.83[13.35, 40.85]21/30 4.28 × 10 2
T ¯ 0.258 ± 0.369[0.120, 0.396]21/30 4.28 × 10 2
W ¯ 54.84 ± 108.28[14.41, 95.27]22/30 1.61 × 10 2
Φ 0.774 ± 1.525[0.205, 1.343]23/30 5.22 × 10 3
χ 4.391 ± 9.493[0.846, 7.935]23/30 5.22 × 10 3
Table 16. Cross-network mean performance against the LSTM baseline on Acosta and Pasubio.
Table 16. Cross-network mean performance against the LSTM baseline on Acosta and Pasubio.
NetworkMethod N out  ↑ T ¯ (min) ↓ W ¯ (s) ↓ Φ  ↓ χ  ↓
AcostaLSTM393.419.206208.892.5938.239
AcostaMamba-PTC455.287.953161.731.8845.191
PasubioLSTM513.757.027220.882.6619.044
PasubioMamba-PTC530.156.808190.412.2226.505
OverallLSTM453.588.117214.892.6278.642
OverallMamba-PTC492.727.381176.072.0535.848
Table 17. Paired-run robustness in the cross-network validation.
Table 17. Paired-run robustness in the cross-network validation.
NetworkMetric Δ Mean ± SD95% CI of Δ Sign-Test p
Acosta N o u t 61.87 ± 37.46[53.25, 70.48] 3.73 × 10 18
Acosta T ¯ 1.253 ± 0.769[1.077, 1.430] 3.73 × 10 18
Acosta W ¯ 47.154 ± 53.300[34.894, 59.415] 1.59 × 10 7
Acosta Φ 0.709 ± 0.763[0.534, 0.885] 1.59 × 10 7
Acosta χ 3.047 ± 3.502[2.242, 3.853] 1.59 × 10 7
Pasubio N o u t 16.40 ± 27.61[10.05, 22.75] 1.08 × 10 3
Pasubio T ¯ 0.219 ± 0.368[0.134, 0.303] 1.08 × 10 3
Pasubio W ¯ 30.465 ± 59.731[16.725, 44.205] 3.56 × 10 1
Pasubio Φ 0.439 ± 0.842[0.245, 0.632] 3.56 × 10 1
Pasubio χ 2.538 ± 5.044[1.378, 3.699] 3.56 × 10 1
Overall N o u t 39.13 ± 39.95[32.69, 45.58] 7.71 × 10 16
Overall T ¯ 0.736 ± 0.794[0.608, 0.864] 7.71 × 10 16
Overall W ¯ 38.810 ± 57.034[29.608, 48.011] 1.23 × 10 5
Overall Φ 0.574 ± 0.812[0.443, 0.705] 1.23 × 10 5
Overall χ 2.793 ± 4.335[2.094, 3.492] 1.23 × 10 5
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ouyang, Z.; Li, C.; Tang, Y.; Shu, Y.; Wang, Z.; Ma, Y.; Ding, T. Transition-Sensitive Congestion Dynamics in Heterogeneous Urban Traffic Networks Under Coordinated Reinforcement Learning. Sustainability 2026, 18, 5561. https://doi.org/10.3390/su18115561

AMA Style

Ouyang Z, Li C, Tang Y, Shu Y, Wang Z, Ma Y, Ding T. Transition-Sensitive Congestion Dynamics in Heterogeneous Urban Traffic Networks Under Coordinated Reinforcement Learning. Sustainability. 2026; 18(11):5561. https://doi.org/10.3390/su18115561

Chicago/Turabian Style

Ouyang, Zhenghan, Chenxin Li, Yifeng Tang, Yuqingyun Shu, Zhiling Wang, Yuhang Ma, and Tongqiang Ding. 2026. "Transition-Sensitive Congestion Dynamics in Heterogeneous Urban Traffic Networks Under Coordinated Reinforcement Learning" Sustainability 18, no. 11: 5561. https://doi.org/10.3390/su18115561

APA Style

Ouyang, Z., Li, C., Tang, Y., Shu, Y., Wang, Z., Ma, Y., & Ding, T. (2026). Transition-Sensitive Congestion Dynamics in Heterogeneous Urban Traffic Networks Under Coordinated Reinforcement Learning. Sustainability, 18(11), 5561. https://doi.org/10.3390/su18115561

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop