Next Article in Journal
Assessment of Gait and Balance in Elderly Individuals with Knee Osteoarthritis Using Inertial Measurement Units
Previous Article in Journal
RGB-D Cameras and Brain–Computer Interfaces for Human Activity Recognition: An Overview
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Physical Layer Security Enhancement in IRS-Assisted Interweave CIoV Networks: A Heterogeneous Multi-Agent Mamba RainbowDQN Method

College of Electrical Engineering and Automation, Fuzhou University, Fuzhou 350108, China
*
Author to whom correspondence should be addressed.
Sensors 2025, 25(20), 6287; https://doi.org/10.3390/s25206287
Submission received: 11 September 2025 / Revised: 4 October 2025 / Accepted: 9 October 2025 / Published: 10 October 2025
(This article belongs to the Section Sensor Networks)

Abstract

The Internet of Vehicles (IoV) relies on Vehicle-to-Everything (V2X) communications to enable cooperative perception among vehicles, infrastructures, and devices, where Vehicle-to-Infrastructure (V2I) links are crucial for reliable transmission. However, the openness of wireless channels exposes IoV to eavesdropping, threatening privacy and security. This paper investigates an Intelligent Reflecting Surface (IRS)-assisted interweave Cognitive IoV (CIoV) network to enhance physical layer security in V2I communications. A non-convex joint optimization problem involving spectrum allocation, transmit power for Vehicle Users (VUs), and IRS phase shifts is formulated. To address this challenge, a heterogeneous multi-agent (HMA) Mamba RainbowDQN algorithm is proposed, where homogeneous VUs and a heterogeneous secondary base station (SBS) act as distinct agents to simplify decision-making. Simulation results show that the proposed method significantly outperform benchmark schemes, achieving a 13.29% improvement in secrecy rate and a 54.2% reduction in secrecy outage probability (SOP). These results confirm the effectiveness of integrating IRS and deep reinforcement learning (DRL) for secure and efficient V2I communications in CIoV networks.

1. Introduction

Intelligent Transportation Systems (ITSs) [1], as an essential component of smart cities, are gradually evolving into an integrated framework that combines sensing, communication, and intelligent decision-making. Within this framework, the Internet of Vehicles (IoV) [2,3] plays a pivotal role by enabling large-scale, low-latency, and high-reliability information exchange, thereby supporting traffic safety, road management, and mobility services [4,5].
Among various communication modes in IoV, Vehicle to Vehicle (V2V) and Vehicle to Infrastructure (V2I) are two fundamental paradigms [6,7]. V2V communication relies on direct links between neighboring vehicles, enabling distributed information sharing. However, its coverage is limited and highly sensitive to vehicle speed and density [8]. In contrast, V2I communication leverages roadside units or base stations as relays, offering wider coverage and more stable connectivity. This makes V2I particularly suitable for critical applications such as traffic signal optimization, road safety alerts, and route planning. Consequently, ensuring the security and reliability of V2I communications under highly dynamic vehicular environments has become a key research challenge in IoV.
With the rapid expansion of IoV, wireless communication systems are facing increasing spectrum scarcity. Fixed spectrum allocation can no longer satisfy the bandwidth demands of massive vehicular access and dynamic traffic services, resulting in inefficient spectrum utilization. To address this issue, cognitive radio (CR) [9,10] has been introduced into IoV, resulting in the formation of the Cognitive Internet of Vehicles (CIoV) [11,12]. In a CR system, primary users (PUs) are licensed users with priority access to the spectrum, while secondary users (SUs) are unlicensed devices that opportunistically access idle spectrum. CR enables environmental sensing and dynamic spectrum access, allowing SUs to opportunistically utilize idle spectrum when it is not occupied by PUs, thereby significantly improving spectrum efficiency. Depending on the access mechanism, CR can operate in three typical modes: overlay, underlay, and interweave. Among them, the interweave mode has attracted particular attention, as it enables efficient spectrum sharing while guaranteeing the protection of PUs [13]. In this mode, SUs are permitted to access the channel only when spectrum holes are detected. This property makes interweave mode well-suited to the requirements of V2I communications in IoV. Compared with the distributed access in V2V, V2I communications under interweave mode, supported by infrastructure, can achieve larger coverage and more stable connectivity. However, the resource allocation must simultaneously address the challenges of spectrum dynamics and secure transmission, which significantly increases the complexity. This serves as the primary motivation of this study on resource allocation for V2I communications in an interweave CIoV network.
The openness and broadcast nature of the CIoV network make it highly vulnerable to physical layer attacks, particularly passive eavesdropping [14,15]. Eavesdroppers (Eves) can threaten legitimate transmissions merely by intercepting V2I signals without introducing active interference. To address this security concern, Intelligent Reflecting Surface (IRS) [16,17], an emerging reconfigurable environment technology, has recently been introduced into physical layer security communications [18,19]. By adaptively adjusting the phase shifts of its reflecting elements, IRS is capable of intelligently reconfiguring the wireless propagation environment. This allows it to enhance the strength of legitimate links while suppressing eavesdropping channels, thereby significantly improving system security with low power consumption and low hardware complexity.

2. Related Work

Recent studies have highlighted the effectiveness of IRS in vehicular communication systems. Yang et al. [20] proposed a weighted performance optimization algorithm using alternating optimization to jointly design IRS phase shifts and beamforming, mitigating resource conflicts among sensing, communication, and computation. Cao et al. [21] introduced an IRS-assisted framework based on statistical CSI, employing the JAPBNB algorithm for joint active and passive beamforming to maximize the sum-rate in multi-vehicle scenarios. Duan et al. [22] further developed an IRS-enhanced V2X framework that jointly optimizes mode selection, spectrum reuse, computation offloading, and beamforming, reducing latency and energy consumption in dynamic vehicular environments.
In CR networks, IRS technology has also received widespread attention. In [23], an IRS-assisted UAV communication framework was proposed, where the UAV trajectory, IRS beamforming, and power allocation are jointly optimized to maximize the SUs throughput under PUs interference constraints. Similarly, ref. [24] studied an IRS-aided MISO cognitive radio system, where the transmit and reflective beamforming are jointly optimized to maximize the SUs rate while satisfying PUs interference and IRS power constraints, demonstrating significant throughput improvements. Moreover, in [25], an IRS-assisted CR system with opportunistic spectrum sharing was considered, supporting both secondary transmission and spectrum sensing. A two-stage IRS phase shift optimization algorithm based on block coordinate descent (BCD) was proposed to maximize the average secondary transmission rate.
Recent studies have shown that IRS effectively mitigated eavesdropping and enhanced physical layer security. In [26], IRSs in underlay cognitive radio networks using selection combining (SC) and maximal ratio combining (MRC) significantly reduced SOP and increased the probability of nonzero secrecy capacity (PNSC). A hybrid IRS with element on–off control and two-stage beamforming–amplification optimization was proposed in [27], improving PLS with lower power consumption. Additionally, ref. [28] optimized IRS reflection under uncertain eavesdropper locations, which enhanced PNSC and ergodic secrecy capacity while reducing SOP.
The aforementioned works mostly rely on conventional resource management and optimization techniques to tackle the inherently non-convex problems in IRS-assisted networks. However, such methods typically incur high computational complexity, require accurate mathematical modeling. In contrast, DRL is model-free and capable of learning near-optimal policies directly through interaction with complex and uncertain environments, which makes it particularly promising for IRS-assisted CIoV systems. Zhao et al. [29] applied DRL to adjust UAV trajectories and IRS phase shifts in real time, maximizing communication rate. Qi et al. [30] employed SAC-based DRL to jointly optimize IRS phase shifts and vehicular resource allocation, minimizing age of information and enhancing payload delivery. Dong et al. [31] used PPO-based DRL to optimize UAV trajectory, beamforming, and IRS phase shifts, improving the secrecy rate and reducing outage under eavesdropper CSI. Qin et al. [32] applied DDPG-based DRL to jointly optimize channel allocation, transmit power, and STAR-IRS coefficients in NOMA systems with dual eavesdroppers, achieving a robust sum secrecy rate. Ju et al. [33,34] recently explored AI-assisted physical layer security in vehicular networks. Specifically, ref. [33] proposed a NOMA-assisted secure offloading scheme for vehicular edge computing using cooperative jamming and asynchronous deep reinforcement learning to reduce energy consumption under latency constraints, while ref. [34] developed a deep recurrent reinforcement learning-based framework for mmWave vehicular networks, jointly optimizing beam allocation, relay/jammer selection, and transmit power to enhance secrecy capacity and energy efficiency. These studies collectively demonstrate the effectiveness of DRL for real-time joint optimization in IRS-assisted networks.
Owing to its efficiency in modeling long-range dependencies and spatial–temporal features, Mamba has increasingly been adopted in reinforcement learning applications. In [35], the authors proposed a hybrid trading framework that integrates Mamba state-space models, temporal convolutional networks, and attention mechanisms for feature extraction, combined with advanced DQN variants for reinforcement learning-based decision optimization, and demonstrated superior returns and risk-adjusted performance in high-frequency cryptocurrency trading. In [36], the authors proposed Mamba-DQN, which integrates the Mamba-SSM encoder and a full state sequence replay strategy into DQN to preserve temporal alignment, and achieved superior stability and sample efficiency over DQN, LSTM-DQN, and Transformer-DQN in dynamic environments.

3. Motivation and Contributions

IRS-assisted secure communications face several challenges in vehicular networks: the high mobility of vehicles causes rapidly varying channels; multiple VUs and potential eavesdroppers increase system complexity; and outdated CSI complicates the joint optimization of spectrum allocation, transmit power, and IRS phase shifts. Existing works typically assume simplified scenarios with static channels or a single eavesdropper, without considering dynamic spectrum sharing or multi-agent resource allocation. To the best of our knowledge, IRS-assisted interweave CIoV networks against malicious eavesdropping attacks have not been studied.
To address these challenges, we propose a multi-agent deep reinforcement learning (MADRL) framework under centralized training and decentralized execution (CTDE), capable of efficiently optimizing multi-user resource allocation and IRS phase shifts in dynamic environments.
The main contributions of this paper are listed in the following.
  • We propose a novel IRS-assisted interweave CIoV network, accounting for outdated CSI, and formulate two optimization problems: maximizing the total minimum secrecy rate and minimizing the average maximum SOP. A heterogeneous MADRL approach is employed to solve these nonlinear and non-convex problems efficiently.
  • We design a heterogeneous multi-agent Mamba RainbowDQN framework, where VUs handle transmit power and channel allocation as homogeneous agents, and the SBS acts as a heterogeneous agent to focus on IRS phase optimization, reducing computational burden and improving system flexibility.
  • We integrate the Mamba module into the heterogeneous MADRL framework to enable the IRS to more effectively assist the interweave CIoV network in enhancing physical layer security. By capturing long-term temporal dependencies and high-dimensional state correlations, the proposed Mamba-enhanced MADRL allows agents to make more accurate and stable decisions under outdated CSI, and simulations show that the proposed framework outperforms baseline methods in system secrecy performance.

4. Paper Organization

The remainder of the paper is organized as follows. Section 5 introduces the IRS-assisted interweave CIoV network model under malicious eavesdropping attacks and formulates the corresponding optimization problem. Section 6 establishes the Markov Decision Process (MDP) framework and proposes the HMA-Mamba RainbowDQN algorithm to solve the optimization problem. Section 7 presents simulation results to evaluate the system performance via our proposed method as compared to benchmark schemes. Finally, Section 8 concludes this paper.

5. System Model and Problem Formulation

5.1. System Model

In this paper, we consider an IRS-assisted interweave CIoV uplink network, as shown in Figure 1, consisting of a primary base station (PBS), N authorized primary users (PUs), a secondary base station (SBS) equipped with a single antenna [37,38], I vehicle users (VUs), M eavesdroppers (Eves), and an IRS with L reflecting elements controlled by the SBS. The PBS provides infrastructure support for PU communications but does not assist secondary systems. VUs independently detect PU activity through spectrum sensing, using conventional methods such as energy detection or cooperative sensing, and adopt an interweave transmission strategy, transmitting only when channels are sensed as idle.
The IRS-assisted secure communication framework we propose is governed by the geometric relationships among vehicles, the IRS, and the base station rather than the underlying road topology. To facilitate geometric modeling and illustrative clarity, a grid-based urban layout is adopted. Each VU randomly selects an initial direction and moves at a constant speed, turning at intersections according to predefined probabilities while maintaining connectivity with the SBS. Although vehicles in real urban scenarios frequently accelerate or decelerate, this simplification enables tractable mobility modeling and focuses the analysis on IRS optimization and secure V2I transmission, providing an indicative upper bound on system performance.
Each PU is allocated a dedicated orthogonal channel to avoid mutual interference so that the number of channels equals the number of PUs. The activity of each PU is modeled as a two-state Markov chain (idle/active), with details given in Appendix A.

5.2. Eavesdropping Model

Unlike traditional physical-layer security studies that often assume a single eavesdropper, in practical vehicular networks, malicious nodes may be distributed at critical locations such as intersections or toll stations. To capture this, we consider a multi-eavesdropper scenario where M eavesdroppers ( M 2 ) are randomly deployed within a circular region of radius 100 m centered at the SBS, with their coordinates denoted by E = ( x m , y m ) m = 1 , 2 , , M .
To rigorously and conservatively evaluate secrecy performance, we adopt a worst-case security assumption [39,40]. Specifically, each Eve is modeled as a powerful passive adversary capable of intercepting uplink transmissions from all VUs. The eavesdroppers act independently and do not share or aggregate intercepted signals. For analytical tractability and to ensure conservative results, each Eve is further assumed to ideally separate and decode the multi-user signals without intrinsic information loss (e.g., via optimal multi-user detection).
Under this assumption, the instantaneous secrecy rate of each VU is defined against the strongest eavesdropper, i.e., with respect to the maximum eavesdropping rate among all Eves. This does not imply cooperation, but rather reflects a worst-case perspective: since the system may be compromised once the most capable Eve succeeds, secrecy performance is conservatively evaluated against the strongest adversary. Moreover, for IRS-assisted transmission design, we assume that the SBS has perfect knowledge of the CSI of all Eves, which enables a robust assessment of system security under the most challenging conditions.

5.3. Communication Model

Let the coordinates of the SBS, PBS, and IRS be ( x SBS , y SBS ) , ( x PBS , y PBS ) , and ( x IRS , y IRS ) , respectively. At the t-th time slot, the coordinates of the i-th VU and the m-th Eve are ( x i , y i ) and ( x m , y m ) , respectively. The channel gain between the i-th VU and the SBS can be defined as      
h i , S = l 0 ϕ ( d i , S ) α g i , S ,
where l 0 denotes the path loss at a reference distance of 1 m, ϕ represents log-normal shadowing, d i , S is the distance between the i-th VU and the SBS, α is the path-loss exponent, and g i , S denotes the small-scale fading component, which follows a Rayleigh distribution.
Similarly, the channel gains between the i-th VU and the IRS, and between the IRS and the SBS, are defined as
h i , R = l 0 ϕ ( d i , R ) α e j 2 π d i , R / λ a AoA ,
h R , S = l 0 ϕ ( d R , S ) α e j 2 π d R , S / λ a AoD ,
where d i , R and d R , S denote the distances between the i-th VU and the IRS, and between the IRS and the SBS, respectively; λ is the carrier wavelength. a AoA and a AoD are the array response vectors for the IRS incident and reflected links, given by
a AoA = [ 1 , e j 2 π d λ sin ( θ AoA ) , , e j 2 π d λ ( L 1 ) sin ( θ AoA ) ] T ,
a AoD = [ 1 , e j 2 π d λ sin ( θ AoD ) , , e j 2 π d λ ( L 1 ) sin ( θ AoD ) ] T .
where θ AoA and θ AoD denote the angle of arrival and departure at the IRS, respectively, and d is the IRS element spacing. Similarly, the channel gains between the m-th Eve and the VU or IRS are denoted by h i , m and h R , m , respectively.
The IRS phase shift matrix is defined as
Θ = diag { ϕ 1 , ϕ 2 , , ϕ L } , ϕ l = e j θ l , θ l [ 0 , 2 π ) , l ,
where ϕ l denotes the reflection coefficient of the l-th IRS element. Here, we assume ideal full-reflection, i.e., all amplitudes β l = 1  [29]. The phase shifts are practically hardware-limited and thus discretized in our action space modeling.
The performance of IRS-assisted communication critically depends on accurate CSI. In this work, the CSI of VUs is obtained via pilot-based channel estimation [41]. In high-mobility V2I scenarios, the rapid movement of vehicles leads to outdated CSI [22,42], which complicates the IRS phase shift design [43]. Therefore, the impact of outdated CSI is explicitly considered for V2I links. By contrast, eavesdropping links are assumed quasi-static due to the near-SBS static deployment of eavesdroppers [44], and outdated CSI is not considered.
Let T s denote the CSI acquisition delay. The outdated CSI can be modeled as
h ( t + T s ) = κ h ˜ ( t ) + 1 κ 2 Δ h ,
where h ˜ ( t ) is the estimated channel, Δ h is the independent estimation error, and
κ = J 0 ( 2 π f D T s )
represents the channel correlation coefficient with f D = v f c / c being the maximum Doppler shift. A larger κ indicates stronger correlation and more accurate CSI, while smaller κ reflects faster channel variation. The actual V2I rate depends on h ( t + T s ) rather than the estimated CSI.
At time slot t, the Signal-to-Interference-plus-Noise Ratio (SINR) of the i-th VU on channel n is given by
γ i n ( t ) = p i | h i , S + h R , S H Θ h i , R | 2 k I , k i δ k n p k | h k , S + h R , S H Θ h k , R | 2 + σ i 2 ,
where p i denotes the transmit power of the i-th VU, σ i 2 is the noise power, and δ k n { 0 , 1 } indicates whether the k-th VU accesses channel n. It should be emphasized that the interference term only accounts for VUs sharing channel n as indicated by δ k n since different channels are assumed orthogonal in cognitive radio systems. Thus, only co-channel users contribute to the interference of the i-th VU.
Accordingly, the achievable data rate is expressed as
C i n ( t ) = b n ( t ) W log 2 1 + γ i n ( t ) ,
where W is the channel bandwidth, and b n ( t ) { 0 , 1 } represents the channel state b n ( t ) = 1 if the channel is idle, and b n ( t ) = 0 if occupied by PUs.
Similarly, the SINR experienced at the m-th Eve when eavesdropping on the i-th VU is given by
γ m , i n ( t ) = p i | h i , m + h R , m H Θ h i , R | 2 k I , k i δ k n p k | h k , m + h R , m H Θ h k , R | 2 + σ m 2 ,
and the corresponding eavesdropping rate is
C m , i n ( t ) = b n ( t ) W log 2 1 + γ m , i n ( t ) .
To evaluate secrecy performance rigorously, we adopt a worst-case assumption: each eavesdropper may intercept transmissions from all VUs, and the instantaneous secrecy rate of the i-th VU is defined using the maximum eavesdropping rate among all eavesdroppers:
R i sec ( t ) = [ C i n ( t ) max m M C m , i n ( t ) ] + ,
where [ z ] + = max ( 0 , z ) . The total secrecy rate of the system at time slot t is
R sec ( t ) = i = 1 I R i sec ( t ) .
While R sec ( t ) measures the instantaneous secrecy of the V2I link, the SOP [31] captures its robustness and is defined as
P out ( t ) = Pr R sec ( t ) < r th ,
where r th is the target secrecy threshold.

5.4. Problem Formulation

In this work, we formulate two optimization problems to enhance the physical layer security of IRS-assisted V2I communications.
(1)
Secrecy Rate Maximization
The first problem aims to maximize the total secrecy rate of the system by jointly optimizing the IRS phase shift matrix Θ , the transmit power vector p = [ p 1 , , p I ] , and the channel allocation matrix δ = [ δ i n ] :  
P 1 : max Θ , p , δ R sum sec ( t ) ,
s . t . R i sec ( t ) R i ( sec , min ) , i I ,
C i n ( t ) C i min , i I , n N ,
0 θ l < 2 π , l L ,
0 p i p max , i I ,
δ i n { 0 , 1 } , i I , n N ,
n = 1 N δ i n 1 , i I .
Constraints C1–C2 guarantee the minimum secrecy and data rate requirements of each VU, C3 restricts the IRS phase shifts, C4 limits the transmit power, and C5–C6 ensure that each VU occupies at most one channel.
(2)
Secrecy Outage Probability Minimization
P 2 : min Θ , p , δ max i I P i out ( t ) ,
s . t . C i n ( t ) r th , i I , n N ,
0 θ l < 2 π , l L ,
0 p i p max , i I ,
δ i n { 0 , 1 } , i I , n N ,
n = 1 N δ i n 1 , i I .
Constraint C1 ensures that the achievable rate of each VU on the allocated channel is no less than the secrecy threshold r th , while the remaining constraints are consistent with those in problem P 1 .
The joint optimization is non-convex due to coupled variables. Conventional methods (e.g., AO and SDR) are complex, require precise CSI, and provide only suboptimal solutions. DRL can learn near-optimal policies in unknown environments, but single-agent MDPs face large state–action spaces and inefficiency. Hence, we adopt a MADRL framework for improved flexibility and efficiency.

6. Deep Reinforcement Learning for Resource Allocation

6.1. Multi-Agent DRL Framework

Most existing works focus on homogeneous MADRL, assuming identical state and action spaces for all agents. However, in IRS-assisted interweave CIoV networks, the power and channel allocation actions of VUs differ from the phase shift actions of IRS, making this assumption unsuitable. Hence, we model the problem as an MDP with I + 1 heterogeneous agents, where VUs and the SBS are treated as two types of agents. Each interacts with the environment in discrete time steps to optimize its policy. This HMA strategy decomposes the large mixed (continuous–discrete) action space into manageable sub-tasks, alleviating the “curse of dimensionality” and improving efficiency and convergence compared with single-agent methods.

6.2. Observation Space

At time slot t, the local observation state of the i-th VU agent is defined as
s i t = { O i ( t ) , h i ( t ) , h m , i ( t ) , G i ( t 1 ) } ,
where O i ( t ) = { O i 1 ( t ) , , O i N ( t ) } denotes the spectrum sensing result set of the VU, and O i n ( t ) { 0 , 1 } indicates the activity of the PU on the n-th channel (1: busy, 0: idle). h i ( t ) = { h i , S , h i , R , h R , S } represents the CSI related to this VU, and h m , i ( t ) = { h m , S , h m , R , h R , S } represents the CSI of all Eves eavesdropping on this VU. G i ( t 1 ) denotes the co-channel interference from other VUs in the previous time slot.
The SBS agent is heterogeneous with respect to the VUs. Its local observation state is defined as
s SBS t = { Θ ( t 1 ) , { h i ( t ) i I } , { h m , i ( t ) i I } } ,
where Θ ( t 1 ) is the IRS phase shift matrix from the previous slot, and { h i ( t ) } and { h m , i ( t ) } denote the CSI sets of all VUs and their corresponding Eves, respectively. To enable IRS optimization, VUs periodically report limited local information to the SBS, forming a lightweight exchange that remains feasible even in high-speed vehicular scenarios.

6.3. Action Space

In time slot t, each VU agent jointly decides on channel selection and power allocation. The transmit power is quantized into A p discrete levels, and one channel c i is selected from N available channels. Accordingly, the action space size is | A VU | = N A p , given by
A VU = { a ( c i , p j ) i = 1 , , N ; j = 1 , , A p } ,
where c i denotes the i-th channel and p j the j-th power level.
For the SBS agent, the action in slot t is the phase shift variation Δ Θ ( t ) , updated as
Θ ( t ) = Θ ( t 1 ) Δ Θ ( t ) ,
where ⊙ denotes element-wise multiplication. Each Δ Θ ( t ) is selected from a subset of the discrete Fourier transform (DFT) basis vectors
v ( ψ ) = [ 1 , e j π ψ / L , , e j π ( L 1 ) ψ / L ] .
Thus, the SBS action space is
A SBS = { v ( ψ 1 ) , v ( ψ 2 ) , , v ( ψ n ) } ,
where ψ k ( k = 1 , , n ) are predefined discrete values (e.g., rational fractions). This design fixes the action space size at n     | A SBS | , independent of the number of IRS reflecting elements L, ensuring stable convergence even for large L.

6.4. Reward Design

To simultaneously maximize the total V2I secrecy rate and minimize the maximum SOP while strictly satisfying QoS constraints, the reward function at time slot t is defined as
r t = μ 1 R sec ( t ) R max sec μ 2 log 1 + P out ( t ) μ 3 i = 1 I Φ i QoS ( t ) + Φ i sec ( t ) ,
where μ 1 , μ 2 , μ 3 are weighting parameters to balance different objectives during training, and R max sec denotes the theoretical maximum secrecy rate.
The QoS penalty for the i-th VU is defined as
Φ i QoS ( t ) = max ( 0 , C i min C i n ( t ) ) 2 , 0 < C i n ( t ) ρ QoS , otherwise
where ρ QoS > 0 is a constant to ensure sufficient punishment when the VU selects a PU-occupied channel, leading to a zero transmission rate.
The secrecy-rate penalty for the i-th VU is defined as
Φ i sec ( t ) = max ( 0 , R i ( sec , min ) R i sec ( t ) ) 2 , 0 < R i sec ( t ) ρ sec , otherwise
where ρ sec > 0 ensures sufficient punishment when the secrecy rate falls below the minimum required threshold.

6.5. HMA-Mamba RainbowDQN Algorithm

As the baseline algorithm, Rainbow DQN integrates multiple enhancements to improve the learning efficiency and stability of traditional DQN [45]. Prioritized experience replay (PER) adjusts sampling probabilities based on the TD error | δ t | , defined as
P sam , t | δ t | ω ,
enabling the model to focus on critical samples. Double DQN (DDQN) decouples action selection from evaluation, with the target value formulated as
Q t a r g e t D D Q N = r t + γ Q θ s t + 1 , arg max a Q θ ( s t + 1 , a ) ,
which effectively mitigates Q-value overestimation. Dueling DQN separates the state value and advantage functions, computed as
Q ( s , a ) = V ( s ) + A ( s , a ) 1 | A | a A ( s , a ) ,
providing a more accurate assessment of state importance. Multi-step Learning incorporates n-step returns, expressed as
Q t a r g e t ( n ) = k = 0 n 1 γ k r t + k + γ n max a Q θ ( s t + n , a ) ,
to balance convergence speed and stability. NoisyNet introduces learnable noise into parameters, such as
W = μ W + σ W ϵ W ,
achieving adaptive exploration. Finally, Distributional RL (C51) learns the full return distribution Z ( s , a ) instead of only its expectation, with the target distribution defined as
T Z ( s t , a t ) = r t + γ Z ( s t + 1 , a * ) , a * = arg max a E [ Z ( s t + 1 , a ) ] ,
which captures uncertainty in value estimates and enhances policy robustness.
In conventional DQN, feature extraction primarily relies on MLPs or CNNs. However, MLPs struggle to capture temporal dependencies, while CNNs, although proficient in local pattern recognition, are limited in representing high-dimensional abstract states, particularly those involving historical trajectory information. This often leads to insufficient state representation and training instability. In contrast, Mamba exhibits strong sequential modeling capability, enabling it to automatically capture dynamic relationships among states with high computational efficiency [46]. Leveraging this advantage, this work incorporates Mamba into the Rainbow DQN framework and proposes the heterogeneous multi-agent Mamba RainbowDQN method as shown in Figure 2.
It should be noted that the original Rainbow incorporates C51 distributional learning; however, in multi-agent CIoV scenarios, this approach may induce environmental non-stationarity and suboptimal solutions. Therefore, in this work, C51 is omitted, and the centralized training with decentralized execution (CTDE) paradigm is adopted to ensure training stability and execution efficiency. Under CTDE, agents learn coordinated policies guided by a shared system-level reward during centralized training, while executing their actions independently based on local observations during deployment.
As shown in Figure 3, the Mamba state feature extraction pipeline projects the raw environment state into a hidden space, serializes it into a sequential structure, and applies state-space modeling before compression into a high-dimensional feature.
Formally, the environment state vector s R d state is first projected as
s proj = W in s + b in ,
where W in R d hidden × d state and b in R d hidden .
To fit the sequential modeling structure of Mamba, s proj is reshaped into a sequence:
s seq = unsqueeze ( s proj , dim = 1 ) ,
resulting in an input of shape ( 1 , d hidden ) . Despite the sequence length being one, Mamba is designed to capture temporal dependencies and perform convolutional expansion along the feature dimension, allowing the extraction of rich, abstract features:
s mamba = Mamba ( s seq ) .
Finally, the output sequence is compressed to yield a fixed-length feature vector
s mamba _ flat = squeeze ( s mamba , dim = 1 ) , s mamba _ flat R d hidden ,
which is then fed into the Dueling architecture.
Next, the Dueling architecture decomposes this representation into state-value and advantage functions. The state-value function is
V ( s t w , β ) = W V s mamba _ flat + b V ,
where W V R 1 × d hidden , b V R . The advantage function is
A ( s t , a t w , α ) = W A s mamba _ flat + b A ,
where W A R | A | × d hidden , b A R | A | . To eliminate action-scale bias, the advantage function is normalized to zero mean. The final Q-value is then expressed as
Q ( s t , a t w , α , β ) = V ( s t w , β ) + A ( s t , a t w , α ) 1 | A | a A A ( s t , a w , α ) .
In Mamba Rainbow, the Q-value target is defined as
y t MambaRainbow = k = 0 n 1 γ k r t + k + γ n Q s t + n , arg max a Q ( s t + n , a ; w , α , β ) ; w , α , β ,
where w , α , β denote the parameters of the target network, and the sampling priority is determined by | δ t | α .
Based on this target, the loss function with PER is
L ( w , α , β ) = E ( s , a , r , s ) D w i · ( y t MambaRainbow Q ( s t , a t w , α , β ) ) 2 ,
where each transition i is sampled with probability
P ( i ) = | δ i | α j | δ j | α , δ i = y t MambaRainbow Q ( s t , a t )
and the importance-sampling weight w i = 1 N 1 P ( i ) β corrects the bias introduced by prioritized sampling.
The target network parameters are updated via soft synchronization:
w τ w + ( 1 τ ) w , α τ α + ( 1 τ ) α , β τ β + ( 1 τ ) β .
The training process of the proposed HMA-Mamba RainbowDQN method is detailed in Algorithm 1.
Algorithm 1 Proposed HMA-Mamba RainbowDQN method for IRS-assisted interweave CIoV network.
  • Input: PU activity status, V2I link channel gain, CSI of VUs and Eve, co-channel interference from previous time slot, IRS phase shift matrix
  • Output: Joint action a t of all agents, including VUs agents’ channel and power selection and SBS agent’s IRS phase adjustment
  •       Initialization: Online network parameters Q ( · ) and target network parameters Q ( · ) , prioritized experience replay buffer D , initial state of Mamba feature extractor h 0 = 0
1:
for each episode do
2:
    Initialize VUs positions and velocities, obtain initial environment state
3:
    for each time step t do
4:
        for each agent u = 1 , 2 , , I + 1  do
5:
           Obtain local observation s t u
6:
           Extract temporal features via Mamba: f t u = Mamba ( s t u ; h t 1 )
7:
           Select action from NoisyNet output: a t u = arg max a Q ( f t u , a ; θ noisy )
8:
        end for
9:
        Execute joint action a t , obtain reward r t
10:
        for each agent u = 1 , 2 , , I + 1  do
11:
           Observe next state s t + 1 u
12:
           Store single-step experience ( s t u , a t u , r t , s t + 1 u ) into buffer D
13:
        end for
14:
    end for
15:
    for each agent u = 1 , 2 , , I + 1  do
16:
        Sample a batch from D via prioritized sampling
17:
        Construct n-step returns and compute target y t MambaRainbow by Equation (40)
18:
        Compute gradients and update network parameters w , α , β by Equation (41)
19:
        Update the priorities of the sampled experiences based on TD-error
20:
        Soft-update target network parameters w , α , β by Equation (43)
21:
    end for
22:
end for

6.6. Computational Complexity Analysis

The computational complexity of the proposed HMA-Mamba RainbowDQN algorithm is primarily determined by the forward and backward propagation of the neural networks of all agents, with the linear scanning mechanism of the Mamba module as the core. For a single agent, the forward pass complexity is T main = z = 0 Z 1 f z f z + 1 + Φ ( d state d hidden + d hidden 2 ) , where Z is the number of fully connected layers (including the value and advantage streams in the Dueling architecture), f z is the number of neurons in the z-th layer, Φ is the number of stacked Mamba modules, and d state and d hidden are the input and hidden dimensions of the Mamba module. Let N agents = I + 1 denote the total number of agents (including I VUs and one SBS). During training, which involves backpropagation, experience replay, and policy updates, the complexity scales linearly with the number of agents and is further multiplied by the number of episodes and steps, giving a total training complexity of O N episodes N steps N agents ( 2 z = 0 Z 1 f z f z + 1 + Φ ( d state d hidden + d hidden 2 ) ) . In contrast, during real-time execution, each agent only performs forward inference with per-step complexity O ( z = 0 Z 1 f z f z + 1 + Φ ( d state d hidden + d hidden 2 ) ) , which scales linearly with the number of agents and is significantly lower than training, ensuring feasible online implementation while maintaining high system performance.

7. Simulation Results

In this section, simulations are conducted to evaluate the physical layer security performance of the proposed HMA-Mamba RainbowDQN algorithm under malicious eavesdropping attacks, with relevant parameters listed in Table 1. The SBS is located at ( 250 ,   150 ) , the PBS at ( 100 ,   300 ) , and the IRS is deployed at ( 200 ,   375 ) . Each VU is equipped with a single antenna and randomly distributed along the lanes. PUs are randomly located within a circular area of 100 m radius centered at the PBS, while malicious Eves are randomly distributed within a circular area of 100 m radius centered at the SBS.
Experiments are conducted in Python 3.10 using PyTorch 2.3.1 on an NVIDIA GeForce RTX 4060 GPU (Nvidia Corporation, Santa Clara, CA, USA). The HMA-Mamba RainbowDQN framework maps environment states to a hidden dimension of d hidden = 256 and uses the Mamba feature extractor with state dimension d state = 16 . The output Q-values are generated by a noisy linear layer. The target network is updated via soft and periodic hard updates, and experience replay uses a buffer of 100,000 with prioritized sampling. Exploration is performed using parameterized Gaussian Noisy Nets. Key training parameters are summarized in Table 2.
To evaluate the effectiveness of the proposed method, it is compared with the following benchmark schemes:
(1)
HMA-RainbowDQN: a heterogeneous multi-agent RL algorithm that excludes the Mamba module, serving to isolate and highlight the contribution of Mamba-based feature extraction.
(2)
HMA-D3QN and HMA-DQN: representative conventional DRL algorithms widely adopted in multi-agent resource allocation, used to demonstrate the relative advantages of the proposed framework over mainstream MADRL approaches.
(3)
Random IRS Random RA: a baseline without intelligent optimization, where both IRS phase shifts and resource allocation are randomly determined.
Figure 4 and Figure 5 illustrate the per-episode minimum total secrecy rate and average maximum SOP during training for different algorithms. As shown in Figure 4, the random scheme consistently achieves the lowest secrecy rate, while the secrecy rate of other methods generally increase over training, with notable performance differences. Both HMA-D3QN and HMA-DQN converge to relatively low levels within the same number of training episodes, whereas the proposed HMA-Mamba RainbowDQN achieves the highest secrecy rate. Once training stabilizes, the proposed method demonstrates a 13.29% secrecy rate improvement over HMA-RainbowDQN. Figure 5 indicates that higher average minimum secrecy rate correspond to lower average maximum SOP, confirming that the proposed method also effectively minimizes system secrecy outages.
Figure 6 and Figure 7 examine the impact of the number of IRS reflecting elements on the minimum total secrecy rate and the average maximum SOP of the V2I link. To highlight the role of the IRS, a baseline scenario without IRS assistance and with randomly allocated VU resources is included for comparison. The results indicate that all IRS-based methods benefit from an increased number of reflecting elements, with the proposed HMA-Mamba RainbowDQN method showing the most significant improvement. This is because additional reflecting elements provide greater degrees of freedom to enhance the signal strength of legitimate channels while suppressing eavesdropping channels, thereby enlarging the difference between the two and improving physical layer security. In contrast, the randomly configured IRS scheme only slightly outperforms the no-IRS case, whose transmission rate remains nearly unchanged as the number of IRS reflecting elements increases, highlighting the importance of proper IRS deployment. Nevertheless, the selection of an appropriate number of reflecting elements is critical: too few elements may fail to satisfy the minimum communication rate requirements for all vehicles, whereas increasing the number can enhance security performance but may also introduce inter-user interference and higher computational complexity.
Figure 8 and Figure 9 illustrate the impact of the number of PUs on the secrecy performance of the V2I link. As the number of PUs increases, all schemes achieve higher secrecy rate and lower SOP. This improvement is primarily due to the increased availability of spectrum: more PUs provide additional orthogonal channels, enabling vehicles to strategically select channels with the lowest eavesdropping risk. Under limited channel resources (number of channels < number of VUs), secrecy performance is severely constrained, as VUs are forced to use high-risk channels. When channel availability becomes sufficient, the rate of secrecy improvement slows, indicating a diminishing marginal effect of additional spectrum on reducing eavesdropping risk. Notably, the proposed HMA-Mamba RainbowDQN consistently outperforms all benchmark algorithms, demonstrating its efficiency in identifying and selecting the safest channels in dynamic spectrum environments and achieving secure spectrum allocation even under persistent external eavesdropping threats.
Figure 10 and Figure 11 illustrate the impact of VU mobility on the secrecy rate and SOP of the V2I link. As vehicle speed increases, the secrecy rate decreases while the SOP rises. This is primarily due to higher speeds increasing coverage distance and path loss, as well as more rapid changes in the environment, which degrade the timeliness of CSI. Additionally, increased speed amplifies Doppler shifts f c , causing fast channel variations and reducing channel estimation accuracy. Specifically, as speed increases from [10, 15] m/s to [40, 45] m/s, the proposed HMA-Mamba RainbowDQN and the HMA-RainbowDQN, HMA-D3QN, and HMA-DQN algorithms experience approximately 10–30% reduction in secrecy rate and 20–40% increase in SOP. Nevertheless, the proposed method consistently outperforms the benchmarks across all speeds, demonstrating its robustness to mobility-induced environmental changes and its ability to maintain stable physical layer security performance.
To enhance the comprehensiveness of the study, we further analyze the impact of increasing the number of VUs on system performance. Figure 12 and Figure 13 show that the secrecy performance of each user improves with the number of IRS elements. However, as the number of VUs grows, the average minimum secrecy rate per VU decreases, while the average SOP increases. This is primarily due to the limited spectrum, power, and IRS reflection resources being shared among more users, reducing the effective channel gain per VU. Meanwhile, increased inter-user interference and channel correlation make it difficult for the IRS phase configuration to accommodate all users, weakening its capability to enhance legitimate links and suppress eavesdropping. Furthermore, under the worst-case criterion, the secrecy performance of each VU is determined by the Eve with the highest achievable rate. As the number of VUs increases, some users are more likely to face excessively strong eavesdropping channels, significantly lowering their secrecy rates and degrading overall system performance. These observations indicate that while IRS deployment can improve system performance, the challenges introduced by a larger number of VUs must be carefully considered.
Table 3 presents the average training runtime of different algorithms under the malicious eavesdropping scenario (evaluated over 100 episodes). As expected, the proposed HMA-Mamba RainbowDQN requires significantly longer training time than the benchmark schemes. Compared with DQN and D3QN, RainbowDQN inherently incurs additional computational overhead due to the integration of modules such as NoisyNet and prioritized experience replay. On top of this, the incorporation of the Mamba architecture introduces further complexity from the state space model (SSM) and selective scanning mechanism, thereby prolonging the overall training process. Nevertheless, as demonstrated by the superior secrecy performance reported earlier, this additional training overhead can be regarded as a reasonable trade-off for enhancing physical layer security.
It is worth noting that the preceding experiments were conducted under the baseline assumption of perfect CSI at the eavesdroppers, a common simplification for analytical tractability. To assess the robustness of the proposed scheme under more practical conditions, we further consider the case of imperfect Eve CSI. In this setting, Eve’s channels are modeled using the outdated CSI assumption together with a norm-bounded error model as in [47], with the detailed formulation provided in Appendix B. As illustrated in Figure 14 and Figure 15, all schemes experience a degradation in secrecy rate and an increase in SOP when Eve’s CSI is imperfect. Nevertheless, the proposed HMA-Mamba RainbowDQN consistently achieves the highest secrecy rate and the lowest SOP among the benchmarks, demonstrating its superiority and robustness against CSI uncertainty at the eavesdroppers.

8. Conclusions

This work investigates enhancing the physical layer security of interweave CIoV network under malicious eavesdropping using IRS. To maximize the total minimum secrecy rate and minimize the average maximum SOP, a non-convex joint optimization problem of IRS phase shifts, VU transmit power, and channel allocation is formulated. To address this problem, a HMA-Mamba RainbowDQN based resource allocation method is proposed, where VUs and the SBS act as heterogeneous agents interacting independently with the environment. Simulation results demonstrate that the proposed method significantly outperforms baseline schemes, achieving higher secrecy rate while effectively reducing secrecy outage risk; compared with the basic HMA-RainbowDQN, the secrecy rate increases by 13.29% and the SOP decreases by 54.2%. The design fully accounts for imperfect CSI, random PUs spectrum occupancy, and the stochastic positions and velocities of VUs, employing a worst-case criterion to counter independent non-cooperative Eves. These results validate the robustness and reliability of the proposed method in dynamic vehicular networks. The HMA-Mamba RainbowDQN demonstrates that integrating IRS with DRL can substantially enhance physical layer security in dynamic vehicular networks. The proposed method can be extended to broader dynamic scenarios, such as UAV networks or smart city IoT systems. For future research, we plan to extend the framework to multi-antenna (MISO/MIMO) BS scenarios and investigate highly mobile eavesdroppers with time-varying channels, as well as the deployment of multiple IRSs to further enhance the security and reliability of communications in large-scale or high-density vehicular networks.

Author Contributions

Conceptualization, R.L.; methodology, S.X. and W.C.; software, S.X.; validation, S.X., W.C. and T.X.; formal analysis, S.X.; investigation, S.X.; resources, R.L.; data curation, S.X., W.C. and T.X.; writing—original draft preparation, S.X.; writing—review and editing, R.L., S.X. and W.C.; visualization, S.X., W.C. and T.X.; supervision, R.L.; project administration, R.L.; funding acquisition, R.L. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported in part by the Major Science and Technology Project of Fuzhou City under Grants No. 2025FZHYJB0111, in part by the Science and Technology Project Fundation of Fuzhou City under Grant 2024-Y-011, and in part by the Special Funding Projects for Promoting High-Qualty Development of Marine and Fishery Industry of Fujian Province under Grants No. FHYF-ZH-2023-06.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflict of interest.

Appendix A

The activity of each PU is characterized by a two-state Markov chain as illustrated in Figure A1. The transition probability matrix for the n-th PU is given by
P n = p 00 p 01 p 10 p 11 ,
where p 01 denotes the probability of transitioning from the idle state to the transmission state, and p 10 denotes the probability of transitioning from the transmission state to the idle state.
Figure A1. Two-state Markov model of PU activity.
Figure A1. Two-state Markov model of PU activity.
Sensors 25 06287 g0a1

Appendix B

Following [47], the imperfect CSI of each Eve is modeled by a norm-bounded error model. Let h ^ i , m and h ^ R , m denote the estimated direct and IRS to Eve channels, respectively. The actual channels are expressed as
h i , m = h ^ i , m + Δ h i , m ,   Δ h i , m 2 ( ε i , m ) 2 ,
h R , m = h ^ R , m + Δ h R , m ,   Δ h R , m 2 ( ε r , m ) 2 ,
where Δ h i , m and Δ h R , m denote bounded estimation errors.

References

  1. Zeadally, S.; Javed, M.A.; Hamida, E.B. Vehicular communications for ITS: Standardization and challenges. IEEE Commun. Stand. Mag. 2020, 4, 11–17. [Google Scholar] [CrossRef]
  2. Chen, S.; Hu, J.; Shi, Y.; Zhao, L.; Li, W. A vision of C-V2X: Technologies, field testing, and challenges with Chinese development. IEEE Internet Things J. 2020, 7, 3872–3881. [Google Scholar] [CrossRef]
  3. Quan, W.; Liu, M.; Cheng, N.; Zhang, X.; Gao, D.; Zhang, H. Cybertwin-driven DRL-based adaptive transmission scheduling for software defined vehicular networks. IEEE Trans. Veh. Technol. 2022, 71, 4607–4619. [Google Scholar] [CrossRef]
  4. Gyawali, S.; Xu, S.; Qian, Y.; Hu, R.Q. Challenges and solutions for cellular based V2X communications. IEEE Commun. Surv. Tutor. 2020, 23, 222–255. [Google Scholar] [CrossRef]
  5. Garcia, M.H.C.; Molina-Galan, A.; Boban, M.; Gozalvez, J.; Coll-Perales, B.; Şahin, T.; Kousaridas, A. A tutorial on 5G NR V2X communications. IEEE Commun. Surv. Tutor. 2021, 23, 1972–2026. [Google Scholar] [CrossRef]
  6. Wang, P.; Wu, W.; Liu, J.; Chai, G.; Feng, L. Joint spectrum and power allocation for V2X communications with imperfect CSI. IEEE Trans. Veh. Technol. 2023, 72, 16338–16353. [Google Scholar] [CrossRef]
  7. Tahir, M.N.; Leviäkangas, P.; Katz, M. Connected vehicles: V2V and V2I road weather and traffic communication using cellular technologies. Sensors 2022, 22, 1142. [Google Scholar] [CrossRef]
  8. Wu, Q.; Zhao, Y.; Fan, Q.; Fan, P.; Wang, J.; Zhang, C. Mobility-aware cooperative caching in vehicular edge computing based on asynchronous federated and deep reinforcement learning. IEEE J. Sel. Top. Signal Process. 2022, 17, 66–81. [Google Scholar] [CrossRef]
  9. Mitola, J. Cognitive radio for flexible mobile multimedia communications. In Proceedings of the 1999 IEEE International Workshop on Mobile Multimedia Communications (MoMuC’99) (Cat. No. 99EX384), San Diego, CA, USA, 15–17 November 1999; IEEE: Piscataway, NJ, USA, 1999; pp. 3–10. [Google Scholar]
  10. Khasawneh, M.; Azab, A.; Alrabaee, S.; Sakkal, H.; Bakhit, H.H. Convergence of IoT and cognitive radio networks: A survey of applications, techniques, and challenges. IEEE Access 2023, 11, 71097–71112. [Google Scholar] [CrossRef]
  11. Chen, G.; Zhu, H.; Xu, Y.; Song, T.; Hu, J. A novel joint spectrum sensing and resource allocation scheme in cognitive internet of vehicles networks. IEEE Trans. Veh. Technol. 2024, 73, 13412–13424. [Google Scholar] [CrossRef]
  12. Tang, F.; Mao, B.; Kato, N.; Gui, G. Comprehensive survey on machine learning in vehicular network: Technology, applications and challenges. IEEE Commun. Surv. Tutor. 2021, 23, 2027–2057. [Google Scholar] [CrossRef]
  13. Hedhly, W.; Amin, O.; Alouini, M.S. Benefits of improper Gaussian signaling in interweave cognitive radio with full and partial CSI. IEEE Trans. Cogn. Commun. Netw. 2020, 6, 1256–1268. [Google Scholar] [CrossRef]
  14. Kaur, R.; Bansal, B.; Majhi, S.; Jain, S.; Huang, C.; Yuen, C. A survey on reconfigurable intelligent surface for physical layer security of next-generation wireless communications. IEEE Open J. Veh. Technol. 2024, 5, 172–199. [Google Scholar] [CrossRef]
  15. Ai, Y.; Felipe, A.; Kong, L.; Cheffena, M.; Chatzinotas, S.; Ottersten, B. Secure vehicular communications through reconfigurable intelligent surfaces. IEEE Trans. Veh. Technol. 2021, 70, 7272–7276. [Google Scholar] [CrossRef]
  16. Zhu, Y.; Mao, B.; Kato, N. Intelligent reflecting surface in 6G vehicular communications: A survey. IEEE Open J. Veh. Technol. 2022, 3, 266–277. [Google Scholar] [CrossRef]
  17. Liu, Y.; Liu, X.; Mu, X.; Hou, T.; Xu, J.; Di Renzo, M.; Al-Dhahir, N. Reconfigurable intelligent surfaces: Principles and opportunities. IEEE Commun. Surv. Tutor. 2021, 23, 1546–1577. [Google Scholar] [CrossRef]
  18. Almohamad, A.; Tahir, A.M.; Al-Kababji, A.; Furqan, H.M.; Khattab, T.; Hasna, M.O.; Arslan, H. Smart and secure wireless communications via reflecting intelligent surfaces: A short survey. IEEE Open J. Commun. Soc. 2020, 1, 1442–1456. [Google Scholar] [CrossRef]
  19. Li, Y.; Khan, F.; Ahmed, M.; Soofi, A.A.; Khan, W.U.; Sheemar, C.K.; Asif, M.; Han, Z. RIS-based Physical Layer Security for Integrated Sensing and Communication: A Comprehensive Survey. IEEE Internet Things J. 2025, 12, 32444–32468. [Google Scholar] [CrossRef]
  20. Yang, R.; Wang, D.; Zhu, S.; Zhu, C.; Bao, J.; Yang, Z.; Huang, C. Beamforming Design for RIS-aided ISCC in Internet of Vehicles Systems. IEEE Internet Things J. 2024, 12, 12153–12165. [Google Scholar] [CrossRef]
  21. Cao, X.; Wang, S.; Ren, X. IRS-Enhanced V2X Communication and Computation Systems: Resource Allocation and Performance Optimization. IEEE Internet Things J. 2024, 12, 9180–9194. [Google Scholar] [CrossRef]
  22. Duan, W.; Gu, X.; Zhang, G.; Wen, M.; Ding, Z.; Ho, P.H. Sum-rate maximization for RIS-IoV: From instantaneous to statistical CSI. IEEE Trans. Wirel. Commun. 2024, 23, 8071–8084. [Google Scholar] [CrossRef]
  23. Yu, Y.; Liu, X.; Liu, Z.; Durrani, T.S. Joint trajectory and resource optimization for RIS assisted UAV cognitive radio. IEEE Trans. Veh. Technol. 2023, 72, 13643–13648. [Google Scholar] [CrossRef]
  24. Yang, S.; Long, R.; Liang, Y.C. Active reconfigurable intelligent surface-aided cognitive radio system. In Proceedings of the ICC 2023-IEEE International Conference on Communications, Rome, Italy, 28 May–1 June 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 3866–3871. [Google Scholar]
  25. Wu, W.; Wang, Z.; Wu, Y.; Zhou, F.; Wang, B.; Wu, Q.; Ng, D.W.K. Joint sensing and transmission optimization for IRS-assisted cognitive radio networks. IEEE Trans. Wirel. Commun. 2023, 22, 5941–5956. [Google Scholar] [CrossRef]
  26. Khoshafa, M.H.; Ngatched, T.M.; Ahmed, M.H. RIS-aided physical layer security improvement in underlay cognitive radio networks. IEEE Syst. J. 2023, 17, 6437–6448. [Google Scholar] [CrossRef]
  27. Chen, Z.; Guo, Y.; Zhang, P.; Jiang, H.; Xiao, Y.; Huang, L. Physical layer security improvement for hybrid RIS-assisted MIMO communications. IEEE Commun. Lett. 2024, 28, 2493–2497. [Google Scholar] [CrossRef]
  28. Gu, X.; Duan, W.; Zhang, G.; Sun, Q.; Wen, M.; Ho, P.-H. Physical layer security for RIS-aided wireless communications with uncertain eavesdropper distributions. IEEE Syst. J. 2022, 17, 848–859. [Google Scholar] [CrossRef]
  29. Zhao, H.; Sun, W.; Ni, Y.; Xia, W.; Gui, G.; Zhu, C. Deep deterministic policy gradient-based rate maximization for RIS-UAV-assisted vehicular communication networks. IEEE Trans. Intell. Transp. Syst. 2024, 25, 15732–15744. [Google Scholar] [CrossRef]
  30. Qi, K.; Wu, Q.; Fan, P.; Cheng, N.; Chen, W.; Wang, J.; Letaief, K.B. Deep-reinforcement-learning-based aoi-aware resource allocation for ris-aided iov networks. IEEE Trans. Veh. Technol. 2024, 74, 1365–1378. [Google Scholar] [CrossRef]
  31. Dong, R.; Wang, B.; Cao, K.; Tian, J.; Cheng, T. Secure transmission design of RIS enabled UAV communication networks exploiting deep reinforcement learning. IEEE Trans. Veh. Technol. 2024, 73, 8404–8419. [Google Scholar] [CrossRef]
  32. Qin, X.; Song, Z.; Wang, J.; Du, S.; Gao, J.; Yu, W.; Sun, X. Deep-Reinforcement-Learning-Based Uplink Security Enhancement for STAR-RIS-Assisted NOMA Systems With Dual Eavesdroppers. IEEE Internet Things J. 2024, 11, 28050–28063. [Google Scholar] [CrossRef]
  33. Ju, Y.; Cao, Z.; Chen, Y.; Liu, L.; Pei, Q.; Mumtaz, S.; Dong, M.; Guizani, M. NOMA-assisted secure offloading for vehicular edge computing networks with asynchronous deep reinforcement learning. IEEE Trans. Intell. Transp. Syst. 2023, 25, 2627–2640. [Google Scholar] [CrossRef]
  34. Ju, Y.; Gao, Z.; Wang, H.; Liu, L.; Pei, Q.; Dong, M.; Mumtaz, S.; Leung, V.C. Energy-efficient cooperative secure communications in mmwave vehicular networks using deep recurrent reinforcement learning. IEEE Trans. Intell. Transp. Syst. 2024, 25, 14460–14475. [Google Scholar] [CrossRef]
  35. Liu, J.C.; Ma, J.C.; Jiang, Z.Q. AlphaSeek FinRL: A Hybrid Deep Learning Architecture for High-Frequency Cryptocurrency Trading. In Proceedings of the 2025 IEEE 11th International Conference on Intelligent Data and Security (IDS), New York, NY, USA, 9–11 May 2025; IEEE: Piscataway, NJ, USA, 2025; pp. 55–57. [Google Scholar]
  36. Ryu, H.; Sohn, C.B.; Kim, D.Y. Latent Mamba-DQN: Improving Temporal Dependency Modeling in Deep Q-Learning via Selective State Summarization. Appl. Sci. 2025, 15, 8956. [Google Scholar] [CrossRef]
  37. Ji, Y.; Wang, Y.; Zhao, H.; Gui, G.; Gacanin, H.; Sari, H.; Adachi, F. Multi-agent reinforcement learning resources allocation method using dueling double deep Q-network in vehicular networks. IEEE Trans. Veh. Technol. 2023, 72, 13447–13460. [Google Scholar] [CrossRef]
  38. Liang, L.; Ye, H.; Li, G.Y. Spectrum sharing in vehicular networks based on multi-agent reinforcement learning. IEEE J. Sel. Areas Commun. 2019, 37, 2282–2292. [Google Scholar] [CrossRef]
  39. Wang, L.; Wu, W.; Zhou, F.; Wu, Q.; Dobre, O.A.; Quek, T.Q. Hybrid hierarchical DRL enabled resource allocation for secure transmission in multi-IRS-assisted sensing-enhanced spectrum sharing networks. IEEE Trans. Wirel. Commun. 2023, 23, 6330–6346. [Google Scholar] [CrossRef]
  40. Guo, X.; Chen, Y.; Wang, Y. Learning-based robust and secure transmission for reconfigurable intelligent surface aided millimeter wave UAV communications. IEEE Wirel. Commun. Lett. 2021, 10, 1795–1799. [Google Scholar] [CrossRef]
  41. Wang, Z.; Liu, L.; Cui, S. Channel estimation for intelligent reflecting surface assisted multiuser communications: Framework, algorithms, and analysis. IEEE Trans. Wirel. Commun. 2020, 19, 6607–6620. [Google Scholar] [CrossRef]
  42. Chen, Y.; Wang, Y.; Jiao, L. Robust transmission for reconfigurable intelligent surface aided millimeter wave vehicular communications with statistical CSI. IEEE Trans. Wirel. Commun. 2021, 21, 928–944. [Google Scholar] [CrossRef]
  43. Huang, Z.; Zheng, B.; Zhang, R. Roadside IRS-aided vehicular communication: Efficient channel estimation and low-complexity beamforming design. IEEE Trans. Wirel. Commun. 2023, 22, 5976–5989. [Google Scholar] [CrossRef]
  44. Chu, Z.; Hao, W.; Xiao, P.; Mi, D.; Liu, Z.; Khalily, M.; Kelly, J.R.; Feresidis, A.P. Secrecy rate optimization for intelligent reflecting surface assisted MIMO system. IEEE Trans. Inf. Forensics Secur. 2020, 16, 1655–1669. [Google Scholar] [CrossRef]
  45. Hessel, M.; Modayil, J.; Van Hasselt, H.; Schaul, T.; Ostrovski, G.; Dabney, W.; Horgan, D.; Piot, B.; Azar, M.; Silver, D. Rainbow: Combining improvements in deep reinforcement learning. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA, 2–7 February 2018; Volume 32. [Google Scholar]
  46. Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar] [CrossRef]
  47. Liu, X.; Yu, Y.; Peng, B.; Zhai, X.B.; Zhu, Q.; Leung, V.C. RIS-UAV enabled worst-case downlink secrecy rate maximization for mobile vehicles. IEEE Trans. Veh. Technol. 2022, 72, 6129–6141. [Google Scholar] [CrossRef]
Figure 1. IRS-assisted interweave CIoV communication network under eavesdropping attacks.
Figure 1. IRS-assisted interweave CIoV communication network under eavesdropping attacks.
Sensors 25 06287 g001
Figure 2. Representation of HMA-Mamba RainbowDQN framework.
Figure 2. Representation of HMA-Mamba RainbowDQN framework.
Sensors 25 06287 g002
Figure 3. Mamba State Feature Extraction Module.
Figure 3. Mamba State Feature Extraction Module.
Sensors 25 06287 g003
Figure 4. Minimum total secrecy rate versus the number of training episodes.
Figure 4. Minimum total secrecy rate versus the number of training episodes.
Sensors 25 06287 g004
Figure 5. Average maximum SOP versus the number of training episodes.
Figure 5. Average maximum SOP versus the number of training episodes.
Sensors 25 06287 g005
Figure 6. Minimum total secrecy rate versus the number of IRS elements.
Figure 6. Minimum total secrecy rate versus the number of IRS elements.
Sensors 25 06287 g006
Figure 7. Average maximum SOP versus the number of IRS elements.
Figure 7. Average maximum SOP versus the number of IRS elements.
Sensors 25 06287 g007
Figure 8. Minimum total secrecy rate versus the number of PUs.
Figure 8. Minimum total secrecy rate versus the number of PUs.
Sensors 25 06287 g008
Figure 9. Average maximum SOP versus the number of PUs.
Figure 9. Average maximum SOP versus the number of PUs.
Sensors 25 06287 g009
Figure 10. Minimum total secrecy rate versus VUs velocity.
Figure 10. Minimum total secrecy rate versus VUs velocity.
Sensors 25 06287 g010
Figure 11. Average maximum SOP versus VUs velocity.
Figure 11. Average maximum SOP versus VUs velocity.
Sensors 25 06287 g011
Figure 12. Average minimum secrecy rate over the number of IRS elements with different number of VUs.
Figure 12. Average minimum secrecy rate over the number of IRS elements with different number of VUs.
Sensors 25 06287 g012
Figure 13. Average SOP over the number of IRS elements with different number of VUs.
Figure 13. Average SOP over the number of IRS elements with different number of VUs.
Sensors 25 06287 g013
Figure 14. Comparison of secrecy rate versus the number of IRS elements under imperfect Eve CSI.
Figure 14. Comparison of secrecy rate versus the number of IRS elements under imperfect Eve CSI.
Sensors 25 06287 g014
Figure 15. Comparison of SOP versus the number of IRS elements under imperfect Eve CSI.
Figure 15. Comparison of SOP versus the number of IRS elements under imperfect Eve CSI.
Sensors 25 06287 g015
Table 1. Simulation parameters.
Table 1. Simulation parameters.
ParameterValue
Number of VUs (I)4
Number of PUs (Q)6
Number of IRS elements (L)18
Carrier frequency2 GHz
Bandwidth1 MHz
V2I transmit power ( p i )[0, 23] dBm
BS and vehicles antenna gains8, 3 dBi
BS and vehicles receiver noise gains5, 11 dBi
Number of discrete power levels9
Vehicle speed range[10, 15] m/s
SBS antenna height25 m
VUs antenna height1.5 m
IRS height25 m
Noise power 114 dBm
V2I link path loss model 128.1 + 37.6 log 10 ( d )
Table 2. Neural network parameters.
Table 2. Neural network parameters.
ParameterValue
Number of episodes3000
Number of iterations per episode100
OptimizerAdam
Learning rate ( μ )0.001
Discount factor ( γ )0.99
Mamba internal activationSiLU
Prioritized experience replay buffer size D100,000
Target network soft update0.005
Prioritization exponent ( ω )0.5
Prioritization typeproportional
Multi-step returns (n)3
Exploration ( ε )0.0
Noisy Nets ( σ 0 )0.5
SSM state dimension ( d state )16
Network hidden dimension ( d hidden )256
Expansion factor (E)2
Table 3. Comparison of running time.
Table 3. Comparison of running time.
AlgorithmTime (s)
HMA-Mamba RainbowDQN60.06
HMA-RainbowDQN45.69
HMA-D3QN24.67
HMA-DQN21.45
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Lin, R.; Xie, S.; Chen, W.; Xu, T. Physical Layer Security Enhancement in IRS-Assisted Interweave CIoV Networks: A Heterogeneous Multi-Agent Mamba RainbowDQN Method. Sensors 2025, 25, 6287. https://doi.org/10.3390/s25206287

AMA Style

Lin R, Xie S, Chen W, Xu T. Physical Layer Security Enhancement in IRS-Assisted Interweave CIoV Networks: A Heterogeneous Multi-Agent Mamba RainbowDQN Method. Sensors. 2025; 25(20):6287. https://doi.org/10.3390/s25206287

Chicago/Turabian Style

Lin, Ruiquan, Shengjie Xie, Wencheng Chen, and Tao Xu. 2025. "Physical Layer Security Enhancement in IRS-Assisted Interweave CIoV Networks: A Heterogeneous Multi-Agent Mamba RainbowDQN Method" Sensors 25, no. 20: 6287. https://doi.org/10.3390/s25206287

APA Style

Lin, R., Xie, S., Chen, W., & Xu, T. (2025). Physical Layer Security Enhancement in IRS-Assisted Interweave CIoV Networks: A Heterogeneous Multi-Agent Mamba RainbowDQN Method. Sensors, 25(20), 6287. https://doi.org/10.3390/s25206287

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop