Next Article in Journal
Artificial Intelligence-Driven Sensing Framework with Multimodal Sensor Importance Learning for Smart Energy Systems
Next Article in Special Issue
A Wide Dynamic Range RF Attenuation Calibration System for 9 kHz to 10 MHz
Previous Article in Journal
An Indoor Mapping Algorithm Fusing LiDAR-IMU Tightly Coupled Fusion and Scan Context: IS-LEGO-LOAM
Previous Article in Special Issue
Wideband MIMO Antenna System Employing Slot and Via Loading Technique for 5G Terminals
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Deep Learning-Based Dynamic Time Division ISAC Beamforming for Vehicular Networks

Department of Electronic Engineering, Sogang University, Seoul 04107, Republic of Korea
*
Author to whom correspondence should be addressed.
Sensors 2026, 26(9), 2790; https://doi.org/10.3390/s26092790
Submission received: 24 March 2026 / Revised: 18 April 2026 / Accepted: 26 April 2026 / Published: 30 April 2026
(This article belongs to the Special Issue Feature Papers in Communications Section 2025–2026)

Abstract

Integrated sensing and communications (ISAC) is a promising key technology for vehicular networks, because it allows roadside units to support both data transmission and radar-like sensing over the same spectrum and hardware platform. In conventional time division ISAC systems, each frame is divided into sensing and communication phases with a fixed ratio, which determines the tradeoff between the sensing accuracy and the communication throughput. However, in high-mobility vehicular environments, a fixed sensing–communication split is often suboptimal due to time-varying channel and intervehicle interference variations. In this paper, we propose a dynamic sensing–communication time division and ISAC beamforming scheme that minimizes the Cramér–Rao lower bound while satisfying the minimum effective communication sum rate. We further develop a deep reinforcement learning framework based on proximal policy optimization to find the optimal time division ratio and beamforming vectors. Simulation results show that the proposed dynamic time division beamforming scheme significantly outperforms the conventional fixed time division beamforming schemes in terms of sensing accuracy and the communication sum rate.

1. Introduction

Rapid advancements in intelligent transportation systems and autonomous driving technologies have led to a significant increase in wireless traffic between vehicles and networks, as well as between vehicles. Therefore, vehicle-to-everything (V2X) technologies are required to simultaneously provide high-data-rate services and accurate situational awareness within limited spectrum resources, even in highly dynamic wireless environments [1,2,3].
Recently, integrated sensing and communications (ISAC) technology has attracted significant attention as a promising solution that enables wireless communication and sensing in the same system. ISAC performs both radar-like sensing and data transmission over the same spectrum and hardware platform, thereby improving the spectral efficiency while reducing the hardware cost and system complexity [4,5,6]. Previous ISAC studies can be broadly categorized into those on time division ISAC (TD-ISAC), resource allocation-based ISAC, and beamforming-oriented ISAC.
Many studies have focused on TD-ISAC systems, which separate sensing and communication into different time intervals [7,8]. In particular, TD-ISAC has attracted considerable attention because it can address challenges such as full-duplex hardware design and interference suppression by separating the sensing and communication phases in the time domain [9,10]. For this reason, TD-ISAC is often regarded as a practical first step toward ISAC implementation. The authors of [7] proposed a TD-ISAC system for connected automated vehicles, in which they considered various sensing and communication time configurations and evaluated the performance under different sensing-to-communication-time ratios. The authors of [11] studied the time allocation problem in TD-ISAC systems and mathematically analyzed the relationship between radar detection accuracy and the achievable data rate for a given fixed time interval. The authors of [12] focused on the design of a cross-layer scheduling policy for a TD-ISAC system with a buffer-equipped base station in a single-input single-output (SISO) scenario, where they used a queuing model to derive the communication performance and assumed that the radar detection accuracy is determined by the allocated time resources. In a TD-ISAC system, the time ratio between sensing and communication determines the tradeoff between the two functions. In high-mobility vehicular environments, it becomes increasingly important to dynamically adjust the sensing-to-communication-time ratio in response to channel variations. In addition, the roadside unit (RSU) should determine the beamforming matrices for sensing and communication [13]. However, most previous TD-ISAC studies have considered a predetermined or enumerated sensing/communication time split and then evaluated performance under this fixed ratio. Although these studies provide useful insights into the sensing–communication tradeoff, they do not dynamically optimize the time split according to channel and mobility variations.
Some studies have addressed time, frequency, and power allocation in ISAC systems [14,15,16]. The authors of [17] analyzed the performance tradeoff between sensing and communication in frequency division ISAC networks, where separate frequency resources are assigned to sensing and data transmission. In particular, they mathematically derived the probability of detection for sensing and the probability of coverage for communication under a given transmit power and bandwidth budget. The authors of [18,19,20,21] investigated joint time–frequency resource allocation. Specifically, the authors of [18,19,20] proposed resource optimization designs that minimize the Cramér–Rao lower bound (CRLB) for delay and Doppler estimation, whereas the authors of [21] focused on suppressing delay Doppler sidelobes. The authors of [22] aimed to jointly optimize resource and power allocation for a monostatic orthogonal frequency division multiplexing (OFDM)–ISAC system while satisfying sensing signal-to-noise ratio (SNR) and communication sum rate requirements. However, although these studies dynamically allocate time–frequency resources, they do not consider beamforming matrices for sensing and data transmission.
ISAC beamforming studies have mainly focused on designing beams for simultaneous sensing and communication, where the beamformer matrix determines the tradeoff between sensing and communication performance [23,24]. A large portion of the ISAC beamforming literature considers simultaneous sensing–communication transmission, where optimization focuses on beam patterns, transmit covariance matrices, or communication precoders, rather than on dynamic frame-level time division. The authors of [24,25,26,27] investigated joint beam pattern design for ISAC, aiming either to improve the sensing accuracy by minimizing the CRLB or to enhance the communication capacity by maximizing the sum rate. In particular, the authors of [24] formulated an ISAC beamforming design for vehicle-mounted transmitters in the uplinks of vehicle-to-infrastructure (V2I) networks, while the authors of [25] proposed sensing-assisted predictive beam tracking in the downlinks of V2I networks. Recently, many researchers have utilized deep learning or reinforcement learning for beamforming design in ISAC networks [28,29]. The authors of [30,31] used deep reinforcement learning (DRL) for dynamic beamforming and power allocation in ISAC systems, and the authors of [32] proposed an energy-efficient DRL-based beamforming scheme for ISAC-enabled V2X networks. However, these ISAC beamforming studies do not formulate the joint TD-ISAC problem with both adaptive time partitioning and multi-user beamforming—that is, they do not dynamically optimize frame-level time division.
In this paper, we propose a dynamic time allocation and beamforming scheme for TD-ISAC systems. The proposed scheme jointly optimizes the sensing-to-communication-time allocation and transmit beamforming vectors for sensing and communication. Our paper differs from prior studies in that it integrates two strongly coupled decisions that are usually treated separately: how much frame time should be assigned to sensing vs. communication and how the corresponding sensing and communication beams should be designed under this time split. This coupling is central to our formulation. For example, a larger sensing ratio improves the effective sensing SNR and reduces the CRLB, but it shortens the communication phase and reduces the effective sum rate. At the same time, the beamforming quality affects both the sensing gain and communication SINR. Because these decisions interact nonlinearly and vary over time with vehicle positions and channels, it is insufficient to optimize them independently. The main contributions of this paper are as follows. First, we propose a joint optimization scheme for time allocation and beamforming vectors in ISAC-aided V2I networks. Specifically, the proposed scheme dynamically adjusts the sensing-to-communication-time ratio and optimizes the beamforming vectors for sensing and communication according to the channel condition. Second, we incorporate a CRLB-oriented sensing objective together with a minimum effective sum rate constraint, thereby explicitly balancing sensing accuracy and communication performance. Third, we develop a proximal policy optimization (PPO) algorithm to solve the optimization problem, where the reward is determined as the weighted sum of the sensing performance and the communication performance.
The rest of this paper is organized as follows. Section 2 describes the system and channel models. Section 3 presents the sensing and communication performance. Section 4 formulates the optimization problem and presents the proposed dynamic time division (DTD)–ISAC beamforming scheme, where the state, action, and reward of DRL are described in detail. Section 5 provides the simulation results, and Section 6 concludes the paper.

2. System Model

2.1. System Description

We consider a downlink of an ISAC-enabled V2I network with a single RSU and K vehicles, as shown in Figure 1a. An RSU is equipped with a uniform linear transmit array of N t antennas and with a uniform linear receive array of N r antennas dedicated to monostatic radar sensing. The RSU serves K vehicles, where each vehicle is equipped with M receive antennas.
The RSU carries out the sensing and communication functions in a time division manner. As shown in Figure 1b, each frame is divided into a sensing phase with a ratio of ρ and a communication phase with a ratio of 1 ρ , where ρ [ 0 , 1 ] . During the sensing phase of ρ Δ T , the RSU transmits a dedicated sensing waveform and receives the reflected echoes to estimate the angle and range of the vehicles. During the communication phase of ( 1 ρ ) Δ T , the RSU transmits independent K data streams to the K vehicles by using a multi-user transmit beamforming matrix.

2.2. Channel Model

The wireless channel between the RSU and the k-th vehicle is characterized by a line-of-sight (LoS) propagation model, suitable for highway and open-road scenarios [2,33]. Let β k ( t ) denote the large-scale path loss coefficient and θ k ( t ) denote the geometric angle of departure (AoD) from the RSU to vehicle k. The downlink channel H k ( t ) C M × N t from the RSU to vehicle k at time t can then be expressed as follows:
H k ( t ) = β k ( t ) a rx , veh ( θ k ( t ) ) a tx H ( θ k ( t ) ) ,
where the vector a tx ( · ) represents the transmit array response at the RSU, and the vector a rx , veh ( · ) denotes the receive array response at the vehicle. Assuming uniform linear arrays (ULAs) with half-wavelength spacing, the transmit and receive array responses are, respectively, given by
a tx ( θ ) = 1 N t 1 , e j π sin θ , , e j π ( N t 1 ) sin θ T ,
a rx , veh ( θ ) = 1 M 1 , e j π sin θ , , e j π ( M 1 ) sin θ T .
The path loss coefficient β k ( t ) follows a standard distance-dependent decay model as follows:
β k ( t ) = β 0 r 0 r k ( t ) α ,
where r k ( t ) is the distance between the RSU and vehicle k at time t, r 0 is a reference distance, β 0 is the path loss at the reference distance, and α is the path loss exponent.

3. Performance Evaluation in a TD-ISAC System

The RSU employs different transmit signals during the sensing and communication phases. The total transmit power of an RSU during a frame is constrained by a budget P tot , which is shared between the sensing and communication phases. In a TD-ISAC system, the frame duration Δ T is divided into a sensing phase and a communication phase. First, during ρ Δ T , an RSU carries out the sensing function. During this sensing phase, the RSU transmits a dedicated probing signal and receives echoes on its sensing array. Then, during the remaining ( 1 ρ ) Δ T , the RSU carries out the communication function. During this communication phase, the RSU serves K vehicles simultaneously using downlink multi-user beamforming.

3.1. Sensing Performance

During the sensing phase, the RSU transmits a dedicated sensing beam f sens C N t × 1 using a probing waveform x sens ( t ) . The RSU then receives the reflected echoes with a sensing array consisting of N r elements. After matched filtering and coherent integration over τ independent echoes within the sensing period ρ Δ T , the RSU forms sufficient statistics for estimating the angle θ k and range r k of each vehicle. The quality of these estimates is captured through the CRLB [34,35], which provides a lower bound on the variance of any unbiased estimator of a parameter.
The effective sensing SNR for target vehicle k after the coherent integration of τ echoes within the sensing period ρ Δ T is proportional to the sensing time, the path loss, and the array gain in the transmit direction, as follows [36,37]:
SNR sens , k ρ · τ β k | a tx H ( θ k ) f sens | 2 σ sens 2 ,
where ρ is the sensing-to-communication-time ratio, β k is the path loss coefficient for vehicle k, a tx ( · ) is the transmit array response at the RSU, f sens is the sensing beamforming vector, and σ sens 2 is the noise power.
The CRLB for angle estimation is inversely proportional to the sensing SNR and the square of the effective aperture D ap . Hence, the CRLB for the angle estimation of vehicle k can then be expressed as follows:
CRLB θ , k 1 SNR sens , k D ap 2 .
Similarly, the CRLB for range estimation is inversely proportional to the sensing SNR and the square of the system bandwidth B; therefore, the CRLB for the range estimation of vehicle k can be expressed as follows:
CRLB r , k 1 SNR sens , k B 2 .
Consequently, the integrated CRLB for the sensing performance can be expressed as a weighted sum of the per-user angle and range CRLBs as follows:
Ω sens = w θ 1 K k = 1 K CRLB θ , k + ( 1 w θ ) 1 K k = 1 K CRLB r , k ,
where CRLB θ , k and CRLB r , k are, respectively, the angle and range CRLBs of vehicle k, and w θ 0 is a system parameter that controls the relative weighting of the angle and range estimations.
In particular, increasing the sensing duration ρ Δ T allows the receiver to accumulate more energy from the reflected signals, thereby linearly increasing the effective sensing SNR. Because the CRLB is inversely proportional to the SNR from (6) and (7), allocating more time resources to sensing directly suppresses the estimation error variance.

3.2. Communication Performance

During the communication phase, the RSU transmits data symbols to K vehicles using a transmit beamforming matrix
F comm = f 1 , , f K C N t × K ,
where f k C N t × 1 is the beamforming vector for vehicle k. Let s k C denote the information symbol intended for vehicle k, with unit average power E [ | s k | 2 ] = 1 . The transmitted signal vector at the RSU is
x comm = k = 1 K f k s k = F comm s ,
where s = [ s 1 , , s K ] T .
The received signal vector at vehicle k is
y k = H k x comm + n k ,
where n k CN ( 0 , σ comm 2 I M ) is additive white Gaussian noise (AWGN). Vehicle k applies a linear receive combiner w k C M × 1 to obtain the decision variable
s ^ k = w k H y k .
Here, we choose common maximum ratio combining (MRC) for w k as follows:
w k = H k f k H k f k ,
which maximizes the instantaneous receive SNR in single-user settings.
The instantaneous signal-to-interference-plus-noise ratio (SINR) at vehicle k under linear receive combining is given by
SINR k = w k H H k f k 2 j k w k H H k f j 2 + σ comm 2 w k 2 .
Assuming Gaussian signaling and ideal coding, the effective sum rate per frame is given by
R comm = ( 1 ρ ) · k = 1 K log 2 1 + SINR k .

4. Proposed DTD-ISAC Beamforming Scheme

4.1. Problem Formulation

Our objective is to minimize the integrated CRLB of Ω sens ( t ) at each decision epoch t, while satisfying the effective sum rate requirement for communication greater than or equal to a threshold of R min . In the proposed DTD-ISAC system, the sensing–communication time split can be dynamically adjusted according to the time-varying channel. Hence, the parameters to be dynamically determined at time t are as follows: the sensing-to-communication-time ratio of ρ t , the sensing beamforming vector of f sens , t , and the communication beamforming vectors of { f k , t } k = 1 K . The optimization problem can then be expressed as follows:
min ρ t , f sens , t , f 1 , t , , f K , t lim sup T 1 T E t = 0 T 1 Ω sens ( t ) s . t . R comm ( t ) R min , ρ min ρ t ρ max , k = 1 K f k , t 2 + f sens , t 2 P tot .
It is difficult to directly solve (16) due to the fact that Ω sens ( t ) and R comm ( t ) have a non-convex dependency on the beamforming vectors, the temporal coupling induced by the channel and mobility dynamics, and the requirement for real-time decision-making. Consequently, we reformulate the problem of (16) using a DRL framework based on the PPO algorithm.

4.2. MDP Modeling

Let π denote a control policy that maps observations o t at decision epoch t to a continuous action vector a t :
a t = π ( o t ) .
To represent the tradeoff between sensing and communication performance as a single scalar measure, an integrated reward is defined as follows:
r t = Ω sens ( t ) λ rate R min R comm ( t ) + + λ exp ( 1 2 | ρ t 0.5 | ) ,
where [ x ] + = max ( x , 0 ) , λ rate is a penalty weight that controls the severity of violations of the communication rate requirement, and λ exp is a weight for an exploration bonus. In the reward of (18), each term is designed to reflect a specific control objective. The first term directly encourages the reduction of the integrated sensing error represented by the CRLB. The second term imposes a penalty when the communication sum rate falls below the minimum required threshold; therefore, λ rate is introduced to enforce the communication rate constraint with sufficient importance during training. The third term provides an exploration incentive for the sensing time ratio ρ t so that the learned policy does not become prematurely biased toward extreme boundary values in the early stage of learning. Accordingly, λ exp is used as a regularization-type parameter to support stable exploration rather than as a dominant performance term. These weighting parameters can be selected empirically so that the reward components have a comparable influence during training and the agent can learn a balanced policy without excessive bias toward only one objective. More specifically, λ rate is chosen to be large enough to strongly discourage rate constraint violation, while λ exp is set to be relatively small so that it assists exploration without overriding the main sensing–communication tradeoff.
Given a policy π , the long-term performance over a horizon of T epi frames is characterized by the expected discounted return
J ( π ) = E π t = 0 T epi 1 γ r t ,
where γ ( 0 , 1 ] is a discount factor. The control objective within the DRL framework is to find a policy π that maximizes the expected return, as follows:
π = arg max π J ( π ) .
The DTD-ISAC control problem can be modeled as a Markov decision process (MDP) described by the tuple ( S , A , P , R , γ ) , where S is the state, A is the action, P is the transition, R is the reward, and γ is the discount factor.

4.2.1. State and Observation

The underlying system state at time t, denoted by s t S , includes the positions, velocities, and channel realizations of all vehicles, as follows:
s t = p t , v t , { H k ( t ) } k = 1 K ,
where p t is the positions of vehicles, v t is the velocities of vehicles, and H k ( t ) is the channel of vehicle k. The policy operates on an observation vector o t = O ( s t ) , where O ( · ) denotes the observation function. Note that performance metrics are normalized using statistics collected from an initial calibration phase to ensure stable training.

4.2.2. Action Space

The RSU controls the high-dimensional beamforming vectors, but learning these vectors directly is computationally complex. Hence, we employ a parameterized action space, where the agent controls the beamforming direction and power allocation. The action vector a t R 2 + 2 K at time t is defined as
a t = ρ t , p sens , t , Δ ϕ 1 , t , , Δ ϕ K , t , p 1 , t , , p K , t ,
where ρ t ( 0 , 1 ) is the sensing-to-communication-time ratio, p sens , t ( 0 , 1 ) is the proportion of the total power allocated to sensing, and p k , t is a power allocation coefficient that determines the power of the communication beam for vehicle k. Additionally, Δ ϕ k , t is an angular offset applied to the estimated line-of-sight (LoS) angle for vehicle k. The communication beam f k , t is constructed as a steering vector pointing toward θ ^ k + Δ ϕ k , t .
This structured action design is adopted to maintain the tractability of the DRL problem in a high-mobility vehicular environment. Directly optimizing full complex-valued sensing and communication beamforming vectors would lead to a much higher-dimensional continuous action space and significantly increase the training difficulty. Therefore, the proposed parameterization captures the dominant beam steering and power allocation decisions relevant to directional vehicular channels while keeping the learning problem manageable. Although this action space does not represent the most general beamforming structure and may not achieve globally optimal unconstrained precoding, it provides a practical and stable formulation for jointly optimizing time division and beam control in TD-ISAC vehicular networks.

4.2.3. Transition and Reward

The state evolves according to a stochastic transition probability P ( s t + 1 s t , a t ) that reflects the mobility and channel dynamics of vehicles. The transition model is realized through a simulation environment that generates vehicle trajectories and channel realizations in a vehicular network.
The reward at time t is given by (18). The reward combines the sensing performance Ω sens ( t ) and the communication performance R comm ( t ) , which encourages the DRL agent to reduce the CRLB-based sensing error while meeting the minimum communication rate requirement.

4.3. PPO-Based Learning

We apply an actor–critic architecture via PPO to learn a stochastic policy π θ ( a o ) , parameterized by θ , where a is the action and o is the observation. PPO is a policy gradient method that employs a clipped surrogate objective to stabilize policy updates [38].
The policy network takes the observation o t as input and outputs the parameters (for example, mean and variance) of a multivariate Gaussian distribution over the continuous action space. The value network has a similar structure but outputs a scalar estimate V ϕ ( o t ) of the state value, with parameters ϕ . In this paper, both networks are implemented using two fully connected hidden layers with 128 units each and nonlinear activation functions.
Let θ old denote the policy parameters before an update. For a batch of trajectories, the PPO objective is given by
L PPO ( θ ) = E t min p t ( θ ) A ^ t , clip p t ( θ ) , 1 ϵ , 1 + ϵ A ^ t ,
where ϵ > 0 is a clipping parameter, A ^ t is an estimate of the advantage function at time step t, clip ( x , a , b ) is a clip function of [38], and p t ( θ ) is the probability ratio as follows:
p t ( θ ) = π θ ( a t o t ) π θ old ( a t o t ) .
The value network parameters ϕ are updated by minimizing the squared value loss as follows:
L V ( ϕ ) = E t V ϕ ( o t ) V ^ t 2 ,
where V ^ t is a target value, such as the empirical return.
In the PPO framework, the training process alternates between collecting trajectories using the current policy and updating the policy and value networks based on the collected data. At time step t, the agent first takes an action a t based on the current observation o t using the policy π . Then, the agent calculates the reward r t of (18) and the next observation o t + 1 . The transition ( o t , a t , r t ) is stored for later processing. After computing the advantage estimates { A ^ t } , the parameters θ are updated using the mini-batch stochastic gradient descent (SGD) method by increasing the PPO objective of (23), and the parameters ϕ are updated by decreasing the value loss of (25). This procedure is repeated until the policy converges or a termination criterion, such as a maximum number of iterations, is reached. The details of the PPO-based DTD-ISAC beamforming scheme are summarized in Algorithm 1.
Algorithm 1 PPO-based DTD-ISAC beamforming
 1:
Initialize the policy parameters θ and value parameters ϕ .
 2:
for  episode = 0 , 1 , , N ep 1  do
 3:
   Reset the simulation environment and obtain the initial observation o 0 .
 4:
   for time step t = 0 , 1 , , N T 1 do
 5:
      Take action a t π θ ( · o t ) .
 6:
      Observe reward r t according to (18).
 7:
      Observe the next observation o t + 1 .
 8:
      Store ( o t , a t , r t ) for later processing.
 9:
   Compute empirical returns and advantage estimates { A ^ t } .
10:
  Update θ by increasing the PPO objective of (23) and update ϕ by decreasing the value loss of (25).
11:
  end for
12:
end for

5. Simulation Results

We evaluate the performance of the proposed DTD-ISAC beamforming scheme in a V2I downlink scenario with a single RSU and K = 3 vehicles, where the vehicles move at a constant speed of v [10 m/s, 12 m/s] on a road. The vehicles are randomly dropped onto different lanes. The simulation parameters are summarized in Table 1. We run N ep = 100 episodes with independent initial positions and velocities, where each episode consists of N T = 100 time steps with duration Δ T = 0.1 s ; therefore, the episode duration is T ep = N T Δ T = 10 s .
The parameters for the PPO algorithm are as follows: the mini-batch size is 256, the discount factor is 0.99, the generalized advantage estimator (GAE) parameter is 0.95, the clipping range is 0.2, the number of hidden layers for the actor and critic networks is 2, the rollout length per update is 100 frames, the policy learning rate is 0.0003, the value learning rate is 0.001, the value loss coefficient is 0.5, the entropy regularization weight is 0.01, the number of PPO epochs per update is 10, and the number of training episodes is 100 per scenario.
We compare the performance of the proposed DTD-ISAC beamforming scheme with that of the conventional fixed time division (FTD)–ISAC beamforming scheme, which has a deterministically fixed sensing-to-communication-time ratio, ρ { 0.1 , 0.3 , 0.5 , 0.7 , 0.9 } . The performance of the ISAC beamforming schemes is evaluated in terms of sensing and communication performance metrics such as the average CRLB for angle estimation, the average CRLB for range estimation, and the average sum rate, as follows:
R ¯ comm = 1 N ep N T e = 1 N ep t = 1 N T R comm ( e ) ( t ) ,
CRLB ¯ θ = 1 N ep N T K e = 1 N ep t = 1 N T k = 1 K CRLB θ , k ( e ) ( t ) ,
CRLB ¯ r = 1 N ep N T K e = 1 N ep t = 1 N T k = 1 K CRLB r , k ( e ) ( t ) .
Figure 2 shows the average CRLB for angle estimation, and Figure 3 shows the average CRLB for range estimation, where the x-axis in both figures represents the average sum rate. In conventional FTD-ISAC beamforming, as the value of ρ decreases, the communication phase lengthens, increasing the average sum rate, but the sensing phase shortens, degrading the sensing performance, namely CRLB ¯ θ and CRLB ¯ r . Additionally, conventional FTD-ISAC beamforming satisfies the communication sum rate requirement only when ρ is less than or equal to 0.3 . However, the proposed DTD-ISAC beamforming scheme dynamically adjusts the sensing-to-communication-time ratio, ρ , depending on the channel environment. Hence, the proposed DTD-ISAC beamforming scheme shows good sensing accuracy while satisfying the minimum communication sum rate, R min . Because the conventional FTD-ISAC beamforming scheme cannot account for changes in wireless channels by fixing the sensing-to-communication-time ratio, the conventional scheme either significantly degrades the sensing accuracy when satisfying the communication data rate requirement or, conversely, significantly degrades the communication data rate when increasing the sensing accuracy. In particular, when ρ is 0.3 in the conventional scheme, the proposed scheme has similar sum rate performance but improves the CRLB by more than 80%. When ρ is 0.7 in the conventional scheme, the proposed scheme has similar CRLB performance but improves the sum rate by more than 45%.
Figure 4 shows the experimental distribution of ρ during the simulation. It should be noted that the distribution of ρ varies depending on the channel environment. As shown in Figure 4, the value of ρ is not constant but varies as the vehicle moves. In this simulation, the average value of ρ is about 0.66 .

6. Conclusions

In this paper, we propose a PPO-based DTD-ISAC beamforming scheme for V2I networks. The proposed scheme dynamically allocates time resources between sensing and data transmission while jointly optimizing the beamforming vectors in order to improve the sensing accuracy under the minimum communication data rate requirement. We developed an optimization problem that determines the sensing-to-communication-time ratio, the sensing beamforming vector, and the communication beamforming vector. However, because it is too difficult to solve the problem directly, we used the PPO algorithm by modeling the integrated reward, which includes the tradeoff between sensing and communication performance, as a single scalar measure. The proposed DRL agent observes the mobility-related states of the vehicles and link conditions, and it outputs in real time both the sensing-to-communication-time ratio and beamforming vectors. In the proposed DTD-ISAC system, the agent effectively increases the sensing time when the vehicle is far away or moving rapidly but reduces the sensing time when the communication data rate decreases to the minimum required rate. The proposed PPO-based DTD-ISAC beamforming scheme shows significantly better performance than the conventional FTD-ISAC beamforming scheme in terms of the sensing accuracy and data rate.
In this work, although the simulation was conducted with a limited number of vehicles, the overall performance trends of the proposed method would remain similar across scenarios, since the main advantage of the proposed DTD-ISAC framework arises from dynamically adjusting the sensing-to-communication-time ratio and beamforming decisions according to the channel and mobility conditions. When there are K vehicles, the action dimension is 2 + 2 K , since the agent determines one sensing-to-communication-time ratio, one sensing power ratio, K steering offsets, and K communication power allocation variables. Accordingly, as K increases, the learning space becomes larger and the computational burden for both training and action selection also increases. Nevertheless, because the proposed action structure scales linearly with K, the framework remains applicable to larger vehicular scenarios, although with increased training time and implementation complexity.
Future research may include performance evaluations under more realistic channel models. Although this work adopts an LoS-dominant channel model, LoS-only channel models cannot fully reflect the various propagation characteristics encountered in practical vehicular environments, where NLoS components, blockage events, and time-varying fading can significantly affect both the sensing and communication performance. Hence, extending the proposed framework to more realistic propagation environments that include NLoS propagation, blockages, Doppler spread, and time-varying fading is an important direction for future research [39]. Moreover, the research would be strengthened by including comparisons with other dynamic or learning-based approaches, such as optimization-based adaptive schemes or alternative DRL algorithms.

Author Contributions

Conceptualization, J.L. and J.S.; software, J.L.; validation, J.S.; investigation, J.S.; writing—original draft preparation, J.S.; writing—review and editing, J.S.; supervision, J.S.; project administration, J.S.; funding acquisition, J.S. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Sogang University Research Grant of 2026 (202615001.01).

Data Availability Statement

Data are contained within the article.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Clancy, J.; Mullins, D.; Deegan, B.; Horgan, J.; Ward, E.; Eising, C.; Denny, P.; Jones, E.; Glavin, M. Wireless access for V2X communications: Research, challenges and opportunities. IEEE Commun. Surv. Tutor. 2024, 26, 2082–2119. [Google Scholar] [CrossRef]
  2. Lee, J.; Kim, H.; So, J. Reinforcement learning-based joint beamwidth and beam alignment interval optimization in V2I communications. Sensors 2024, 24, 837. [Google Scholar] [CrossRef]
  3. Yang, Q.; Liu, K.; Chen, L.; Zhang, Z.; Zhao, J. Adaptive beam prediction for enhancing mmWave V2I communication performance in complex real-world scenarios. IEEE Internet Things J. 2025, 12, 35717–35730. [Google Scholar] [CrossRef]
  4. Liu, F.; Cui, Y.; Masouros, C.; Xu, J.; Han, T.X.; Eldar, Y.C.; Buzzi, S. Integrated sensing and communications: Toward dual-functional wireless networks for 6G and beyond. IEEE J. Sel. Areas Commun. 2022, 40, 1728–1767. [Google Scholar] [CrossRef]
  5. Wei, Z.; Qu, H.; Wang, Y.; Yuan, X.; Wu, H.; Du, Y.; Han, K.; Zhang, N.; Feng, Z. Integrated sensing and communication signals toward 5G-A and 6G: A survey. IEEE Internet Things J. 2023, 10, 11068–11092. [Google Scholar] [CrossRef]
  6. Mollahosseini, P.; Shen, R.; Ahmed Khan, T.; Ghasempour, Y. Integrated sensing and communication: Research advances and industry outlook. IEEE J. Sel. Top. Electromagn. Antennas Propag. 2025, 1, 375–392. [Google Scholar] [CrossRef]
  7. Zhang, Q.; Sun, H.; Gao, X.; Wang, X.; Feng, Z. Time-division ISAC enabled connected automated vehicles cooperation algorithm design and performance evaluation. IEEE J. Sel. Areas Commun. 2022, 40, 2206–2218. [Google Scholar] [CrossRef]
  8. Zhang, Q.; Wang, X.; Li, Z.; Wei, Z. Design and performance evaluation of joint sensing and communication integrated system for 5G mmWave enabled CAVs. IEEE J. Sel. Top. Signal Process. 2021, 15, 1500–1514. [Google Scholar] [CrossRef]
  9. Zhang, J.A.; Rahman, M.L.; Wu, K.; Huang, X.; Guo, Y.J.; Chen, S.; Yuan, J. Enabling joint communication and radar sensing in mobile networks—A survey. IEEE Commun. Surv. Tutor. 2022, 24, 306–345. [Google Scholar]
  10. Cheng, X.; Duan, D.; Gao, S.; Yang, L. Integrated sensing and communications (ISAC) for vehicular communication networks (VCN). IEEE Internet Things J. 2022, 9, 23441–23451. [Google Scholar] [CrossRef]
  11. Chen, Y.; Gu, X. Time allocation for integrated bi-static radar and communication systems. IEEE Commun. Lett. 2021, 25, 1033–1036. [Google Scholar]
  12. Xie, Z.; Li, R.; Jiang, Z.; Zhu, J.; She, Z.; Chen, P. Optimal scheduling policy for time-division joint radar and communication systems: Cross-Layer design and sensing for free. IEEE Internet Things J. 2023, 10, 20746–20760. [Google Scholar]
  13. Xue, Q.; Ji, C.; Ma, S.; Guo, J.; Xu, Y.; Chen, Q.; Zhang, W. A survey of beam management for mmWave and THz communications towards 6G. IEEE Commun. Surv. Tutor. 2024, 26, 1520–1559. [Google Scholar]
  14. Chang, B.; Tang, W.; Yan, X.; Tong, X.; Chen, Z. Integrated scheduling of sensing, communication, and control for mmWave/THz communications in cellular connected UAV networks. IEEE J. Sel. Areas Commun. 2022, 40, 2103–2113. [Google Scholar]
  15. Chu, N.H.; Nguyen, D.N.; Hoang, D.T.; Pham, Q.V.; Phan, K.T.; Hwang, W.J.; Dutkiewicz, E. AI-enabled mm-Waveform configuration for autonomous vehicles with integrated communication and sensing. IEEE Internet Things J. 2023, 10, 16727–16743. [Google Scholar]
  16. Zhao, C.; Liu, X.; Wu, J.; Liu, X. Integrated sensing and communication task optimization for cellular-connected UAV. IEEE Trans. Veh. Technol. 2026, 75, 5215–5220. [Google Scholar]
  17. Li, X.; Guo, S.; Li, T.; Zou, X.; Li, D. On the performance trade-off of distributed integrated sensing and communication networks. IEEE Wirel. Commun. Lett. 2023, 12, 2033–2037. [Google Scholar] [CrossRef]
  18. Ni, Z.; Zhang, J.A.; Huang, X.; Liu, R.P. Frequency-time resource allocation for multiuser uplink ISAC systems. IEEE Trans. Veh. Technol. 2024, 73, 18893–18906. [Google Scholar] [CrossRef]
  19. Mura, S.; Tagliaferri, D.; Mizmizi, M.; Spagnolini, U.; Petropulu, A. Waveform design for OFDM-based ISAC systems under resource occupancy constraint. In Proceedings of the IEEE Radar Conf. (RadarConf), Denver, CO, USA, 6–10 May 2024; pp. 1–6. [Google Scholar]
  20. Mura, S.; Tagliaferri, D.; Mizmizi, M.; Spagnolini, U.; Petropulu, A. Optimized waveform design for OFDM-based ISAC systems under limited resource occupancy. IEEE Trans. Wirel. Commun. 2025, 24, 5241–5254. [Google Scholar] [CrossRef]
  21. Zhang, F.; Mao, T.; Liu, R.; Han, Z.; Chen, S.; Wang, Z. Crossdomain dual-functional OFDM waveform design for accurate sensing/positioning. IEEE J. Sel. Areas Commun. 2024, 42, 2259–2274. [Google Scholar]
  22. Li, P.; Li, M.; Liu, R.; Liu, Q.; Swindlehurst, A.L. Sensing-oriented adaptive resource allocation designs for OFDM-ISAC systems. IEEE Trans. Signal Process. 2025, 73, 5121–5135. [Google Scholar] [CrossRef]
  23. Liao, B.; Ngo, H.Q.; Matthaiou, M.; Smith, P.J. Power allocation for massive MIMO-ISAC systems. IEEE Trans. Wirel. Commun. 2024, 23, 14232–14248. [Google Scholar] [CrossRef]
  24. Cong, D.; Guo, S.; Dang, S.; Zhang, H. Vehicular behavior-aware beamforming design for integrated sensing and communication systems. IEEE Trans. Intell. Transp. Syst. 2023, 24, 5923–5935. [Google Scholar] [CrossRef]
  25. Du, Z.; Liu, F.; Yuan, W.; Masouros, C.; Zhang, Z.; Xia, S.; Caire, G. Integrated sensing and communications for V2I networks: Dynamic predictive beamforming for extended vehicle targets. IEEE Trans. Wirel. Commun. 2023, 22, 3612–3627. [Google Scholar] [CrossRef]
  26. Zhao, Z.; Zhang, L.; Jiang, R.; Zhang, X.P.; Tang, X.; Dong, Y. Joint beamforming scheme for ISAC systems via robust Cramér–Rao bound optimization. IEEE Wirel. Commun. Lett. 2024, 13, 889–893. [Google Scholar] [CrossRef]
  27. Zhang, X.; Yuan, W.; Liu, C.; Wu, J.; Ng, D.W.K. Predictive beamforming for vehicles with complex behaviors in ISAC systems: A deep learning approach. IEEE J. Sel. Top. Signal Process. 2024, 18, 828–841. [Google Scholar] [CrossRef]
  28. Liu, C.; Yuan, W.; Li, S.; Liu, X.; Li, H.; Ng, D.W.K.; Li, Y. Learning-based predictive beamforming for integrated sensing and communication in vehicular networks. IEEE J. Sel. Areas Commun. 2022, 40, 2317–2334. [Google Scholar] [CrossRef]
  29. Xia, F.; Fei, Z.; Huang, J.; Wang, X.; Wang, R.; Yuan, W.; Ng, D.W.K. Sensing-enabled predictive beamforming design for RIS-assisted V2I systems: A deep learning approach. IEEE Trans. Wirel. Commun. 2024, 23, 5571–5586. [Google Scholar] [CrossRef]
  30. Liu, Y.; Zhang, S.; Li, X.; Huang, Y.; Fang, Y. Deep reinforcement learning-based beamforming design in ISAC-assisted vehicular networks. In Proceedings of the IEEE Wireless Communications and Networking Conference (WCNC), Dubai, United Arab Emirates, 21–24 April 2024; pp. 1–6. [Google Scholar]
  31. Dao, D.N.; Kokkeler, A.B.J.; Zhang, H.; Miao, Y. Dynamic beamforming and power allocation in ISAC via deep reinforcement learning. In Proceedings of the IEEE Global Communications Conference (GLOBECOM), Taipei, Taiwan, 8–12 December 2025; pp. 3542–3548. [Google Scholar]
  32. Shang, C.; Yu, J.; Hoang, D.T. Energy-efficient learning-based beamforming for ISAC-enabled V2X networks. In Proceedings of the IEEE Global Communications Conference (GLOBECOM), Taipei, Taiwan, 8–12 December 2025; pp. 925–930. [Google Scholar]
  33. Va, V.; Shimizu, T.; Bansal, G.; Heath, R., Jr. Beam design for beam switching based millimeter wave vehicle-to-infrastructure communications. In Proceedings of the IEEE International Conference on Communications (ICC), Kuala Lumpur, Malaysia, 23–27 May 2016; pp. 1–6. [Google Scholar]
  34. Tang, X.; Tang, J.; He, Q.; Wan, S.; Tang, B.; Sun, P.; Zhang, N. Cramer–Rao bounds and coherence performance analysis for next generation radar. Sensors 2013, 13, 5347–5367. [Google Scholar] [CrossRef]
  35. Das, P.; Vilà-Valls, J.; Vincent, F.; Davain, L.; Chaumette, E. A new compact delay, Doppler stretch and phase estimation CRB with a band-limited signal for generic remote sensing applications. Remote Sens. 2020, 12, 2913. [Google Scholar] [CrossRef]
  36. O’Donnell, R.M. Detection of Targets in Noise and Pulse Compression Techniques. MIT Lincoln Laboratory Radar Course Slides. 2001. Available online: https://www.ll.mit.edu/sites/default/files/outreach/doc/2018-07/lecture%205.pdf (accessed on 18 April 2026).
  37. Richards, M.A. Fundamentals of Radar Signal Processing, 2nd ed.; McGraw-Hill: New York, NY, USA, 2014. [Google Scholar]
  38. Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef]
  39. Deng, Q.; Ge, Y.; Ding, Z. A unifying view of OTFS and its many variants. IEEE Commun. Surv. Tutor. 2025, 27, 3561–3586. [Google Scholar] [CrossRef]
Figure 1. The system model.
Figure 1. The system model.
Sensors 26 02790 g001
Figure 2. Average CRLB for angle estimation vs. average sum rate.
Figure 2. Average CRLB for angle estimation vs. average sum rate.
Sensors 26 02790 g002
Figure 3. Average CRLB for range estimation vs. average sum rate.
Figure 3. Average CRLB for range estimation vs. average sum rate.
Sensors 26 02790 g003
Figure 4. Distribution of ρ .
Figure 4. Distribution of ρ .
Sensors 26 02790 g004
Table 1. Simulation parameters.
Table 1. Simulation parameters.
ParameterSymbolValue
Number of transmit antennas at RSU N t 8
Number of receive antennas at RSU N r 4
Number of receive antennas at vehicleM4
Number of vehiclesK3
Carrier frequency f c 28 GHz
Episode duration T ep 10 s
Frame duration Δ T 0.1 s
Total transmit power of RSU P tot 1.0 W
Communication noise power σ comm 2 10 10 W
Path loss exponent α 2.5
Minimum sum rate requirement R min 2.4 bps/Hz
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Lim, J.; So, J. Deep Learning-Based Dynamic Time Division ISAC Beamforming for Vehicular Networks. Sensors 2026, 26, 2790. https://doi.org/10.3390/s26092790

AMA Style

Lim J, So J. Deep Learning-Based Dynamic Time Division ISAC Beamforming for Vehicular Networks. Sensors. 2026; 26(9):2790. https://doi.org/10.3390/s26092790

Chicago/Turabian Style

Lim, Junseok, and Jaewoo So. 2026. "Deep Learning-Based Dynamic Time Division ISAC Beamforming for Vehicular Networks" Sensors 26, no. 9: 2790. https://doi.org/10.3390/s26092790

APA Style

Lim, J., & So, J. (2026). Deep Learning-Based Dynamic Time Division ISAC Beamforming for Vehicular Networks. Sensors, 26(9), 2790. https://doi.org/10.3390/s26092790

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop