Skip to Content
TechnologiesTechnologies
  • Article
  • Open Access

9 September 2026

Energy-Efficient Anti-Jamming over Time-Varying Fading Channels via DQN-Based Joint Channel Selection and Power Control

,
and
1
School of Electronic Information Engineering, Nanjing University of Information Science & Technology, Nanjing 210007, China
2
The Sixty-Third Research Institute, National University of Defense Technology, Nanjing 210007, China
*
Author to whom correspondence should be addressed.

Abstract

Addressing the dual threats of malicious jamming and time-varying fading faced by wireless communication links in complex dynamic electromagnetic adversarial environments, existing intelligent anti-jamming methods predominantly focus on single-dimensional resource optimization under quasi-static channels. This focus neglects the nonlinear superposition effects of multi-path deep fading and dynamic strong jamming in the time-frequency domain, making it challenging for systems to balance transmission reliability and system energy efficiency in physical environments where fading and suppression coexist. To address this issue, this study proposes a joint intelligent anti-jamming method for channel switching and transmit power control based on a Deep Q-Network (DQN). Initially, a composite communication environment model incorporating Markov time-varying fading and jamming is constructed. Subsequently, the joint resource scheduling problem is formulated as a Markov Decision Process. The environment state space is reconstructed by integrating continuous channel state estimation and jamming observation features, accompanied by the design of a highly aggregated two-dimensional discrete action space for both channel and power. Finally, a composite reward function evaluating both communication success rates and power consumption costs is proposed to guide the agent in multi-dimensional resource joint optimization. Simulation results demonstrate that the proposed algorithm effectively extracts implicit features under the composite state of fading and jamming. When encountering extreme deep fading or full-band blocking, the agent strategically triggers a silent mechanism to avoid exorbitant invalid energy consumption penalties, while precisely matching interference-free channels with the minimum effective transmit power during favorable communication windows. Simulation results show that compared with traditional xx algorithms, the proposed method significantly improves the dynamic successful transmission rate and system energy efficiency in complex, highly dynamic scenarios, achieving an effective optimization of anti-jamming reliability and low power overhead.

1. Introduction

Wireless communication systems have become critical information infrastructure in modern defense and civilian sectors due to their inherent flexibility and convenience. However, the open nature of wireless channels exposes these systems to increasingly severe electromagnetic security threats. Particularly in malicious adversarial environments, wireless communication links must possess the capability to transmit information reliably and effectively under complex dynamic electromagnetic jamming [1,2]. Malicious jammers specifically disrupt the demodulation process of wireless communication systems by strategically configuring various jamming patterns, leading to interruptions in communication transmission [3]. With the introduction of cognitive radio and artificial intelligence technologies, malicious jamming in modern electromagnetic adversarial environments exhibits characteristics such as high intensity, multi-channel coverage, and high-frequency time-varying switching. These features pose severe challenges to fixed-rule anti-jamming technologies, including traditional spread spectrum and frequency hopping. To counter this severe jamming environment, intelligent communication anti-jamming technologies equipped with environmental perception and dynamic decision-making capabilities have gradually become research frontiers. Intelligent anti-jamming research has extensively explored single-dimensional resource optimization, where classical Q-learning algorithms were widely applied in dynamic spectrum access and channel selection due to their model-free learning nature; cooperative extensions further combined wideband spectrum sensing with greedy Q-learning policies to proactively evade hostile jamming [4]. Regarding power control, earlier studies proposed reinforcement learning-based anti-jamming power control schemes aimed at minimizing energy consumption while maintaining communication connections [5]. Additionally, research focusing on the time dimension of anti-jamming proposed time-domain evasion algorithms against random pulse jamming [6]. Controlling a single domain (frequency, power, or time) struggles to cope with complex and dynamic jamming variations, making joint multi-domain resource anti-jamming a prominent new trend. A two-dimensional joint anti-jamming communication game model utilizing reinforcement learning algorithms to jointly schedule channel evasion and transmit power was previously proposed [7]. This achieved the global optimization goal of maximizing system transmission rates under dynamic malicious jamming threats while significantly reducing equipment transmit power and switching overhead. Traditional tabular reinforcement learning algorithms like Q-learning rely heavily on discrete tabular mechanisms to store state-action value functions. When addressing multi-domain joint optimization, the dimensionality of the state space expands exponentially, triggering the curse of dimensionality and rendering discrete representation and tabular storage methods infeasible. Deep reinforcement learning has been introduced into the anti-jamming decision-making domain to overcome this computation and storage bottleneck. Deep reinforcement learning achieves an end-to-end mapping from complex continuous environmental observations to optimal anti-jamming actions by utilizing the powerful high-dimensional feature extraction and nonlinear generalization capabilities of deep neural networks. Some studies introduced concealed mechanisms to enhance policy robustness by weakening the predictive capabilities of reactive jammers through feature compression [8]. Further research explored the application of federated learning in 5G heterogeneous networks to solve collaborative anti-jamming problems among multiple nodes [9]. Deep Q-networks have been utilized to extract jamming features directly from spectrum waterfall plots, achieving intelligent evasion of dynamic sweep jamming by relying on the strong feature extraction capabilities of convolutional neural networks [10]. A joint frequency and power anti-jamming algorithm based on prioritized experience replay was designed to improve the learning efficiency of joint policies [11]. Deep dueling neural networks combined with backscatter technology realized anti-jamming communication under extremely low signal-to-interference-plus-noise ratios [12]. The advantages of multi-domain joint decision-making in enhancing system throughput under jamming environments were verified in unknown information environments and multi-agent collaborative scenarios, respectively [13,14].
Despite achieving significant results in theory and simulation, these studies still possess limitations when facing real physical channels. Most research is based on quasi-static or ideal Rayleigh channel assumptions, failing to effectively address the complex coupling of time-varying fading and malicious jamming in the time-frequency domain. Under deep fading conditions, the deterioration of received signal strength originates from the channel’s inherent characteristics rather than an increase in jamming power. Severe policy mismatches will likely occur if the wireless communication system cannot decouple and distinguish these two signal impairments originating from different physical sources based on environmental observations. For instance, the multi-path effect in frequency-selective fading channels causes severe fluctuations in signal amplitude, greatly degrading the signal-to-interference-plus-noise ratio at the receiver when superimposed with artificial jamming. This rapid time-varying channel characteristic is particularly evident in highly dynamic scenarios such as unmanned aerial vehicle communications [15,16]. Blindly increasing transmit power when mistakenly attributing deep fading to jamming fails to improve the bit error rate and instead causes massive invalid energy consumption. Consequently, a critical bottleneck in the current intelligent anti-jamming field is designing a robust joint decision-making framework for channel and power to maximize the system’s dynamic energy efficiency while ensuring reliable communication link transmission in a dynamic environment where time-varying fading and jamming are coupled.
Addressing the limitations of existing research regarding fading channels and multi-dimensional resource joint mechanisms, this paper conducts an in-depth study on joint channel switching and power control over fading channels. The core contributions of this study are as follows:
  • Constructing a system model with coexisting time-varying fading and jamming: Targeting the inadequate consideration of channel time-variance in existing research, a communication environment model incorporating time-varying fading channels is established to characterize the nonlinear superposition effects of channel gain and jamming power at the signal-to-interference-plus-noise ratio level.
  • Proposing a joint design method for state space and anti-jamming reward function aimed at decoupling fading and jamming, utilizing the DQN framework for anti-jamming decision-making. Designing a high-dimensional state space incorporating both channel gain and jamming observations to distinguish the causes of communication performance impairment. Constructing a composite action space encompassing transmission status, channel selection, and power levels alongside a reward function, guiding the agent to learn a multi-modal joint policy. This policy involves remaining silent during deep fading, switching channels during jamming, and reducing transmit power during favorable channel conditions, thereby enhancing energy efficiency while ensuring communication quality.
The remainder of this paper is organized as follows: Section 2 establishes the system and time-varying channel fading models. Section 3 formulates the joint channel selection and power control optimization problem as a Markov Decision Process. Section 4 presents the proposed DQN-based joint anti-jamming algorithm along with its computational complexity analysis. Section 5 provides numerical simulation results, comparative analyses, and robustness evaluations. Finally, Section 6 concludes the paper.

2. System Model

As illustrated in Figure 1, this paper considers a wireless communication system operating in a dynamic electromagnetic environment. The system comprises three main components: a wireless communication transmitter, a receiver, and a malicious jammer.
Figure 1. Diagram of the wireless communication system under dynamic jamming.
The model construction is based on the following core assumptions and system definitions:
  • The available spectrum of the wireless communication system is divided into K non-overlapping channels, denoted by the channel set C = { f 1 , f 2 , , f K } .
  • The transmit power P set is defined as P = { 0 , p 1 , p 2 , , p M } . The receiver is responsible for demodulating the received mixed signals and measuring state information, such as the instantaneous signal-to-interference-plus-noise ratio of the current channel in real time.
  • At any discrete time slot, the transmitter dynamically selects an optimal sub-channel c ( t ) C and matches the corresponding transmit power P ( t ) P for data transmission. This decision is executed by an internally integrated deep reinforcement learning agent based on the channel state information and jamming observation results fed back by the receiver.
  • To focus on the anti-jamming action problem at the transmitting end, this paper assumes that the necessary control information is fed back to the transmitter over a dedicated low-rate control channel that is physically separated from the data channel in the time-frequency domain. Within the operational timescale considered in this paper, this control channel is treated as error-free, which is a standard assumption in physical-layer anti-jamming literature [17]. We emphasize that only the data channel is subject to malicious jamming, whereas the feedback assumption isolates the contribution of the proposed joint channel-and-power scheduling scheme. The jammer acts as a non-cooperative party attempting to reduce the communication link’s S I N R by transmitting suppressive jamming signals to block communication.
Considering the multi-path effects caused by dense spatial scatterers in complex urban or suburban non-line-of-sight transmission environments, the physical layer wireless channel exhibits typical highly dynamic time-varying characteristics. This paper models the physical channel as a quasi-static flat block fading channel, indicating the channel gain remains constant within a single time slot but dynamically changes over time across different slots. For the K - t h sub-channel, the complex channel gain at time slot t is denoted as h k ( t ) . Due to the abundance of scatterers and the lack of a direct line-of-sight path in the environment, the channel envelope | h k ( t ) | follows a Rayleigh distribution. The probability density function is given by:
f | h | ( x ) = 2 x Ω exp x 2 Ω ,   x 0
where Ω represents the average channel power gain.
To accurately describe the dynamic evolution characteristics of the channel over time, a first-order autoregressive process is employed to model the Markovian nature of the complex channel gain:
h k ( t ) = ρ h k ( t 1 ) + 1 ρ 2 z k ( t )
where z k ( t ) is an independent and identically distributed complex Gaussian white noise process representing the innovation of channel variation. The parameter ρ [ 0 , 1 ] is the channel time correlation coefficient. It is worth emphasizing that the AR(1) recursion in Equation (2) is mathematically consistent with the Rayleigh envelope assumption in Equation (1). The complex baseband channel gain h k ( t ) is a zero-mean circularly symmetric complex Gaussian process, whose real and imaginary parts are independent Gaussian variables with equal variance; consequently, its magnitude h k t is Rayleigh-distributed at every time slot. Since the innovation term w(t) in Equation (2) is independent and identically distributed (i.i.d.) zero-mean unit-variance circularly symmetric complex Gaussian noise, the recursively generated h k ( t ) remains zero-mean circularly symmetric complex Gaussian at any slot t, so its envelope continues to follow the same Rayleigh distribution. In other words, Equation (2) only introduces temporal correlation between successive slots and does not alter the marginal distribution of the channel envelope at any individual time slot. According to the Jakes fading model, its physical relationship with moving speed can be expressed as:
ρ = J 0 ( 2 π f d τ )
where J 0 ( ) is the zeroth-order Bessel function of the first kind, f d is the maximum Doppler shift, and τ is the time slot length. For scenarios representing slow-moving vehicular communication terminals or low-altitude unmanned aerial vehicle communications, utilizing a relative moving speed of approximately 8.75 m/s at a 2.4 GHz carrier frequency yields a maximum Doppler shift of about 70 Hz. Under these conditions, the higher-order statistical properties of the channel change minimally within a time slot, making the modeling error of approximating the Jakes channel with a first-order autoregressive model acceptable [18]. Equation (3) indicates that as the Doppler shift increases, the time correlation coefficient decays in a step-like manner, exacerbating the channel’s time decorrelation effect and significantly increasing the difficulty and signal processing complexity for the receiver when decoupling multi-path features.
To comprehensively evaluate the robustness of the intelligent anti-jamming algorithm, this paper assumes the jammer possesses multiple jamming modes capable of dynamically distributing its jamming power across the time-frequency domain based on specific strategies, albeit constrained by a total peak transmit power P max . Let J k ( t ) denote the jamming power exerted by the jammer on channel k at time slot t .

3. Problem Modeling

3.1. Description of the Joint Resource Optimization Problem

In discrete time slots, the intelligent transmitter adapts to the dynamic electromagnetic environment by jointly optimizing channel selection and transmit power. The system’s joint transmission policy is defined as a t = { c t , p t } , where c t C represents the working channel index, C = { 1 , , K } and p t { P = 0 , P 1 , , P M } represents the discrete transmit power level.
To ensure the physical feasibility and service quality of the wireless communication system, the transmission policy must satisfy the following constraints: (1) the communication transmitter can only implement exclusive single-sub-channel occupation or strategic silence within any discrete time slot, satisfying the mutual exclusivity of frequency domain access. (2) The transmit power must be non-negative and not exceed the peak power P max . (3) To simplify the analysis, this paper adopts a hard-decision threshold model. The data packet in the current time slot is considered successfully demodulated and received only when the instantaneous signal-to-interference-plus-noise ratio γ t at the receiving end exceeds the demodulation threshold Γ t h ; otherwise, the current link is deemed to have an instantaneous outage.
This paper aims to design a robust policy that maximizes throughput while minimizing overhead. The system’s instantaneous utility function U ( a t , s t ) is defined as:
U ( a t , s t ) = λ s u c I ( γ t Γ t h ) λ p w r p t P max λ s w I ( c t c t 1 )
where I ( ) is the indicator function, I ( γ t Γ t h ) is the transmission success indicator function. It takes the value 1 if the instantaneous signal-to-interference-plus-noise ratio at the receiver meets the demodulation threshold, indicating successful data packet transmission, and 0 for transmission failure. The term P t P max is the normalized transmit power used to characterize the radio frequency energy consumption overhead of the communication equipment in the current time slot, where P t is the actual executed transmit power and P max is the maximum transmit power supported by the hardware. The variable I s w i t c h is the channel switching indicator function, taking the value 1 if channel switching occurs in the current time slot, yielding radio frequency switching delay and overhead, and 0 if the original channel is maintained for transmission. The coefficients λ s u c , λ p w r , λ s w are non-negative weighting factors for throughput, energy consumption, and switching stability, respectively.
Based on this formulation, the joint resource optimization problem can be formalized as finding the optimal policy π * to maximize the cumulative expected discounted utility over the time domain:
( P 1 ) : max π V ( π ) = E π t = 0 δ t · U ( a t , s t )
s . t .   C 1 :   c t C { } ,   t
C 2 :   p t P ,   0 p t P max ,   t
C 3 :   γ t = p t | h c t ( t ) | 2 J c t ( t ) + N 0 ,   t
where δ [ 0 , 1 ) is the discount factor used to balance current rewards against long-term returns. The constraints C 1 represent channel access limits (where 0 represents silence) and transmit power hardware constraints C 2 , C 3 revealing the underlying physical evolution mechanism where the receiver’s signal-to-interference-plus-noise ratio is suppressed by the coupling of highly dynamic fading gains and non-cooperative jamming. Given that the dynamic distributions of channel gain | h | 2 and jamming power J in the constraints are unknown a priori, and the objective involves non-convex discrete variables, this paper transforms the problem ( P 1 ) into a Markov Decision Process for resolution. The decision process is described by constructing a standard four-tuple M = ( S , A , R , δ ) .

3.2. State Space

To distinguish between fading and jamming, the system state s t S at time t is defined as a composite vector comprising channel state information, jamming observation features, and historical actions:
s t = [ g e s t ( t ) , J o b s ( t ) T , c t 1 ] T
Constrained by the objective physical limitations of the communication node’s half-duplex radio frequency front-end or non-full-band sensing hardware, the receiver typically cannot simultaneously perform high-precision channel fading detection on all candidate channels. g e s t ( t ) indicates that the receiver can only conduct precise channel c t 1 estimation on the current communication path utilizing orthogonal pilot signal sequences. The state space contains only the instantaneous power gain estimation value of the channel occupied in the previous time slot rather than full-band information. Utilizing this localized fading feature as a state anchor provides the agent with an intuitive physical baseline to judge whether the current channel has fallen into a deep fading trap, effectively preventing blind power compensation caused by unknown channel states. Although high-precision channel estimation is restricted to the operating frequency band, modern communication receivers are typically equipped with wideband energy detection bypasses capable of sweeping full-band energy at a lower resolution. The state vector J o b s ( t ) R K incorporates the full-band jamming power observation vector and the previous time slot’s channel index c t 1 , which is used to calculate the switching cost. Through the aforementioned refined dimensionality reduction and feature decoupling, the composite state vector defined in this study effectively demystifies the black box of environmental information. This enables the deep reinforcement learning model to extract clear nonlinear mapping boundaries within the complex field where fading and jamming are coupled.

3.3. Joint Action Space

To resolve the coupled control problem, a discrete joint action space A is designed. The agent outputs an action index a t { 0 , 1 , , K M } mapped to corresponding physical parameters F ( a t ) :
( c t , p t ) = F ( a t ) = ( , 0 ) , if   a t = 0 ( k , P m ) , if   a t > 0
where the a t = 0 mapping encompasses a silent shutdown state for the radio frequency transmit link. This action is primarily utilized to strategically avoid invalid energy consumption when the system encounters extremely harsh physical environments, such as when all candidate channels are in deep fading or full-band blocking jamming prevents meeting the demodulation threshold. When a t > 0 , the action index is restored to specific channel K indices and power levels P M via the mapping function. This encoding paradigm reduces the dimensionality of the high-dimensional concurrent frequency and power allocation problem to a single scalar output space. It avoids the action space decoupling errors inherent in traditional dual-branch networks while strictly ensuring the physical execution synchronization of frequency domain evasion and power compensation.
In addition to its energy-saving role, the silent action ( a t = 0 ) also serves as a safety fallback when the most recent feedback becomes unavailable or unreliable: (i) the agent can deliberately suspend transmission to avoid physical collisions with the jammer; (ii) the historical residence-frequency features already embedded in the state space allow the agent to remember safe channel windows even in the absence of instantaneous feedback; and (iii) because the jamming patterns considered (multi-path sweep and intelligent blocking) are structurally periodic, the learned policy implicitly encodes the jammer’s transition rule rather than depending solely on instantaneous observations. These three layers jointly preserve the operational integrity of the closed-loop system under complete feedback corruption.

3.4. Reward Function

The reward function r t R directly maps the optimization objective ( P 1 ) , adopting a segmented design:
r t = η w a i t , if   a t = 0 R l i n k λ p w r p t P max λ s w I ( c t c t 1 ) , if   a t > 0
When executing the silent action ( a t = 0 ) , the system passively imposes a static queueing delay penalty η w a i t to prevent the agent from falling into a local optimum of passive evasion out of fear of failure deductions. Conversely, in the active transmission state a t > 0 , the reward consists of link state benefits R l i n k and cost penalty terms. If the instantaneous demodulation threshold γ t Γ t h is met, a positive transmission reward R l i n k = R s u c is obtained; otherwise, a severe link disconnection penalty R l i n k = R f a i l is faced. Simultaneously, the formulation introduces a power penalty λ p w r positively correlated with the current power and a switching penalty triggered by a Boolean indicator function λ s w . These penalties aim to eliminate the agent’s detrimental policy tendencies of blindly increasing power and frequently hopping frequencies. This reward function effectively guides the agent to abandon short-sighted behaviors that solely pursue link connection rates, thereby precisely balancing radio frequency energy consumption and physical execution stability while ensuring communication reliability.

3.5. State-Action Value Function

The agent’s objective is to discover the optimal policy that maximizes the long-term expected cumulative return after executing an action a t from the current state s t , Q ( s , a ) known as the state-action value function. According to the Bellman optimality equation, the function Q * satisfies a specific recursive relationship:
Q * ( s , a ) = E s r ( s , a ) + δ max a A Q * ( s , a ) s , a
where s denotes the state at the next time step, and the expectation operator E s is defined with respect to the state transition probability of the environment. Since the considered state space encompasses continuous channel gains and jamming intensities, and the environmental dynamic model is not known a priori, traditional dynamic programming methods cannot solve this equation directly. Furthermore, the continuity of the state space introduces a certain smoothness in the Q-values between adjacent states, making it suitable to introduce the Deep Q-Network algorithm to approximate the aforementioned function Q * using deep neural networks.

4. DQN-Based Joint Channel Selection and Power Control Algorithm

To enhance the efficiency of anti-jamming resource scheduling in complex electromagnetic environments, this paper proposes a rapid anti-jamming algorithm based on joint channel switching and power synergy over fading channels. The algorithm’s fundamental structure adopts the classic architecture of the conventional Deep Q-Network. It comprehensively accounts for the lag in the physical process of state perception in communication anti-jamming problems and the delayed feedback mechanism of economic rewards. This enables the agent to collect interaction samples with strict causal timing when facing coupled deep fading and jamming, ultimately accelerating the learning process of multi-modal anti-jamming policies. The specific execution steps are detailed as follows in Algorithm 1:
Algorithm 1. DQN-Based Joint Channel Selection and Power Control
Input: Composite state vector s t = [ | h c t 1 ( t ) | 2 , J ( t ) , c t 1 ] discount factor γ ; learning rate α ; replay memory capacity N c a p ; mini-batch size B ; target network update period C ; exploration parameters ϵ init , ϵ end , β ; maximum episodes M and slot horizon T max .
Output: Optimal joint transmission action a t = ( c t , P t )   ( channel   selection   c t K   and   transmit   power   level   P t P optimal Q-network parameters θ .
1:  Randomly initialize the training network parameters θ   and   initialize   the   target   network   parameters   to   θ θ
2:  Initialize the experience replay buffer D   with   a   capacity   of   N c a p
3:for each training round, episode = 1 to M  do
4:  Reset the environment and observe the initial state s 0
5:  For each time slot t = 1   to   T m a x  do
6:    Generate a random r a n d   number   between   [ 0 , 1 ]
7:    if r a n d < ϵ then randomly select an action a t A ;
8:    else, select the optimal action a t = arg max a Q ( s t , a ; θ )
9:    Execute an action a t ,   observe   the   immediate   reward   r t ,   and   the   next   state   s t + 1
10:     Store   the   quadruplet   ( s t , a t , r t , s t + 1 ) in the buffer D
11:      if the number of samples in the buffer D > B  then
12:      Randomly sample B samples from the set D
13:       Calculate   the   target   value   y j = r j + δ max a Q ( s j + 1 , a ; θ )
14:      Perform gradient descent updates θ   to   minimize   the   loss   L = 1 B ( y j Q ( s j , a j ; θ ) ) 2
15:    end if
16:    if t mod C = = 0 then Update Target Network θ θ
17:     Update   Status   s t s t + 1
18:  end for
19:end for
To facilitate a step-by-step understanding of Algorithm 1, consider a representative time slot t with K = 3 candidate channels, power levels { 0 , 0.2 , 0.6 , 1.0 }   W , and demodulation threshold γ t h = 2.5 . State Construction: The agent receives the previous channel gain, full-band jamming observations J = [ 0.8 , 0.5 , 0.0 ]   W , and past residence features, assembling state s t . Action Selection: With exploration probability ϵ = 0.2 , the agent explores randomly; with probability 0.8, it exploits   a * =   a r g   m a x a Q s t ,   a ;   θ Since channel 3 is free of jamming while channel 1 is heavily suppressed, the policy selects action a * = ( c t = 3 , P t = 0.2   W ) . Feedback and Update: The transmitter switches to channel 3 with 0.2   W . Because the unjammed channel yields S I N R > γ t h , transmission succeeds. The agent receives a positive reward minus the power cost ( 0.2   W ) and switching overhead. The transition ( s t , a * , r t , s t + 1 ) is pushed into the buffer D ; once samples exceed mini-batch size B = 256 , a batch is sampled to update network weights θ via gradient descent. This concrete execution loop directly demonstrates the algorithmic workflow.
The proposed algorithm operates as an online learning algorithm. Prior to transmission termination, the algorithm iterates once per communication time slot, gradually approaching the optimal dual-domain joint transmission policy for fading and jamming. The three components of spectrum sensing, learning and decision-making, and communication transmission operate in parallel within the same time slot. The output of the current time slot’s spectrum sensing serves as the input state for the next time slot’s learning and decision-making, while the action instruction output from the current learning and decision-making is executed by the communication transmission physical hardware in the subsequent time slot. This parallel architecture assists intelligent nodes in acquiring more continuous spectrum observation states and allows ample time for learning and decision-making. Simultaneously, the perception, learning, and decision processes do not occupy valuable communication transmission time. Compared to the serial and partially parallel structures commonly used in intelligent anti-jamming algorithms, the parallel time slot structure achieves greater system throughput at the same transmission rate.
The computational complexity of the proposed scheme is governed by three operational parts. (i) Environment interaction, reward calculation, and metric evaluation incur O ( 1 ) time complexity per slot. (ii) DQN forward inference requires multiply-accumulate (MAC) operations per decision slot, where N = 22 is the composite state dimension, L = 128 is the hidden-layer width, and M = 31 is the cardinal size of the discrete joint action space. This results in approximately 2.3 × 10 4 MAC operations, which consume less than 0.03   ms on standard embedded processors and effortlessly satisfy the 1   ms real-time execution constraint. (iii) Offline network training via backpropagation takes O ( B ( N L + L 2 + L M ) ) operations per gradient step with mini-batch size B = 256 , yielding around 1.4 × 10 7 MAC operations executed asynchronously in the background.
In contrast, while conventional tabular Q-learning has a minimal per-step lookup time of O ( | A | ) , its memory and convergence complexity scale with O ( | S | | A | ) . Even under coarse discretization of the continuous channel gains and jamming features across K = 10 channels, the state cardinality | S | explodes combinatorially, triggering the prohibitive curse of dimensionality. Furthermore, regarding scalability, as the channel count K increases, the state space of tabular methods grows exponentially as | O ( C K ) , rendering memory storage and policy convergence impossible. Conversely, the proposed DQN scales gracefully: the input and output dimensions grow strictly linearly with K (state dimension ~ 2 K + 2 and action space ~ 3 K + 1 ), increasing inference FLOPs marginally while retaining robust continuous-state representation.
Regarding scalability, as the channel count $K$ expands, tabular Q-learning suffers from exponential explosion O ( N J K ) , rendering state-table storage intractable. In contrast, the proposed DQN model scales linearly: expanding K from 10 to 30 merely increases network weights by approximately 50 % .

5. Algorithm Simulation

5.1. Parameter Settings

The algorithm simulation in this paper is implemented utilizing the Python3.8 platform and the PyTorch 2.0 deep learning framework. The detailed simulation parameter configurations are summarized in Table 1. Both the evaluation network and the target network of the DQN employ fully connected neural networks with two hidden layers, each containing 128 nodes. To address the limitations of the Partially Observable Markov Decision Process caused by intelligent blocking jamming, the input layer dimension of the network is expanded to 22. This dimension encompasses precise channel gains, full-band jamming observations, self-historical residence frequencies, and the previous time slot’s channel index. The output layer dimension is 31, corresponding to one silent evasion action and 30 joint channel-and-power actions. The maximum capacity of the sample memory is 20,000, and the mini-batch size is 256.
Table 1. Simulation parameters.
The channel time slot duration is Δ t = 1   ms . With a carrier frequency f c = 2.4   GHz and a terminal velocity v 8.75   m / s , the maximum Doppler shift is f d = v f c c 70   Hz . Under Jakes’ model, the theoretical correlation coefficient between adjacent time slots is 0.95. Hence, setting ρ = 0.95 accurately matches the physical channel variation between consecutive 1 ms slots.
The channel coherence time can be approximated as T c 0.423 f d 0.423 70 6.04   ms . In reinforcement learning, the effective planning horizon over which cumulative rewards are weighed is given by H e f f = 1 1 γ For γ = 0.9 . The effective horizon is exactly 10 ms.
This horizon closely matches the channel coherence span. If a higher discount factor were selected, the agent would evaluate policies based on channel predictions far beyond coherence time, leading to severe policy bias in non-stationary environments. Conversely, a smaller γ would make the agent myopic. Therefore, γ = 0.9 achieves an optimal balance between Markovian long-term planning and the physical non-stationarity of the fading channel.
The algorithm’s exploration and exploitation mechanism employs the epsilon-greedy method, where epsilon denotes the probability of exploration. This involves randomly selecting actions with a certain probability during training to explore more states and policies. Online anti-jamming algorithms require a smooth transition from exploration to exploitation. During the initial training phase, extensive exploration is necessary, requiring a larger epsilon value. In the later training stages, the algorithm must increasingly utilize the learned anti-jamming policies, necessitating a gradual reduction in the epsilon value. The simulation calculates the exploration rate for each time slot utilizing an exponential decay function:
ϵ ( T ) = ϵ e n d + ( ϵ s t a r t ϵ e n d ) e T ϵ d e c a y  
In this context, ϵ s t a r t is set to 1 as the initial value to ensure sufficient exploration at the start of training; ϵ e n d is set to 0.01 as the termination value to ensure full exploitation in the later stages of training; T = 1 , 2 , T max represents the number of time slots; and ϵ d e c a y represents the decay rate, set to 80. The larger the value of ϵ d e c a y , the more steps the algorithm requires to transition from exploration to exploitation.
The simulation considers two typical dynamic jamming patterns:
  • Multi-path sweep jamming: Employs three narrow-band jamming signals to perform periodic linear sweeping sequentially across various channels within the target frequency band. This jamming method is widely applied in practical communication confrontation scenarios due to its high jamming efficiency and ease of generation.
  • Intelligent blocking jamming: Also known as intelligent responsive jamming, this is one of the widely researched intelligent jamming methods in existing works. Its jamming attacks are based on a perception-action loop capable of selecting the three channels with the highest occupancy rates to apply two strong and one weak narrow-band jams, based on the legitimate signal’s channel occupancy over the past four time slots.
To comprehensively verify the superiority of the proposed multi-domain joint intelligent anti-jamming algorithm, this paper compares it against three typical intelligent communication anti-jamming baseline algorithms:
  • Conventional DQN algorithm: The algorithm processes, network structures, and parameter settings align entirely with the proposed algorithm, except that its state space lacks the capability to perceive time-varying channel fading features.
  • Q-learning algorithm considering channel fading: Transmission action decisions are based on the widely applied traditional tabular Q-learning algorithm. To prevent the curse of dimensionality triggered by high-dimensional states, its state space is substantially reduced, containing only quantized channel fading identifiers (normal or deep fading).
  • Ordinary Q-learning algorithm: Adopts a tabular Q-learning mechanism where its state space completely disregards physical channel fading features, making purely reactive table lookup decisions based solely on discretized primary jamming locations. This algorithm represents the performance lower bound of the system lacking multi-domain fusion perception conditions.

5.2. Algorithm Analysis

Figure 2 illustrates the converged communication time-frequency state of the proposed algorithm under a multi-path sweep jamming environment. The figure captures genuine physical communication time slots after the algorithm has undergone sufficient iterations to reach a stable convergence phase. It is clearly observable that the dark (strong jamming) and light (weak jamming) blocks representing jamming signals exhibit a strict periodic step-like scanning pattern. The communication node driven by the proposed algorithm accurately anticipates the frequency-hopping trajectory of the multi-path sweep jamming, completely and precisely evading all narrow-band jamming signals in the domain. The residence frequencies of the communication nodes perfectly intersperse within the interference-free white safe communication time slots without any physical collisions. This visually corroborates, in the time-frequency dimension, the physical mechanism enabling the proposed algorithm’s normalized throughput to approach the optimal value.
Figure 2. Time-frequency plots of algorithm convergence under multi-channel sweep jamming.
Figure 3 presents the converged time-frequency state distribution of the proposed algorithm under an intelligent blocking jamming environment. Against intelligent jammers implementing dynamic tracking based on the historical residence frequencies of legitimate nodes, the proposed algorithm constructs a stable periodic frequency-hopping evasion sequence within the restricted spectrum resources. A typical local spectrum induction game feature can be observed from the figure: the communication node tends to restrict its frequency-hopping actions to a specific subset of sub-channels for cyclic switching, effectively concentrating and pinning the jammer’s suppressive energy within a localized frequency band. Benefiting from the algorithm state space’s effective integration of historical temporal features, the communication node accurately evaluates the jammer’s transfer rules, maintaining low-power transmission on interference-free safe time-frequency resource blocks and avoiding the frequent blind collisions typical of traditional memoryless algorithms. These time-frequency distribution characteristics thoroughly validate the effectiveness and robustness of the proposed multi-domain joint intelligent anti-jamming algorithm in countering non-stationary tracking jamming.
Figure 3. Time-frequency plots of algorithm convergence under intelligent blocking jamming.
Figure 4 illustrates the normalized throughput performance comparison between the proposed algorithm, conventional DQN, QL(considering fading), and ordinary QL algorithms under a multi-path sweep jamming environment. Due to the presence of multi-path sweep jamming covering 30% of the frequency band (two strong jams and one weak jam) in the initial environment, the normalized throughput starting points for all algorithms in the initial blind exploration phase reasonably land around 0.4. As iterations progress, the proposed algorithm rapidly escalates after training for approximately 200 communication time slots, stably converging to the optimal policy approaching 0.9. Under identical parameter settings, the convergence speed of the conventional DQN algorithm is roughly equivalent to the proposed algorithm, but its final normalized throughput only converges to approximately 0.75. The fundamental reason for this performance gap is that the conventional DQN lacks joint regulation capabilities in the power domain, consistently transmitting anti-jamming signals at maximum transmit power. This results in exorbitant effective energy consumption penalties despite successfully evading jamming. In contrast, the proposed algorithm accurately matches interference-free channels with the minimum effective transmit power, immensely improving energy efficiency. Simultaneously, constrained by the tabular method’s dimensionality reduction compromise for high-dimensional states (causing blindness to weak jamming), both the QL (considering fading) and ordinary QL algorithms easily fall into secondary jamming traps, with their normalized throughput ultimately converging merely to around 0.68. The causal origin of this gap is twofold. First, the conventional DQN’s state space omits the fading features, so the agent cannot tell whether a low SINR stems from deep fading or from jamming; it therefore always transmits at maximum power, incurring the maximum energy penalty even on clean channels. Second, the tabular Q-learning baselines sacrifice state resolution to control the curse of dimensionality, which makes them blind to weak jamming and prone to secondary jamming traps. In contrast, the proposed algorithm decouples these two impairment sources through its composite state space and jointly selects channel and power, which is exactly why it can match interference-free channels with the minimum effective power.
Figure 4. Comparison of the normalized throughput performance of various algorithms under multi-channel frequency-sweep Jamming.
Figure 5 demonstrates the normalized throughput performance comparison of various algorithms under the more challenging intelligent blocking jamming environment. In this highly dynamic game scenario where the jammer possesses historical tracking capabilities, the performance disparities among algorithms are further amplified. The proposed algorithm still exhibits exceptional robustness. By relying on the self-historical residence-frequency features introduced in the state space, it effectively shatters the limitations of the Partially Observable Markov Decision Process, converging to an optimal policy with an average normalized throughput greater than 0.93 in just about 300 time slots. In comparison, the performance ceiling of the conventional DQN is still firmly suppressed below 0.75 by fixed high power consumption. Due to the lack of memory capabilities for temporal features and the generalization support of deep neural networks, traditional QL and ordinary QL algorithms completely lose their evasion capabilities when facing tracking-style intelligent blocking jamming equipped with sensing loops, leading to frequent physical collisions and causing their normalized throughput to stagnate in a trough around 0.45.
Figure 5. Comparison of the normalized throughput performance of various algorithms under Intelligent Blocking Jamming.
To intuitively measure the system’s energy consumption efficiency during the anti-jamming process, this paper introduces system energy efficiency as a core evaluation metric, defined as the effective amount of data successfully transmitted per unit joule of energy. The calculation formula is:
E E = R P t x + P c
where R { 0 , 1 } represents the instantaneous effective throughput of the current time slot. Specifically, assuming a normalized target transmission rate of 1 bit/s/Hz, R takes the value of 1 (bit/s/Hz) if the receiver’s signal-to-interference-plus-noise ratio satisfies the demodulation threshold indicating successful transmission, and 0 otherwise. P t x denotes the actual radio frequency transmit power executed by the communication node at the current moment, while P c signifies the fixed static circuit power consumption required to maintain the operation of the radio frequency hardware link. By implicitly incorporating the time slot duration and bandwidth normalization into the numerator R , the dimension of this metric rigorously aligns with bits/J/Hz. This metric thoroughly reflects the algorithm’s success rate in evading jamming and embodies the communication node’s power scheduling cost in complex dynamic games. The proposed algorithm’s advantage, however, is not unconditional. When the maximum Doppler shift grows to several hundred hertz so that the channel coherence time falls below the decision interval, the learned policy becomes stale before the next action is executed, and a recurrent or meta-learning architecture would then be required. Likewise, when the jamming power saturates the entire band so that no sub-channel offers a usable SINR margin, the algorithm correctly falls back to the silent action, which caps the achievable normalized throughput at zero—a graceful degradation rather than a hard failure. These boundary conditions delimit the applicability of the proposed scheme.
Figure 6 illustrates the system energy efficiency comparison of various algorithms under a multi-path sweep jamming environment. As depicted, the conventional DQN algorithm, lacking perception of time-varying fading features in its state space, fails to avoid link interruptions caused by deep fading when attempting energy-saving transmission at low power, causing its energy efficiency convergence ceiling to stagnate around 2.3 bits/J/Hz. The QL (considering fading) and ordinary QL algorithms, restricted by the curse of dimensionality of the tabular method in processing continuous high-dimensional states, converge slowly and easily fall into sub-optimal policies, maintaining an energy efficiency chronically below 1.7 bits/J/Hz. Conversely, relying on a decoupled composite state space, the proposed algorithm accurately identifies high-quality resources free of jamming and with favorable channel gains, precisely matching the lowest effective transmit power of 0.2 W. Its energy efficiency rapidly escalates after approximately 200 training episodes of exploration, stabilizing above 3.2 bits/J/Hz and achieving a remarkable improvement over baseline algorithms.
Figure 6. Comparison of the performance and efficiency of various algorithms under multi-channel sweep jamming.
Figure 7 demonstrates the system energy efficiency performance of various algorithms under an intelligent blocking jamming environment. In this highly dynamic game scenario equipped with historical tracking capabilities, the difficulty of jamming evasion further intensifies, causing the energy efficiency ceilings of the baseline algorithms to experience varying degrees of suppression and intensified local oscillations. Due to a lack of fading perception capabilities, the conventional DQN algorithm’s energy efficiency further declines to approximately 2.1 bits/J/Hz. Traditional QL and ordinary QL algorithms experience frequent physical collisions as they fail to cope with non-stationary tracking jamming, resulting in rock-bottom energy efficiency. The proposed algorithm still demonstrates strong robustness in this harsh environment, effectively anticipating jamming transfer rules relying on the state space fused with historical residence frequencies, and flexibly adjusting power for transmission amidst dynamic jamming. The proposed joint channel-and-power multi-domain mechanism effectively overcomes the exorbitant energy consumption penalty incurred by single frequency domain evasion, realizing a comprehensive optimization of system transmission reliability and energy efficiency.
Figure 7. Comparison of energy efficiency of various algorithms under intelligent blocking jamming.
Furthermore, to address potential control plane congestion, we analytically evaluate the algorithmic resilience against random feedback latency τ d . In slotted communication frameworks ( Δ t = 1   ms ), any propagation and queuing latency within the subframe guard time ( τ d < 1   ms ) incurs zero decision mismatch. When severe control channel congestion introduces a delay of one discrete slot ( τ d = 1 ), the agent executes decision-making conditioned on the delayed observation s t 1 . Benefiting from the high temporal correlation coefficient ( ρ = 0.95 ), the channel state persistence remains exceptionally high, with the expected prediction error bounded by E [ | h ( t ) h ( t 1 ) | 2 ] = ( 1 ρ 2 ) σ h 2 0.0975 . The observed historical state still preserves over 90 % of the mutual information with the instantaneous channel. Moreover, because the multi-path sweep jamming follows continuous time-frequency trajectories, the short latency does not disrupt the agent’s spatial-temporal evasion capability. In the rare event of prolonged burst congestion ( τ d 2 ), the transmitter autonomously reverts to the strategic silent action ( a = 0 ), preventing catastrophic transmission collisions and providing graceful degradation without link collapse.

6. Conclusions

Targeting the dual threats of malicious multi-mode suppressive jamming and time-varying multi-path fading faced by wireless communication links in complex dynamic electromagnetic adversarial environments, single-dimensional resource scheduling struggles to balance system transmission reliability and energy efficiency. This paper proposes a joint intelligent anti-jamming method for channel switching and transmit power over fading channels based on deep reinforcement learning. The algorithm transforms joint resource scheduling into a Markov Decision Process, reconstructing a composite environment state space by fusing continuous channel gains, jamming observations, and historical residence frequencies. This mechanism effectively decouples physical layer impairment sources and resolves partially observable blind spots. Concurrently, a highly aggregated two-dimensional joint discrete action space for frequency and power is designed, along with a composite reward function balancing communication success rates and power costs, guiding the agent to learn joint optimization policies for multi-dimensional resources end-to-end.
Simulation validations indicate that under highly dynamic jamming environments such as multi-path sweep and intelligent blocking, the proposed algorithm demonstrates effective environmental adaptation and joint resource-selection capabilities. Compared to traditional DQN and conventional Q-learning algorithms, the proposed method not only exhibits no jamming collisions within the presented time-frequency examples but also selects the minimum effective power during favorable channel windows to achieve high-efficiency transmission. This is accomplished by relying on the introduced joint channel-and-power multi-domain mechanism, thereby mitigating the short-sighted strategy of blindly elevating transmit power. Furthermore, the proposed algorithm substantially enhances the system’s dynamic energy efficiency and normalized throughput, genuinely realizing an optimal balance between highly reliable anti-jamming communication and low power overhead. It provides robust theoretical support and methodological references for the design of intelligent communication equipment in complex electromagnetic adversarial environments.

Author Contributions

Conceptualization, Y.W., Y.Z. and Y.N.; methodology, Y.W.; validation, Y.W., Y.Z. and Y.N.; formal analysis, Y.N.; investigation, Y.W.; data curation, Y.Z.; writing—original draft preparation, Y.W.; writing—review and editing, Y.W.; supervision, Y.W.; funding acquisition, Y.Z. and Y.N. All authors have read and agreed to the published version of the manuscript.

Funding

The work was supported in part by the Research Program of the National University of Defense Technology under Grant No. ZK24-58 and in part by the Research Program of the National Key Laboratory of Wireless Communication under Grant No. 2024-kgr-JJ-08. This work was supported by the National Science Foundation of China grant number 62371461.

Data Availability Statement

Due to institutional data privacy requirements, our data are unavailable.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Aref, M.A.; Jayaweera, S.K.; Yepez, E. Survey on cognitive anti-jamming communications. IET Commun. 2020, 14, 3110–3127. [Google Scholar] [CrossRef] [Scilit]
  2. Illi, E.; Bazzi, A.; Qaraqe, M.; Ghrayeb, A. On the Secrecy-Sensing Optimization of RIS-assisted Full-Duplex Integrated Sensing and Communication Network. IEEE Trans. Wirel. Commun. 2026, 25, 9530–9547. [Google Scholar] [CrossRef] [Scilit]
  3. Yao, F. Communication Anti-Jamming Engineering and Practice, 3rd ed.; Electronic Industry Press: Beijing, China, 2025. [Google Scholar]
  4. Slimeni, F.; Chtourou, Z.; Scheers, B.; Le Nir, V.; Attia, R. Cooperative Q-learning based channel selection for cognitive radio networks. Wirel. Netw. 2019, 25, 4161–4171. [Google Scholar] [CrossRef] [Scilit]
  5. Xiao, L.; Li, Y.; Liu, J.; Zhao, Y. Power control with reinforcement learning in cooperative cognitive radio networks against jamming. J. Supercomput. 2015, 71, 3237–3257. [Google Scholar] [CrossRef] [Scilit]
  6. Zhou, Q.; Li, Y.; Niu, Y. A countermeasure against random pulse jamming in time domain based on reinforcement learning. IEEE Access 2020, 8, 97164–97174. [Google Scholar] [CrossRef] [Scilit]
  7. Xiao, L.; Jiang, D.; Xu, D.; Zhu, H.; Zhang, Y.; Poor, H.V. Two-dimensional anti-jamming mobile communication game based on reinforcement learning. IEEE Trans. Commun. 2017, 65, 4319–4329. [Google Scholar]
  8. Wang, Y.; Liu, X.; Wang, M.; Yu, Y. A hidden anti-jamming method based on deep reinforcement learning. KSII Trans. Internet Inf. Syst. 2021, 15, 3444–3457. [Google Scholar] [CrossRef] [Scilit]
  9. Sharma, H.; Kumar, N.; Tekchandani, R. Mitigating jamming attack in 5G heterogeneous networks: A federated deep reinforcement learning approach. IEEE Trans. Veh. Technol. 2023, 72, 2439–2452. [Google Scholar] [CrossRef] [Scilit]
  10. Liu, X.; Xu, Y.; Jia, L.; Wu, Q.; Anpalagan, A. Anti-jamming communications using spectrum waterfall: A deep reinforcement learning approach. IEEE Commun. Lett. 2018, 22, 998–1001. [Google Scholar] [CrossRef] [Scilit]
  11. Wan, B.; Niu, Y.; Chen, C.; Zhou, Z.; Xiang, P. A novel algorithm of joint frequency-power domain anti-jamming based on PER-DQN. Neural Comput. Appl. 2025, 36, 7823–7840. [Google Scholar] [CrossRef] [Scilit]
  12. Van Huynh, N.; Nguyen, D.N.; Hoang, D.T.; Dutkiewicz, E. “Jam me if you can”: Defeating jammer with deep dueling neural network architecture and ambient backscattering communications. IEEE J. Sel. Areas Commun. 2019, 37, 2603–2620. [Google Scholar] [CrossRef] [Scilit]
  13. Li, Y.; Wang, J.; Gao, Z. Learning-based multi-domain anti-jamming communication with unknown information. Electronics 2023, 12, 3901. [Google Scholar] [CrossRef] [Scilit]
  14. Yao, F.; Jia, L. A collaborative multi-agent reinforcement learning anti-jamming algorithm in wireless networks. IEEE Wirel. Commun. Lett. 2019, 8, 1024–1027. [Google Scholar] [CrossRef] [Scilit]
  15. Wu, T.M. A suboptimal maximum-likelihood receiver for FFH/BFSK systems with multitone jamming over frequency-selective Rayleigh-fading channels. IEEE Trans. Veh. Technol. 2008, 57, 1316–1322. [Google Scholar] [CrossRef] [Scilit]
  16. Hu, L.; Shao, Y.; Qian, Y.; Du, F.; Li, J.; Lin, Y.; Wang, Z. Meta-reinforcement learning in time-varying UAV communications: Adaptive anti-jamming channel selection. Radioengineering 2024, 33, 417–431. [Google Scholar] [CrossRef] [Scilit]
  17. Li, X.; Chen, J.; Ling, X.; Wu, T. Deep reinforcement learning-based anti-jamming algorithm using dual action network. IEEE Trans. Wirel. Commun. 2023, 22, 4625–4637. [Google Scholar] [CrossRef] [Scilit]
  18. Pärlin, K.; Byman, A.; Meriläinen, T.; Riihonen, T. Known-interference cancellation over time-varying Rayleigh-fading channels. In Proceedings of the 2025 IEEE Military Communications Conference (MILCOM); IEEE: New York, NY, USA, 2025; pp. 350–355. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.