Figure 1.
Conceptual framework of the Extractor for Adaptive Tradeoff Between Exploration and Exploitation (EATBEE) algorithm for correlation analysis and data quality enhancement within the latent adjacency matrix space. Black solid arrows represent the agent’s real trajectory transitions from the initial to the end state. Black dashed arrows denote the mapping and inverse mapping between the state and latent state spaces. Teal and red dashed arrows indicate latent state directions, with the red dashed arrow pointing to the interaction point between the real and ideal trajectories.
Figure 1.
Conceptual framework of the Extractor for Adaptive Tradeoff Between Exploration and Exploitation (EATBEE) algorithm for correlation analysis and data quality enhancement within the latent adjacency matrix space. Black solid arrows represent the agent’s real trajectory transitions from the initial to the end state. Black dashed arrows denote the mapping and inverse mapping between the state and latent state spaces. Teal and red dashed arrows indicate latent state directions, with the red dashed arrow pointing to the interaction point between the real and ideal trajectories.
Figure 2.
The optimal policy in a trajectory . Black solid arrows denote state transitions induced by the original policy . Red solid arrows denote state transitions induced by the optimal policy . The red dashed arrow represents the skipped segment of the original trajectory, which is replaced by the optimal sub-trajectory. The original trajectory is induced by the original sub-policy , whereas the optimal sub-trajectory is induced by the optimal sub-policy , which skips the sub-sequence from to .
Figure 2.
The optimal policy in a trajectory . Black solid arrows denote state transitions induced by the original policy . Red solid arrows denote state transitions induced by the optimal policy . The red dashed arrow represents the skipped segment of the original trajectory, which is replaced by the optimal sub-trajectory. The original trajectory is induced by the original sub-policy , whereas the optimal sub-trajectory is induced by the optimal sub-policy , which skips the sub-sequence from to .
Figure 3.
Backpropagation mechanism of EATBEE. Within the latent state space, EATBEE identifies a potentially optimal trajectory and updates the policy accordingly. The robot icon represents the agent interacting with the environment, and the red star denotes the target end state. Orange dots are original states in the state and latent spaces, green dots are newly identified potentially optimal states added via backpropagation, and the red dot is the central target state in the latent space. Black solid arrows indicate state transitions in the original trajectory, while black dashed lines separate different processing stages. Red dashed arrows represent the encoding mapping and decoding mapping between the state space and latent state space. Orange dashed arrows show the temporal correspondence between trajectory states and the correlation matrix, and green curved arrows denote the backpropagation process that propagates optimal state information to update the policy. The color bar on the right represents the correlation coefficient between latent states (ranging from 0 to 1, with deeper blue indicating higher correlation).
Figure 3.
Backpropagation mechanism of EATBEE. Within the latent state space, EATBEE identifies a potentially optimal trajectory and updates the policy accordingly. The robot icon represents the agent interacting with the environment, and the red star denotes the target end state. Orange dots are original states in the state and latent spaces, green dots are newly identified potentially optimal states added via backpropagation, and the red dot is the central target state in the latent space. Black solid arrows indicate state transitions in the original trajectory, while black dashed lines separate different processing stages. Red dashed arrows represent the encoding mapping and decoding mapping between the state space and latent state space. Orange dashed arrows show the temporal correspondence between trajectory states and the correlation matrix, and green curved arrows denote the backpropagation process that propagates optimal state information to update the policy. The color bar on the right represents the correlation coefficient between latent states (ranging from 0 to 1, with deeper blue indicating higher correlation).
![Mathematics 14 01624 g003 Mathematics 14 01624 g003]()
Figure 4.
The statistical mechanism of EATBEE. The light blue bars represent the original data distribution, and the light orange bars represent the optimized best data distribution. The brown bars denote the overlapping region between the two distributions, corresponding to the retained valid data that balances exploration and exploitation in the light exploration interval. The light blue and light orange smooth curves are the fitted density curves of the original and optimized distributions, respectively. The blue-labeled regions indicate overexploration segments where data deviates from the task goal. The red-bracketed regions mark the light exploration segments that balance exploration and exploitation, while the orange-bracketed region highlights the exploitation-focused data distribution driven by the task objective. When the agent explores excessively, EATBEE extracts exploitative transitions to improve policy learning.
Figure 4.
The statistical mechanism of EATBEE. The light blue bars represent the original data distribution, and the light orange bars represent the optimized best data distribution. The brown bars denote the overlapping region between the two distributions, corresponding to the retained valid data that balances exploration and exploitation in the light exploration interval. The light blue and light orange smooth curves are the fitted density curves of the original and optimized distributions, respectively. The blue-labeled regions indicate overexploration segments where data deviates from the task goal. The red-bracketed regions mark the light exploration segments that balance exploration and exploitation, while the orange-bracketed region highlights the exploitation-focused data distribution driven by the task objective. When the agent explores excessively, EATBEE extracts exploitative transitions to improve policy learning.
Figure 5.
The return for an excluded state . Black solid arrows denote the standard mapping from the original return to the filtered return for normally retained states. Black dashed arrows indicate the correspondence between original and filtered returns. Red solid and dashed arrows represent the special handling for the filtered state : the return is forced to 0, and its experience is excluded from the return calculation. The orange dashed box highlights the processing unit of the filtered state . When is filtered out by EATBEE, is set to 0.
Figure 5.
The return for an excluded state . Black solid arrows denote the standard mapping from the original return to the filtered return for normally retained states. Black dashed arrows indicate the correspondence between original and filtered returns. Red solid and dashed arrows represent the special handling for the filtered state : the return is forced to 0, and its experience is excluded from the return calculation. The orange dashed box highlights the processing unit of the filtered state . When is filtered out by EATBEE, is set to 0.
Figure 6.
The operational mechanism of EATBEE. The blue horizontal line separates the two trajectory groups: the top row shows the optimal trajectory refined by EATBEE, while the bottom row shows the raw execution of the original trajectory. Black curved arrows labeled “Optimal Path” denote the progression of the optimized trajectory across steps. Green solid arrows labeled “Efficient Backpropagation” indicate the direction of the efficient backpropagation process that propagates optimal information backward. Orange bidirectional arrows show the correspondence between steps in the original and optimized trajectories. Red dashed boxes highlight redundant steps in the original trajectory that are eliminated by EATBEE. In each game frame, the blue triangle represents the agent, the light blue symbols represent the key and door respectively, and the green square indicates the terminal state. The original trajectory comprises 225 steps, which is significantly reduced to 27 steps following EATBEE optimization. Both the current and total step counts are displayed in the upper-left corner of each frame.
Figure 6.
The operational mechanism of EATBEE. The blue horizontal line separates the two trajectory groups: the top row shows the optimal trajectory refined by EATBEE, while the bottom row shows the raw execution of the original trajectory. Black curved arrows labeled “Optimal Path” denote the progression of the optimized trajectory across steps. Green solid arrows labeled “Efficient Backpropagation” indicate the direction of the efficient backpropagation process that propagates optimal information backward. Orange bidirectional arrows show the correspondence between steps in the original and optimized trajectories. Red dashed boxes highlight redundant steps in the original trajectory that are eliminated by EATBEE. In each game frame, the blue triangle represents the agent, the light blue symbols represent the key and door respectively, and the green square indicates the terminal state. The original trajectory comprises 225 steps, which is significantly reduced to 27 steps following EATBEE optimization. Both the current and total step counts are displayed in the upper-left corner of each frame.
![Mathematics 14 01624 g006 Mathematics 14 01624 g006]()
Figure 7.
Consistency and monotonicity analysis of value functions: EATBEE vs. original trajectories. Black curved arrows labeled “Optimal Path” denote the progression of the optimized trajectory across steps. Green solid arrows labeled “Efficient Backpropagation” indicate the direction of the efficient backpropagation process that propagates optimal information backward. Orange bidirectional arrows show the correspondence between steps in the original and optimized trajectories. In each game frame, the blue triangle represents the agent, the light blue symbols represent the key and door respectively, and the green square indicates the terminal state. In the original sequence (bottom), the agent transitions through numerous redundant states (e.g., from step 2 to 20). EATBEE identifies these as low-information segments and refines them into an optimized trajectory (top). Crucially, the labels and indicate that state-value estimates remain stable and monotonically increasing despite the removal of intermediate steps. This confirms that the extractor effectively preserves the task-relevant knowledge distribution.
Figure 7.
Consistency and monotonicity analysis of value functions: EATBEE vs. original trajectories. Black curved arrows labeled “Optimal Path” denote the progression of the optimized trajectory across steps. Green solid arrows labeled “Efficient Backpropagation” indicate the direction of the efficient backpropagation process that propagates optimal information backward. Orange bidirectional arrows show the correspondence between steps in the original and optimized trajectories. In each game frame, the blue triangle represents the agent, the light blue symbols represent the key and door respectively, and the green square indicates the terminal state. In the original sequence (bottom), the agent transitions through numerous redundant states (e.g., from step 2 to 20). EATBEE identifies these as low-information segments and refines them into an optimized trajectory (top). Crucially, the labels and indicate that state-value estimates remain stable and monotonically increasing despite the removal of intermediate steps. This confirms that the extractor effectively preserves the task-relevant knowledge distribution.
![Mathematics 14 01624 g007 Mathematics 14 01624 g007]()
Figure 8.
Environments with discrete actions. (a) MiniGrid-DoorKey-8 × 8: In each game frame, the red triangle represents the agent, the yellow symbols represent the key and door respectively, and the green square indicates the terminal state. A small grid-world environment with sparse rewards, where the agent must collect a key, open the door, and reach the goal state. (b) MiniGrid-DoorKey-: A larger-scale variant of the DoorKey task with increased complexity. (c) LunarLander: A continuous control environment with dense rewards, where the agent controls a lander to safely land between two flags.
Figure 8.
Environments with discrete actions. (a) MiniGrid-DoorKey-8 × 8: In each game frame, the red triangle represents the agent, the yellow symbols represent the key and door respectively, and the green square indicates the terminal state. A small grid-world environment with sparse rewards, where the agent must collect a key, open the door, and reach the goal state. (b) MiniGrid-DoorKey-: A larger-scale variant of the DoorKey task with increased complexity. (c) LunarLander: A continuous control environment with dense rewards, where the agent controls a lander to safely land between two flags.
Figure 9.
Environments with continuous actions. The agent needs to keep walking and must not fall down. (a) Hopper: A single-leg hopping environment where the agent controls a single articulated leg to maintain balance and move forward. (b) Walker2d: A two-legged walking environment where the agent controls two legs (purple for the left leg, brown for the right leg) to achieve stable forward locomotion. (c) Ant: A four-legged ant-like environment with higher-dimensional continuous action space, requiring coordinated control of all legs to move forward.
Figure 9.
Environments with continuous actions. The agent needs to keep walking and must not fall down. (a) Hopper: A single-leg hopping environment where the agent controls a single articulated leg to maintain balance and move forward. (b) Walker2d: A two-legged walking environment where the agent controls two legs (purple for the left leg, brown for the right leg) to achieve stable forward locomotion. (c) Ant: A four-legged ant-like environment with higher-dimensional continuous action space, requiring coordinated control of all legs to move forward.
Figure 10.
Visualization of data distribution for different trajectory types. (a) Type 1: exploration with task failure. (b) Type 2: exploration with task success. (c) Type 3: exploitation with task success. The horizontal axis represents the state embedding after Principal Component Analysis (PCA) projection (from to ), and the vertical axis represents density. The blue bars and corresponding blue solid Kernel Density Estimation (KDE) curve show the distribution of the original trajectory , while the orange bars and corresponding orange solid KDE curve show the distribution of the EATBEE-filtered trajectory .
Figure 10.
Visualization of data distribution for different trajectory types. (a) Type 1: exploration with task failure. (b) Type 2: exploration with task success. (c) Type 3: exploitation with task success. The horizontal axis represents the state embedding after Principal Component Analysis (PCA) projection (from to ), and the vertical axis represents density. The blue bars and corresponding blue solid Kernel Density Estimation (KDE) curve show the distribution of the original trajectory , while the orange bars and corresponding orange solid KDE curve show the distribution of the EATBEE-filtered trajectory .
Figure 11.
Evaluations in the MiniGrid-DoorKey-8 × 8 environment. (
a) The EATBEE performance. (
b) The Ratio between exploration and exploitation. Detailed hyperparameter configurations for both Proximal Policy Optimization (PPO) and Deep Q-Network (DQN) are summarized in
Table 1. In our benchmarks, PPO + Long Short-Term Memory (LSTM) denotes the PPO algorithm integrated with an LSTM network, while PPO + Intrinsic Curiosity Module (ICM) and PPO + EATBEE refer to PPO augmented with the ICM and our proposed EATBEE module, respectively. Regarding the DQN baselines, DQN + Experience Replay (ER) represents the vanilla DQN equipped with ER and a target network, whereas DQN + Prioritized Experience Replay (PER) denotes the variant incorporating PER.
Figure 11.
Evaluations in the MiniGrid-DoorKey-8 × 8 environment. (
a) The EATBEE performance. (
b) The Ratio between exploration and exploitation. Detailed hyperparameter configurations for both Proximal Policy Optimization (PPO) and Deep Q-Network (DQN) are summarized in
Table 1. In our benchmarks, PPO + Long Short-Term Memory (LSTM) denotes the PPO algorithm integrated with an LSTM network, while PPO + Intrinsic Curiosity Module (ICM) and PPO + EATBEE refer to PPO augmented with the ICM and our proposed EATBEE module, respectively. Regarding the DQN baselines, DQN + Experience Replay (ER) represents the vanilla DQN equipped with ER and a target network, whereas DQN + Prioritized Experience Replay (PER) denotes the variant incorporating PER.
Figure 12.
EATBEE performance in the MiniGrid-DoorKey-16 × 16 environment. (a) The EATBEE performance. (b) The Ratio between exploration and exploitation.
Figure 12.
EATBEE performance in the MiniGrid-DoorKey-16 × 16 environment. (a) The EATBEE performance. (b) The Ratio between exploration and exploitation.
Figure 13.
EATBEE performance in the LunarLander environment. (a) The EATBEE performance. (b) The Ratio between exploration and exploitation.
Figure 13.
EATBEE performance in the LunarLander environment. (a) The EATBEE performance. (b) The Ratio between exploration and exploitation.
Figure 14.
EATBEE performance in MuJoCo environment. (a) Hopper. (b) Walker2d. (c) Ant.
Figure 14.
EATBEE performance in MuJoCo environment. (a) Hopper. (b) Walker2d. (c) Ant.
Figure 15.
The analysis of transfer learning using EATBEE in the MiniGrid-DoorKey environment. (a) The transfer learning in MiniGrid-DoorKey-8 × 8 for EATBEE. (b) The transfer learning in MiniGrid-DoorKey-16 × 16 for EATBEE. The EATBEE approach does not focus on a specific policy but rather on the data distribution of the task, making it highly adaptable to similar or identical tasks. The following symbols are explained: 8 × 8 PPO refers to the MiniGrid-DoorKey-8 × 8 environment using PPO; 8 × 8 EATBEE refers to the MiniGrid-DoorKey-8 × 8 environment using PPO with EATBEE; 8 × 8 5 × 5-no refers to using the converged PPO with EATBEE from the 5 × 5 environment without updating the EATBEE network in the MiniGrid-DoorKey-8 × 8 environment; and, finally, 8 × 8 5 × 5-adj refers to using the converged PPO with the EATBEE module from the 5 × 5 environment while updating the EATBEE network in an 8 × 8 MiniGrid-DoorKey setting. In a MiniGrid-DoorKey-16 × 16 scenario, transfer learning with EATBEE carries the same significance.
Figure 15.
The analysis of transfer learning using EATBEE in the MiniGrid-DoorKey environment. (a) The transfer learning in MiniGrid-DoorKey-8 × 8 for EATBEE. (b) The transfer learning in MiniGrid-DoorKey-16 × 16 for EATBEE. The EATBEE approach does not focus on a specific policy but rather on the data distribution of the task, making it highly adaptable to similar or identical tasks. The following symbols are explained: 8 × 8 PPO refers to the MiniGrid-DoorKey-8 × 8 environment using PPO; 8 × 8 EATBEE refers to the MiniGrid-DoorKey-8 × 8 environment using PPO with EATBEE; 8 × 8 5 × 5-no refers to using the converged PPO with EATBEE from the 5 × 5 environment without updating the EATBEE network in the MiniGrid-DoorKey-8 × 8 environment; and, finally, 8 × 8 5 × 5-adj refers to using the converged PPO with the EATBEE module from the 5 × 5 environment while updating the EATBEE network in an 8 × 8 MiniGrid-DoorKey setting. In a MiniGrid-DoorKey-16 × 16 scenario, transfer learning with EATBEE carries the same significance.
![Mathematics 14 01624 g015 Mathematics 14 01624 g015]()
Figure 16.
Ablation experiments on state-value functions. (a) Number of state-value functions. (b) The Ratio between exploration and exploitation. Symbol explanation: V2 denotes the use of two state-value functions. As the number of state-value functions increases, EATBEE exhibits enhanced capability in data filtering, leading to more stable overall policy performance. For instance, PPO-EATBEE-V5 and -V10 demonstrate smaller standard deviations, with the final Ratio metric approaching 1.
Figure 16.
Ablation experiments on state-value functions. (a) Number of state-value functions. (b) The Ratio between exploration and exploitation. Symbol explanation: V2 denotes the use of two state-value functions. As the number of state-value functions increases, EATBEE exhibits enhanced capability in data filtering, leading to more stable overall policy performance. For instance, PPO-EATBEE-V5 and -V10 demonstrate smaller standard deviations, with the final Ratio metric approaching 1.
Figure 17.
Ablation experiments on the knowledge coefficient. (a) Size of knowledge coefficient . (b) The Ratio between exploration and exploitation. As this coefficient increases, EATBEE consistently guarantees stable convergence of RL algorithms. The Ratio results reveal that a larger knowledge coefficient enables stricter capture of environmental knowledge distribution. For tasks requiring precise modeling of environmental knowledge patterns, a relatively large value of is recommended.
Figure 17.
Ablation experiments on the knowledge coefficient. (a) Size of knowledge coefficient . (b) The Ratio between exploration and exploitation. As this coefficient increases, EATBEE consistently guarantees stable convergence of RL algorithms. The Ratio results reveal that a larger knowledge coefficient enables stricter capture of environmental knowledge distribution. For tasks requiring precise modeling of environmental knowledge patterns, a relatively large value of is recommended.
Figure 18.
Ablation experiments on reachability. (a) Latent Space and Reachability Effects. (b) The Ratio between exploration and exploitation. When reachability within the latent space is compromised, EATBEE loses rigorous constraints during data filtering. This issue introduces abrupt jumps in pruned trajectory data, weakening the model’s capacity to capture environmental knowledge. Consequently, the RL algorithm fails to learn from valid, high-quality samples.
Figure 18.
Ablation experiments on reachability. (a) Latent Space and Reachability Effects. (b) The Ratio between exploration and exploitation. When reachability within the latent space is compromised, EATBEE loses rigorous constraints during data filtering. This issue introduces abrupt jumps in pruned trajectory data, weakening the model’s capacity to capture environmental knowledge. Consequently, the RL algorithm fails to learn from valid, high-quality samples.
Table 1.
Hyperparameter details.
Table 1.
Hyperparameter details.
| | PPO | DQN |
|---|
| Optimizer | Adam [49] | Adam |
| learning rate | | |
| discount | | |
| Replay Buffer Size | None | |
| Number of Hidden Layers | 2 | 2 |
| Number of Hidden Units Per Layer | 64 | 64 |
| Number of Samples Per Minibatch | None | 256 |
| Nonlinearity | ReLU | ReLU |
Table 2.
The performance comparison in the MiniGrid-DoorKey-8 × 8 environment.
Table 2.
The performance comparison in the MiniGrid-DoorKey-8 × 8 environment.
| | Min Step | Max Step | 10% MS | 20% MS | 30% MS | 40% MS | 50% MS |
|---|
| PPO | | 500 | | | | | |
| PPO + LSTM | | 500 | | | | | |
| PPO + ICM | | 500 | | | | | |
| PPO + EATBEE | | 500 | | | | | |
| PPO + LSTM + EATBEE | | 500 | | | | | |
| DQN + ER | | 500 | | | | | |
| DQN + PER | | 500 | | | | | |
Table 3.
The performance comparison in the MiniGrid-DoorKey-16 × 16 environment.
Table 3.
The performance comparison in the MiniGrid-DoorKey-16 × 16 environment.
| | Min Step | Max Step | 10% MS | 20% MS | 30% MS | 40% MS | 50% MS |
|---|
| PPO | 339.4 ± 198.52 | 500 | 496.47 ± 4.63 | 496.89 ± 3.28 | 496.77 ± 2.55 | 497.42 ± 2.05 | 497.79 ± 1.71 |
| PPO + LSTM | | 500 | | | | | |
| PPO + ICM | | 500 | | | | | |
| PPO + EATBEE | | 500 | | | | | |
| PPO + LSTM + EATBEE | | 500 | | | | | |
| DQN + ER | 426.2 ± 147.6 | 500 | 499.86 ± 0.28 | 500.0 ± 0.0 | 500 ± 0.0 | 500 ± 0.0 | 500 ± 0.0 |
| DQN + PER | 448.2 ± 103.60 | 500 | 499.98 ± 0.05 | 499.99 ± 0.03 | 499.99 ± 0.02 | 499.99 ± 0.02 | 499.99 ± 0.01 |
Table 4.
The performance comparison in the LunarLander environment.
Table 4.
The performance comparison in the LunarLander environment.
| | Min Reward | Max Reward | 10% MR | 20% MR | 30% MR | 40% MR | 50% MR |
|---|
| PPO | −314.42 ± 101.27 | 290.33 ± 23.22 | 230.98 ± 42.04 | 231.10 ± 42.77 | 231.19 ± 43.22 | 230.71 ± 43.78 | 229.61 ± 43.92 |
| PPO + LSTM | | | | | | | |
| PPO + EATBEE | | | | | | | |
| PPO + LSTM + EATBEE | | | | | | | |
| DQN + ER | −538.21 ± 1313.68 | 291.78 ± 9.14 | 108.86 ± 93.92 | 120.17 ± 99.12 | 141.25 ± 86.89 | 155.61 ± 72.13 | 156.20 ± 57.56 |
| DQN + PER | −787.09 ± 1244.59 | 287.51 ± 13.34 | −78.10 ± 111.93 | −57.22 ± 67.68 | −54.24 ± 49.29 | −47.92 ± 39.76 | −37.41 ± 36.66 |
Table 5.
The performance comparison in MuJoCo.
Table 5.
The performance comparison in MuJoCo.
| | Min Reward | Max Reward | 10% MR | 20% MR | 30% MR | 40% MR | 50% MR |
|---|
| PPO-Hopper | 10.42 ± 3.41 | 1001.86 ± 981.75 | 555.89 ± 353.09 | 520.14 ± 302.63 | 506.22 ± 282.70 | 484.76 ± 260.26 | 463.50 ± 231.01 |
| PPO + EATBEE-Hopper | | | | | | | |
| PPO-Walker | −7.71 ± 9.17 | 612.83 ± 450.23 | 291.29 ± 41.78 | 289.34 ± 43.31 | 289.58 ± 43.20 | 290.96 ± 43.87 | 290.74 ± 46.67 |
| PPO + EATBEE-Walker | | | | | | | |
| PPO-Ant | −1183.82 ± 872.45 | 66.14 ± 46.72 | −1103.86 ± 975.20 | −982.08 ± 898.34 | −897.62 ± 883.52 | −857.43 ± 888.60 | −834.28 ± 894.73 |
| PPO + EATBEE-Ant | | | | | | | |