Dynamic Task Planning for Heterogeneous Platforms via Spatio-Temporal and Capability Dual-Driven Framework
Abstract
1. Introduction
- Spatio-Temporal Coverage Integrity: Maximizing continuous coverage duration for both stationary and dynamic targets represents a fundamental objective.
- Real-Time Performance Constraints: Minimizing computational overhead in large-scale mission scheduling remains imperative.
- Resource Elasticity and Reconfiguration: Supporting dynamic strategy adaptation and priority reassignment for emergent high-value targets is essential.
- A multi-granularity task decomposition algorithm architecture incorporating platform dynamics is designed, enabling adaptive task allocation for heterogeneous platforms.
- A curvature-compensated accessibility–visibility joint evaluation algorithm is proposed, significantly improving detection accuracy in complex environments.
- A multi-objective optimized DRL algorithm module is constructed, achieving coordinated optimization of task coverage and planning efficiency.
- A simulation environment and task dataset for land–sea–air–space heterogeneous platform cooperative observation have been established, designed with three distinct scenarios to validate the algorithm’s effectiveness across small-, medium- and large-scale targets.
2. Related Work
2.1. Research Progress on Multi-Agent DRL Scheduling Methods
2.2. Research Progress on Heterogeneous Platform Task Planning
2.3. Autonomous Collaborative Mission Planning: Real-World Motivation for the Proposed Model
- Unified Description of Heterogeneous Resources: A vehicle fleet may comprise vehicles from different brands, with varying sensor configurations and performance capabilities (acceleration, braking), analogous to the capability diversity of land, sea, air, and space platforms.
- Modeling of Spatio-Temporally Coupled Constraints: Vehicle movement must strictly adhere to dynamic constraints (speed, acceleration), traffic rules (routes, signals), and spatio-temporal relationships (collision avoidance, platoon formation maintenance). This necessitates a refined spatio-temporal and capability model similar to the one established in Section 3.2.1.
- Dynamic Real-Time Decision-Making: In the face of dynamic events such as sudden road conditions, pedestrian interference, or vehicle failures, the cooperative system must possess millisecond-level real-time re-planning and re-scheduling capabilities. This is precisely the shortcoming of traditional meta-heuristic algorithms (e.g., Genetic Algorithm, Ant Colony Optimization) and forms the fundamental reason for introducing the DRL.
3. Materials and Methods
3.1. Dynamic Task Preprocessing Design for Heterogeneous Platforms in Complex Application Scenarios
- First, identify detection tasks executable by satellites;
- Second, rationally allocate tasks that cannot be completed by satellites.
3.1.1. Spatiotemporal Segmentation Description of Detection Tasks
- Space platform Detection Capability Description
- 2.
- Detection Time Window Partitioning
- Treat time windows detectable by space platforms as independent detection periods;
- Consider discontinuous periods within required detection time ranges, segmented by space platform-available windows, as separate independent periods;
- Regard periodic or intermittent measurement periods as independent segments.
3.1.2. Complex Detection Task Decomposition
- Decomposition Mode Classification
- Pre-decomposition Oriented to Task Feasibility
- Task Decomposition Adapted to Actual Situations
- 2.
- Task Domain Decomposition Algorithm Design
- Grid Partitioning of Surveillance Area
- Calculation of Effective Detection Positions
3.2. Spatiotemporal-Capability Dual-Driven Accessibility–Visibility Characterization for Heterogeneous Platforms
3.2.1. Detection Performance Characterization Model Design
3.2.2. Accessibility Calculation for Heterogeneous Platforms
- Feasible Detection Boundary Calculation
- Detection Boundary Calculation for Land and sea platforms
- Detection Boundary Calculation for Air Platforms
- Calculation of Detection Range
- Optimal Path Calculation
3.3. Design of a Spatiotemporal-Capability Dual-Driven Dynamic Task Planning Algorithm for Heterogeneous Platforms
- Task Allocation Completion Rate
- Task Time Overlap Rate
- Residual Detection Capability
- Design of Comprehensive Optimization Metric Function
- Implementation of Task Planning Strategy (Priority Handling)
3.4. Design of Task Planning Optimization Algorithm Based on DPL
- State (at time )
- Action (at time )
- Reward Function (at time )
- DRL under Sufficient Computational Resources
4. Result
5. Discussion
5.1. Analysis of Algorithm Results
- Target Coverage Rate
- Planning Time Consumption
- Resource Utilization Rate
5.2. Analysis of Algorithm Principles
- Greedy Algorithm: Rule-Driven Myopic Decision-Making Paradigm
- Genetic Algorithm: Search-Driven Random Exploration Paradigm
- STAS-Net: Learning-Driven Graph Reasoning Decision-Making Paradigm
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| A | surveillance area |
| AEOS | Agile Earth Observing Satellites |
| three-dimensional grid with index | |
| Action parameter in STAS-Net model | |
| receiver bandwidth | |
| book | flag vector, with a size equal to the number of nodes |
| CLDE | Centralized Training with Decentralized Execution |
| DRL | Deep Reinforcement Learning |
| DMASA | Distributed Multi-Agent Scheduling Architecture |
| maximum detection distance of the i-th category | |
| -th detection unit | |
| receiver noise figure | |
| FJSP | Flexible Job Shop Scheduling Problem |
| GE-HetGNN | Graph Embedding–Heterogeneous Graph Neural Network |
| GCN | Graph Convolutional Network |
| GNN | Graph Neural Networks |
| antenna gain in the given direction | |
| receiving antenna gain | |
| Boltzmann’s constant | |
| grid side length | |
| MAPPO | Multi-Agent Proximal Policy Optimization |
| MAS | Multi-Agent Systems |
| the -th type of detection task | |
| the -th task, which can correspond to different types of tasks | |
| mindis | A set stores the distances from the starting point to all other points |
| category of detection unit | |
| total number of various detection tasks | |
| Earth’s center | |
| A location point of detection unit, including longitude, latitude and height | |
| A location point of target, including longitude, latitude and height | |
| set of vertices for which the shortest path has been determined | |
| positional center of the -th target area | |
| echo power | |
| transmission power | |
| location of the -th detection platform | |
| distance between the radar and the target | |
| stores the sequence of nodes traversed from the starting point | |
| RH-MAPPO | Rolling-Horizon-based Multi-Agent Proximal Policy Optimization |
| Current Detection Platform Coverage Rate | |
| proportion of the total detection area that can be effectively covered by unallocated resources | |
| STAS-Net | Spatio-Temporal Adaptive Scheduling Network |
| the output signal-to-noise ratio of the receiver | |
| State parameter in STAS-Net model | |
| T | observation time |
| set of grid indices of the area effectively covered by the | |
| temperature | |
| UAV | Unmanned Aerial Vehicle |
| V2V | vehicle-to-vehicle |
| V2X | vehicle-to-everything |
| set of unvisited points | |
| α | a proportionality coefficient (typically 1/3) |
| the expected coverage performance of the area | |
| target scattering cross-sectional area | |
| radar wavelength | |
| the weight value | |
| expected detectable probability of the area | |
| expected positioning accuracy | |
| the fixed movement distance per step for each detection unit | |
| a given threshold angle for ground clutter interference |
References
- Feng, H.L. Research on Mission Planning Methods for Space-Air Cooperative Observation. Master’s Thesis, Harbin Engineering University, Harbin, China, 2022. [Google Scholar]
- Hou, R.; Cheng, Y.T.; Li, H. Intelligent Mission Planning Technology for Cooperative Monitoring of Marine Multi-Platform and Multi-Sensor. Mar. Inform. 2020, 35, 11–19. [Google Scholar]
- Huang, W.; Qin, Q. Mission Planning Method of Heterogeneous Reconnaissance System Based on Simulated Annealing Algorithm. J. Phys. Conf. Ser. 2024, 1290, 012034. [Google Scholar] [CrossRef]
- Park, K.; Moon, I. Multi-agent deep reinforcement learning approach for EV charging scheduling in a smart grid. Appl. Energy 2022, 328, 120111. [Google Scholar] [CrossRef]
- Jing, X.; Yao, X.; Liu, M.; Zhou, J. Multi-agent reinforcement learning based on graph convolutional network for flexible job shop scheduling. J. Intell. Manuf. 2024, 35, 75–93. [Google Scholar] [CrossRef]
- Pu, Y.; Li, F.; Rahimifard, S. Multi-Agent Reinforcement Learning for Job Shop Scheduling in Dynamic Environments. Sustainability 2024, 16, 3234. [Google Scholar] [CrossRef]
- Zhu, X.; Xu, J.; Ge, J.; Wang, Y.; Xie, Z. Multi-Task Multi-Agent Reinforcement Learning for Real-Time Scheduling of a Dual-Resource Flexible Job Shop with Robots. Processes 2023, 11, 267. [Google Scholar] [CrossRef]
- Zhao, R.; Tao, S.; Li, P. Safety-efficiency integrated assembly: The next-stage adaptive task allocation and planning framework for human–robot collaboration. Robot. Comput.-Integr. Manuf. 2025, 94, 102942. [Google Scholar] [CrossRef]
- Wang, Y.; Xiang, B.; Huang, S. SCRIMP: Scalable Communication for Reinforcement- and Imitation-Learning-Based Multi-Agent Pathfinding; IEEE: Detroit, MI, USA, 2023; pp. 9301–9308. [Google Scholar]
- Liu, R.; Piplani, R.; Toro, C. A deep multi-agent reinforcement learning approach to solve dynamic job shop scheduling problem. Comput. Oper. Res. 2023, 159, 106294. [Google Scholar] [CrossRef]
- Zhao, X.; Wu, C. Large-Scale Machine Learning Cluster Scheduling via Multi-Agent Graph Reinforcement Learning. IEEE Trans. Netw. Serv. Manag. 2022, 19, 4962–4974. [Google Scholar] [CrossRef]
- Jung, S.; Yun, W.J.; Shin, M.; Kim, J.; Kim, J.H. Orchestrated Scheduling and Multi-Agent Deep Reinforcement Learning for Cloud-Assisted Multi-UAV Charging Systems. IEEE Trans. Veh. Technol. 2021, 70, 5362–5377. [Google Scholar] [CrossRef]
- Jayanetti, A.; Halgamuge, S.; Buyya, R. Multi-agent deep reinforcement learning framework for renewable energy-aware workflow scheduling on distributed cloud data centers. IEEE Trans. Parallel Distrib. Syst. 2024, 35, 604–615. [Google Scholar] [CrossRef]
- Shen, W.; Lin, W.; Wu, W. Reinforcement learning-based task scheduling for heterogeneous computing in end-edge-cloud environment. Clust. Comput. 2025, 28, 179. [Google Scholar] [CrossRef]
- Zhang, X.; Wang, Q.; Yu, J.; Sun, Q.; Hu, H.; Liu, X. A Multi-Agent Deep-Reinforcement-Learning-Based Strategy for Safe Distributed Energy Resource Scheduling in Energy Hubs. Electronics 2023, 12, 4763. [Google Scholar] [CrossRef]
- Kaewdornhan, N.; Srithapon, C.; Liemthong, R.; Chatthaworn, R. Real-Time Multi-Home Energy Management with EV Charging Scheduling Using Multi-Agent Deep Reinforcement Learning Optimization. Energies 2023, 16, 2357. [Google Scholar] [CrossRef]
- Yang, H.Y. Research on Large-Scale Mission Planning Methods for Giant Remote Sensing LEO Satellite Clusters. Master’s Thesis, Harbin Engineering University, Harbin, China, 2024. [Google Scholar]
- Wang, J.F.; Jia, G.W.; Guo, Z. A Survey of Research on Multi-UAV Cooperative Mission Planning Methods. Syst. Eng. Electron. 2024, 46, 3437–3450. [Google Scholar]
- Liu, Y.; Cong, J.Y. Research Review on UAV Cooperative Mission Planning. Ship Electron. Eng. 2025, 45, 22–27. [Google Scholar]
- Hall, N.G.; Magazine, M.J. Maximizing the value of a space mission. Eur. J. Oper. Res. 1994, 78, 224–241. [Google Scholar] [CrossRef]
- Harrison, S.A.; Price, M.E.; Philpott, M.S. Task Scheduling for Satellite Based Imagery; Durham University: Durham, UK, 1999; pp. 64–78. [Google Scholar]
- Iacopino, C.; Palmer, P.; Brewer, A. EO Constellation MPS Based on Ant Colony Optimization Algorithms; IEEE: Istanbul, Turkey, 2013; pp. 159–164. [Google Scholar]
- He, R.J. Research on Scheduling Problems for Imaging Reconnaissance Satellites. Ph.D. Thesis, National University of Defense Technology, Changsha, China, 2005. [Google Scholar]
- Wang, X.; Zhao, F.; Shi, Z.; Jin, Z. Deep Reinforcement Learning-Based Periodic Earth Observation Scheduling for Agile Satellite Constellation. J. Aerosp. Inf. Syst. 2023, 20, 508–519. [Google Scholar] [CrossRef]
- Li, Z.; Zhu, X.; Liu, C.; Song, J.Y.; Liu, Y.; Yin, C. Dynamic task scheduling optimization by rolling horizon deep reinforcement learning for distributed satellite system. Expert Syst. Appl. 2025, 289, 128350. [Google Scholar] [CrossRef]
- Xue, P.; Pi, Y. Collaborative planning and control of heterogeneous multi-ground unmanned platforms. Eng. Appl. Artif. Intell. 2024, 136, 108968. [Google Scholar] [CrossRef]
- Wiseman, Y. Autonomous Vehicles. In Encyclopedia of Information Science and Technology; IGI Global Scientific Publishing: Hershey, PA, USA, 2020; Volume 1, pp. 1–11. [Google Scholar]
- Lu, J. Research on Multi-Platform Joint Mission Planning Method for Maritime Target Search. Master’s Thesis, National University of Defense Technology, Changsha, China, 2020. [Google Scholar]
- Liu, C.L.; Xu, J.F.; Peng, J.X. Construction of Space-Air Integrated Unmanned Intelligent Detection System and Its Technical Prospects. Natl. Def. Sci. Technol. 2025, 46, 89–95+142. [Google Scholar]
- Bai, S.; Jiang, N.; Qin, H. Research on Multi-Platform Cooperative Search Task Planning Method Based on Cuckoo Search Algorithm. Shipboard Electron. Countermeas. 2025, 48, 71–75+79. [Google Scholar]
- Zhao, W.D.; Wang, T.J. Research on New Technical Means of UAV Detection. Digit. Commun. World 2021, 15–16+24. [Google Scholar]
- Xu, C.; Tang, B. Digital Twin-Driven Collaborative Scheduling for Heterogeneous Task and Edge-End Resource via Multi-Agent Deep Reinforcement Learning. IEEE J. Sel. Areas Commun. 2023, 41, 3120–3135. [Google Scholar] [CrossRef]
- Hu, H.F.; Wu, A.D.; Han, B. Research on Intelligent Planning Algorithm for Arctic Optimal Routes Based on Deep Reinforcement Learning. J. Glaciol. Geocryol. 2025, 47, 587–598. [Google Scholar]
- Huang, G.Q. Implementation Scheme of Mission Planning Framework for Mobile Reconnaissance Platforms. In Proceedings of the 10th China Command and Control Conference, Beijing, China, 7–9 July 2022; Volume I, pp. 29–36. [Google Scholar]
- Zhao, D.M.; Xiong, J. A dynamic planning method for satellite imaging mission based on improved genetic algorithm. Appl. Math. Nonlinear Sci. 2024, 9, 20241526. [Google Scholar] [CrossRef]
- Liu, C.R. A Multi-Satellite Multi-Target Observation Task Planning and Replanning Method Based on DQN. Sensors 2025, 25, 1123. [Google Scholar]
- Hu, C. Research on Key Technologies of Multi-Agent Collaboration Based on Deep Reinforcement Learning. Ph.D. Thesis, Beijing University of Posts and Telecommunications, Beijing, China, 2025. [Google Scholar]
- Huang, C.L. Research on Networking Methods for Integrated Air-Ground Networks Based on Deep Reinforcement Learning. Ph.D. Thesis, Beijing University of Posts and Telecommunications, Beijing, China, 2025. [Google Scholar]










| No. | Parameter | Description |
|---|---|---|
| 1 | Detection Type | Categorizes detection units |
| 2 | Operating Mode | Active, passive, frequency band, mode, etc. |
| 3 | Detection Envelope | Envelope shape of detectable area |
| 4 | Transmission Power | Used for calculating received signal power |
| 5 | Antenna Area | Receiver sensitivity |
| 6 | Detection Capability | For comprehensive assessment of platform detection performance |
| No. | Parameter | Description |
|---|---|---|
| 1 | Platform Type | Four types: land, sea, air, and space platforms |
| 2 | Operating Altitude | Altitude range within which the detection unit can operate effectively |
| 3 | Operational Area | Ground projection of the area where a given detection unit can operate, considering various constraints |
| 4 | Mobility Range | Maneuvering range of the detection unit |
| 5 | Mobility Speed | Maximum speed at which the detection unit can move; 0 indicates no mobility capability |
| 6 | Payload Type | Types and quantities of detectable equipment that can be carried; a single carrier can accommodate different types of detection units |
| No. | Parameter | Description |
|---|---|---|
| 1 | Target Type | Point, area, trajectory, and their combinations |
| 2 | Task Location | Effective description of points, areas, or trajectories |
| 3 | Detection Task Mode | Operating modes such as passive, active, scanning, frequency monitoring, etc., supporting concurrent description of multiple tasks |
| 4 | Detection Task Parameters | Specific parameters for the employed detection mode, such as frequency band, frequency point, period, etc. |
| 5 | Task Weight | |
| 6 | Task Start Time | |
| 7 | Task End Time |
| Scenarios | Algorithm | Target Coverage Rate | Planning Time Consumption (s) | Resource Utilization Rate |
|---|---|---|---|---|
| Small | Greedy Algorithm | 94.4% | 7.8 | 62.8% |
| Genetic Algorithm | 98.7% | 20.4 | 68.1% | |
| STAS-Net | 98.4% | 9.5 | 66.5% | |
| Medium | Greedy Algorithm | 86.1% | 17.3 | 66.3% |
| Genetic Algorithm | 88.3% | 53.0 | 71.4% | |
| STAS-Net | 92.8% | 14.2 | 82.5% | |
| Large | Greedy Algorithm | 78.2% | 39.3 | 72.5% |
| Genetic Algorithm | 81.7% | 286.5 | 85.3% | |
| STAS-Net | 89.7% | 18.3 | 86.8% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhu, G.; Wang, G.; Fu, W.; Han, C. Dynamic Task Planning for Heterogeneous Platforms via Spatio-Temporal and Capability Dual-Driven Framework. Electronics 2026, 15, 202. https://doi.org/10.3390/electronics15010202
Zhu G, Wang G, Fu W, Han C. Dynamic Task Planning for Heterogeneous Platforms via Spatio-Temporal and Capability Dual-Driven Framework. Electronics. 2026; 15(1):202. https://doi.org/10.3390/electronics15010202
Chicago/Turabian StyleZhu, Guangxi, Gang Wang, Wei Fu, and Changxing Han. 2026. "Dynamic Task Planning for Heterogeneous Platforms via Spatio-Temporal and Capability Dual-Driven Framework" Electronics 15, no. 1: 202. https://doi.org/10.3390/electronics15010202
APA StyleZhu, G., Wang, G., Fu, W., & Han, C. (2026). Dynamic Task Planning for Heterogeneous Platforms via Spatio-Temporal and Capability Dual-Driven Framework. Electronics, 15(1), 202. https://doi.org/10.3390/electronics15010202

