CCTD-MARL: Coupled Communication-Task Decoupling Framework for Multi-Agent Systems Under Partial Observability
Abstract
1. Introduction
- 1.
- State Dependence: The state of subtask at time t, denoted as , is a function of the state of other subtasks and the joint action executed by all agents at time , i.e., , where is a task-specific transition function.
- 2.
- Reward Coupling: The cumulative reward of the system depends on the combined performance of all subtasks, with the immediate reward being a non-separable function of the individual subtask rewards , where g denotes the reward aggregation mechanism.
- We propose a dynamic state compensation framework that enables agents to reconstruct accurate global states from partial observations and incomplete communication information.
- We design a multi-agent hierarchical architecture that effectively decomposes complex, coupled tasks while maintaining inter-task dependencies and achieving efficient coordination.
- We provide a flexible integration mechanism that allows for flexible combination with existing MARL algorithms and specific task solutions.
2. Background
2.1. Multi-Agent Reinforcement Learning
2.2. Hierarchical Deep Reinforcement Learning
2.3. Multi-Agent Communication Architectures
3. Methodology
3.1. Hierarchical Reinforcement Learning
- Hyperweight Generation: A shared Hypernetwork takes each agent’s spatiotemporal feature (position, velocity, orientation) as input, generating customized weights . This adapts to individual agent states while maintaining homogeneity.
- Invariant Aggregation: The input layer computes latent embeddings via symmetric summation: . Symmetric summation ensures e is unchanged by agent input permutations, eliminating redundant computations.
- Equivariant Mapping: For entity-specific actions (e.g., assigning delivery targets to UAVs), submodular weights are remapped to the original agent order, Achieving Permutation Equivariance (APE). This preserves action-agent correspondence while retaining invariant aggregation benefits.
| Algorithm 1 Training Procedure of HLC Based on Pre-trained LLC (MDP Formulation) |
|
- Low-Level Controller
- High-level controller
3.2. State Compensation Framework
| Algorithm 2 Distributed State Alignment and Position Prediction |
Input: —Set of agents; t—Target timestamp (global Cartesian coordinate system); D—Historical trajectory buffer (stores for each agent, where = position, = timestamp, = velocity); —Delay threshold (unit: s); N—Length of historical sequence for RNN input; d—Dimension of agent state vector (e.g., for x/y/z coordinates, x/y/z velocity) Output: —Time-aligned global state matrix of size (global Cartesian coordinate system)
|
4. Experiments
4.1. Environments
- Reward Settings
4.2. Baseline
- MAPPO [28]: A MARL algorithm based on PPO (Proximal Policy Optimization), belonging to the policy gradient method. It is suited for complex policy and continuous action space scenarios. As a non-communication, non-hierarchical baseline, it serves to verify the necessity of communication optimization and task decomposition in coupled multi-task scenarios.
- QMix [20]: A value-decomposition MARL algorithm, designed for complex team tasks in discrete action spaces. It adopts a centralized training and decentralized execution paradigm but lacks explicit communication mechanisms and hierarchical task management, acting as a baseline for validating the effectiveness of our proposed communication compensation and hierarchical architecture.
- QMix-communicate MARL: An improved version of QMix with an explicit communication mechanism based on CommNet [27]. Each agent encodes its local observation (position, velocity, orientation, and partial environmental information) into a 32-dimensional feature vector, which is explicitly transmitted to all other agents through a fully connected communication network. The received messages are aggregated with the agent’s own observation via element-wise summation, enhancing the global state awareness of each agent. This baseline is used to compare the performance of different communication optimization strategies.
- QMix-Hierarchical MARL: A hierarchical extension of QMix, built on the Hierarchical Deep Q-Network (h-DQN) [23] architecture. It decomposes the complex coupled task into two independent subtasks (search and delivery) at the high level, with a meta-controller selecting subtasks and a low-level controller executing specific actions for each subtask. The meta-controller and low-level controllers are both trained using the QMix value-decomposition framework, but lack mechanisms to preserve subtask interdependencies and handle communication latency. This baseline is used to verify the advantages of our CCTD framework in maintaining task coupling and adapting to communication-constrained environments.
4.3. Experimental Results
4.4. Ablation Studies
- CCTD w/o com: Using CCTD without the communication compensation module
- CCTD w/o hir: Using CCTD without hierarchical architecture
5. Conclusions and Discussion
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Canese, L.; Cardarilli, G.C.; Di Nunzio, L.; Fazzolari, R.; Giardino, D.; Re, M.; Spanò, S. Multi-agent reinforcement learning: A review of challenges and applications. Appl. Sci. 2021, 11, 4948. [Google Scholar] [CrossRef] [Scilit]
- Zhang, K.; Yang, Z.; Başar, T. Multi-agent reinforcement learning: A selective overview of theories and algorithms. In Handbook of Reinforcement Learning and Control; Springer: Berlin/Heidelberg, Germany, 2021; pp. 321–384. [Google Scholar]
- Wen, M.; Kuba, J.; Lin, R.; Zhang, W.; Wen, Y.; Wang, J.; Yang, Y. Multi-agent reinforcement learning is a sequence modeling problem. Adv. Neural Inf. Process. Syst. 2022, 35, 16509–16521. [Google Scholar]
- Agrawal, A.; Won, S.J.; Sharma, T.; Deshpande, M.; McComb, C. A multi-agent reinforcement learning framework for intelligent manufacturing with autonomous mobile robots. Proc. Des. Soc. 2021, 1, 161–170. [Google Scholar] [CrossRef] [Scilit]
- Zheng, Z.; Gu, S. Safe multi-agent reinforcement learning with bilevel optimization in autonomous driving. IEEE Trans. Artif. Intell. 2024, 6, 829–842. [Google Scholar] [CrossRef] [Scilit]
- Du, W.; Ding, S. A survey on multi-agent deep reinforcement learning: From the perspective of challenges and applications. Artif. Intell. Rev. 2021, 54, 3215–3238. [Google Scholar] [CrossRef] [Scilit]
- Nekoei, H.; Badrinaaraayanan, A.; Sinha, A.; Amini, M.; Rajendran, J.; Mahajan, A.; Chandar, S. Dealing with non-stationarity in decentralized cooperative multi-agent deep reinforcement learning via multi-timescale learning. In Proceedings of the Conference on Lifelong Learning Agents, Montreal, QC, Canada, 22–25 August 2023; pp. 376–398. [Google Scholar]
- Amato, C. An introduction to centralized training for decentralized execution in cooperative multi-agent reinforcement learning. arXiv 2024, arXiv:2409.03052. [Google Scholar] [CrossRef] [Scilit]
- Phan, T.; Ritz, F.; Altmann, P.; Zorn, M.; Nüßlein, J.; Kölle, M.; Gabor, T.; Linnhoff-Popien, C. Attention-based recurrence for multi-agent reinforcement learning under stochastic partial observability. In Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July 2023; pp. 27840–27853. [Google Scholar]
- Liu, X.; Jin, B. Information-Theoretic Multi-Agent Algorithm Based on the CTDE Framework. In 2024 9th International Conference on Electronic Technology and Information Science (ICETIS); IEEE: Piscataway, NJ, USA, 2024; pp. 511–516. [Google Scholar]
- Hu, S.; Shen, L.; Zhang, Y.; Tao, D. Learning multi-agent communication from graph modeling perspective. arXiv 2024, arXiv:2405.08550. [Google Scholar] [CrossRef] [Scilit]
- Li, C.; Dong, S.; Yang, S.; Hu, Y.; Ding, T.; Li, W.; Gao, Y. Multi-task multi-agent reinforcement learning with interaction and task representations. IEEE Trans. Neural Netw. Learn. Syst. 2024, 36, 13431–13445. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhao, N.; Ye, Z.; Pei, Y.; Liang, Y.C.; Niyato, D. Multi-agent deep reinforcement learning for task offloading in UAV-assisted mobile edge computing. IEEE Trans. Wirel. Commun. 2022, 21, 6949–6960. [Google Scholar] [CrossRef] [Scilit]
- Zhu, X.; Xu, J.; Ge, J.; Wang, Y.; Xie, Z. Multi-task multi-agent reinforcement learning for real-time scheduling of a dual-resource flexible job shop with robots. Processes 2023, 11, 267. [Google Scholar] [CrossRef] [Scilit]
- Pateria, S.; Subagdja, B.; Tan, A.H.; Quek, C. Hierarchical reinforcement learning: A comprehensive survey. ACM Comput. Surv. (CSUR) 2021, 54, 1–35. [Google Scholar] [CrossRef] [Scilit]
- Hutsebaut-Buysse, M.; Mets, K.; Latré, S. Hierarchical reinforcement learning: A survey and open research challenges. Mach. Learn. Knowl. Extr. 2022, 4, 172–221. [Google Scholar] [CrossRef] [Scilit]
- Al-Emran, M. Hierarchical reinforcement learning: A survey. Int. J. Comput. Digit. Syst. 2015, 4, 137–143. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Williams, R.J. Reinforcement Learning and Markov Decision Processes; CSG220; Spring: Boston, MA, USA, 2007; Available online: https://ccs.neu.edu/home/rjw/com3480/lectures/reinforcement.pdf (accessed on 15 October 2025).
- Yao, Z.; Xia, S.; Li, Y.; Wu, G. Cooperative task offloading and service caching for digital twin edge networks: A graph attention multi-agent reinforcement learning approach. IEEE J. Sel. Areas Commun. 2023, 41, 3401–3413. [Google Scholar] [CrossRef] [Scilit]
- Rashid, T.; Samvelyan, M.; De Witt, C.S.; Farquhar, G.; Foerster, J.; Whiteson, S. Monotonic value function factorisation for deep multi-agent reinforcement learning. J. Mach. Learn. Res. 2020, 21, 1–51. [Google Scholar]
- Hao, X.; Wang, W.; Mao, H.; Yang, Y.; Li, D.; Zheng, Y.; Wang, Z.; Hao, J. API: Boosting multi-agent reinforcement learning via agent-permutation-invariant networks. arXiv 2022, arXiv:2203.05285. [Google Scholar]
- Zaheer, M.; Kottur, S.; Ravanbakhsh, S.; Poczos, B.; Salakhutdinov, R.R.; Smola, A.J. Deep sets. Adv. Neural Inf. Process. Syst. 2017, 30, 3394–3404. [Google Scholar]
- Skrynnik, A.; Staroverov, A.; Aitygulov, E.; Aksenov, K.; Davydov, V.; Panov, A.I. Hierarchical deep q-network from imperfect demonstrations in minecraft. Cogn. Syst. Res. 2021, 65, 74–78. [Google Scholar] [CrossRef] [Scilit]
- Andrychowicz, M.; Wolski, F.; Ray, A.; Schneider, J.; Fong, R.; Welinder, P.; McGrew, B.; Tobin, J.; Pieter Abbeel, O.; Zaremba, W. Hindsight experience replay. Adv. Neural Inf. Process. Syst. 2017, 30, 5055–5065. [Google Scholar]
- Röder, F.; Eppe, M.; Nguyen, P.D.; Wermter, S. Curious hierarchical actor-critic reinforcement learning. In International Conference on Artificial Neural Networks; Springer: Cham, Switzerland, 2020; pp. 408–419. [Google Scholar]
- Foerster, J.; Assael, I.A.; De Freitas, N.; Whiteson, S. Learning to communicate with deep multi-agent reinforcement learning. Adv. Neural Inf. Process. Syst. 2016, 29, 2145–2153. [Google Scholar]
- Sukhbaatar, S.; Fergus, R. Learning multiagent communication with backpropagation. Adv. Neural Inf. Process. Syst. 2016, 29, 2252–2260. [Google Scholar]
- Yu, C.; Velu, A.; Vinitsky, E.; Gao, J.; Wang, Y.; Bayen, A.; Wu, Y. The surprising effectiveness of ppo in cooperative multi-agent games. Adv. Neural Inf. Process. Syst. 2022, 35, 24611–24624. [Google Scholar]
- Wang, S. Real operational labeled data of air handling units from office, auditorium, and hospital buildings. Sci. Data 2025, 12, 1481. [Google Scholar] [CrossRef] [Scilit] [PubMed]






| Parameter | Value | Parameter | Value |
|---|---|---|---|
| Map Size | m | Discount Factor () | 0.99 |
| Number of UAVs | 12 | Batch Size | 32 |
| Communication Range | 200 m | RNN Hidden Units | 128 |
| Delay Dist. () | s | Learning Rate | |
| Reward (Search) | per area | Reward (Delivery) | |
| Step Penalty | Optimizer | Adam |
| Algorithm | Inference Time (ms) | Train Time (s/1000 Steps) |
|---|---|---|
| MAPPO | 4.5 | 12.4 |
| QMix | 3.2 | 8.5 |
| QMix-com | 4.6 | 13.8 |
| QMix-Hir | 5.8 | 15.2 |
| CCTD (Ours) | 7.2 | 19.8 |
| Method | 0–5 s Delay Error | 5–10 s Delay Error |
|---|---|---|
| RNN-based | ≈19 m | ≈21 m |
| Historical-location | ≈27 m | ≈32 m |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, K.; Wang, Z.; Tang, X.; You, H.; Hu, L.; Xie, H.; Chen, M. CCTD-MARL: Coupled Communication-Task Decoupling Framework for Multi-Agent Systems Under Partial Observability. Big Data Cogn. Comput. 2026, 10, 52. https://doi.org/10.3390/bdcc10020052
Li K, Wang Z, Tang X, You H, Hu L, Xie H, Chen M. CCTD-MARL: Coupled Communication-Task Decoupling Framework for Multi-Agent Systems Under Partial Observability. Big Data and Cognitive Computing. 2026; 10(2):52. https://doi.org/10.3390/bdcc10020052
Chicago/Turabian StyleLi, Kehan, Zhenya Wang, Xin Tang, Heng You, Long Hu, Haidong Xie, and Min Chen. 2026. "CCTD-MARL: Coupled Communication-Task Decoupling Framework for Multi-Agent Systems Under Partial Observability" Big Data and Cognitive Computing 10, no. 2: 52. https://doi.org/10.3390/bdcc10020052
APA StyleLi, K., Wang, Z., Tang, X., You, H., Hu, L., Xie, H., & Chen, M. (2026). CCTD-MARL: Coupled Communication-Task Decoupling Framework for Multi-Agent Systems Under Partial Observability. Big Data and Cognitive Computing, 10(2), 52. https://doi.org/10.3390/bdcc10020052
