DEDMAC: Disentangling Environment and Decision Messages for Multi-Agent Communication
Abstract
1. Introduction
- We propose DEDMAC, a communication framework that explicitly separates environment and decision information. It applies distinct processing pathways tailored to the specific characteristics of each, effectively addressing both partial observability and environmental non-stationarity in one unified framework.
- We design a differentiated processing mechanism that integrates environment messages into recurrent states for long-term modeling while utilizing decision messages as instantaneous biases for real-time policy alignment.
- We introduce a perception-first recalibration logic that utilizes aggregated global information to correct the temporal bias of asynchronous intents, ensuring consistency between signaled messages and final actions.
- We implement an independent encoding and explicit disentangling strategy to ensure semantic purity in communication. By employing specialized losses, DEDMAC disentangles the message streams. This mechanism prevents semantic ambiguity and enables precise feature extraction for robust multi-agent coordination.
2. Preliminary
2.1. Problem Formulation
2.2. Deep Q-Learning
3. Method
3.1. Overall Architecture of the DEDMAC Agent
3.2. Message Generator
3.3. Decision Generator
3.3.1. Environment Message Integration and Long-Term Memory Update
- Feature Aggregation and State Reconstruction: The system utilizes a self-attention mechanism [24] to capture the spatial dependencies among disparate environment messages. Subsequently, through flattening and Multi-Layer Perceptron (MLP) processing, the global state representation is reconstructed. This process is formalized in (4):
- Memory State Update: The reconstructed global state representation and the temporary hidden state (containing initial local cognition) are fed into a Gated Recurrent Unit (GRU). Through its internal gating mechanism, the GRU consolidates the shared environment information into the agent’s final hidden state , as defined in (5):This path achieves a “long-term update” of cognition, ensuring that the environmental context provided by peers can guide the agent’s subsequent decisions across multiple timesteps.
3.3.2. Decision Message Integration and Instantaneous Intent Bias
3.4. Loss Functions of DEDMAC
3.4.1. Temporal Difference Loss
3.4.2. Global State Reconstruction Loss
3.4.3. Intent Consistency Loss
3.4.4. Message Semantic Disentangling Loss
4. Experiments
4.1. Experimental Settings
4.1.1. Baselines
- QMIX [9]: A state-of-the-art non-communicative MARL algorithm that utilizes a monotonic value decomposition mechanism. It employs a mixing network conditioned on the global state to integrate individual Q-functions into a joint Q-function during training. This approach implements the Centralized Training with Decentralized Execution (CTDE) paradigm and effectively addresses the credit assignment problem while maintaining scalability.
- MASIA [15]: This method adopts a self-supervised learning pathway to extract abstract representations of the global state by aggregating observations from all agents and performing future predictions. Each agent subsequently extracts decision-relevant information from this global representation.
- TarMAC [17]: This algorithm introduces an attention mechanism to process communication messages. It allows agents to selectively extract information most critical to their own decision-making from the incoming message stream via a “signature-query” mechanism.
- MAIC [16]: This method facilitates explicit collaboration by generating incentive messages that directly influence the Q-functions of other agents. It utilizes target teammate models to customize message content and enhances communication efficiency through a sparse communication mechanism.
- T2MAC [28]: A recent communication framework where agents generate “evidence messages” and learn to selectively participate in communication. It improves information integration capabilities in complex environments through an evidence-driven fusion mechanism.
4.1.2. Experimental Benchmarks
- StarCraft Multi-Agent Challenge (SMAC) [31]: This is an authoritative collaborative MARL benchmark based on the game StarCraft II. We select two maps characterized by highly asymmetric information: 1o2r_vs_4r and 1o10b_vs_1r [29]. In these maps, the agent team consists of several combat units and an Overseer. The Overseer possesses superior reconnaissance vision but lacks combat capabilities, whereas combat units have restricted sight ranges. This setup creates an intrinsic asymmetric information structure, requiring the Overseer to extract and accurately communicate critical environment information to guide the combat units’ decision-making. We utilize StarCraft II version 4.6.2.6923.
- Hallway [29]: This environment is specifically designed to evaluate temporal coordination under extreme information asymmetry. In this task, multiple agents are randomly placed on Markov chains of varying lengths. Each agent’s perception is strictly localized, knowing only its own position within its hallway. Since agents cannot observe the states or progress of their teammates, effective communication is the sole means to complete collaborative tasks. The task requires all agents belonging to the same group to reach the goal state g simultaneously at the same timestep. Due to differing path lengths and random initial positions, agents must utilize communication to align their progress in real-time. In extended multi-group settings (referencing MASIA [15]), different subgroups are required to reach g at different time points; touching the goal simultaneously results in collision penalties. For instance, Hallway: 3x5-4x6x10 describes a complex task involving a two-agent group (paths of length 3 and 5) and a three-agent group (paths of length 4, 6, and 10). Here, agents must not only achieve internal synchrony but also negotiate the “passing order” across groups, providing an ideal scenario to validate DEDMAC’s semantic decoupling.
- Level-Based Foraging (LBF) [32]: LBF is a grid-world task that demands high levels of collaboration and local perception. Each agent and food item is randomly assigned a level. The success of a “load” action depends on the sum of the participating agents’ levels: a collection succeeds only if the sum of the levels of all agents adjacent to a food item who simultaneously perform the “load” action is greater than or equal to the food item’s level. This rule compels agents to flexibly seek partners based on level disparities. We test two difficulty configurations: for example, LBF: 11x11-6p-4f-s1 denotes an map with six agents (p) and four food items (f), where each agent’s sight range (s) is limited to one. In this configuration, task allocation—determining which agents should converge on which food item—rigorously tests the efficiency of communication algorithms in sharing both environment and intentional information.
- Traffic Junction (TJ) [33]: This benchmark simulates urban traffic flow where agents must navigate pre-defined intersecting routes while avoiding collisions. We evaluate DEDMAC on the Medium ( grid with 10 agents) and Hard ( grid with 20 agents) configurations. The action space is discrete, consisting of two operations: GAS (moving one step forward) and BRAKE (remaining in the current cell). A critical constraint is the extreme partial observability (vision = 0), meaning agents can only perceive their own identity and current cell. To encourage both efficiency and safety, the reward function issues a constant penalty for each timestep () and a severe penalty for each collision (). Furthermore, the environment implements a curriculum learning mechanism that gradually increases the agent arrival rate (add_rate) over training epochs, significantly raising traffic density and the frequency of potential conflicts. This setup rigorously tests the architectural efficacy of DEDMAC in filtering high-value decision intent and maintaining a stable environment state representation under increasingly congested and safety-critical conditions.
4.1.3. Implementation Details
- Input Encoder: Utilizes a single-layer linear transformation to project raw observations into a 64-dimensional latent space, followed by a ReLU activation function and a 64-unit GRU.
- Message Generator: Composed of a three-layer MLP (), where the environment message dimension is set to 6 and the decision message dimension is set to 4. This network generates the mean and variance for both environment and decision messages in parallel.
- Self-Attention Aggregation Module: Employs a single self-attention head with an embedding dimension of 16. A linear layer then maps the flattened output of the self-attention module to a vector with a dimension eight times the number of agents, followed by a 64-unit GRU.
- Differential Attention Module: Consists of two cross-attention heads to perform the subtraction operation, with an embedding dimension of 16.
- Policy Network: Implemented as an MLP that maps inputs to a 64-dimensional hidden representation, followed by a ReLU activation function before being projected into the action space.
- Auxiliary Evaluation Networks: The variational intent estimators used exclusively during training are designed as MLPs with a single hidden layer of 64 units and ReLU activation.
- QMIX [9]: As a non-communicative baseline, its communication dimension is 0.
- MASIA [15]: Following its self-supervised representation learning design, the message dimension is set to match the agent’s raw observation dimension to facilitate full feature transmission.
- TarMAC [17]: Referring to the original settings, this method employs a 16-dimensional signature/query pair coupled with a 32-dimensional transmission message, resulting in a fixed total communication overhead of 48 dimensions per timestep.
- DEDMAC (Ours): We restrict the message space to a minimal fixed dimension of 10 in total. Specifically, six dimensions are allocated for environment semantic encoding, while four dimensions are reserved for decision intent representation. We set to cover the entropy of the maximum action space (), while is chosen based on empirical results showing diminishing returns with higher dimensions.
- Mixing Network: We employ the standard QMIX monotonic constraint network with a mixing embedding dimension of 32.
- Hypernetworks: The hypernetworks responsible for generating the weights of the mixing network utilize a two-layer MLP architecture (including one hidden layer), with the internal hypernet embedding dimension standardized to 64.
4.2. Comparison with SOTA Methods
4.3. Ablation Studies and Structural Analysis
4.4. Visualization Analysis of Messages
4.5. Study on the Effectiveness of the Intent Filtering and Bias Correction Module
4.6. Communication Robustness Analysis
5. Discussion
6. Conclusions
- Disentanglement Is Essential: Explicitly severing the mutual information between and prevents “semantic leakage,” leading to more robust coordination in partially observable environments.
- Differentiated Processing Is Key: The effectiveness of DEDMAC stems from aligning processing pathways with the distinct natures of information. Stable, spatial environment messages are integrated into long-term memory to maintain a consistent world model, while transient, intent-driven decision messages are used for real-time strategic correction. This asymmetric design ensures that heterogeneous information is processed according to its specific functional role.
- Superior Efficiency: DEDMAC outperforms SOTA baselines in win rate and convergence speed across SMAC, Hallway, and LBF benchmarks, all while maintaining a minimal communication overhead of only 10 dimensions.
Author Contributions
Funding
Institutional Review Board Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Appendix A. Computational Complexity Analysis
| Environment | Scenario | MASIA (Baseline) | DEDMAC (Ours) | ||
|---|---|---|---|---|---|
| Params (K) | FLOPs (M) | Params (K) | FLOPs (M) | ||
| SMAC | 1o10b vs. 1r | 66.991 | 0.741 | 93.558 | 0.925 |
| 1o2r vs. 4r | 48.778 | 0.148 | 64.572 | 0.197 | |
| LBF | 20x20-10p-6f | 56.734 | 0.571 | 86.748 | 0.795 |
| 11x11-6p-4f | 48.094 | 0.291 | 70.908 | 0.419 | |
| Hallway | 4x6x10 | 39.491 | 0.120 | 60.142 | 0.183 |
| 3x5-4x6x10 | 41.259 | 0.208 | 65.502 | 0.327 | |
| TJ | Medium | 59.594 | 0.600 | 86.804 | 0.795 |
| Hard | 102.154 | 2.049 | 147.076 | 2.125 | |
- Equivalent Complexity: Both models share an time complexity for self-attention during decentralized execution; DEDMAC introduces no higher-order complexity.
- Excellent Scalability: In the most demanding TJ: Hard scenario (20 agents), the relative FLOPs increase is marginal (only 3.7%). The absolute FLOPs of DEDMAC are ≈2.125 M, ensuring millisecond-level real-time inference even in high-density systems.
- Justified Trade-off: The modest computational cost is a necessary trade-off for the substantial gains in coordination performance and model interpretability.

Appendix B. Hyperparameter Analysis
| Hyperparameter | Role | Value | Win Rate (%) |
|---|---|---|---|
| Global State Reconstruction Loss | 0.01 | 81.5 ± 5.2 | |
| 0.1 | 83.4 ± 4.3 | ||
| 0.5 | 87.4 ± 6.4 | ||
| 1.0 | 84.3 ± 6.7 | ||
| Intent Consistency Loss | 0.01 | 82.2 ± 5.9 | |
| 0.1 | 87.4 ± 6.4 | ||
| 0.5 | 83.7 ± 8.2 | ||
| 1.0 | 84.0 ± 7.0 | ||
| Message Semantic Disentangling Loss | 0.001 | 83.6 ± 6.7 | |
| 0.01 | 87.4 ± 6.4 | ||
| 0.1 | 81.9 ± 5.5 | ||
| 0.5 | 77.4 ± 7.6 |
References
- Hüttenrauch, M.; Šošić, A.; Neumann, G. Guided deep reinforcement learning for swarm systems. arXiv 2017, arXiv:1709.06011. [Google Scholar] [CrossRef] [Scilit]
- Cao, Y.; Yu, W.; Ren, W.; Chen, G. An overview of recent progress in the study of distributed multi-agent coordination. IEEE Trans. Ind. Inform. 2012, 9, 427–438. [Google Scholar] [CrossRef] [Scilit]
- Du, X.; Wang, J.; Chen, S.; Liu, Z. Multi-agent deep reinforcement learning with spatio-temporal feature fusion for traffic signal control. In Proceedings of the Machine Learning and Knowledge Discovery in Databases, Applied Data Science Track: European Conference, ECML PKDD 2021, Bilbao, Spain, 13–17 September 2021; Proceedings, Part IV 21; Springer: Berlin/Heidelberg, Germany, 2021; pp. 470–485. [Google Scholar]
- Xue, K.; Xu, J.; Yuan, L.; Li, M.; Qian, C.; Zhang, Z.; Yu, Y. Multi-agent dynamic algorithm configuration. In Advances in Neural Information Processing Systems; ACM Digital Library: New York, NY, USA, 2022; Volume 35, pp. 20147–20161. [Google Scholar]
- Osband, I.; Blundell, C.; Pritzel, A.; Van Roy, B. Deep exploration via bootstrapped DQN. In Advances in Neural Information Processing Systems; ACM Digital Library: New York, NY, USA, 2016; Volume 29, pp. 4026–4034. [Google Scholar]
- Mao, W.; Zhang, K.; Miehling, E.; Başar, T. Information state embedding in partially observable cooperative multi-agent reinforcement learning. In Proceedings of the 2020 59th IEEE Conference on Decision and Control (CDC); IEEE: New York, NY, USA, 2020; pp. 6124–6131. [Google Scholar]
- Papoudakis, G.; Christianos, F.; Rahman, A.; Albrecht, S.V. Dealing with non-stationarity in multi-agent deep reinforcement learning. arXiv 2019, arXiv:1906.04737. [Google Scholar]
- Lowe, R.; Wu, Y.I.; Tamar, A.; Harb, J.; Pieter Abbeel, O.; Mordatch, I. Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems; ACM Digital Library: New York, NY, USA, 2017; Volume 30. [Google Scholar]
- Rashid, T.; Samvelyan, M.; De Witt, C.S.; Farquhar, G.; Foerster, J.; Whiteson, S. Monotonic value function factorisation for deep multi-agent reinforcement learning. J. Mach. Learn. Res. 2020, 21, 7234–7284. [Google Scholar]
- Albrecht, S.V.; Stone, P. Autonomous agents modelling other agents: A comprehensive survey and open problems. Artif. Intell. 2018, 258, 66–95. [Google Scholar] [CrossRef] [Scilit]
- He, H.; Boyd-Graber, J.; Kwok, K.; Daumé, H., III. Opponent modeling in deep reinforcement learning. In Proceedings of the International Conference on Machine Learning, PMLR, New York, NY, USA, 20–22 June 2016; Association for Computing Machinery: New York, NY, USA, 2016; pp. 1804–1813. [Google Scholar]
- Zhang, K.; Yang, Z.; Başar, T. Multi-agent reinforcement learning: A selective overview of theories and algorithms. In Handbook of Reinforcement Learning and Control; Springer: Berlin/Heidelberg, Germany, 2021; pp. 321–384. [Google Scholar]
- Foerster, J.; Assael, I.A.; De Freitas, N.; Whiteson, S. Learning to communicate with deep multi-agent reinforcement learning. In Advances in Neural Information Processing Systems; ACM Digital Library: New York, NY, USA, 2016; Volume 29, pp. 2137–2145. [Google Scholar]
- Sukhbaatar, S.; Szlam, A.; Fergus, R. Learning multiagent communication with backpropagation. In Advances in Neural Information Processing Systems; ACM Digital Library: New York, NY, USA, 2016; Volume 29, pp. 2244–2252. [Google Scholar]
- Guan, C.; Chen, F.; Yuan, L.; Wang, C.; Yin, H.; Zhang, Z.; Yu, Y. Efficient multi-agent communication via self-supervised information aggregation. In Advances in Neural Information Processing Systems; ACM Digital Library: New York, NY, USA, 2022; Volume 35, pp. 1020–1033. [Google Scholar]
- Yuan, L.; Wang, J.; Zhang, F.; Wang, C.; Zhang, Z.; Yu, Y.; Zhang, C. Multi-agent incentive communication via decentralized teammate modeling. In Proceedings of the AAAI Conference on Artificial Intelligence; The AAAI Press: Palo Alto, CA, USA, 2022; Volume 36, pp. 9466–9474. [Google Scholar]
- Das, A.; Gervet, T.; Romoff, J.; Batra, D.; Parikh, D.; Rabbat, M.; Pineau, J. Tarmac: Targeted multi-agent communication. In Proceedings of the International Conference on Machine Learning, PMLR, Long Beach, CA, USA, 9–15 June 2019; Association for Computing Machinery: New York, NY, USA, 2019; pp. 1538–1546. [Google Scholar]
- Chung, J.; Gulcehre, C.; Cho, K.; Bengio, Y. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv 2014, arXiv:1412.3555. [Google Scholar] [CrossRef] [Scilit]
- Cheng, P.; Hao, W.; Dai, S.; Liu, J.; Gan, Z.; Carin, L. Club: A contrastive log-ratio upper bound of mutual information. In Proceedings of the International Conference on Machine Learning, PMLR, Virtual, 13–18 July 2020; Association for Computing Machinery: New York, NY, USA, 2020; pp. 1779–1788. [Google Scholar]
- Oliehoek, F.A.; Amato, C. A Concise Introduction to Decentralized POMDPs; Springer: Berlin/Heidelberg, Germany, 2016; Volume 1. [Google Scholar]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sunehag, P.; Lever, G.; Gruslys, A.; Czarnecki, W.M.; Zambaldi, V.; Jaderberg, M.; Lanctot, M.; Sonnerat, N.; Leibo, J.Z.; Tuyls, K.; et al. Value-decomposition networks for cooperative multi-agent learning. arXiv 2017, arXiv:1706.05296. [Google Scholar]
- Wang, J.; Ren, Z.; Liu, T.; Yu, Y.; Zhang, C. Qplex: Duplex dueling multi-agent q-learning. arXiv 2020, arXiv:2008.01062. [Google Scholar]
- Bahdanau, D.; Cho, K.; Bengio, Y. Neural machine translation by jointly learning to align and translate. arXiv 2014, arXiv:1409.0473. [Google Scholar]
- Ye, T.; Dong, L.; Xia, Y.; Sun, Y.; Zhu, Y.; Huang, G.; Wei, F. Differential transformer. arXiv 2024, arXiv:2410.05258. [Google Scholar]
- Wainwright, M.J.; Jordan, M.I. Graphical models, exponential families, and variational inference. In Foundations and Trends® in Machine Learning; Now Publishers Inc.: Hanover, MA, USA, 2008; Volume 1, pp. 1–305. [Google Scholar]
- Alemi, A.A.; Fischer, I.; Dillon, J.V.; Murphy, K. Deep variational information bottleneck. arXiv 2016, arXiv:1612.00410. [Google Scholar]
- Sun, C.; Zang, Z.; Li, J.; Li, J.; Xu, X.; Wang, R.; Zheng, C. T2mac: Targeted and trusted multi-agent communication through selective engagement and evidence-driven integration. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20–27 February 2024; AAAI Press: Washington, DC, USA, 2024; Volume 38, pp. 15154–15163. [Google Scholar]
- Wang, T.; Wang, J.; Zheng, C.; Zhang, C. Learning nearly decomposable value functions via communication minimization. arXiv 2019, arXiv:1910.05366. [Google Scholar]
- Hu, J.; Jiang, S.; Harding, S.A.; Wu, H.; Liao, S.-w. Rethinking the implementation tricks and monotonicity constraint in cooperative multi-agent reinforcement learning. arXiv 2021, arXiv:2102.03479. [Google Scholar]
- Samvelyan, M.; Rashid, T.; de Witt, C.S.; Farquhar, G.; Nardelli, N.; Rudner, T.G.J.; Hung, C.M.; Torr, P.H.S.; Foerster, J.; Whiteson, S. The StarCraft Multi-Agent Challenge. arXiv 2019, arXiv:1902.04043. [Google Scholar]
- Papoudakis, G.; Christianos, F.; Schäfer, L.; Albrecht, S.V. Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks. arXiv 2020, arXiv:2006.07869. [Google Scholar]
- Singh, A.; Jain, T.; Sukhbaatar, S. Learning when to Communicate at Scale in Multiagent Cooperative and Competitive Tasks. arXiv 2018, arXiv:1812.09755. [Google Scholar] [CrossRef] [Scilit]
- Maaten, L.v.d.; Hinton, G. Visualizing data using t-SNE. J. Mach. Learn. Res. 2008, 9, 2579–2605. [Google Scholar]





| Scenario | QMIX | MASIA | TarMAC | MAIC | T2MAC | DEDMAC (Ours) |
|---|---|---|---|---|---|---|
| SMAC | ||||||
| 1o10b_vs_1r | 0 | 84 | 48 | 7 | 7 | 10 (6 + 4) |
| 1o2r_vs_4r | 0 | 49 | 48 | 10 | 10 | 10 (6 + 4) |
| Hallway | ||||||
| 4x6x10 | 0 | 1 | 48 | 3 | 3 | 10 (6 + 4) |
| 3x5-4x6x10 | 0 | 2 | 48 | 3 | 3 | 10 (6 + 4) |
| LBF | ||||||
| 11x11-6p-4f-s1 | 0 | 30 | 48 | 6 | 6 | 10 (6 + 4) |
| 20x20-10p-6f-s1 | 0 | 48 | 48 | 6 | 6 | 10 (6 + 4) |
| TJ | ||||||
| Medium | 0 | 61 | 48 | 2 | 2 | 10 (6 + 4) |
| Hard | 0 | 149 | 48 | 2 | 2 | 10 (6 + 4) |
| Category | Parameter | Symbol | Value |
|---|---|---|---|
| General | Batch size | - | 32 episodes |
| Buffer size | - | 5000 episodes | |
| Target update interval | - | 200 episodes | |
| Optimizer | Learning rate | ||
| Grad norm clip | - | 10.0 | |
| Discount factor | 0.99 | ||
| Loss Weights | State reconstruction | 0.5 | |
| Intent consistency | 0.1 | ||
| Semantic disentanglement | 0.01 |
| Variant | Final Win Rate (%) | Performance Drop () |
|---|---|---|
| DEDMAC (Full) | 100.0 | – |
| Optimization Objectives (Losses) | ||
| w/o l_a (Intent Consistency) | 39.9 | −60.1 |
| w/o l_s (Global Reconstruction) | 62.6 | −37.4 |
| w/o l_msg (Semantic Disentangling) | 91.9 | −8.1 |
| Structural Components | ||
| w/o Env-Msg (Environment Messages) | 77.4 | −22.6 |
| w/o Dec-Msg (Decision Messages) | 93.8 | −6.2 |
| w/o Env-Proc (Long-term Memory Update) | 96.2 | −3.8 |
| w/o Dec-Proc (Policy Bias) | 82.4 | −17.6 |
| w/o Diff-Attn (Differential Attention) | 71.7 | −28.3 |
| Algorithm | Packet Loss Rate | |||
|---|---|---|---|---|
| 0% (Baseline) | 10% | 20% | 30% | |
| TarMAC | 82.0 | 65.7 (−16.3) | 54.0 (−28.0) | 42.8 (−39.2) |
| T2MAC | 10.6 | 16.2 (+5.6) | 8.1 (−2.5) | 7.6 (−3.0) |
| MAIC | 55.6 | 54.1 (−1.5) | 52 (−3.6) | 46.7 (−8.9) |
| MASIA | 78.2 | 71.8 (−6.4) | 69.4 (−8.8) | 60.1 (−18.1) |
| DEDMAC (Ours) | 86.2 | 81.9 (−4.3) | 78.7 (−7.5) | 71.5 (−14.7) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Liang, Y.; Li, J. DEDMAC: Disentangling Environment and Decision Messages for Multi-Agent Communication. Information 2026, 17, 332. https://doi.org/10.3390/info17040332
Liang Y, Li J. DEDMAC: Disentangling Environment and Decision Messages for Multi-Agent Communication. Information. 2026; 17(4):332. https://doi.org/10.3390/info17040332
Chicago/Turabian StyleLiang, Yihan, and Jinlong Li. 2026. "DEDMAC: Disentangling Environment and Decision Messages for Multi-Agent Communication" Information 17, no. 4: 332. https://doi.org/10.3390/info17040332
APA StyleLiang, Y., & Li, J. (2026). DEDMAC: Disentangling Environment and Decision Messages for Multi-Agent Communication. Information, 17(4), 332. https://doi.org/10.3390/info17040332
