Adversarial Distributed Multi-Task Meta-Inverse Reinforcement Learning with Theory of Mind and Mean-Field Method
Abstract
1. Introduction
- (1)
- To enhance knowledge transfer between tasks, this paper proposes a meta-adversarial IRL framework based on theory of mind, which uses the theory of mind to make better use of the relationships between partially heterogeneous tasks and capture the representational information between tasks.
- (2)
- To reduce the computational complexity of reward and strategy optimization, the mean-field theory is introduced into TMMF-MTAIRL, which can transform interactions between complex tasks into the interaction between the main task and the rest of the average tasks.
- (3)
- By introducing additional latent variables to characterize modal information within mixed expert demonstrations, the TMMF-MTAIRL enables rapid adaptation to novel tasks during meta-adversarial training of the discriminator and generator, even with limited expert demonstration data.
- (4)
- The effectiveness of TMMF-MTAIRL is evaluated on both the maze-point benchmark and rolling bearing fault diagnosis tasks. Experimental results indicate that TMMF-MTAIRL outperforms conventional multi-task IRL methods, achieving state-of-the-art performance in both reward learning and policy optimization.
2. Related Works
3. Preliminaries
3.1. Markov Games
3.2. Adversarial Inverse Reinforcement Learning
4. Multi-Task Meta-Adversarial Mean-Field IRL Based on Theory of Mind
4.1. Adversarial Mean-Field IRL Based on Probabilistic Context Variable
4.2. The Embedding of Theory of Mind for Improved MMF-MTAIRL
4.3. TMMF-MTAIRL with Mutual Information Regularization
| Algorithm 1 Multi-task meta-adversarial mean-field IRL based on theory of mind |
|
5. Experiments and Simulations
5.1. Point-Maze Environments
5.2. Rolling Bearing Fault Diagnosis Experiments
6. Conclusions and Future Works
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Appendix A
Appendix B
Appendix C
| Notation | Meaning |
|---|---|
| State space | |
| Action space | |
| R | Reward space |
| Discount factor | |
| State transition distribution | |
| Expert demonstrations | |
| Feature function of each reward | |
| Partition function | |
| State transition probability | |
| M | The number of expert demonstration trajectories |
| Expert strategy | |
| Expert demonstration trajectories | |
| Discriminator | |
| Value space | |
| Reward functions of expert different tasks | |
| Corresponding reward functions for different tasks | |
| Parametrized reward function | |
| -conditional mean-field flow | |
| -conditional mind flow | |
| Policy flow | |
| Marginal distribution of expert policy | |
| A prior distribution | |
| Policies | |
| State of mind | |
| Tasks | |
| Average task | |
| Average state of mind | |
| The mind-based context-conditional trajectory distribution induced by a policy | |
| Context variable inference model | |
| Conditional distribution | |
| Posterior distribution | |
| Feature function of each reward | |
| K | The number of inner updates |
| T | Trajectory length |
Appendix D
| Contribution | Novelty/Key Idea |
|---|---|
| Embedding ToM for enhanced multi-task knowledge transfer | Introduces ToM-based task representations to explicitly model inter-task relationships and better leverage environmental/task interaction information, enabling effective knowledge transfer across tasks |
| Mean-field modeling of task interactions to reduce computational complexity | Approximates multi-task coupling by modeling interactions between the main task and the average of remaining tasks (mean-field approximation), significantly reducing complexity caused by pair-wise task interactions |
| Latent task variable inference for task-specific adaptation under the maximum entropy framework | Accurately infers latent task variables from sampled trajectories of novel tasks and derives task-specific adaptive reward functions and optimal policies under the maximum entropy IRL formulation |
| Continuously updatable adversarial generative learning framework | Constructs an iterative adversarial training procedure that continuously updates and optimizes the generator and discriminator, improving the quality and utility of mixed expert demonstrations |
References
- Zhang, X.Y.; Cai, X.Y.; Liu, B.; Huang, W.D.; Zhu, S.C.; Qi, S.Y.; Yang, Y.D. Differentiable Information Enhanced Model-Based Reinforcement Learning. In Proceedings of the 39th Annual AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA, 25 February–4 March 2025; pp. 1–9. [Google Scholar]
- Mohsin, M.A.; Rizwan, H.; Umer, M.; Bhattacharya, S.; Bilal, A.; Cioffi, J.M. Hierarchical Deep Reinforcement Learning for Adaptive Resource Management in Integrated Terrestrial and Non-Terrestrial Networks. In Proceedings of the 39th Annual AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA, 25 February–4 March 2025; pp. 1–8. [Google Scholar]
- Li, D.D.; Xie, L.; Wang, Z.; Yang, H. Brain Emotion Perception Inspired EEG Emotion Recognition with Deep Reinforcement Learning. IEEE Trans. Neural Netw. Learn. Syst. 2024, 35, 12979–12992. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lv, C.L.; Lv, X.L.; Wang, Z.Y.; Zhao, T.Q.; Tian, W.; Zhou, Q.Q.; Zeng, L.; Wan, M.; Liu, C.G. A focal quotient gradient system method for deep neural network training. Appl. Soft Comput. 2025, 184, 113704. [Google Scholar] [CrossRef] [Scilit]
- Wu, X.; Zhang, Y.T.; Lai, C.K.; Yang, M.Z.; Yang, G.L.; Wang, H.H. A Novel Centralized Federated Deep Fuzzy Neural Network with Multiobjectives Neural Architecture Search for Epistatic Detection. IEEE Trans. Fuzzy Syst. 2025, 33, 94–107. [Google Scholar] [CrossRef] [Scilit]
- Li, W.L.; Qiu, F.; Li, L.X.; Zhang, Y.N.; Wang, K. Simulation of Vehicle Interaction Behavior in Merging Scenarios: A Deep Maximum Entropy-Inverse Reinforcement Learning Method Combined with Game Theory. IEEE Trans. Intell. Veh. 2024, 9, 1079–1093. [Google Scholar] [CrossRef] [Scilit]
- Nan, J.F.; Deng, W.W.; Zhang, R.Z.; Zhao, R.; Wang, Y.; Ding, J. Car-Following Behavior Modeling with Maximum Entropy Deep Inverse Reinforcement Learning. IEEE Trans. Intell. Veh. 2024, 9, 3998–4010. [Google Scholar] [CrossRef] [Scilit]
- Li, J.H.; Wu, H.; He, Q.; Zhao, Y.J.; Wang, X. Dynamic QoS Prediction with Intelligent Route Estimation Via Inverse Reinforcement Learning. IEEE Trans. Serv. Comput. 2024, 17, 509–523. [Google Scholar] [CrossRef] [Scilit]
- Song, L.; Li, D.Z.; Wang, X.; Xu, X. AdaBoost Maximum Entropy Deep Inverse Reinforcement Learning with Truncated Gradient. Inform. Sci. 2022, 602, 328–350. [Google Scholar] [CrossRef] [Scilit]
- Wang, P.; Li, H.H.; Chan, C.Y. Meta-adversarial Inverse Reinforcement Learning for Decision-making Tasks. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation, Xi’an, China, 30 May–5 June 2021; pp. 12632–12638. [Google Scholar]
- Chen, J.B.; Yu, T.; Pan, Z.N.; Zhang, M.Y.; Lu, G.H.; Zhu, K.D. Stochastic Dynamic Power Dispatch with Human Knowledge Transfer Using Graph-GAN Assisted Inverse Reinforcement Learning. IEEE Trans. Smart Grid 2024, 15, 3303–3315. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.J.; Niu, Y.C.; Zhu, W.Y.; Chen, W.Q.; Li, Q.; Wang, T. Predicting Pedestrian Crossing Behavior at Unsignalized Mid-Block Crosswalks Using Maximum Entropy Deep Inverse Reinforcement Learning. IEEE Trans. Intell. Transp. Syst. 2024, 25, 3685–3698. [Google Scholar] [CrossRef] [Scilit]
- Bighashdel, A.; Meletis, P.; Jancura, P.; Dubbelman, G. Deep Adaptive Multi-intention Inverse Reinforcement Learning. In Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Bilbao, Spain, 13–17 September 2021; pp. 206–221. [Google Scholar]
- Krishnan, S.; Garg, A.; Liaw, R.; Miller, L.; Pokorny, F.T.; Goldberg, K. Hirl: Hierarchical Inverse Reinforcement Learning for Long-horizon Tasks with Delayed Rewards. arXiv 2016, arXiv:1604.06508. [Google Scholar] [CrossRef] [Scilit]
- Zhang, N.; Zhao, Y.J.; Yang, M.G.; Dai, S.L. Hierarchical Reinforcement Learning with Demonstration for Long-Horizon Robotic Manipulation. Arab. J. Sci. Eng. 2025. [Google Scholar] [CrossRef] [Scilit]
- Glazer, N.; Navon, A.; Shamsian, A.; Fetaya, E. Multi Task Inverse Reinforcement Learning for Common Sense Reward. In Proceedings of the Thirteenth International Conference on Learning Representations, Singapore, 24–28 April 2025; pp. 1–13. [Google Scholar]
- Tang, Q.H.; Guo, H.Y.; Chen, Q.X. Bidding Strategy Evolution Analysis Based on Multi-task Inverse Reinforcement Learning. Electr. Power Syst. Res. 2022, 212, 108286. [Google Scholar] [CrossRef] [Scilit]
- Yang, B.; Lu, Y.N.; Wan, R.; Hu, H.Y.; Yang, C.C.; Ni, R.R. Meta-IRLSOT++: A meta-Inverse Reinforcement Learning Method for Fast Adaptation of Trajectory Prediction Networks. Expert Syst. Appl. 2024, 240, 122499. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.Y.; Tamboli, D.; Lan, T.; Aggarwal, V. Multi-task Hierarchical Adversarial Inverse Reinforcement Learning. In Proceedings of the 40th International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July 2023; pp. 4895–4920. [Google Scholar]
- Chen, Y.; Lin, X.; Yan, B.; Zhang, L.B.; Liu, J.M.; Tan, N.O.; Witbrock, M. Meta-Inverse Reinforcement Learning for Mean Field Games via Probabilistic Context Variables. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20–27 February 2024; pp. 11407–11415. [Google Scholar]
- Ghasemipour, S.K.S.; Gu, S.X.; Zemel, R. SMILe: Scalable Meta Inverse Reinforcement Learning through Context-conditional Policies. In Proceedings of the 33rd Conference on Neural Information Processing Systems, Vancouver, BC, Canada, 8–14 December 2019; pp. 1–11. [Google Scholar]
- Yu, L.T.; Yu, T.H.; Finn, C.; Ermon, S. Meta-Inverse Reinforcement Learning with Probabilistic Context Variables. In Proceedings of the 33rd Conference on Neural Information Processing Systems, Vancouver, BC, Canada, 8–14 December 2019; pp. 11772–11783. [Google Scholar]
- Yoo, S.W.; Seo, S.W. Learning Multi-Task Transferable Rewards via Variational Inverse Reinforcement Learning. In Proceedings of the 2022 International Conference on Robotics and Automation, Pennsylvania, PA, USA, 23–27 May 2022; pp. 434–440. [Google Scholar]
- Gleave, A.; Habryka, O. Multi-task Maximum Entropy Inverse Reinforcement Learning. arXiv 2018, arXiv:1805.08882. [Google Scholar] [CrossRef] [Scilit]
- Baert, M.; Mazzaglia, P.; Leroux, S.; Simoens, P. Maximum Causal Entropy Inverse Constrained Reinforcement Learning. Mach. Learn. 2025, 114, 103. [Google Scholar] [CrossRef] [Scilit]
- Cheng, G.R.; Dong, L.; Cai, W.Z.; Sun, C.Y. Multi-Task Reinforcement Learning with Attention-Based Mixture of Experts. IEEE Robot Autom. Lett. 2023, 8, 3812–3819. [Google Scholar] [CrossRef] [Scilit]
- Vazquez-Chanlatte, M.; Seshia, S.A. Maximum Causal Entropy Specification Inference from Demonstrations. arXiv 2020, arXiv:1907.11792. [Google Scholar] [CrossRef] [Scilit]
- Nishi, K.; Shimosaka, M. Fine-Grained Driving Behavior Prediction via Context-Aware Multi-Task Inverse Reinforcement Learning. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation, Paris, France, 31 May–4 June 2020; pp. 2281–2287. [Google Scholar]
- Song, L.; Li, D.Z.; Xu, X. Adaptive Generative Adversarial Maximum Entropy Inverse Reinforcement Learning. Inform. Sci. 2025, 695, 121712. [Google Scholar] [CrossRef] [Scilit]
- Wu, K.Y.; Wu, F.G.; Lin, Y.J.; Zhao, J.S. Stable Control Policy and Transferable Reward Function via Inverse Reinforcement Learning. In Proceedings of the 2023 9th International Conference on Computing and Artificial Intelligence, Tianjin, China, 17–20 March 2023; pp. 733–742. [Google Scholar]
- Finn, C.; Levine, S.; Abbeel, P. Guided cost learning: Deep Inverse Optimal Control via Policy Optimization. In Proceedings of the 33rd International Conference on International Conference on Machine Learning, New York, NY, USA, 19–24 June 2016; pp. 49–58. [Google Scholar]
- Venuto, D.; Chakravorty, J.; Boussioux, L.; Wang, J.H.; McCracken, G.; Precup, D. Oirl: Robust Adversarial Inverse Reinforcement Learning with Temporally Extended Actions. arXiv 2002, arXiv:2002.09043. [Google Scholar]
- Wang, P.; Liu, D.P.; Chen, J.Y.; Li, H.H.; Chan, C.Y. Decision Making for Autonomous Driving via Augmented Adversarial Inverse Reinforcement Learning. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation, Xi’an, China, 30 May–5 June 2021; pp. 1036–1042. [Google Scholar]
- Sun, J.K.; Yu, L.T.; Dong, P.Q.; Lu, B.; Zhou, B.L. Adversarial Inverse Reinforcement Learning with Self-attention Dynamics Model. IEEE Rob. Autom. Lett. 2021, 6, 1880–1886. [Google Scholar] [CrossRef] [Scilit]
- Sestini, A.; Kuhnle, A.; Bagdanov, A.D. Demonstration Efficient Inverse Reinforcement Learning in Procedurally Generated Environments. arXiv 2020, arXiv:2012.02527. [Google Scholar] [CrossRef] [Scilit]
- Yang, Y.D.; Luo, R.; Li, M.; Zhou, M.; Zhang, W.N.; Wang, J. Mean Field Multi-Agent Reinforcement Learning. In Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018; pp. 5571–5580. [Google Scholar]
- Wang, X.Q.; Ke, L.J.; Zhang, G.W.; Zhu, D.P. Adaptive Mean Field Multi-agent Reinforcement Learning. Inform. Sci. 2024, 669, 120560. [Google Scholar] [CrossRef] [Scilit]
- Yu, C. Hierarchical Mean-Field Deep Reinforcement Learning for Large-Scale Multiagent Systems. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20–27 February, 2023; pp. 11744–11752. [Google Scholar]
- Hao, Q.Y.; Huang, W.Z.; Feng, T.; Yuan, J.; Li, Y. GAT-MF: Graph Attention Mean Field for Very Large Scale Multi-Agent Reinforcement Learning. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, New York, NY, USA, 6–10 August 2023; pp. 685–697. [Google Scholar]
- Chen, Y.; Zhang, L.B.; Liu, J.M.; Witbrock, M. Adversarial Inverse Reinforcement Learning for MeanField Games. In Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems, London, UK, 29 May–2 June 2023; pp. 1088–1096. [Google Scholar]
- Wu, H.C.; Sequeira, P.; Pynadath, D.V. Multiagent Inverse Reinforcement Learning via Theory of Mind Reasoning. In Proceedings of the 22nd International Conference on Autonomous Agents and Multiagent Systems, London, UK, 29 May–2 June 2023; pp. 708–716. [Google Scholar]
- Wei, R.; Zeng, S.; Li, C.L.; Garcia, A.; McDonald, A.; Hong, M.Y. A Bayesian Approach to Robust Inverse Reinforcement Learning. In Proceedings of the 7th Conference on Robot Learning, Atlanta, GA, USA, 6–9 November 2023; pp. 1–19. [Google Scholar]
- Oguntola, I.; Campbell, J.; Stepputtis, S.; Sycara, K. Theory of Mind as Intrinsic Motivation for Multi-Agent Reinforcement Learning. In Proceedings of the First Workshop on Theory of Mind in Communicating Agents, Honolulu, HI, USA, 10 July 2023; pp. 1–8. [Google Scholar]
- Wang, Y.F.; Zhong, F.W.; Xu, J.; Wang, Y.Z. Tom2c: Target-oriented Multi-agent Communication and Cooperation with Theory of Mind. In Proceedings of the International Conference on Learning Representations, Virtual Event, 25–29 April 2022; pp. 1–17. [Google Scholar]
- Shi, H.B.; Li, J.C.; Chen, S.C.; Hwang, K.S. A behavior Fusion Method based on Inverse Reinforcement Learning. Inform. Sci. 2022, 609, 429444. [Google Scholar] [CrossRef] [Scilit]
- Alsaleh, R.; Sayed, T. Markov-game Modeling of Cyclist-pedestrian Interactions in Shared Spaces: A Multi-agent Adversarial Inverse Reinforcement Learning Approach. Transp. Res. Part C Emerg. Technol. 2021, 128, 103191. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y.; Zhang, L.B.; Liu, J.M.; Hu, S.Y. Individual-Level Inverse Reinforcement Learning for MeanField Games. In Proceedings of the 21st International Conference on Autonomous Agents and Multiagent Systems, Virtual Event, 9–16 May 2022; pp. 253–262. [Google Scholar]
- Li, Y.Z.; Song, J.M.; Ermon, S. InfoGAIL: Interpretable Imitation Learning From Visual Demonstrations. In Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA, 3–9 December 2017; pp. 3815–3825. [Google Scholar]
- Fu, J.; Kumar, A.; Nachum, O.; Tucker, G.; Levine, S. D4RL: Datasets for Deep Data-Driven Reinforcement Learning. arXiv 2021, arXiv:2004.07219. [Google Scholar]
- Fan, S.; Zhang, X.M.; Song, Z.H. Imbalanced Sample Selection with Deep Reinforcement Learning for Fault Diagnosis. IEEE Trans. Ind. Inf. 2022, 18, 2518–2527. [Google Scholar] [CrossRef] [Scilit]
- Mismar, F.B.; Evans, B.L. Deep Q-Learning for Self-organizing Networks Fault Management and Radio Performance Improvement. In Proceedings of the 52nd Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, USA, 28–31 October 2018; pp. 1457–1461. [Google Scholar]
- Zhou, H.; Aral, A.; Brandić, I.; Erol-Kantarc, M. Multiagent Bayesian Deep Reinforcement Learning for Microgrid Energy Management under Communication Failures. IEEE Internet Things J. 2022, 9, 11685–11698. [Google Scholar] [CrossRef] [Scilit]









| Methods | Faced Problems | Advantages | References |
|---|---|---|---|
| Behavioral fusion IRL | Difficult to adapt to dynamic environments | Estimations based on trajectory distributions have small differences | [45] |
| Model-based IRL | [34] | ||
| Semantic-based augmented adversarial IRL | [33] | ||
| Hierarchical multi-task adversarial IRL | Difficult to adapt to multi-task scenarios | Capturing information among multiple tasks in real time | [19] |
| Meta-adversarial IRL | [20,21,22] | ||
| Variational multi-task IRL | [23] | ||
| Causal entropy-based multi-task IRL | [27] | ||
| Mean-field actor–critic algorithms | Complexity of large-scale interactive computing and modeling | Reducing computational complexity, improving generalization ability | [36] |
| GAT-MF | [39] | ||
| HMF-based DRL | [46] | ||
| ToM-based IRL | Modeling of multi-agent interactions is difficult; knowledge transfer for heterogeneous tasks is challenging | Enhanced cognitive depth and generalization ability of agents | |
| Bayesian MoT-based IRL | [41] | ||
| Intrinsically motivated ToM-based IRL | [42] | ||
| Goal-oriented IRL based on ToM | [43] |
| Control Parameters | Values |
|---|---|
| Number of iterations | 2000 |
| Batch size | 20,000 |
| Maximum path length | 500 |
| Discount factor | 0.99 |
| Step size | 0.01 |
| Control Parameters | Values |
|---|---|
| Number of iterations | 3000 |
| Batch size | 16 |
| Meta batch size | 50 |
| Maximum path length | 100 |
| Discount factor | 0.99 |
| Step size | 0.01 |
| Number of pre-training epochs | 1000 |
| Entropy weight | 1.0 |
| Information coefficient | 0.1 |
| Environments | Methods | TaskRmean | Mulinfo | MeanKL | dLoss |
|---|---|---|---|---|---|
| Point-maze | PEMIRL | ||||
| left-v0 | Meta-IRLSOT | ||||
| TMMF-MTAIRL | |||||
| Point-maze | PEMIRL | ||||
| right-v0 | Meta-IRLSOT | ||||
| TMMF-MTAIRL | |||||
| Point-maze | PEMIRL | ||||
| cont-v0 | Meta-IRLSOT | ||||
| TMMF-MTAIRL |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Song, L.; Yang, K.; Chen, C. Adversarial Distributed Multi-Task Meta-Inverse Reinforcement Learning with Theory of Mind and Mean-Field Method. Mathematics 2026, 14, 691. https://doi.org/10.3390/math14040691
Song L, Yang K, Chen C. Adversarial Distributed Multi-Task Meta-Inverse Reinforcement Learning with Theory of Mind and Mean-Field Method. Mathematics. 2026; 14(4):691. https://doi.org/10.3390/math14040691
Chicago/Turabian StyleSong, Li, Kun Yang, and Chao Chen. 2026. "Adversarial Distributed Multi-Task Meta-Inverse Reinforcement Learning with Theory of Mind and Mean-Field Method" Mathematics 14, no. 4: 691. https://doi.org/10.3390/math14040691
APA StyleSong, L., Yang, K., & Chen, C. (2026). Adversarial Distributed Multi-Task Meta-Inverse Reinforcement Learning with Theory of Mind and Mean-Field Method. Mathematics, 14(4), 691. https://doi.org/10.3390/math14040691

