Enhancing Multi-Agent Reinforcement Learning via Knowledge-Embedded Modular Framework for Online Basketball Games
Abstract
1. Introduction
- (1)
- Situation-Specific Modular Reinforcement Learning: By partitioning various game states in team sports into individual sub-models for training, we can reduce the complexity of high-dimensional state-action space and facilitate the learning of optimal policies for each situation. This approach mitigates the impact of high sample complexity and enables specialized decision making in different contexts.
- (2)
- Knowledge-Embedded Observation Layers: Through multistage feature engineering of the raw game data, we can derive a high-dimensional semantic context—such as shot success probability—and incorporate it into the agent’s modularized observation space. This domain-specific state representation can help to counter low sample efficiency by allowing the agent to learn more effectively with limited data.
- (3)
- Dynamic and Dense Reward Scheme: To complement the sparse rewards offered by simple metrics (such as scoring or winning), we can assign more finely grained rewards based on the tactical indicators extracted in (2). This enables the agent to receive immediate feedback from the outset of training, thereby improving the sample efficiency and aligning more closely with knowledge-based data for stable and rapid policy formation.
2. Background
2.1. Sample Complexity in Reinforcement Learning and Sports Game Environment
2.2. Multi-Agent Proximal Policy Optimization (MAPPO)

2.3. Modular Reinforcement Learning
2.4. Knowledge-Based Feature Engineering
3. Methods
3.1. Modular Architecture for Context-Specific Policies
3.2. Knowledge-Embedded Observation Layers
Feature Estimation for Knowledge-Embedded Observation
3.3. Dynamic Dense Reward Design
4. Experimental Results
4.1. Experimental Environment
4.2. Setting
| Algorithm 1. Knowledge-Embedded Modular MAPPO Training |
| Require: (1) Env. interface (states, actions, rewards); (2) MAPPO models for each module (Offense, Defense, Loose-ball); (3) γ, λ, ε; (4) Teams: {0, 1} each with N agents. |
|
1: Initialize: Set seeds; Create buffers 0, 1 Load Knowledge-embedded modular MAPPO params; if available; Reset trackers; 2: for episode e = 1, … do 3: Reset env; Clear 0, 1 4: if (e mod 200) < 100 then 5: Update team 0, freeze team 1 6: else 7: Update team 1, freeze team 0 8: end if 9: while not done do 10: Get state s from 11: for team ∈ {0, 1} do 12: if team is frozen then 13: // Only forward pass, no param update 14: end if 15: Determine module M_team (Offense/Defense/Loose-ball) 16: Build local obs oi for each agent; masks mi 17: Global state g_team 18: (ai, log π(ai|oi)) ← MAPPO_{M_team}(oi) 19: Exec ai in env; Store transition (oi, ai, ri, …) in _team 20: end for 21: Env step; Compute rewards rt; done ← (time = 0) 22: end while 23: for team ∈ {0, 1} do 24: if team not frozen then 25: Group transitions in _team by trajectory 26: for each traj. τ = {(st, ot, rt)}T−1_{t=0} do 27: for t = T − 1 down to 0 do 28: δt = rt + γV_φ(st+1) − V_φ(st) 29: if t = T − 1 then 30: At = δt 31: else 32: At = δt + γλ At+1 33: end if 34: end for 35: for t = 0 to T − 1 do 36: Rt = At + V_φ(st) 37: end for 38: end for 39: PPO Update: 40: for minibatch in _team do 41: for sample t in minibatch do 42: r(θ) ← π_θ(at|st) / π_{θ_old}(at|st) 43: L^{CLIP}_t = min(r(θ) At, clip(r(θ), 1−ε, 1+ε) At) 44: L^{VF}_t = (Rt − V_φ(st))2 45: end for 46: L^{CLIP} = mean(L^{CLIP}_t); L^{VF} = mean(L^{VF}_t) 47: L(θ, φ) = −L^{CLIP} + c1 L^{VF} + c2 Entropy(π_θ) 48: θ, φ ← θ, φ − η∇L(θ, φ) 49: end for 50: end if 51: end for 52: Save model if training 53: end for |
4.3. Comparison Experiment with Competing Methods
4.4. Ablation Study
- (a)
- Baseline (No Proposed Methods): This setting mirrored the training regime of our previous model and did not incorporate any newly proposed techniques. The agent received raw game state observation data and relied solely on sparse rewards using a single-model approach.
- (b)
- Knowledge-Embedded Observation and Dynamic Dense Reward (No Modular): For this variant, we embedded the domain knowledge into the agent’s observation data, as described in Section 3.2, and applied our dense reward strategy. However, the modular policy component was omitted; therefore, the agent employed a single unified policy across all contexts.
- (c)
- Modular Policy with Raw Game State Observation and Sparse Reward: This setting adopted a modular RL design while retaining the original raw game state observation and sparse reward used in the baseline approach. This allowed us to isolate the impact of modular decomposition on the learning performance without the additional influence of knowledge-embedded features or denser reward signals.
- (d)
- Full KEMF (All Proposed Methods): Finally, we evaluated the complete KEMF where the knowledge-embedded observation, dynamic dense reward, and modular architectures were employed. This served as a reference point for comparison with the other three variations.
4.5. Experiment on Real-World Live Services
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| DRL | Deep reinforcement learning |
| FSM | Finite state machine |
| GAE | Generalized advantage estimation |
| IPPO | Independent proximal policy optimization |
| KEMF | Knowledge-embedded modular framework |
| LLM | Large language model |
| MAPPO | Multi-agent proximal policy optimization |
| MARL | Multi-agent reinforcement learning |
| MDP | Markov decision process |
| PPO | Proximal policy optimization |
| RL | Reinforcement learning |
| RNN | Recurrent neural network |
References
- Vinyals, O.; Babuschkin, I.; Czarnecki, W.M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D.H.; Powell, R.; Ewalds, T.; Georgiev, P.; et al. Grandmaster Level in StarCraft II Using Multi-Agent Reinforcement Learning. Nature 2019, 575, 350–354. [Google Scholar] [CrossRef] [Scilit]
- Hu, Y.-J.; Lin, S.-J. Deep Reinforcement Learning for Optimizing Finance Portfolio Management. In Proceedings of the Amity International Conference, Artificial Intelligence (AICAI), Dubai, United Arab Emirates, 4–6 February 2019; pp. 14–20. [Google Scholar]
- Kiran, B.R.; Sobh, I.; Talpaert, V.; Mannion, P.; Sallab, A.A.A.; Yogamani, S.; Perez, P. Deep Reinforcement Learning for Autonomous Driving: A Survey. IEEE Trans. Intell. Transp. Syst. 2022, 23, 4909–4926. [Google Scholar] [CrossRef] [Scilit]
- Han, D.; Mulyana, B.; Stankovic, V.; Cheng, S. A Survey on Deep Reinforcement Learning Algorithms for Robotic Manipulation. Sensors 2023, 23, 3762. [Google Scholar] [CrossRef] [Scilit]
- Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction, 2nd ed.; MIT Press: Cambridge, MA, USA, 2018. [Google Scholar]
- Souchleris, K.; Sidiropoulos, G.K.; Papakostas, G.A. Reinforcement Learning in Game Industry Review, Prospects and Challenges. Appl. Sci. 2023, 13, 2443. [Google Scholar] [CrossRef] [Scilit]
- Berner, C.; Brockman, G.; Chan, B.; Cheung, V.; Dębiak, P.; Dennison, C.; Farhi, D.; Fischer, Q.; Hashme, S.; Hesse, C.; et al. Dota 2 with Large-Scale Deep Reinforcement Learning. arXiv 2019, arXiv:1912.06680. [Google Scholar]
- Xenou, K.; Chalkiadakis, G.; Afantenos, S. Deep Reinforcement Learning in Strategic Board Game Environments. In Proceedings of the European Conference on Multi-Agent Systems, Bergen, Norway, 6–7 December 2018; pp. 233–248. [Google Scholar]
- Piergigli, D.; Ripamonti, L.A.; Maggiorini, D.; Gadia, D. Deep Reinforcement Learning to Train Agents in a Multiplayer First-Person Shooter: Some Preliminary Results. In Proceedings of the IEEE Conference, Games (CoG), London, UK, 20–23 August 2019; pp. 1–8. [Google Scholar]
- Schut, L.; Tomašev, N.; McGrath, T.; Hassabis, D.; Paquet, U.; Kim, B. Bridging the Human–AI Knowledge Gap Through Concept Discovery and Transfer in AlphaZero. Proc. Natl. Acad. Sci. USA 2025, 122, e2406675122. [Google Scholar] [CrossRef] [Scilit]
- Kang, J.; Yoon, J.S.; Lee, B. How AI-Based Training Affected the Performance of Professional Go Players. In Proceedings of the CHI Conference on Human Factors in Computing Systems, New Orleans, LA, USA, 29 April–5 May 2022; ACM: New York, NY, USA, 2022; pp. 1–12. [Google Scholar] [CrossRef] [Scilit]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-Level Control through Deep Reinforcement Learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit]
- Zhang, K.; Yang, Z.; Başar, T. Multi-agent reinforcement learning: A selective overview of theories and algorithms. In Handbook of Reinforcement Learning and Control; Springer: Cham, Switzerland, 2021; pp. 321–384. [Google Scholar]
- Yu, C.; Velu, A.; Vinitsky, E.; Gao, J.; Wang, Y.; Bayen, A.; Wu, Y. The Surprising Effectiveness of PPO in Cooperative Multi-agent Games. Adv. Neural Inf. Process. Syst. 2022, 35, 24611–24624. [Google Scholar]
- Song, Y.; Jiang, H.; Tian, Z.; Zhang, H.; Zhang, Y.; Zhu, J.; Dai, Z.; Zhang, W.; Wang, J. An empirical study on Google research football multi-agent scenarios. Mach. Intell. Res. 2024, 21, 549–570. [Google Scholar] [CrossRef] [Scilit]
- Business Research. Insights. “Sport Games Market Size, Share, Growth, and Industry Analysis by Type” (Client Type, Web Game Type) by Application (PC, Mobile, Tablet & Others) Regional Insights And Forecast From 2026 To 2035. Available online: https://www.businessresearchinsights.com/market-reports/sport-games-market-103860 (accessed on 30 August 2025).
- Zhao, Y.; Borovikov, I.; Rupert, J.; Somers, C.; Beirami, A. On Multi-agent Learning in Team Sports Games. arXiv 2019, arXiv:1906.10124. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Janner, M.; Li, Q.; Levine, S. Offline Reinforcement Learning as One Big Sequence Modeling Problem. Adv. Neural Inf. Process. Syst. 2022, 34, 1273–1286. [Google Scholar]
- Jeon, M.; Lee, J.; Ko, S.-K. Modular Reinforcement Learning for Playing the Game of Tron. IEEE Access 2022, 10, 63394–63402. [Google Scholar] [CrossRef] [Scilit]
- Lee, D.; Tang, H.; Zhang, J.; Xu, H.; Darrell, T.; Abbeel, P. Modular Architecture for StarCraft II with Deep Reinforcement Learning. In Proceedings of the AAAI Conference on Artificial Intelligence & Interactive Digital Entertainment, Edmonton, AB, Canada, 13–17 November 2018; Volume 14, pp. 187–193. [Google Scholar] [CrossRef] [Scilit]
- Kovalchik, S.A. Player tracking data in sports. Annu. Rev. De Stat. Appliquée 2023, 10, 677–697. [Google Scholar] [CrossRef] [Scilit]
- Cervone, D.; D’Amour, A.; Bornn, L.; Goldsberry, K. A Multiresolution Stochastic Process Model for Predicting Basketball Possession Outcomes. J. Am. Stat. Assoc. 2016, 111, 585–599. [Google Scholar] [CrossRef] [Scilit]
- Franks, A.; Miller, A.; Bornn, L.; Goldsberry, K. Characterizing the Spatial Structure of Defensive Skill in Professional Basketball. Ann. Appl. Stat. 2015, 9, 94–121. [Google Scholar] [CrossRef] [Scilit]
- Mannion, P.; Devlin, S.; Duggan, J.; Howley, E. Reward Shaping for Knowledge-Based Multi-objective Multi-agent Reinforcement Learning. Knowl. Eng. Rev. 2018, 33, e23. [Google Scholar] [CrossRef] [Scilit]
- Ng, A.Y.; Harada, D.; Russell, S.J. Policy Invariance under Reward Transformations: Theory and Application to Reward Shaping. In Proceedings of the International Conference, Machine Learning (ICML), Bled, Slovenia, 27–30 June 1999; pp. 278–287. [Google Scholar]
- Humphrys, M. Action Selection Methods Using Reinforcement Learning. Anim. Animat. 1996, 4, 135–144. [Google Scholar] [CrossRef] [Scilit]
- Abe, T.; Orihara, R.; Sei, Y.; Tahara, Y.; Ohsuga, A. Step-by-Step Acquisition of Cooperative Behavior in Soccer Task. J. Adv. Inf. Technol. 2022, 13, 147–154. [Google Scholar] [CrossRef] [Scilit]
- Liu, R.-Z.; Pang, Z.-J.; Meng, Z.-Y.; Wang, W.; Yu, Y.; Lu, T. On Efficient Reinforcement Learning for Full-Length Game of StarCraft II. J. Artif. Intell. Res. 2022, 75, 213–260. [Google Scholar] [CrossRef] [Scilit]
- Kakade, S.M. On the Sample Complexity of Reinforcement Learning; University College London: London, UK, 2003. [Google Scholar]
- Azar, M.G.; Munos, R.; Kappen, B. On the Sample Complexity of Reinforcement Learning with a Generative Model. arXiv 2012, arXiv:1206.6461. [Google Scholar] [CrossRef] [Scilit]
- Zhang, K.; Kakade, S.M.; Basar, T.; Yang, L.F. Model-Based Multi-agent RL in Zero-Sum Markov Games with Near-Optimal Sample Complexity. J. Mach. Learn. Res. 2023, 24, 1–53. [Google Scholar]
- Zhang, G.; Jia, Y.; Yang, L.F.; Jiang, N. Settling the Sample Complexity of Online Reinforcement Learning. In Proceedings of the Thirty Seventh Annual Conference on Learning Theory, Edmonton, AB, Canada, 30 June–3 July 2024; pp. 2447–2448. [Google Scholar]
- Shi, L.; Chi, J. Distributionally Robust Model-Based Offline Reinforcement Learning with Near-Optimal Sample Complexity. J. Mach. Learn. Res. 2024, 25, 1–52. [Google Scholar]
- Azar, M.G.; Munos, R.; Kappen, B. Minimax PAC Bounds on the Sample Complexity of Reinforcement Learning with a Generative Model. Mach. Learn. 2013, 91, 325–349. [Google Scholar] [CrossRef] [Scilit]
- Kurach, K.; Raichuk, A.; Stańczyk, P.; Zając, M.; Bachem, O.; Espeholt, L.; Riquelme, C.; Vincent, D.; Michalski, M.; Bousquet, O.; et al. Google Research Football: A Novel Reinforcement Learning Environment. In Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA, 7–12 February 2020; Volume 34, pp. 4501–4510. [Google Scholar] [CrossRef] [Scilit]
- Zhao, J.; Lin, J.; Zhang, X.; Li, Y.; Zhou, X.; Sun, Y. From Mimic to Counteract: A Two-Stage Reinforcement Learning Algorithm for Google Research Football. Neural Comput. Appl. 2024, 36, 7203–7219. [Google Scholar] [CrossRef] [Scilit]
- Scott, A.; Fujii, K.; Onishi, M. How Does AI Play Football? An Analysis of RL and Real-World Football Strategies. arXiv 2021, arXiv:2111.12340. [Google Scholar] [CrossRef] [Scilit]
- Chen, X.; Jiang, J.Y.; Jin, K.; Zhou, Y.; Liu, M.; Brantingham, P.J.; Wang, W. Reliable: Offline reinforcement learning for tactical strategies in professional basketball games. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, Atlanta, GA, USA, 17–21 October 2022; pp. 3023–3032. [Google Scholar]
- Jia, H.; Ren, C.; Hu, Y.; Chen, Y.; Lv, T.; Fan, C.; Tang, H.; Hao, J. Mastering basketball with deep reinforcement learning: An integrated curriculum training approach. In Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems, Auckland, New Zealand, 9–13 May 2020; pp. 1872–1874. [Google Scholar]
- Choi, T.; Cho, K.; Sung, Y. Approaches That Use Domain-Specific Expertise: Behavioral-Cloning-Based Advantage Actor-Critic in Basketball Games. Mathematics 2023, 11, 1110. [Google Scholar] [CrossRef] [Scilit]
- Lee, S.; Lee, G.; Kim, W.; Kim, J.; Park, J.; Cho, K. Human Strategy Learning-Based Multi-agent Deep Reinforcement Learning for Online Team Sports Game. IEEE Access 2025, 13, 15437–15452. [Google Scholar] [CrossRef] [Scilit]
- Jia, H.; Hu, Y.; Chen, Y.; Ren, C.; Lv, T.; Fan, C.; Zhang, C. Fever Basketball: A Complex, Flexible, and Asynchronized Sports Game Environment for Multi-agent Reinforcement Learning. arXiv 2020, arXiv:2012.03204. [Google Scholar] [CrossRef] [Scilit]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef] [Scilit]
- Wijmans, E.; Kadian, A.; Morcos, A.; Lee, S.; Essa, I.; Parikh, D.; Savva, M.; Batra, D. DD-PPO: Learning Near-Perfect Pointgoal Navigators from 2.5 Billion Frames. arXiv 2019, arXiv:1911.00357. [Google Scholar]
- Zheng, R.; Dou, S.; Gao, S.; Hua, Y.; Shen, W.; Wang, B.; Liu, Y.; Jin, S.; Liu, Q.; Zhou, Y.; et al. Secrets of RLHF in Large Language Models Part I: PPO. arXiv 2023, arXiv:2307.04964. [Google Scholar] [CrossRef] [Scilit]
- Liang, Z.; Chen, H.; Zhu, J.; Jiang, K.; Li, Y. Adversarial Deep Reinforcement Learning in Portfolio Management. arXiv 2018, arXiv:1808.09940. [Google Scholar] [CrossRef] [Scilit]
- Wan, Z.; Wang, X.; Liu, C.; Alam, S.; Zheng, Y.; Liu, J.; Qu, Z.; Yan, S.; Zhu, Y.; Zhang, Q.; et al. Efficient Large Language Models: A Survey. arXiv 2024, arXiv:2312.03863. [Google Scholar]
- Li, T.; Sahu, A.K.; Zaheer, M.; Sanjabi, M.; Talwalkar, A.; Smith, V. Federated Optimization in Heterogeneous Networks. In Proceedings of the Machine Learning and Systems (MLSys), Austin, TX, USA, 2–4 March 2020; pp. 429–450. [Google Scholar]
- Russell, S.J.; Zimdars, A. Q-Decomposition for Reinforcement Learning Agents. In Proceedings of the International Conference, Machine Learning (ICML), Washington, DC, USA, 21–23 August 2003; pp. 656–663. [Google Scholar]
- Gupta, V.; Anand, D.; Paruchuri, P.; Kumar, A. Action Selection for Composable Modular Deep Reinforcement Learning. In Proceedings of the International Conference, Autonomous Agents & Multi-Agent Systems (AAMAS), London, UK, 3–7 May 2021; pp. 565–573. [Google Scholar] [CrossRef] [Scilit]
- Andreas, J.; Klein, D.; Levine, S. Modular Multitask Reinforcement Learning with Policy Sketches. In Proceedings of the International Conference, Machine Learning (ICML), Sydney, Australia, 6–11 August 2017; pp. 166–175. [Google Scholar]
- Yu, B.; Lee, T. Modular Reinforcement Learning for a Quadrotor UAV with Decoupled Yaw Control. IEEE Robot. Autom. Lett. 2025, 10, 572–579. [Google Scholar] [CrossRef] [Scilit]
- Shi, H.; Di, Y.; Ruan, X.; Liao, M.; Zhang, Q.; Ma, R.; Guan, H. Efficient federated recommender system with adaptive model pruning and momentum-based batch adjustment. ACM Trans. Recomm. Syst. 2025, 3, 1–23. [Google Scholar] [CrossRef] [Scilit]
- Deng, H.; Hua, Y.; Song, T.; Xue, Z.; Ma, R.; Robertson, N.; Guan, H. Reinforcing Neural Network Stability with Attractor Dynamics. In Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA, 9–11 February 2020; Volume 34, pp. 3765–3772. [Google Scholar] [CrossRef] [Scilit]
- Yarats, D.; Zhang, A.; Kostrikov, I.; Amos, B.; Pineau, J.; Fergus, R. Improving sample efficiency in model-free reinforcement learning from images. In Proceedings of the AAAI Conference on Artificial Intelligence, Online, 2–9 February 2021; Volume 35, pp. 10674–10681. [Google Scholar]
- Zahavy, T.; Haroush, M.; Merlis, N.; Mankowitz, D.J.; Mannor, S. Learn What Not to Learn: Action Elimination with Deep Reinforcement Learning. Adv. Neural Inf. Process. Syst. 2018, 31, 3566–3577. [Google Scholar]
- Liu, B.; Pu, Z.; Zhang, T.; Wang, H.; Yi, J.; Mi, J. Learning to Play Football from Sports Domain Perspective: A Knowledge-Embedded Deep Reinforcement Learning Framework. IEEE Trans. Games 2023, 15, 648–657. [Google Scholar] [CrossRef] [Scilit]
- Song, Y.; Jiang, H.; Zhang, H.; Tian, Z.; Zhang, W.; Wang, J. Boosting Studies of Multi-agent Reinforcement Learning on Google Research Football Environment: The past, Present, and Future. arXiv 2023, arXiv:2309.12951. [Google Scholar] [CrossRef] [Scilit]
- Baker, B.; Kanitscheider, I.; Markov, T.; Wu, Y.; Powell, G.; McGrew, B.; Mordatch, I. Emergent Tool Use from Multi-agent Autocurricula. In Proceedings of the International Conference on Learning Representations (ICLR), Addis Ababa, Ethiopia, 26–30 April 2020. [Google Scholar]
- Zhu, J.; Kuang, M.; Zhou, W.; Shi, H.; Zhu, J.; Han, X. Mastering Air Combat Game with Deep Reinforcement Learning. Def. Technol. 2024, 34, 295–312. [Google Scholar] [CrossRef] [Scilit]
- Tang, M.; Chen, R.; Zhu, J. A Multi-agent Cooperative Group Game Model Based on Intention-Strategy Optimization. Algorithms 2026, 19, 22. [Google Scholar] [CrossRef] [Scilit]








| Situation | Knowledge-Embedded Observation Data |
|---|---|
| Offense | Start flag |
| Ball possession state | |
| Distance, angle to the closest opponent agent | |
| Distance, angle to the rim | |
| Action, state of the closest opponent agent | |
| Personal, team average shoot success rate | |
| Distance, angle to the closest team agent | |
| Highest, lowest shoot success rate agent in team | |
| Defense | Start flag |
| Ball possession state of the mark target | |
| Distance, angle to the mark target | |
| Action, state of the mark target | |
| Shoot success rate of the mark target | |
| Personal, team average defensive accuracy | |
| Steal action available | |
| Block action available | |
| Loose-ball | Start flag |
| Distance to the ball | |
| Angle to the ball | |
| Ball state | |
| Ball’s height | |
| Closest agent to the ball in team | |
| Rebound action available |
| Situation | Action | Condition | Reward |
|---|---|---|---|
| Offense | Movement | Only for movement actions | (Previous team average shoot success rate–Current team average shoot success rate)/Max step |
| Pass | Only when possessing the ball and lowest shoot success rate | Average team shoot success rate × Scale factor | |
| Shoot | Only when possessing the ball and highest shoot success rate | Shoot success rate × Scale factor | |
| Defense | Movement | Only for movement actions | (Previous defensive accuracy rate–Current defensive accuracy rate)/Max step |
| Face Up | Only when distance to mark target is lower than defense effective range | Defensive accuracy rate × Scale factor | |
| Steal | Only when distance to mark target is lower than defense effective range | Defensive accuracy rate × Scale factor | |
| Block | Only when distance to mark target is lower than defense effective range | Defensive accuracy rate × Scale factor | |
| Loose-ball | Movement | Only for movement actions and when closest to the ball on team | (Previous distance to the ball–Current distance to the ball)/Max step |
| Only for movement actions and when not closest to the ball on team | |||
| Rebound | Only when ball height is higher than minimum height to rebound | Distance to the ball × Scale factor |
| Situation | Actions |
|---|---|
| Common | Movement (8 directions), Stop |
| Offense | Shoot, Pass, Breakthrough |
| Defense | Steal, Block, Face Up |
| Loose-ball | Rebound, Diving Catch |
| Hyperparameter | Value |
|---|---|
| Learning rate | 0.00001 |
| Gamma | 0.99933 |
| Mini batch size | 64 |
| GAE lambda | 0.95 |
| Clip ratio | 0.2 |
| Gradient clipping max norm | 0.5 |
| Hidden layer size | 64 |
| Layer | 2 fully connected layers |
| Activation | Hyperbolic tangent |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kim, J.; Park, J.; Cho, K. Enhancing Multi-Agent Reinforcement Learning via Knowledge-Embedded Modular Framework for Online Basketball Games. Mathematics 2026, 14, 419. https://doi.org/10.3390/math14030419
Kim J, Park J, Cho K. Enhancing Multi-Agent Reinforcement Learning via Knowledge-Embedded Modular Framework for Online Basketball Games. Mathematics. 2026; 14(3):419. https://doi.org/10.3390/math14030419
Chicago/Turabian StyleKim, Junhyuk, Jisun Park, and Kyungeun Cho. 2026. "Enhancing Multi-Agent Reinforcement Learning via Knowledge-Embedded Modular Framework for Online Basketball Games" Mathematics 14, no. 3: 419. https://doi.org/10.3390/math14030419
APA StyleKim, J., Park, J., & Cho, K. (2026). Enhancing Multi-Agent Reinforcement Learning via Knowledge-Embedded Modular Framework for Online Basketball Games. Mathematics, 14(3), 419. https://doi.org/10.3390/math14030419

