Mapping the Knowledge Structure of Physical Artificial Intelligence: A Data-Driven Systematic Review
Abstract
1. Introduction
2. Methods
2.1. Article Collection and Corpus Creation (Step 1)
2.2. Informetric Network Analysis (Step 2)
2.3. Topic Modeling (Step 3)
3. Results
3.1. Results of Bibliometric Networks
3.1.1. Descriptive Statistics
3.1.2. Visualization of the Keyword Co-Occurrence Network
3.1.3. Visualization of the Co-Citation Network of Cited References
3.2. Results of LDA Topic Modeling and Discovery
3.2.1. Layer 1: Physical Interaction & Infrastructure
Topic 1: Embodied AI
Topic 2: Human–Robot Interaction
3.2.2. Layer 2: Policy Learning & Control
Topic 4: Simulation-Based Learning
3.2.3. Layer 3: Cognitive Integration & Multimodal Reasoning
Topic 6: Language-Grounded Action
Topic 7: Multimodal Learning
4. Discussion
4.1. Three-Layered Hierarchical Architecture of PAI
4.2. Educational, Research Policy, and Societal Implications of PAI
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Wu, E.-H.; Liu, Y.-Q.; Xu, T.-C.; Ren, L.-X.; Qin, Y.-M.; Wei, M.-Y.; He, X.-W.; Yuan, D.-Y.; Hou, W.-C.; Ma, Z.-W.; et al. Physical AI: Evolution, Progress, Challenges, and Prospects. J. Comput. Sci. Technol. 2026, 41, 271–288. [Google Scholar] [CrossRef] [Scilit]
- Abou Ali, M.; Dornaika, F.; Charafeddine, J. Agentic AI: A comprehensive survey of architectures, applications, and future directions. Artif. Intell. Rev. 2025, 59, 11. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Chen, W.; Bai, Y.; Liang, X.; Li, G.; Gao, W.; Lin, L. Aligning Cyber Space With Physical World: A Comprehensive Survey on Embodied AI. IEEE/ASME Trans. Mechatron. 2025, 30, 7253–7274. [Google Scholar] [CrossRef] [Scilit]
- Salehi, V. Fundamentals of Physical AI. J. Intell. Syst. Syst. Lifecycle Manag. 2025, 2. [Google Scholar] [CrossRef] [Scilit]
- Bousetouane, F. Physical AI agents: Integrating cognitive intelligence with real-world action. arXiv 2025, arXiv:2501.08944. Available online: https://arxiv.org/abs/2501.08944 (accessed on 18 October 2025).
- Miriyev, A.; Kovač, M. Skills for physical artificial intelligence. Nat. Mach. Intell. 2020, 2, 658–660. [Google Scholar] [CrossRef] [Scilit]
- Thakur, A.; Kaipa, K.; Banerjee, A.G.; Cappelleri, D.J.; Krovi, V.N.; Gupta, S.K. Physical Artificial Intelligence for Powering the Next Revolution in Robotics. J. Comput. Inf. Sci. Eng. 2025, 25, 120809. [Google Scholar] [CrossRef] [Scilit]
- Cheng, X.; Shen, Z.; Zhang, Y. Bioinspired 3D flexible devices and functional systems. Natl. Sci. Rev. 2024, 11, nwad314. [Google Scholar] [CrossRef] [Scilit]
- Sørensen, L.; Sagen Johannesen, D.T.; Melkas, H.; Johnsen, H.M. User Acceptance of a Home Robotic Assistant for Individuals With Physical Disabilities: Explorative Qualitative Study. JMIR Rehabil. Assist. Technol. 2025, 12, e63641. [Google Scholar] [CrossRef] [Scilit]
- Bajestani, M.S.; Kim, C.; Lee, K.-C.; Kim, D.B. Self-X-based secure human-cyber-physical system (SSHCPS) for autonomous manufacturing in the era of industry 5.0. Adv. Eng. Inform. 2025, 69, 104054. [Google Scholar] [CrossRef] [Scilit]
- Sharma, A.; Bhowmik, B. Autonomous agentic AI with policy adaptation for physics-informed spectral learning in Structural Health Monitoring. Adv. Eng. Inform. 2026, 70, 104224. [Google Scholar] [CrossRef] [Scilit]
- Balasubramani, M.; Chen, J.; Chang, R.; Shieh, J.-S. Development of a Human-Centric Autonomous Heating, Ventilation, and Air Conditioning Control System Enhanced for Industry 5.0 Chemical Fiber Manufacturing. Machines 2025, 13, 421. [Google Scholar] [CrossRef] [Scilit]
- Ohueri, C.C.; Seghier, T.E.; Jing, K.T.; Esa, M. AI-powered adaptive exoskeletons for long-term musculoskeletal disorder prevention in dynamic construction environments. Adv. Eng. Inform. 2025, 69, 104042. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Li, Z.; Duan, Y.; Spulber, A.-B. Physical artificial intelligence (PAI): The next-generation artificial intelligence. Front. Inf. Technol. Electron. Eng. 2023, 24, 1231–1238. [Google Scholar] [CrossRef] [Scilit]
- Ray, P.P. Physical AI: Bridging the sim-to-real divide toward embodied, ethical, and autonomous intelligence. Mach. Learn. Comput. Sci. Eng. 2026, 2, 1. [Google Scholar] [CrossRef] [Scilit]
- Wu, J.; You, H.; Du, J. AI generations: From AI 1.0 to AI 4.0. Front. Artif. Intell. 2025, 8, 1585629. [Google Scholar] [CrossRef] [Scilit]
- Agarwal, N.; Ali, A.; Bala, M.; Balaji, Y.; Barker, E.; Cai, T.; Chattopadhyay, P.; Chen, Y.; Cui, Y.; Ding, Y.; et al. Cosmos world foundation model platform for physical ai. arXiv 2025, arXiv:2501.03575. Available online: https://arxiv.org/abs/2501.03575 (accessed on 13 September 2026).
- Booth, A.; Sutton, A.; Clowes, M.; Martyn-St James, M. Systematic Approaches to a Successful Literature Review, 3rd ed.; SAGE Publications: London, UK, 2021. [Google Scholar]
- Lee, C.-H.; Liu, C.-L.; Trappey, A.J.; Mo, J.P.T.; Desouza, K.C. Understanding digital transformation in advanced manufacturing and engineering: A bibliometric analysis, topic modeling and research trend discovery. Adv. Eng. Inform. 2021, 50, 101428. [Google Scholar] [CrossRef] [Scilit]
- Pranckutė, R. Web of Science (WoS) and Scopus: The titans of bibliographic information in today’s academic world. Publications 2021, 9, 12. [Google Scholar] [CrossRef] [Scilit]
- Donthu, N.; Kumar, S.; Mukherjee, D.; Pandey, N.; Lim, W.M. How to conduct a bibliometric analysis: An overview and guidelines. J. Bus. Res. 2021, 133, 285–296. [Google Scholar] [CrossRef] [Scilit]
- Archambault, É.; Campbell, D.; Gingras, Y.; Larivière, V. Comparing bibliometric statistics obtained from the Web of Science and Scopus. J. Am. Soc. Inf. Sci. Technol. 2009, 60, 1320–1326. [Google Scholar] [CrossRef] [Scilit]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. Br. Med. J. 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit]
- Wang, B.; Liu, S.; Ding, K.; Liu, Z.; Xu, J. Identifying technological topics and institution-topic distribution probability for patent competitive intelligence analysis: A case study in LTE technology. Scientometrics 2014, 101, 685–704. [Google Scholar] [CrossRef] [Scilit]
- Penning de Vries, B.B.L.; van Smeden, M.; Rosendaal, F.R.; Groenwold, R.H.H. Title, abstract, and keyword searching resulted in poor recovery of articles in systematic reviews of epidemiologic practice. J. Clin. Epidemiol. 2020, 121, 55–61. [Google Scholar] [CrossRef] [Scilit]
- Callon, M.; Courtial, J.-P.; Turner, W.A.; Bauin, S. From translations to problematic networks: An introduction to co-word analysis. Soc. Sci. Inf. 1983, 22, 191–235. [Google Scholar] [CrossRef] [Scilit]
- Blei, D.; Ng, A.; Jordan, M. Latent Dirichlet Allocation. J. Mach. Learn. Res. 2003, 3, 993–1022. [Google Scholar]
- Nikolenko, S.I.; Koltcov, S.; Koltsova, O. Topic modelling for qualitative studies. J. Inf. Sci. 2016, 43, 88–102. [Google Scholar] [CrossRef] [Scilit]
- Jacobi, C.; van Atteveldt, W.; Welbers, K. Quantitative analysis of large amounts of journalistic texts using topic modelling. In Rethinking Research Methods in an Age of Digital Journalism; Karlsson, M., Sjøvaag, H., Eds.; Routledge: London, UK, 2018; pp. 89–106. [Google Scholar] [CrossRef] [Scilit]
- Griffiths, T.L.; Steyvers, M. Finding scientific topics. Proc. Natl. Acad. Sci. USA 2004, 101, 5228–5235. [Google Scholar] [CrossRef] [Scilit]
- Bianchi, F.; Terragni, S.; Hovy, D. Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 759–766. [Google Scholar] [CrossRef] [Scilit]
- Angelov, D. Top2vec: Distributed representations of topics. arXiv 2020, arXiv:2008.09470. Available online: https://arxiv.org/abs/2008.09470 (accessed on 13 September 2026).
- Hoyle, A.; Goel, P.; Hian-Cheong, A.; Peskov, D.; Boyd-Graber, J.; Resnik, P. Is Automated Topic Model Evaluation Broken? The Incoherence of Coherence. Adv. Neural Inf. Process. Syst. 2021, 34, 2018–2033. [Google Scholar]
- Antons, D.; Breidbach, C.F. Big Data, Big Insights? Advancing Service Innovation and Design with Machine Learning. J. Serv. Res. 2017, 21, 17–39. [Google Scholar] [CrossRef] [Scilit]
- Zupic, I.; Čater, T. Bibliometric Methods in Management and Organization. Organ. Res. Methods 2015, 18, 429–472. [Google Scholar] [CrossRef] [Scilit]
- Savva, M.; Kadian, A.; Maksymets, O.; Zhao, Y.; Wijmans, E.; Jain, B.; Straub, J.; Liu, J.; Koltun, V.; Malik, J.; et al. Habitat: A platform for embodied ai research. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 9339–9347. [Google Scholar]
- Chaplot, D.S.; Gandhi, D.P.; Gupta, A.; Salakhutdinov, R.R. Object goal navigation using goal-oriented semantic exploration. Adv. Neural Inf. Process. Syst. 2020, 33, 4247–4258. [Google Scholar]
- Wijmans, E.; Kadian, A.; Morcos, A.; Lee, S.; Essa, I.; Parikh, D.; Savva, M.; Batra, D. Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames. arXiv 2019, arXiv:1911.00357. Available online: https://arxiv.org/abs/1911.00357 (accessed on 13 September 2026).
- Gupta, S.; Davidson, J.; Levine, S.; Sukthankar, R.; Malik, J. Cognitive mapping and planning for visual navigation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2616–2625. [Google Scholar]
- Zhu, Y.; Mottaghi, R.; Kolve, E.; Lim, J.J.; Gupta, A.; Fei-Fei, L.; Farhadi, A. Target-driven visual navigation in indoor scenes using deep reinforcement learning. In Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore, 29 May–3 June 2017; IEEE: New York, NY, USA, 2017; pp. 3357–3364. [Google Scholar]
- Kolve, E.; Mottaghi, R.; Han, W.; VanderBilt, E.; Weihs, L.; Herrasti, A.; Deitke, M.; Ehsani, K.; Gordon, D.; Zhu, Y.; et al. Ai2-thor: An interactive 3d environment for visual ai. arXiv 2017, arXiv:1712.05474. Available online: https://arxiv.org/abs/1712.05474 (accessed on 13 September 2026).
- Xia, F.; Zamir, A.R.; He, Z.; Sax, A.; Malik, J.; Savarese, S. Gibson env: Real-world perception for embodied agents. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 9068–9079. [Google Scholar]
- Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft COCO: Common Objects in Context. In Proceedings of the Computer Vision–ECCV 2014; Springer: Cham, Switzerland, 2014; Volume 8693, pp. 740–755. [Google Scholar] [CrossRef] [Scilit]
- Szot, A.; Clegg, A.; Undersander, E.; Wijmans, E.; Zhao, Y.; Turner, J.; Maestre, N.; Mukadam, M.; Chaplot, D.S.; Maksymets, O.; et al. Habitat 2.0: Training Home Assistants to Rearrange their Habitat. Adv. Neural Inf. Process. Syst. 2021, 34, 251–266. [Google Scholar]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. Available online: https://arxiv.org/abs/1707.06347 (accessed on 13 September 2026).
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning, Virtual, 18–24 July 2021; Volume 139, pp. 8748–8763. Available online: https://proceedings.mlr.press/v139/radford21a (accessed on 13 September 2026).
- Graves, A. Long Short-Term Memory. In Supervised Sequence Labelling with Recurrent Neural Networks; Springer: Berlin/Heidelberg, Germany, 2012; pp. 37–45. [Google Scholar] [CrossRef] [Scilit]
- Dai, A.; Chang, A.X.; Savva, M.; Halber, M.; Funkhouser, T.; Nießner, M. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 5828–5839. [Google Scholar]
- Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, 2–7 June 2019; pp. 4171–4186. [Google Scholar] [CrossRef] [Scilit]
- Fried, D.; Hu, R.; Cirik, V.; Rohrbach, A.; Andreas, J.; Morency, L.P.; Berg-Kirkpatrick, T.; Saenko, K.; Klein, D.; Darrell, T. Speaker-follower models for vision-and-language navigation. Adv. Neural Inf. Process. Syst. 2018, 31, 3314–3325. [Google Scholar]
- Qi, Y.; Wu, Q.; Anderson, P.; Wang, X.; Wang, W.Y.; Shen, C.; Hengel, A. Reverie: Remote embodied visual referring expression in real indoor environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 14–19 June 2020; pp. 9982–9991. [Google Scholar]
- Tan, H.; Bansal, M. LXMERT: Learning Cross-Modality Encoder Representations from Transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, Hong Kong, China, 3–7 November 2019; pp. 5100–5111. [Google Scholar] [CrossRef] [Scilit]
- Anderson, P.; Wu, Q.; Teney, D.; Bruce, J.; Johnson, M.; Sünderhauf, N.; Reid, I.; Gould, S.; van den Hengel, A. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 3674–3683. [Google Scholar]
- Röder, M.; Both, A.; Hinneburg, A. Exploring the Space of Topic Coherence Measures. In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining, Shanghai, China, 2–6 February 2015; pp. 399–408. [Google Scholar] [CrossRef] [Scilit]
- Wan, Z.; Du, Y.; Ibrahim, M.; Zhao, Y.; Krishna, T.; Raychowdhury, A. Thinking and moving: An efficient computing approach for integrated task and motion planning in cooperative embodied ai systems. In Proceedings of the 43rd IEEE/ACM International Conference on Computer-Aided Design, New York, NY, USA, 27–31 October 2024; pp. 1–7. [Google Scholar]
- Hu, D.; Lan, D.; Liu, Y.; Ning, J.; Wang, J.; Yang, Y. Embodied AI Through Cloud-Fog Computing: A Framework for Everywhere Intelligence. In Proceedings of the 2024 IEEE 33rd International Symposium on Industrial Electronics (ISIE), Ulsan, Republic of Korea, 18–21 June 2024; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
- Foglia, L.; Wilson, R. Embodied cognition. Wiley Interdiscip. Rev. Cogn. Sci. 2013, 4, 319–325. [Google Scholar] [CrossRef] [Scilit]
- Pfeifer, R.; Bongard, J. How the Body Shapes the Way We Think: A New View of Intelligence; MIT Press: Cambridge, MA, USA, 2006. [Google Scholar]
- Kwon, W.; Baek, S.; Baek, J.; Shin, W.; Gwak, M.; Park, P.; Lee, S. Reinforced Intelligence Through Active Interaction in Real World: A Survey on Embodied AI. Int. J. Control Autom. Syst. 2025, 23, 1597–1612. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Chen, K.; Chen, S.; Zheng, Y.; Huang, T.; Yu, Z. Spikegs: 3d gaussian splatting from spike streams with high-speed camera motion. In Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, VIC, Australia, 28 October–1 November 2024; pp. 9194–9203. [Google Scholar]
- Zitkovich, B.; Yu, T.; Xu, S.; Xu, P.; Xiao, T.; Xia, F.; Wu, J.; Wohlhart, P.; Welker, S.; Wahid, A.; et al. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. In Proceedings of the 7th Conference on Robot Learning, Atlanta, GA, USA, 6–9 November 2023; Volume 229, pp. 2165–2183. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. Available online: https://arxiv.org/abs/2312.00752 (accessed on 13 September 2026).
- Hamburg, S.; Jimenez Rodriguez, A.; Htet, A.; Di Nuovo, A. Active Inference for Learning and Development in Embodied Neuromorphic Agents. Entropy 2024, 26, 582. [Google Scholar] [CrossRef] [Scilit]
- Friston, K. The free-energy principle: A unified brain theory? Nat. Rev. Neurosci. 2010, 11, 127–138. [Google Scholar] [CrossRef] [Scilit]
- You, H.; Zhou, T.; Zhu, Q.; Ye, Y.; Du, E.J. Embodied AI for dexterity-capable construction Robots: DEXBOT framework. Adv. Eng. Inform. 2024, 62, 102572. [Google Scholar] [CrossRef] [Scilit]
- Zhao, M.; Xia, J.; Hou, K.; Liu, Y.; Xia, S.; Jiang, X. FlexiFly: Interfacing the Physical World with Foundation Models Empowered by Reconfigurable Drone Systems. In Proceedings of the 23rd ACM Conference on Embedded Networked Sensor Systems, Irvine, CA, USA, 6–9 May 2025; pp. 463–476. [Google Scholar] [CrossRef] [Scilit]
- De Haro, L. Using Embodied Artificial Intelligence Agents to Automate Biorisk Management Tasks in High-Containment Laboratories. Appl. Biosaf. 2025, 30, 314–325. [Google Scholar] [CrossRef] [Scilit]
- Dennett, D.C. The Intentional Stance; Mit Press: Cambridge, MA, USA, 1987. [Google Scholar]
- Sini, R. Does Saudi robot citizen have more rights than women? BBC News, 26 October 2017. Available online: https://www.bbc.com/news/blogs-trending-41761856 (accessed on 13 September 2026).
- Tiku, N. The Google engineer who thinks the company’s AI has come to life. The Washington Post, 11 June 2022. Available online: https://www.washingtonpost.com/technology/2022/06/11/google-ai-lamda-blake-lemoine/ (accessed on 17 July 2026).
- Gray, H.; Gray, K.; Wegner, D. Dimensions of Mind Perception. Science 2007, 315, 619. [Google Scholar] [CrossRef] [Scilit]
- Zhou, X.; Zhou, Y.; Gong, Y.; Cai, Z.; Qiu, A.; Xiao, Q. Bee and I need diversity! Break Filter Bubbles in Recommendation Systems through Embodied AI Learning. In Proceedings of the 23rd Annual ACM Interaction Design and Children Conference, Delft, The Netherlands, 17–20 June 2024; pp. 44–61. [Google Scholar] [CrossRef] [Scilit]
- Balazadeh, K.; Kajonius, P. Exploring Intimacy with Artificial Intelligence: Validation of Robot Intimacy Receptivity Scale (RIRS). Int. J. Soc. Robot. 2025, 17, 1453–1465. [Google Scholar] [CrossRef] [Scilit]
- Cheetham, M.; Pedroni, A.F.; Antley, A.; Slater, M.; Jäncke, L. Virtual milgram: Empathic concern or personal distress? Evidence from functional MRI and dispositional measures. Front. Hum. Neurosci. 2009, 3, 29. [Google Scholar] [CrossRef] [Scilit]
- Cheetham, M.; Suter, P.; Jancke, L. Perceptual discrimination difficulty and familiarity in the Uncanny Valley: More like a “Happy Valley”. Front. Psychol. 2014, 5, 1219. [Google Scholar] [CrossRef] [Scilit]
- Marquardt, M.; Graf, P.; Jansen, E.; Hillmann, S.; Voigt-Antons, J.N. Situativität, Funktionalität und Vertrauen: Ergebnisse einer szenariobasierten Interviewstudie zur Erklärbarkeit von KI in der Medizin. TATuP-Z. Tech. Theor. Prax. 2024, 33, 41–47. [Google Scholar] [CrossRef] [Scilit]
- Cheon, E.; Zaga, C.; Lee, H.; Lupetti, M.; Dombrowski, L.; Jung, M. Human-Machine Partnerships in the Future of Work: Exploring the Role of Emerging Technologies in Future Workplaces. In Companion Publication of the 2021 Conference on Computer Supported Cooperative Work and Social Computing; Association for Computing Machinery: New York, NY, USA, 2021; pp. 323–326. [Google Scholar] [CrossRef] [Scilit]
- Zubala, A.; Pease, A.; Lyszkiewicz, K.; Hackett, S. Art psychotherapy meets creative AI: An integrative review positioning the role of creative AI in art therapy process. Front. Psychol. 2025, 16, 154839. [Google Scholar] [CrossRef] [Scilit]
- Song, X.; Liu, C.; Xu, L.; Gao, B.; Lu, Z.; Zhang, Y. Affective computing methods for multimodal embodied AI human–computer interaction. Aslib J. Inf. Manag. 2025, 77, 1–25. [Google Scholar] [CrossRef] [Scilit]
- Batra, D.; Gokaslan, A.; Kembhavi, A.; Maksymets, O.; Mottaghi, R.; Savva, M.; Toshev, A.; Wijmans, E. Objectnav revisited: On evaluation of embodied agents navigating to objects. arXiv 2020, arXiv:2006.13171. Available online: https://arxiv.org/abs/2006.13171 (accessed on 13 September 2026).
- Anderson, P.; Chang, A.; Chaplot, D.S.; Dosovitskiy, A.; Gupta, S.; Koltun, V.; Kosecka, J.; Malik, J.; Mottaghi, R.; Savva, M.; et al. On evaluation of embodied navigation agents. arXiv 2018, arXiv:1807.06757. Available online: https://arxiv.org/abs/1807.06757 (accessed on 13 September 2026).
- Armeni, I.; He, Z.Y.; Gwak, J.; Zamir, A.R.; Fischer, M.; Malik, J.; Fischer, M.; Savarese, S. 3D scene graph: A structure for unified semantics, 3d space, and camera. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 5664–5673. [Google Scholar]
- Locatello, F.; Weissenborn, D.; Unterthiner, T.; Mahendran, A.; Heigold, G.; Uszkoreit, J.; Dosovitskiy, A.; Kipf, T. Object-centric learning with slot attention. Adv. Neural Inf. Process. Syst. 2020, 33, 11525–11538. [Google Scholar]
- Pal, A.; Qiu, Y.; Christensen, H. Learning hierarchical relationships for object-goal navigation. In Proceedings of the 2020 Conference on Robot Learning, Virtual, 16–18 November 2020; Volume 155, pp. 517–528. [Google Scholar]
- Seymour, Z.; Thopalli, K.; Mithun, N.; Chiu, H.P.; Samarasekera, S.; Kumar, R. Maast: Map attention with semantic transformers for efficient visual navigation. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi’an, China, 30 May–5 June 2021; IEEE: New York, NY, USA, 2021; pp. 13223–13230. [Google Scholar]
- Li, L.; Chu, W.; Langford, J.; Schapire, R.E. A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th International Conference on World Wide Web, Raleigh, NC, USA, 26–30 April 2010; pp. 661–670. [Google Scholar]
- Rudra, S.; Goel, S.; Santara, A.; Gentile, C.; Perron, L.; Xia, F.; Sindhwani, V.; Parada, C.; Aggarwal, G. A contextual bandit approach for learning to plan in environments with probabilistic goal configurations. arXiv 2022, arXiv:2211.16309. Available online: https://arxiv.org/abs/2211.16309 (accessed on 13 September 2026).
- Zeng, H.; Song, X.; Jiang, S. Multi-object navigation using potential target position policy function. IEEE Trans. Image Process. 2023, 32, 2608–2619. [Google Scholar] [CrossRef] [Scilit]
- Blum, A.; Chalasani, P.; Coppersmith, D.; Pulleyblank, B.; Raghavan, P.; Sudan, M. The minimum latency problem. In Proceedings of the Twenty-Sixth Annual ACM Symposium on Theory of Computing, Montréal, QC, Canada, 23–25 May 1994; pp. 163–171. [Google Scholar]
- Höfer, S.; Bekris, K.; Handa, A.; Gamboa, J.C.; Mozifian, M.; Golemo, F.; Atkeson, C.; Fox, D.; Goldberg, K.; Leonard, J.; et al. Sim2real in robotics and automation: Applications and challenges. IEEE Trans. Autom. Sci. Eng. 2021, 18, 398–400. [Google Scholar] [CrossRef] [Scilit]
- Wu, S.C.; Wald, J.; Tateno, K.; Navab, N.; Tombari, F. Scenegraphfusion: Incremental 3d scene graph prediction from rgb-d sequences. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtually, 19–25 June 2021; pp. 7515–7525. [Google Scholar]
- Wani, S.; Patel, S.; Jain, U.; Chang, A.; Savva, M. Multion: Benchmarking semantic map memory using multi-object navigation. Adv. Neural Inf. Process. Syst. 2020, 33, 9700–9712. [Google Scholar]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level Control through Deep Reinforcement Learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit]
- Petrenko, A.; Wijmans, E.; Shacklett, B.; Koltun, V. Megaverse: Simulating embodied agents at one million experiences per second. In Proceedings of the 38th International Conference on Machine Learning, Virtual, 18–24 July 2021; Volume 139, pp. 8556–8566. [Google Scholar]
- Chang, A.; Dai, A.; Funkhouser, T.; Halber, M.; Niessner, M.; Savva, M.; Song, S.; Zeng, A.; Zhang, Y. Matterport3d: Learning from rgb-d data in indoor environments. arXiv 2017, arXiv:1709.06158. Available online: https://arxiv.org/abs/1709.06158 (accessed on 13 September 2026).
- Kadian, A.; Truong, J.; Gokaslan, A.; Clegg, A.; Wijmans, E.; Lee, S.; Savva, M.; Chernova, S.; Batra, D. Sim2Real Predictivity: Does Evaluation in Simulation Predict Real-World Performance? IEEE Robot. Autom. Lett. 2020, 5, 6670–6677. [Google Scholar] [CrossRef] [Scilit]
- Jain, U.; Liu, I.J.; Lazebnik, S.; Kembhavi, A.; Weihs, L.; Schwing, A.G. Gridtopix: Training embodied agents with minimal supervision. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Virtually, 10–17 October 2021; pp. 15141–15151. [Google Scholar]
- Ramakrishnan, S.K.; Gokaslan, A.; Wijmans, E.; Maksymets, O.; Clegg, A.; Turner, J.; Undersander, E.; Galuba, W.; Westbury, A.; Chang, A.X.; et al. Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai. arXiv 2021, arXiv:2109.08238. Available online: https://arxiv.org/abs/2109.08238 (accessed on 13 September 2026).
- Chao, Y.W.; Paxton, C.; Xiang, Y.; Yang, W.; Sundaralingam, B.; Chen, T.; Murali, A.; Cakmak, M.; Fox, D. Handoversim: A simulation framework and benchmark for human-to-robot object handovers. In Proceedings of the 2022 International Conference on Robotics and Automation (ICRA), Philadelphia, PA, USA, 23–27 May 2022; IEEE: New York, NY, USA, 2022; pp. 6941–6947. [Google Scholar]
- Christen, S.; Yang, W.; Pérez-D’Arpino, C.; Hilliges, O.; Fox, D.; Chao, Y.W. Learning human-to-robot handovers from point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 9654–9664. [Google Scholar]
- Kumar, A.; Fu, Z.; Pathak, D.; Malik, J. Rma: Rapid motor adaptation for legged robots. arXiv 2021, arXiv:2107.04034. Available online: https://arxiv.org/abs/2107.04034 (accessed on 13 September 2026).
- Liang, Y.; Ellis, K.; Henriques, J. Rapid motor adaptation for robotic manipulator arms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 16404–16413. [Google Scholar]
- Zhu, L.; Jia, K.; Zhao, Y.; Qi, Y.; Wang, L.; Huang, H. Spikenerf: Learning neural radiance fields from continuous spike stream. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 6285–6295. [Google Scholar]
- Jiang, Y.; Guo, M.; Li, J.; Exarchos, I.; Wu, J.; Liu, C.K. Dash: Modularized human manipulation simulation with vision and language for embodied ai. In Proceedings of the ACM SIGGRAPH/Eurographics Symposium on Computer Animation, Virtual, 6–9 September 2021; pp. 1–12. [Google Scholar]
- Lu, H.; Tan, X.; Chen, M.; Zhang, Z.; Zhang, X.; Chen, J.; Wei, X.; Zhao, T. Cross-Modal Haptic Compression Inspired by Embodied AI for Haptic Communications. IEEE Trans. Multimed. 2025, 27, 4996–5008. [Google Scholar] [CrossRef] [Scilit]
- Puig, X.; Shu, T.; Tenenbaum, J.B.; Torralba, A. Nopa: Neurally-guided online probabilistic assistance for building socially intelligent home assistants. arXiv 2023, arXiv:2301.05223. Available online: https://arxiv.org/abs/2301.05223 (accessed on 13 September 2026).
- Jain, V.; Magalhaes, G.; Ku, A.; Vaswani, A.; Ie, E.; Baldridge, J. Stay on the path: Instruction fidelity in vision-and-language navigation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 1862–1872. [Google Scholar]
- Lin, B.; Nie, Y.; Wei, Z.; Chen, J.; Ma, S.; Han, J.; Xu, H.; Chang, X.; Liang, X. NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 5945–5957. [Google Scholar] [CrossRef] [Scilit]
- Min, S.Y.; Puig, X.; Chaplot, D.S.; Yang, T.Y.; Rai, A.; Parashar, P.; Salakhutdinov, R.; Bisk, Y.; Mottaghi, R. Situated instruction following. In European Conference on Computer Vision; Springer Nature: Cham, Switzerland, 2024; pp. 202–228. [Google Scholar]
- Gupta, S.K. Embodied ai for smart robotic cells in manufacturing applications. Proc. AAAI Conf. Artif. Intell. 2025, 39, 28630–28636. [Google Scholar] [CrossRef] [Scilit]
- Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q.V.; Zhou, D. Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inf. Process. Syst. 2022, 35, 24824–24837. [Google Scholar] [CrossRef] [Scilit]
- Shridhar, M.; Thomason, J.; Gordon, D.; Bisk, Y.; Han, W.; Mottaghi, R.; Zettlemoyer, L.; Fox, D. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 10740–10749. [Google Scholar]
- Chen, C.; Cong, Y.; Kan, Z. Worldafford: Affordance grounding based on natural language instructions. In Proceedings of the 2024 IEEE 36th International Conference on Tools with Artificial Intelligence (ICTAI), Herndon, VA, USA, 28–30 October 2024; IEEE: New York, NY, USA, 2024; pp. 822–828. [Google Scholar]
- Xu, W.; Wang, M.; Zhou, W.; Li, H. P-RAG: Progressive retrieval augmented generation for planning on embodied everyday task. In Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, VIC, Australia, 28 October–1 November 2024; pp. 6969–6978. [Google Scholar]
- Gao, F.; Shi, L.; Tang, J.; Wang, J.; Li, S.; Ma, S.; Yu, J. Visual and Textual Commonsense-Enhanced Layout Learning for Vision-and-Language Navigation. IEEE Trans. Autom. Sci. Eng. 2025, 22, 21311–21324. [Google Scholar] [CrossRef] [Scilit]
- Rutar, D.; Markelius, A.; Schellaert, W.; Hernández-Orallo, J.; Cheke, L. General interaction battery: Simple object navigation and affordances (GIBSONA). Cogn. Syst. Res. 2025, 94, 101411. [Google Scholar] [CrossRef] [Scilit]
- Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.-T.; Rocktäschel, T.; et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Adv. Neural Inf. Process. Syst. 2020, 33, 9459–9474. [Google Scholar]
- Johnson-Laird, P.N. Mental models and human reasoning. Proc. Natl. Acad. Sci. USA 2010, 107, 18243–18250. [Google Scholar] [CrossRef] [Scilit]
- Wang, R.; Xu, P.; Shi, H.; Schumann, E.; Liu, C.K. FürElise: Capturing and physically synthesizing hand motion of piano performance. In SIGGRAPH Asia 2024 Conference Papers; Association for Computing Machinery: New York, NY, USA, 2024; Article 77; pp. 1–11. [Google Scholar] [CrossRef] [Scilit]
- Alayrac, J.B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. Flamingo: A visual language model for few-shot learning. Adv. Neural Inf. Process. Syst. 2022, 35, 23716–23736. [Google Scholar] [CrossRef] [Scilit]
- Hong, Y.; Zhen, H.; Chen, P.; Zheng, S.; Du, Y.; Chen, Z.; Gan, C. 3d-llm: Injecting the 3d world into large language models. Adv. Neural Inf. Process. Syst. 2023, 36, 20482–20494. [Google Scholar] [CrossRef] [Scilit]
- Ma, X.; Yong, S.; Zheng, Z.; Li, Q.; Liang, Y.; Zhu, S.C.; Huang, S. Sqa3d: Situated question answering in 3d scenes. arXiv 2022, arXiv:2210.07474. Available online: https://arxiv.org/abs/2210.07474 (accessed on 13 September 2026).
- Zhang, R.; Zhao, C.; Du, H.; Niyato, D.; Wang, J.; Sawadsitang, S.; Shen, X.; Kim, D.I. Embodied AI-enhanced vehicular networks: An integrated vision language models and reinforcement learning method. IEEE Trans. Mob. Comput. 2025, 24, 11494–11510. [Google Scholar] [CrossRef] [Scilit]
- Zhou, J.; Wang, J.; Zhang, J.; Sun, W.; Zhang, J.; Birchfield, S.; Guo, D.; Kong, L.; Wang, M.; Zhong, Y. Audio–visual segmentation. In European Conference on Computer Vision; Springer Nature: Cham, Switzerland, 2022; pp. 386–403. [Google Scholar]
- Wang, S.; Huang, X.; Chen, C.; Wu, L.; Li, J. Reform: Error-aware few-shot knowledge graph completion. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, Virtually, 1–5 November 2021; pp. 1979–1988. [Google Scholar]
- Wen, C.; Liang, J.; Yuan, S.; Huang, H.; Bethala, G.C.R.; Liu, Y.S.; Wang, M.; Tzes, A.; Fang, Y. How secure are large language models (llms) for navigation in urban environments? arXiv 2024, arXiv:2402.09546. Available online: https://arxiv.org/abs/2402.09546 (accessed on 13 September 2026).
- Li, K.; Geng, Q.; Wan, M.; Cao, X.; Zhou, Z. Context and Spatial Feature Calibration for Real-Time Semantic Segmentation. IEEE Trans. Image Process. 2023, 32, 5465–5477. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]







| References | Description | Domains | Scope and Limitations |
|---|---|---|---|
| Miriyev and Kovač [6] | PAI concerns both the conceptual foundations and practical development of physical systems that can perform tasks normally linked to intelligent living beings. | Robotics | Emphasizes physical embodiment and morphology, but provides relatively limited consideration of distributed intelligence and higher-level cognition. |
| Li et al. [14] | PAI can be understood as a multidisciplinary field focused on nature-inspired intelligent robots, highlighting the integration of software-based intelligence with hardware elements such as materials and mechanics. | Computer science | Broadens PAI through the integration of software intelligence with hardware, materials, and mechanics, but provides limited detail on higher-level cognition and action mechanisms. |
| Balasubramani et al. [12] | PAI enables systems to learn from both real-time data and simulated environments, supporting adaptive control and ongoing improvement to maintain operational stability in complex industrial contexts. | Manufacturing | Emphasizes adaptive control and continuous learning in industrial environments, but gives relatively limited attention to material and morphological aspects of embodiment. |
| Wu et al. [16] | AI 3.0, or PAI, expands intelligence into the physical world by combining robotics, autonomous vehicles, and sensor-integrated control systems to operate under uncertain real-world conditions. | Intelligent construction | Provides a clear emphasis on sensing, actuation, and operation in uncertain physical environments, but discusses reasoning and learning mechanisms in less detail. |
| Bousetouane [5] | PAI agents are embodied intelligent systems developed to engage directly with the physical environment. | Healthcare, logistics, autonomous vehicles | Emphasizes embodiment and direct interaction with the physical environment, but provides limited detail on adaptive learning and higher-level cognitive reasoning. |
| Agarwal et al. [17] | PAI refers to AI systems equipped with sensors and actuators, where sensors enable environmental perception and actuators enable physical interaction and modification of the environment. | Robotic manipulation, Virtual worlds | Clearly defines PAI through sensing and actuation, but places less emphasis on planning, reasoning, cognitive architecture, and material embodiment. |
| Rank | Source | Frequency | Relative Frequency (%) |
|---|---|---|---|
| 1 | IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) | 26 | 8.20 |
| 2 | CHI Conference on Human Factors in Computing Systems (CHI) | 20 | 6.31 |
| 3 | Advances in Neural Information Processing Systems (NeurIPS) | 17 | 5.36 |
| 4 | European Conference on Computer Vision (ECCV) | 12 | 3.79 |
| 5 | IEEE Robotics and Automation Letters (RA-L) | 11 | 3.47 |
| 6 | IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) | 9 | 2.84 |
| 7 | IEEE International Conference on Robotics and Automation (ICRA) | 9 | 2.84 |
| 8 | ACM International Conference on Multimedia (MM) | 8 | 2.52 |
| 9 | IEEE/CVF International Conference on Computer Vision (ICCV) | 7 | 2.21 |
| 10 | Conference on Robot Learning | 6 | 1.89 |
| 11 | International Conference on Information Fusion (FUSION) | 6 | 1.89 |
| 12 | International Joint Conference on Artificial Intelligence | 4 | 1.26 |
| 13 | AAAI Conference on Artificial Intelligence | 4 | 1.26 |
| 14 | Artificial Life | 4 | 1.26 |
| 15 | Sensors | 3 | 0.95 |
| 16 | Advanced Engineering Informatics | 3 | 0.95 |
| 17 | IEEE Transactions on Emerging Topics in Computational Intelligence | 3 | 0.95 |
| 18 | Frontiers in Neurorobotics | 3 | 0.95 |
| 19 | ACM/IEEE International Conference on Human-Robot Interaction | 3 | 0.95 |
| 20 | Journal of Field Robotics | 3 | 0.95 |
| Cited Reference | Cluster | TLS | Cross-Cluster Strength | Cross-Cluster Ratio |
|---|---|---|---|---|
| Savva et al. [36] | 1 | 972 | 578 | 59.5% |
| Anderson et al. [53] | 2 | 861 | 547 | 63.5% |
| Radford et al. [46] | 3 | 590 | 451 | 76.4% |
| Xia et al. [42] | 4 | 521 | 398 | 76.4% |
| Kolve et al. [41] | 4 | 392 | 311 | 79.3% |
| Layer | Topic | Keywords | Proportion |
|---|---|---|---|
| Physical Interaction & Infrastructure | 1. Embodied AI | system, robot, physical, intelligence, interaction, framework, sensor, control, autonomous, platform | 19.1% |
| 2. Human–Robot Interaction | human, interaction, user, social, trust, experience, communication, perception, assistant, emotional | 13.4% | |
| Policy Learning & Control | 3. Semantic Navigation | navigation, object, scene, goal, map, representation, spatial, obstacle, semantic, exploration | 13.6% |
| 4. Simulation-Based Learning | simulation, simulator, training, policy, state, benchmark, performance, visual, data, control | 20.0% | |
| Cognitive Integration & Multimodal Reasoning | 5. Vision-Language Navigation | vln, vision, language, instruction, egocentric, communication, goal, dataset, behavior, action | 5.0% |
| 6. Language-Grounded Action | language, action, instruction, reasoning, planning, llm, knowledge, navigation, performance, prediction | 21.8% | |
| 7. Multimodal Learning | visual, language, llm, knowledge, representation, feature, perception, decision, multi, cross | 7.1% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Maeng, K.; Jin, H.; Kim, M. Mapping the Knowledge Structure of Physical Artificial Intelligence: A Data-Driven Systematic Review. Appl. Sci. 2026, 16, 9306. https://doi.org/10.3390/app16189306
Maeng K, Jin H, Kim M. Mapping the Knowledge Structure of Physical Artificial Intelligence: A Data-Driven Systematic Review. Applied Sciences. 2026; 16(18):9306. https://doi.org/10.3390/app16189306
Chicago/Turabian StyleMaeng, Kyuho, Hyeonjun Jin, and Minjun Kim. 2026. "Mapping the Knowledge Structure of Physical Artificial Intelligence: A Data-Driven Systematic Review" Applied Sciences 16, no. 18: 9306. https://doi.org/10.3390/app16189306
APA StyleMaeng, K., Jin, H., & Kim, M. (2026). Mapping the Knowledge Structure of Physical Artificial Intelligence: A Data-Driven Systematic Review. Applied Sciences, 16(18), 9306. https://doi.org/10.3390/app16189306

