Next Article in Journal
Sensitivity Enhancement of Weak Reflection Signals in Reflectometry via Concurrent Pulse Superposition
Previous Article in Journal
In-Situ Stress Measurement of Lower Rock Mass and Rockburst Prediction of Underground Caverns in Zhumadian Compressed Air Energy Storage Power Station
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

Mapping the Knowledge Structure of Physical Artificial Intelligence: A Data-Driven Systematic Review

1
Graduate School of Technology and Innovation Management, Hanyang University, 222 Wangsimni-ro, Seongdong-gu, Seoul 04763, Republic of Korea
2
Department of Industrial Engineering, Kumoh National Institute of Technology, 60 Daehak-ro, Gumi 39177, Republic of Korea
3
Department of Industrial Engineering, Ajou University, 206 World Cup-ro, Yeongtong-gu, Suwon 16499, Republic of Korea
*
Author to whom correspondence should be addressed.
These authors contributed equally to this work.
Appl. Sci. 2026, 16(18), 9306; https://doi.org/10.3390/app16189306 (registering DOI)
Submission received: 18 August 2026 / Revised: 13 September 2026 / Accepted: 14 September 2026 / Published: 19 September 2026
(This article belongs to the Section Computing and Artificial Intelligence)

Abstract

Physical artificial intelligence (PAI) has emerged as a transformative paradigm that integrates AI into physical entities, enabling direct interactions with real-world environments. However, despite rapid expansion across diverse domains, PAI research has remained highly fragmented and failed to provide a comprehensive understanding of its overarching knowledge structure. To address this gap, this study conducted a data-driven systematic review of 317 publications indexed in the Web of Science between September 2020 and October 2025. For the analysis of annual publication volume, the growth trend was assessed using complete calendar-year observations from 2021 to 2024, while the 2025 publication count was reported separately as a partial-year observation through October. Combining bibliometric network analysis with latent Dirichlet allocation topic modeling, we identified seven latent research topics. We integrated these fragmented topics into a unified, three-layered hierarchical architecture encompassing (1) physical interaction and infrastructure, (2) policy learning and control, and (3) cognitive integration and multimodal reasoning. The temporal analysis revealed a distinct evolutionary trajectory, indicating a structural shift from simulation-based, navigation-centric studies toward greater cognitive and multimodal integration and the practical implementation of embodied physical systems. This study provides a quantitative and structural mapping of PAI, offering a foundational framework to inform future interdisciplinary research and technological convergence.

1. Introduction

In recent years, artificial intelligence (AI) has evolved beyond cognitive and generative models to agentic systems [1,2], and its scope has rapidly expanded into physical AI (PAI), which integrates AI technologies into physical entities and enables direct interactions with real-world environments [3,4]. This transition indicates that AI is progressing beyond its initial capacity for information processing and decision-making support, evolving into intelligent systems that can perform integrated perception, reasoning, and action within the physical world [5]. Generally, PAI refers to systems in which AI is embodied within physical entities, enabling them to perceive the real world through sensors and actuators and autonomously make decisions and act, thereby interacting organically with their environment [6,7]. Rather than a singular technology or specific algorithm, PAI is a complex, multilayered paradigm that integrates elements from robotics, cyber–physical systems, autonomous systems, embodied AI, and soft robotics [8,9,10,11], and research on PAI covers a range of application domains, including manufacturing systems, autonomous driving, medical and service robots, and logistics automation [5,12,13].
Although these paradigms overlap, they emphasize different aspects of intelligence in the physical world. Cyber–physical systems (CPSs) emphasize the integration of computational and control functions with physical processes [10], whereas Embodied AI focuses on physical embodiment and agent–environment interaction as foundations for perception, cognition, and action [3]. Autonomous systems emphasize autonomous decision-making and control [11], while soft robotics highlights adaptive physical interaction through flexible and compliant physical structures [6,8,14]. PAI can be understood more broadly as a system-level paradigm that integrates perception, learning and reasoning, decision-making, control, and physical action under real-world physical constraints [6,14,15]. Accordingly, PAI provides a broader system-level perspective that connects embodied interaction, cyber–physical integration, autonomous decision-making, and adaptive physical systems within a unified framework for intelligence in the physical world [6,14,15]. In the present study, however, this broader conceptual framing should be distinguished from the operational scope of the empirical corpus, which comprises publications explicitly associated with the terms “Physical AI” or “Embodied AI.” The resulting knowledge structure therefore characterizes this explicitly identified literature rather than the full range of adjacent research that may contribute to the broader conceptual domain of PAI.
However, due to the inherently multidisciplinary and application-oriented nature of PAI, the literature can be located at the intersection of multiple academic fields including computer science, mechanical engineering, control engineering, and human–robot interactions [9,14]. Due to this diversity, PAI research has accumulated around individual subproblems or specific application scenarios, limiting an integrated perspective of the entire research landscape and a systematic classification framework [15]. As a result, PAI-related research topics and related terminologies are dispersed across various fields, with identical or analogous concepts often reappearing in different research contexts under varying terms [7]. Table 1 summarizes the evolving definitions and primary themes of PAI as discussed in the recent literature.
Recent studies have increasingly attempted to define the scope and structure of Embodied AI and PAI. Liu et al. [3] emphasize embodied perception, interaction, agents, and sim-to-real adaptation, while Ray [15] presents PAI as an integrated system combining perception, reasoning, and actuation. These perspectives are broadly consistent with our finding that PAI comprises interconnected layers of physical interaction, policy learning and control, and cognitive integration. However, unlike these primarily conceptual and technology-oriented reviews, the present study derives the field’s knowledge structure empirically from 317 publications using bibliometric analysis and LDA topic modeling. This data-driven approach complements existing taxonomies by quantifying the relative prominence and temporal variation in major PAI research themes and by synthesizing the identified topics into an empirically informed higher-order framework.
Despite these efforts, the fragmented and dispersed nature presents practical constraints on a systematic understanding of the field’s overall knowledge structure and technological trends. Furthermore, despite a partial understanding of individual domains, meta-level analyses using quantitative and structural methods to examine the thematic architecture and interconceptual relations remain scarce [15]. These limitations have become increasingly critical as PAI rapidly expands into diverse industrial sectors such as manufacturing, robotics, services, healthcare, and logistics. With the acceleration of technological development and application, a comprehensive understanding of the overall structure and flow of research that extends beyond individual achievements is needed to set future research directions, explore possibilities for interdisciplinary convergence, and support strategic decision-making.
To address these gaps, this study aims to systematically map the knowledge structure and technological evolution of PAI research through a data-driven literature review. To this end, a corpus of 317 PAI-related studies was constructed from the Web of Science (WoS) database through a systematic screening process. Building on this dataset, this study aims to address three primary research objectives: identifying the key topic structure of PAI research (RQ1), systematizing core concepts into higher-level technological dimensions (RQ2), and investigating the evolutionary patterns of these structures (RQ3).
This study offers three main contributions. First, we comprehensively identified the latent knowledge structure of PAI research, providing a quantitative and structural understanding of the field. Second, we proposed a unified framework for PAI and presented a systematic multilayered technological structure by integrating fragmented topics and concepts into three high-level dimensions (physical interaction, policy learning, and cognitive integration). Third, we empirically clarified the evolutionary trajectory of PAI research by analyzing temporal changes in topics, demonstrating a transition from a simulation- and navigation-centric phase to integrated embedded system implementation.
The rest of this paper is structured as follows. Section 2 presents the research methodology, including the literature collection process and the dual-analysis framework. Section 3 presents the results of the study, first analyzing the bibliometric knowledge landscape of PAI through publication trends, keyword co-occurrence networks, and document co-citation networks, and then reporting the results of LDA topic modeling to identify the seven core research topics and their integration into a three-layered hierarchical architecture. Section 4 discusses these findings by reconceptualizing PAI as a dynamically integrated, multilayered system in which physical implementation, policy learning, and cognitive integration co-evolve, and outlines implications for future research. Finally, Section 5 synthesizes the results and presents concluding remarks.

2. Methods

To systematically examine the knowledge structure and research trends of PAI, we adopted a data-driven methodology combining bibliometric network analysis and topic modeling. This review was not registered in a public database, and post hoc registration was not pursued. The literature identification and selection process is presented in Figure 1. Figure 2 presents the study’s methodological framework consisting of three steps. First, an exploratory literature review was conducted to identify key terms that were conceptually linked to PAI, which served as the basis for the construction of the corpus from academic databases. Second, bibliometric network analysis was performed on the data to analyze keyword co-occurrence and co-citation structures. Third, topic modeling was applied to obtain key research topic structures and reorganize them from a high-level technological architecture perspective.

2.1. Article Collection and Corpus Creation (Step 1)

Relevant literature was collected from the WoS database. The search period spanned from September 2020 to October 2025. This timeframe was selected because the term “PAI” gained significant academic traction from late 2020 onward. The WoS database is highly suitable for bibliometric reviews and provides rigorous journal selection criteria, refined citation metadata, and an optimal structure for citation network analysis [18,19].
PAI currently lacks a universally agreed-upon conceptual framework and remains an emerging field spanning several adjacent research streams. A preliminary literature review was conducted to identify three recurring conceptual characteristics: interaction with the physical environment, embodied intelligence, and integrated perception–action processes. Based on these characteristics, this study selected “Physical AI” and “Embodied AI” as representative search terms. WoS was selected because it provides standardized bibliographic and citation metadata and systematically indexed scholarly records that are well suited to bibliometric and co-citation analyses [20]. Moreover, previous methodological studies suggest that using a single established database can reduce inconsistencies arising from cross-database integration [21], while comparative research has reported broadly consistent bibliometric patterns between WoS and Scopus [22]. To maintain a focused conceptual scope and ensure reproducibility, the complete WoS search query was defined as TS = (“Embodied AI” OR “Physical AI”) AND DOP = (1 September 2020/31 October 2025), where TS denotes the Topic field and DOP specifies the publication-date range. The search results were further restricted to the document types “Article” and “Proceedings Paper,” and no other filters were applied. To maintain a focused conceptual scope, the final search query was defined as (‘Embodied AI’ OR ‘Physical AI’). The identification and selection of records were reported in accordance with PRISMA 2020 [23], and the completed checklist is provided as S1 Checklist. The review was not registered, and no protocol was prospectively prepared. One reviewer assessed records for corpus inclusion by applying the predefined eligibility criteria to the bibliographic metadata, titles, and abstracts retrieved from Web of Science. Full texts were not retrieved for a separate eligibility assessment because the criteria used to construct the analytical corpus could be determined from these record-level data. Bibliographic metadata, titles, abstracts, author keywords, and cited references were exported directly from Web of Science.
The search retrieved 341 records. After patents and nonacademic literature were excluded, 333 journal and conference papers remained. Finally, 16 review articles were excluded to prevent secondary syntheses from disproportionately influencing the bibliometric and topic-modeling results, yielding a final analytical corpus of 317 records.
For both the informetric network analysis and LDA topic modeling, the corpus was constructed using only titles and abstracts. This approach was adopted because collecting full texts poses consistent challenges due to heterogeneity in formats across databases, accessibility constraints, and limited integration with bibliometric analysis tools. Existing bibliometric and topic modeling studies on emerging technologies thus frequently adopt a title- and abstract-centric approach [24,25].

2.2. Informetric Network Analysis (Step 2)

Initially, descriptive statistics such as publication year and thematic distribution were analyzed. Subsequently, a keyword co-occurrence network was constructed to examine structural relationships among core concepts. This network represented each keyword as a node. A co-occurrence relationship was established when two keywords appeared within the same document. The strength of this relationship was determined by co-occurrence frequency across publications, which reflects the conceptual centrality of the keywords within the field [26].
A two-dimensional co-word matrix was constructed to quantify these relationships by recording the co-occurrence frequency between all keyword pairs. VOSviewer (version 1.6.20; Centre for Science and Technology Studies, Leiden University, Leiden, The Netherlands) was used to identify keyword clusters and examine the structural grouping of concepts. These clusters were reviewed against source journals and disciplinary distributions to clarify thematic alignments.

2.3. Topic Modeling (Step 3)

Topic modeling is a text mining technique that automatically extracts latent topics from large-scale unstructured text. It is highly effective for understanding research trends in a literature collection [27,28,29]. This study adopted latent Dirichlet allocation (LDA), which explicitly provides topic–word probability distributions, ensuring comparability and interpretability across topics [27]. LDA is a probabilistic generative model that assumes each document contains a finite number of topic mixtures and that each topic is defined by a probability distribution over words. Model estimation relies on Gibbs sampling–based inference to approximate the posterior distribution of latent variables [30].
Recent developments in topic modeling include BERT-based approaches and large language models (LLMs) [31]. However, these approaches require significant computational resources and lack probabilistically specified topic–word distributions, limiting cross-topic comparability [32,33]. In particular, LLM-based methods often exhibit inconsistent results depending on prompts and generation settings, thereby reducing methodological transparency [33]. LDA thus remains the standard for bibliometric reviews of long academic texts and is highly suitable because of its emphasis on interpretability and reproducibility [34,35]. Topic modeling was performed on the final corpus of 317 PAI-related papers, and the top 10 words for each topic were extracted and visualized for interpretation.
All topic-modeling analyses were conducted in Python (version 3.13.15). LDA modeling and vectorization were performed using scikit-learn (version 1.6.1), coherence scores were calculated using Gensim (version 4.4.0), and text preprocessing was performed using NLTK (version 3.9.1). NumPy (version 2.1.3), SciPy (version 1.16.3), pandas (version 2.2.3), and Matplotlib (version 3.10.0) were used for numerical computation, scientific computing, data processing, and visualization, respectively.

3. Results

Although topic modeling constitutes the core analytical contribution of this study, a bibliometric network analysis was conducted to provide an empirical overview of the field’s structural landscape. This preliminary mapping enables the subsequent topic modeling results to be interpreted within an established bibliometric context.

3.1. Results of Bibliometric Networks

3.1.1. Descriptive Statistics

To map the temporal and geographical distribution of the reviewed literature, this study analyzed annual publication trends along with country-specific research activities and collaboration structures. The annual publication volume from 2020 to 2025 exhibits a clear upward trajectory, recorded as 5 → 26 → 35 → 53 → 99 → 99. During the complete calendar years from 2021 to 2024, annual publication counts increased from 26 to 99, with 35 publications in 2022 and 53 in 2023. The growth trend was therefore assessed using these complete-year observations. Notably, the number of publications increased substantially in 2024, reaching 99, approximately 1.9 times the 2023 total. By October 2025, the publication count had already reached 99, matching the full-year total for 2024 and indicating that research activity remained high during the partial 2025 observation period. Because the 2025 value represents only January-October, it was not used for a direct year-on-year growth comparison. Overall, these publication patterns indicate a marked expansion of PAI-related research activity during the study period, particularly from 2023 onward.
In addition to the publication trends, country- and institution-level collaboration networks were examined to characterize the broader research landscape. The country-level network showed that international collaboration was centered primarily on the United States and China, with connections extending to several other research-active countries. The institution-level network likewise revealed multiple interconnected research groups, although collaboration was concentrated around a relatively limited set of institutions. Because the primary focus of this study is the knowledge and thematic structure of PAI, detailed collaboration-network visualizations are provided in Supplementary Figures S1 and S2.
Table 2 demonstrates that PAI studies are not confined to a single academic discipline. Instead, the literature is widely distributed across major international conferences and journals in computer vision, human–computer interactions, machine learning, and robotics. This distribution highlights that PAI represents a multidisciplinary research field, integrating diverse technologies such as visual perception, large-scale language models, Human–Robot Interaction (HRI), and autonomous robotics.

3.1.2. Visualization of the Keyword Co-Occurrence Network

Within this network (Figure 3), specific thematic clusters emerge to define the core components of the field. Concepts such as “navigation,” “reinforcement learning” and “training” form a distinct cluster, highlighting the critical importance of policy learning and path optimization within physical environments. Conversely, keywords like “large language model,” “transformers,” “computer vision,” and “vision-and-language navigation (VLN)” maintain strong ties to the center despite their relatively lower individual frequency. This pattern indicates that language- and vision-based models are increasingly serving as a pivotal integrative layer, facilitating the integration of advanced cognitive processing with behavioral control. This is further complemented by the presence of terms such as “deep learning,” “dataset,” “IoT,” “automation,” and “trust” within the AI cluster, which connects PAI discussions to broader infrastructure and real-world implementation. Notably, the emergence of “trust” reflects a growing academic focus on the safety and reliability of embodied systems, while the term “physical AI” remains directly linked to the AI cluster and adjacent to “embodied AI,” reaffirming its alignment with embedded systems and contemporary AI research trends.
Overall, the structural delineation of the network confirms that PAI is advancing across three interconnected knowledge layers that constitute a tightly integrated stack, marking a transition from conceptual exploration into a phase of substantive empirical integration. This hierarchical progression begins with the foundational Physical Interaction & Infrastructure layer, which serves as the digital and physical bedrock for the field. This layer focuses on the high-fidelity replication of three-dimensional environments and the modeling of physical spaces into learnable digital forms. Its primary role is to provide the necessary environmental representation and experimental infrastructure—such as sensors and simulation platforms—ensuring that agents can interact within a world that accurately reflects real-world physical constraints. Furthermore, this foundation establishes the essential framework for trust and safety, which are critical prerequisites for the eventual transition from simulated environments to real-world deployment.
Building directly upon this foundational base, the functional Policy Learning & Control layer acts as the operational engine that translates environmental perception into purposeful action. This layer is dominated by reinforcement learning algorithms, policy optimization theories, and semantic reasoning frameworks that enable agents to acquire robust behavioral skills. It addresses the core challenges of navigation and path planning, ensuring that an agent’s movements are not only physically feasible but also functionally optimized for goal-directed exploration. The structural integration observed at this level suggests that the acquisition of effective behavioral policies is intrinsically tied to the agent’s ability to reason about its spatial surroundings and the fidelity of the underlying training infrastructure.
Ultimately, the stack is capped by the Cognitive Integration & Multimodal Reasoning layer, which provides the high-level intelligence required for complex decision-making. This layer integrates advanced cognitive processing with behavioral control by facilitating the alignment between linguistic instructions and visual data. Driven by the advancement of LLMs and Transformers, it focuses on language-grounded intelligence and cross-modal reasoning, allowing agents to execute sophisticated tasks through a holistic understanding of vision, audio, and other sensory inputs. These layers do not operate in isolation but function as a cohesive, hierarchical architecture with “embodied AI” acting as the central mediator. This structural convergence confirms that PAI research has moved into a highly integrated phase, where physical interaction, policy optimization, and cognitive expansion are seamlessly combined into a unified, multilayered technological paradigm.

3.1.3. Visualization of the Co-Citation Network of Cited References

To systematically map the intellectual structure and technological evolution of PAI, document co-citation analysis was conducted with a citation threshold of 10. The resulting network (Figure 4) reflects the core knowledge base and research lineage across the field, identifying four main clusters that converge into the hierarchical architecture established in the keyword analysis.
The red cluster represents the foundational Physical Interaction & Infrastructure layer, centered on the Habitat platform [36]. This cluster encompasses seminal works on high-precision three-dimensional environmental reproduction, simulation-based experimental setups, and sim-to-real approaches. By modeling physical spaces as digital environments to create learnable forms [37,38,39,40], these studies confirm that environmental representation and simulation infrastructure serve as the essential bedrock for PAI research. This foundation establishes the physical parameters in which intelligent agents must operate, providing the necessary sensory and spatial constraints that define the limits of autonomous interaction.
The yellow and green clusters together constitute the Policy Learning & Control layer, addressing how agents execute tasks within those environments. The yellow cluster focuses on simulation environments, visual datasets, and learning infrastructures that support the training and execution of realistic behavioral policies [41,42,43,44], addressing task-level challenges such as goal-directed movement and language-guided exploration. Complementing this, the green cluster centers on policy optimization and representation learning theory, specifically reinforcement learning-based approaches [45,46,47,48]. The structural overlap and close proximity between these two clusters indicate that the integration of environmental action learning with robust policy optimization algorithms serves as a functional pillar of PAI research, enabling agents to acquire optimized behavioral skills.
The blue cluster signifies the Cognitive Integration & Multimodal Reasoning layer, driven by pretrained language models and multimodal representation learning [49,50,51,52]. This cluster focuses on large-scale, data-driven cross-modal alignment, reflecting efforts to construct a shared representational space for linguistic and visual information. Rather than directly generating action policies, this layer provides the foundational mechanisms for semantic representations and multimodal alignment, enabling higher-level reasoning.
The cluster cohesion analysis showed a weighted modularity of 0.166 and an overall silhouette score of 0.339, indicating partial cluster separation alongside substantial inter-cluster connectivity. To further examine the integrative structure of the co-citation network, cross-cluster link strength was analyzed to identify bridging references (Table 3). Savva et al. [36], Anderson et al. [53], Radford et al. [46], Xia et al. [42], and Kolve et al. [41] showed high cross-cluster connectivity, with cross-cluster ratios ranging from 59.5% to 79.3%. These results indicate that major references extend across cluster boundaries, suggesting that the co-citation clusters form an interconnected rather than structurally isolated knowledge structure.
A comprehensive evaluation of these cluster interconnections reveals that the knowledge-level citation structures correspond directly to the three technological layers identified in the keyword co-occurrence network (Section 3.1.2). The alignment between the conceptual discourse and the actual citation network confirms that PAI is advancing as a tightly integrated stack, where physical interaction and infrastructure, policy learning and control, and cognitive integration and multimodal reasoning converge into a single, multilayered technological paradigm. This hierarchical architecture reinforces PAI’s maturity as a convergent field that draws on the systematic synergy between physical environment modeling, learning-based control, and cognitive model integration.
To further quantify the core intellectual structure of the co-citation network, citation counts, number of links, total link strength (TLS), and cluster membership were examined for the most strongly connected cited references. As shown in Table S1 (Supplementary Section), Savva et al. [36] exhibited the highest TLS (972), followed by Anderson et al. [53] (TLS = 861) and Radford et al. [46] (TLS = 590). Chaplot et al. [37] and Xia et al. [42] also showed high TLS values of 523 and 521, respectively. These quantitative results support the network-based interpretation by showing that simulation environments and embodied navigation constitute a central intellectual foundation of PAI research, while the strong connectivity of Radford et al. [46] reflects the structural importance of vision–language representation within the cognitive and multimodal dimension.

3.2. Results of LDA Topic Modeling and Discovery

Following the data construction process detailed in Section 2.1, the LDA analysis was performed on the finalized corpus of 317 titles and abstracts to ensure consistency with the preceding network analysis. To determine an appropriate number of topics, coherence scores were used as an initial quantitative criterion rather than as the sole basis for model selection. Coherence scores quantitatively assess semantic relevance among the most frequent keywords constituting each topic, offering a useful measure of topic interpretability relative to probability-based metrics such as perplexity [54]. To assess the sensitivity of the topic-number selection, coherence scores were calculated across K = 3–30 (Figure 5). The results showed that coherence varied across different topic solutions. Although some lower-K solutions showed higher coherence scores, these solutions tended to combine conceptually distinct research streams into broader themes, thereby reducing thematic resolution. The seven-topic solution provided a more differentiated representation of the major PAI research areas while maintaining favorable coherence relative to adjacent solutions. In particular, the coherence scores for K = 6, K = 7, and K = 8 were 0.3460, 0.3641, and 0.3571, respectively. Accordingly, K = 7 was selected based on a balance between quantitative coherence and substantive interpretability rather than on coherence maximization alone. Table 4 presents the results of the topic modeling analysis conducted with K = 7.
The identification of these seven distinct research subareas reveals a clear technological hierarchy within PAI. Rather than existing as fragmented domains, these topics organically converge into three higher-level knowledge layers that reflect the field’s substantive operational flow. The analysis first identifies Topic 1 (Embodied AI) and Topic 2 (Human–Robot Interaction) as the foundational pillars of the field. These topics, which focus on the “Sensors-Think-Act” cycle and the social dimensions of trust and perception, collectively constitute the Physical Interaction & Infrastructure layer. This finding confirms that PAI research is fundamentally rooted in the development of real-world physical systems and the hardware-based interaction frameworks necessary for autonomous operation.
Building upon this physical foundation, a functional cluster emerges through Topic 3 (Semantic Navigation) and Topic 4 (Simulation-Based Learning). Topic 4, in particular, holds a significant research proportion, reflecting the critical role of high-performance simulators in accelerating the training of behavioral policies. Together with Topic 3, which addresses spatial reasoning and object-centric mapping, these areas form the Policy Learning & Control layer. This synthesis indicates that PAI inherently requires a robust decision-making engine to translate environmental understanding into optimized action policies across both simulated and physical spaces.
The field’s expansion into advanced intelligence is further characterized by Topic 5 (Vision-Language Navigation), Topic 6 (Language-Grounded Action), and Topic 7 (Multimodal Learning). Notably, Topic 6 accounts for the largest share of the research, highlighting the current dominance of LLM-based reasoning and hierarchical planning in the literature. By aligning linguistic instructions and diverse sensory inputs with physical behavior, these topics establish the Cognitive Integration & Multimodal Reasoning layer. This layer acts as the integrative cap of the PAI stack, facilitating sophisticated goal-directed reasoning and complex environmental understanding.
To quantitatively examine changes in the thematic composition of PAI research, annual topic prevalence was calculated as the mean document-level topic proportion across publications in each year from 2021 to 2025 (Figure 6). The 2025 values represent partial-year data covering publications indexed through October. Topic 4 (Simulation-Based Learning) showed a high prevalence in 2021–2022 (25.3–27.0%) and subsequently declined to 14.7% in 2025. Similarly, Topic 3 (Semantic Navigation) peaked at 22.4% in 2023 before declining to 11.0% in 2024 and 9.4% in 2025. In contrast, Topic 6 (Language-Grounded Action) maintained a relatively high prevalence throughout the study period, while Topic 7 (Multimodal Learning) showed higher proportions in 2024–2025 than in preceding years. Notably, Topic 1 (Embodied AI) increased from 16.8% in 2024 to 28.9% in 2025. Overall, these patterns suggest a shift in thematic emphasis: simulation- and navigation-oriented topics became relatively less prominent in the later years, whereas multimodal learning and embodied-system research gained greater representation. Given the partial coverage of 2025, these findings are interpreted as descriptive temporal patterns rather than as evidence of a definitive structural transition.
The temporal patterns observed in the topic-prevalence analysis also coincided with several important technological developments in embodied and multimodal AI, including the emergence of multimodal foundation models, vision-language-action models such as RT-2, and advanced embodied simulation platforms such as Habitat 3.0. Although the present analysis does not establish causal relationships between these developments and changes in topic prevalence, their timing provides relevant technological context for the observed shifts in research emphasis.
In summary, the transition from these seven specific research topics to a unified hierarchical stack transcends simple statistical clustering to reflect the convergent technological paradigm of the field. As illustrated in Figure 7, this structure provides a comprehensive map of how physical interaction, learning-based control, and cognitive intelligence are systematically integrated into a single, multilayered architecture. The following subsections provide a detailed examination of the key research themes and challenges identified within each of these functional layers.

3.2.1. Layer 1: Physical Interaction & Infrastructure

Topic 1: Embodied AI
Embodied AI transcends virtual information processing, integrating perception, cognition, and action through physical interactions with the real environment [55,56]. Grounded in the embodied cognition hypothesis, it posits that intelligence stems from bodily experiences and sensorimotor loops, emphasizing an agent’s capacity to adapt and execute goal-directed behaviors [57,58]. Recently, foundation models leveraging LLMs and visual-language models have successfully bridged complex natural language instructions with real-world physical control loops [15,59].
Modern embodied AI systems often employ a hierarchical, modular framework that mimics human cognitive structures by organically integrating perception, memory, communication, planning, and execution modules [55,60]. Technologically, robotic transformer models such as RT-1 and RT-2 have internalized multimodal data to achieve generalized control across diverse environments [61]. Furthermore, the integration of state-space models like Mamba facilitates real-time inference and efficient motion control on resource-constrained edge devices [59,62]. This is complemented by the active inference framework, which enhances autonomy by allowing agents to explore environments and update internal models while minimizing uncertainty [63,64].
These architectures have yielded tangible results across various industrial applications. For instance, the DEXBOT framework executes precise assembly tasks in unstructured construction environments, while the FlexiFly drone system provides reconfigurable sensing for smart homes and laboratories [65,66]. Humanoid agents have also been deployed for decontamination in high-risk bio-laboratories to improve operational safety and efficiency [67]. To ensure real-time computational efficiency, architectures such as the cloud-fog embedded framework process resource-intensive computations in the cloud while delegating immediate control responses to the edge-proximate fog layer [56]. Ultimately, embodied AI establishes a technological foundation for autonomous agents that can collaborate with humans within the physical constraints of the real world [15,55].
Topic 2: Human–Robot Interaction
HRI represents an ontological transition within the modern technological ecosystem from a “design stance” of mere machine control to an “intentional stance” that attributes intentions and beliefs to robots [68]. The increasing ambiguity of robots’ social roles is illustrated by incidents involving the humanoid robot Sophia and public debates regarding AI sentience [69,70]. These developments suggest that as robots become more physically and cognitively integrated into human environments, the nature of the relationship shifts from functional tool usage to social partnership.
The underlying mechanisms of these interactions depend on two core dimensions of perception: agency and experience [71]. Humans readily attribute agency to robots—perceiving them as entities capable of planning and action—but remain hesitant to attribute the capacity for experience, such as the ability to feel pain or pleasure. For example, children interacting with augmented reality applications like BeeTrap often understand the system’s behavioral logic without perceiving it as an emotional entity [72]. Conversely, the formation of deep emotional attachments between users and social chatbots, such as Replika, suggests that intimacy with AI is generating novel interaction dimensions that challenge traditional boundaries [73].
Despite its growing scalability, HRI faces inherent psychological and structural limitations. The cognitive gap defining robots as entities that act but do not feel creates a “paradox of moral agency,” where robots may face disproportionate moral expectations without possessing corresponding rights [71]. Furthermore, cognitive friction arises when robots blur the human–machine boundary, leading to delayed responses and reduced interaction efficiency [74,75]. This friction is particularly evident in medical settings, where professionals often struggle to classify AI diagnostic systems as either peer advice or mere tools, complicating the process of trust-building [76].
Structurally, HRI introduces risks of power asymmetry and deskilling [77]. In fields such as medicine or art therapy, over-reliance on AI as a co-creative partner may threaten human professional competency and agency [76,78]. Additionally, affective computing systems that collect emotional data raise critical ethical concerns regarding privacy and potential data misuse [79]. Therefore, addressing these psychological gaps and social inequalities is essential to advance HRI and foster genuine trustworthiness in human–robot partnerships [79].

3.2.2. Layer 2: Policy Learning & Control

Topic 3: Semantic Navigation
With rapid advancements in embedded AI, spatial navigation is transitioning from geometric coordinate-based SLAM toward semantic navigation, which comprehends object–environment relationships [80]. Unlike conventional systems constrained by obstacle avoidance and shortest-path heuristics, semantic navigation reconstructs environments as relational networks to execute complex natural language instructions [81]. Scene graphs and object-centric world models play central roles in this framework, enabling agents to selectively extract and infer task-critical information across expansive, multiroom settings [82,83]. Methodologically, frameworks such as hierarchical relational object navigation utilize spatial hierarchies encompassing rooms, furniture, and objects [84]. These systems store topological relationships and employ task-driven attention mechanisms to filter extraneous data [85], while contextual bandit algorithms dynamically optimize paths via commonsense probability estimates [86].
These architectures, while sophisticated, present critical limitations that impede real-world mastery. The stochastic nature of physical habitats, where objects frequently shift due to human intervention, introduces substantial uncertainty [87]. Additionally, multiobject navigation suffers from local optimization inefficiency, often generating duplicate paths [88]. From a computational perspective, determining optimal visitation sequences scales into an NP-hard weighted minimum latency problem, which significantly hinders real-time responsiveness [89]. Finally, a persistent sim-to-real gap remains a major hurdle [90]; real-world variations in lighting and occlusion often discourage the flawless semantic segmentation that is easily acquired in simulated environments [91]. Consequently, despite the significance of semantic navigation in developing context-aware agents, ensuring adaptability and computational efficiency in dynamic environments remains a crucial challenge for future research [92].
Topic 4: Simulation-Based Learning
Due to physical constraints and safety risks associated with real-world training, simulation-based learning (SBL) serves as the primary developmental foundation for intelligent agents [93,94]. High-performance engines such as Megaverse and photorealistic simulators such as Habitat enable the processing of billions of frames, leveraging 3D-scanned real-world data to accelerate visual and navigational training [36,95].
However, the sim-to-real gap between idealized virtual models and stochastic physical environments severely limits practical deployment [96]. Agents frequently learn to exploit simulator imperfections, such as oversimplified friction models, leading to systematic failures upon physical transfer. While initiatives such as LoCoNav mitigate this limitation by enforcing rigorous collision constraints, the structural complexity of reward shaping constitutes a more fundamental bottleneck [38]. Designing granular, step-by-step rewards is functionally analogous to manual feature engineering, whereas scalable terminal reward structures suffer from low convergence rates due to nonstationarity in visually complex tasks [97].
Furthermore, current SBL frameworks predominantly use static environmental datasets, limiting the development of dynamic social navigation [42,98]. As demonstrated by HandoverSim, simple rigid-body representations fail to adequately model the nuances of human–robot object handovers [99,100]. Addressing these deficiencies requires the integration of mechanisms such as rapid motor adaptation to instantly adjust to environmental variability [101,102]. The use of ultra-high-speed sensory inputs, such as SpikeGS camera data, has emerged as a critical methodology to accelerate learning and enhance 3D scene reconstruction, thereby bridging the divide between simulated training and physical execution [60,103].

3.2.3. Layer 3: Cognitive Integration & Multimodal Reasoning

Topic 5: Vision-Language Navigation
Transcending mere object recognition, VLN demands precise visual-language grounding to map natural language instructions onto physical environments. Recent advancements have used LLMs for complex planning and state tracking, integrating proprioceptive and tactile data to execute sophisticated operations [104,105]. Notably, online goal inference algorithms allow agents to dynamically interpret ambiguous commands by synthesizing behavioral histories with contextual cues [106].
More specifically, Topic 5 is characterized by a navigation-oriented objective, in which linguistic instructions are grounded in visual and spatial observations to identify destinations, interpret spatial relations, and determine navigation paths [53,107,108]. Anderson et al. [53] established Vision-and-Language Navigation as the task of following visually grounded natural-language instructions in real environments, while Jain et al. [107] further emphasized instruction fidelity during navigation. More recently, NavCoT [108] incorporated explicit reasoning into vision-and-language navigation to improve navigation decisions. Accordingly, Topic 5 is distinguished by its primary focus on determining where an agent should move and how it should reach a spatial goal, rather than on executing a broader sequence of physical actions [53,107,108].
However, critical barriers to real-world deployment have been identified. Agents trained on synthetic instructions struggle to resolve the inherent ambiguity and contextual omissions in human communication [109]. Furthermore, situated reasoning, defined as the capacity to adapt to dynamic environmental changes, remains underdeveloped [106]. Visuomotor bottlenecks severely degrade performance as minor segmentation errors often cascade into catastrophic planning failures [104,109]. Additionally, high-precision industrial applications require physics-informed AI frameworks, as standard probabilistic models cannot guarantee the strict error tolerances needed to ensure physical safety [110].
Topic 6: Language-Grounded Action
Language-grounded action elevates command execution to hierarchical semantic manipulation of the physical world [53], aligning abstract linguistic directives with visual observations and physical constraints [107]. State-of-the-art frameworks use LLMs as cognitive world models to decompose long-horizon tasks, such as those in the ALFRED benchmark, into executable subtask sequences [111,112]. Advanced systems such as NavCoT use “future imagination” to infer obscured spatial layouts [108] while WorldAfford dynamically identifies functionally equivalent tool substitutes during environmental interactions [113].
In contrast to Topic 5, Topic 6 is characterized by action-oriented grounding, in which linguistic instructions are translated into executable physical behaviors and structured sequences of actions [108,112,114]. Its primary concern is therefore not only where an agent should move, but what actions it should perform and how those actions should be organized to accomplish a physical task [108,112]. ALFRED [112], for example, requires embodied agents to interpret natural-language instructions and execute multi-step household tasks involving navigation and object interaction, whereas WorldAfford [113] grounds natural-language instructions in object affordances to identify physically feasible interactions. P-RAG [114] further addresses planning for embodied everyday tasks by incorporating retrieved knowledge into action planning. Thus, although navigation may constitute an intermediate step, Topic 6 is distinguished from Topic 5 by its broader emphasis on transforming linguistic intent into executable task sequences that interact with or modify the physical environment [108,112,114].
However, these advanced capabilities have profound limitations, primarily stemming from the inherent gap between abstract linguistic symbols and the physical environment. While general-purpose LLMs are trained on vast textual corpora, they fundamentally lack the physical senses and spatiotemporal context of real-world 3D environments [108]. As a result, agents frequently generate action plans that violate physical constraints or fail to incorporate the task-specific knowledge unique to simulated environments [114]. Compounding this domain gap is the semantic loss and incomplete visual anchoring during cross-modal information transfer [46]. Translating visual observations into text often distorts critical details, leading to inaccurate decision-making. For instance, ambiguous spatial commands such as “go to a table at the end of the hallway” frequently result in erroneous target selection among multiple distractors [115]. Furthermore, current models lack real-time empirical updating functions. Systems trained on fixed datasets struggle to adapt to dynamic environmental changes [116]; although interaction-driven knowledge accumulation has been proposed, a comprehensive solution remains elusive due to performance saturation of LLM planners and physical constraint violations [114], as well as the exploitation of task-specific shortcuts and limited transferability outside specific simulations [116].
In conclusion, language-grounded action strives for true embodied intelligence, which seamlessly synchronizes abstract language with physical reality [53]. Realizing this vision requires the development of incremental knowledge acquisition architectures that organically integrate the foundational common sense of large-scale models with real-time interaction data [114,117]. Developing rigorous world models that precisely reflect physical constraints and spatiotemporal context would provide vital groundwork for intelligent agents that can freely manipulate the physical world through natural language [107,118,119].
Topic 7: Multimodal Learning
Modern AI research is rapidly evolving toward multimodal learning with the aim of building integrated cognitive systems that are analogous to the integration of human senses [120]. Central to this transition is the alignment of asymmetrical data modalities within a unified embedding space to facilitate context-aware inference, which allows agents to independently determine their 3D spatial orientation and position based on verbal prompts [121,122]. In terms of practical applications, vehicle-deployed vision-language models that use the LLAVA architecture have improved data transmission efficiency by over 90% by isolating essential semantic features from raw imagery [123]. Concurrently, audiovisual integration frameworks such as ALOHA dynamically calibrate spatiotemporal weights to synchronize signals, which has proved crucial for precise localization of noise sources and object segmentation in complex environments [124].
Unlike Topics 5 and 6, Topic 7 is characterized primarily by multimodal representation learning and cross-modal integration rather than by the completion of a specific navigation or physical-action task [120,121,122,124]. Its central concern is how heterogeneous information from vision, language, 3D spatial representations, audio, and other modalities can be aligned or fused to support broader perception and reasoning [120,121,122,124]. Flamingo [120], for example, integrates visual and linguistic information within a general-purpose multimodal learning framework, while 3D-LLM [121] incorporates three-dimensional environmental information into large language models to support spatial understanding. SQA3D [122] further demonstrates multimodal reasoning through situated question answering in 3D environments, while audiovisual segmentation [124] illustrates the integration of visual and auditory information for perceptual tasks. Accordingly, Topic 7 is distinguished by its emphasis on learning shared or integrated multimodal representations that can subsequently support navigation or action, rather than directly producing navigation trajectories or executable physical-action sequences [120,121,122,124].
Nevertheless, structural imbalances between global context comprehension and local detail extraction remain a critical roadblock [124]. Fusion mechanisms often misclassify fine-grained pixel details, such as confounding vocalizing focal subjects with their background during pixel-level segmentation [124]. Moreover, real-world long-tailed attribute distributions impede the recognition of rare object configurations, significantly restricting the scalability of multimodal knowledge graphs [124]. Crucially, the integration of heterogeneous inputs exponentially expands the system’s attack surface, thereby exacerbating semantic vulnerabilities [125]. Sophisticated prompt injection attacks, such as destination hijacking instructions embedded in text that strategically conflict with camera data, can bypass standard defenses and orchestrate physical accidents [126]. Ultimately, robust multimodal intelligence can be acquired through adaptive architectures that dynamically assign spatiotemporal weights to securely integrate external knowledge while rigorously preserving local contextual details [125,127].

4. Discussion

4.1. Three-Layered Hierarchical Architecture of PAI

Although prevailing definitions of PAI emphasize embodied intelligence, perception–action loops, and the integration of cognitive models such as LLMs, these elements have often been treated as loosely connected components or fragmented concepts across disparate domains. Specifically, the conventional perception–cognition–action framework, while effective in functionally distinguishing environmental sensing from reasoning and execution, often fails to capture the interdependent relationships among these components or the profound influence of physical constraints on the overall system. This study thus reconceptualizes PAI not as a single functional pipeline but as a multilayered, integrated system in which physical implementation, policy learning, and cognitive integration interact. By mapping the identified research topics onto this structure, we systematically clarify their structural relationships through a distribution that is empirically grounded in the research corpus.
At the foundational layer, which comprises Physical Interaction & Infrastructure (Topics 1 and 2), intelligence is grounded in material embodiment and real-world constraints. Here, agents perceive the environment through sensorimotor loops and generate executable actions in physical space, while HRI establishes the normative and social constraints—including trust, safety, and acceptance—necessary for operation within a broader sociotechnical context. This foundational intelligence is subsequently processed by the intermediate layer, Policy Learning & Control (Topics 3 and 4, 33.57%), which functions as a computational abstraction layer. By utilizing semantic navigation and SBL, this layer transforms embodied constraints into executable policies, operating as a translational mechanism that converts high-level cognitive objectives into physically realizable behaviors. Finally, the upper layer, Cognitive Integration & Multimodal Reasoning (Topics 5, 6, and 7, 33.89%), operates at the level of symbolic abstraction and cross-modal alignment. This layer performs the critical task of symbol–action alignment, linking abstract semantic systems—driven by LLMs and multimodal learning—to the concrete constraints of the physical world.
Although presented as analytically distinguishable, these layers operate in a mutually recursive manner rather than a strictly linear hierarchy. Embodiment constrains the design of simulations, and simulation, in turn, accelerates the optimization of behavioral policies. Subsequently, policy execution generates real-world interaction data that reshapes high-level cognitive representations, while social interaction continuously informs planning and decision-making processes. Consequently, PAI should be understood as a dynamically integrated intelligence system in which physical implementation, policy learning, and cognitive integration co-evolve across layers. This interconnected and recursive structure, as illustrated in Figure 7, provides a comprehensive framework for future research aiming to achieve seamless embodied intelligence by integrating hardware-based systems, behavioral policies, and advanced reasoning.

4.2. Educational, Research Policy, and Societal Implications of PAI

The three-layer framework proposed in this study can also be applied to curriculum development in PAI-related fields. Because PAI is an interdisciplinary research domain that integrates artificial intelligence, robotics, embedded systems, control, human–robot interaction, and multimodal learning, curricula should emphasize the connections among these areas rather than treating them as isolated technologies. The Physical Interaction & Infrastructure layer can cover competencies related to robotics, sensing, embedded systems, and human–robot interaction, while the Policy Learning & Control layer can address behavioral learning and decision-making capabilities such as simulation, reinforcement learning, and control. The Cognitive Integration & Multimodal Reasoning layer can incorporate higher-level cognitive capabilities, including large language models, vision–language models, and multimodal reasoning. Such an integrated approach can help students understand how perception, reasoning, and action are connected in real-world physical environments and how these competencies interact across different layers of PAI systems.
From a research policy perspective, funding priorities should focus not only on improving the performance of individual models or algorithms but also on addressing key technical challenges that connect different PAI layers. In particular, sim-to-real transfer, real-world adaptation, safe physical control, and the integration of multimodal reasoning with physical action are important challenges for enabling PAI systems to operate effectively in real environments. Future research support could therefore place greater emphasis on cross-layer research that integrates simulation, control, multimodal reasoning, and physical embodiment at the system level.
Finally, the development of PAI should be considered alongside its broader ethical and societal implications. Large-scale simulation and foundation model training may require substantial computational resources and energy, making computational efficiency and environmental sustainability important considerations. The deployment of PAI in manufacturing, logistics, healthcare, and service sectors may also reshape labor markets and job structures, creating a need for reskilling and institutional support for human–AI collaboration. In addition, unequal access to high-performance computing infrastructure and robotic hardware may widen technological disparities between the Global North and Global South. Because PAI systems can act directly in physical environments, issues of safety, accountability, privacy, human oversight, and regulatory frameworks should also be considered from the early stages of system design. These considerations extend the technical framework proposed in this study toward the broader conditions required for responsible PAI development and deployment.

5. Conclusions

This study elucidated the knowledge structure of PAI through a systematic review of 317 papers published in the WoS database from 2020 to 2025. Based on the analysis, we identify seven key research topics and reorganize them into three higher-level knowledge axes: (1) physical interaction and infrastructure, (2) policy learning and control, and (3) cognitive integration and multimodal reasoning. By comprehensively analyzing the keyword co-occurrence network, co-citation structure, and LDA-based topic structure, we empirically demonstrate that rather than acting as a single domain, PAI is an integrated, multilayered paradigm formed through the convergence of embodied interaction, learning-based control, and cognitive models.
Existing reviews of PAI have primarily adopted descriptive approaches, focusing on specific technical components, application cases, or conceptual paradigms such as soft robotics, cyber–physical AI, and AI 3.0. In contrast, this study systematically maps the distribution of research topics, structural connections, and their evolution based on rigorous quantitative data analysis. Notably, the convergence of keyword and co-citation networks onto the three identified axes reinforces the structural consistency and validity of our findings. These axes are not merely descriptive categories; they reflect structural interdependencies across different levels of abstraction, necessitating their reinterpretation as three interdependent layers—physical interaction and infrastructure, policy learning and control, and cognitive integration and multimodal reasoning.
Our analysis reveals that while early research on PAI was dominated by simulation environments and navigation-focused policy learning, recent studies have increasingly shifted toward language-grounded action control and the implementation of integrated embedded systems. In particular, the surge in research on embedded AI observed in 2025 suggests a transition beyond algorithmic experimentation toward the realization of integrated physical hardware and intelligence. This shift serves as a structural signal that the field is moving from exploratory expansion toward early structural consolidation.
The hierarchical architecture illustrated in Figure 7 conceptually integrates these analytical findings into a multilayered framework. The physical execution layer forms the foundation, the policy learning layer abstracts the environment into executable policies, and the multimodal cognitive layer engages in high-level task understanding and symbol–action alignment. Importantly, these layers do not constitute a linear hierarchy but a dynamically integrated system forming a mutual feedback loop. This hierarchical abstraction structure provides a crucial theoretical perspective for understanding PAI as an integrated intelligent system rather than a fragmented set of technologies.
Despite its contributions, this study has several limitations. First, the data were limited to publications indexed in the WoS database, which may not fully reflect rapidly emerging research disseminated through non-indexed conferences, preprint servers, or other scholarly outlets. Moreover, because WoS predominantly indexes English-language publications, the corpus may be subject to language bias, potentially underrepresenting relevant PAI research published in other languages or regional outlets. Although the corpus includes WoS-indexed proceedings from major AI and robotics conferences, such as CVPR, NeurIPS, ECCV, and ICRA, arXiv-only preprints and conference papers not indexed in WoS may be underrepresented. Given the rapid development of PAI, this database coverage may lead to the underrepresentation or delayed capture of newly emerging research topics and technological trends. Future research could mitigate this limitation by incorporating complementary sources such as Scopus and arXiv and systematically comparing whether the resulting topic and network structure remain robust across different bibliographic sources. Second, the corpus primarily comprised titles and abstracts, which may limit in-depth semantic analysis compared to full-text reviews. Third, although keyword-based filtering and preliminary reviews were performed, the inclusion of some documents that do not fully align with the conceptual categories of PAI remains possible, potentially affecting the network structure. Finally, while LDA offers advantages in interpretability and reproducibility, its semantic sophistication may be limited compared to more recent context-based topic models. In addition, the latent topic structure generated by LDA is not directly observable, and the interpretation and labeling of topics inevitably involve researcher judgment. Therefore, the naming of individual topics and their organization into higher-level layers may include a degree of interpretive subjectivity. Notably, although Topic 6 (Language-Grounded Action) was the most prevalent topic, its lexical stability across repeated LDA runs was moderate (approximately 0.30), indicating that its representative keywords may vary to some extent across runs and that detailed keyword-level interpretations should therefore be made cautiously.
The generalizability of the findings is also limited by the PAI-specific scope of the corpus. Because the search focused primarily on “Physical AI” and “Embodied AI,” adjacent research streams such as general robotics, autonomous systems, AI safety, and control may be underrepresented when they do not explicitly use these terms. Accordingly, the identified topic structure and hierarchical framework should primarily be interpreted as characterizing the literature explicitly associated with Physical AI and Embodied AI, rather than the broader AI and robotics research landscape. Future research could address these gaps by incorporating full-text data from multiple bibliographic and preprint sources and utilizing advanced transformer-based topic modeling techniques to assess whether the identified knowledge structure remains consistent across broader corpora and alternative modeling approaches.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/app16189306/s1. Reference [128] is cited in the Supplementary Materials.

Author Contributions

Conceptualization, K.M., H.J. and M.K.; methodology, K.M. and H.J.; software, K.M. and H.J.; validation, M.K.; formal analysis, K.M. and H.J.; investigation, K.M. and H.J.; data curation, K.M. and H.J.; writing—original draft preparation, K.M. and H.J.; writing—review and editing, M.K.; visualization, K.M. and H.J.; supervision, M.K.; project administration, M.K. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The bibliographic data analyzed in this study were obtained from the Web of Science Core Collection under institutional license and are subject to the database provider’s licensing restrictions. The data supporting the findings of this study are available from the corresponding author upon reasonable request.

Acknowledgments

During the preparation of this manuscript, the authors used ChatGPT (GPT-5.2 Instant; OpenAI; accessed on 21 February 2026) to assist with the creation of specific icons used in Figure 7 and to improve language readability. The authors reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Wu, E.-H.; Liu, Y.-Q.; Xu, T.-C.; Ren, L.-X.; Qin, Y.-M.; Wei, M.-Y.; He, X.-W.; Yuan, D.-Y.; Hou, W.-C.; Ma, Z.-W.; et al. Physical AI: Evolution, Progress, Challenges, and Prospects. J. Comput. Sci. Technol. 2026, 41, 271–288. [Google Scholar] [CrossRef] [Scilit]
  2. Abou Ali, M.; Dornaika, F.; Charafeddine, J. Agentic AI: A comprehensive survey of architectures, applications, and future directions. Artif. Intell. Rev. 2025, 59, 11. [Google Scholar] [CrossRef] [Scilit]
  3. Liu, Y.; Chen, W.; Bai, Y.; Liang, X.; Li, G.; Gao, W.; Lin, L. Aligning Cyber Space With Physical World: A Comprehensive Survey on Embodied AI. IEEE/ASME Trans. Mechatron. 2025, 30, 7253–7274. [Google Scholar] [CrossRef] [Scilit]
  4. Salehi, V. Fundamentals of Physical AI. J. Intell. Syst. Syst. Lifecycle Manag. 2025, 2. [Google Scholar] [CrossRef] [Scilit]
  5. Bousetouane, F. Physical AI agents: Integrating cognitive intelligence with real-world action. arXiv 2025, arXiv:2501.08944. Available online: https://arxiv.org/abs/2501.08944 (accessed on 18 October 2025).
  6. Miriyev, A.; Kovač, M. Skills for physical artificial intelligence. Nat. Mach. Intell. 2020, 2, 658–660. [Google Scholar] [CrossRef] [Scilit]
  7. Thakur, A.; Kaipa, K.; Banerjee, A.G.; Cappelleri, D.J.; Krovi, V.N.; Gupta, S.K. Physical Artificial Intelligence for Powering the Next Revolution in Robotics. J. Comput. Inf. Sci. Eng. 2025, 25, 120809. [Google Scholar] [CrossRef] [Scilit]
  8. Cheng, X.; Shen, Z.; Zhang, Y. Bioinspired 3D flexible devices and functional systems. Natl. Sci. Rev. 2024, 11, nwad314. [Google Scholar] [CrossRef] [Scilit]
  9. Sørensen, L.; Sagen Johannesen, D.T.; Melkas, H.; Johnsen, H.M. User Acceptance of a Home Robotic Assistant for Individuals With Physical Disabilities: Explorative Qualitative Study. JMIR Rehabil. Assist. Technol. 2025, 12, e63641. [Google Scholar] [CrossRef] [Scilit]
  10. Bajestani, M.S.; Kim, C.; Lee, K.-C.; Kim, D.B. Self-X-based secure human-cyber-physical system (SSHCPS) for autonomous manufacturing in the era of industry 5.0. Adv. Eng. Inform. 2025, 69, 104054. [Google Scholar] [CrossRef] [Scilit]
  11. Sharma, A.; Bhowmik, B. Autonomous agentic AI with policy adaptation for physics-informed spectral learning in Structural Health Monitoring. Adv. Eng. Inform. 2026, 70, 104224. [Google Scholar] [CrossRef] [Scilit]
  12. Balasubramani, M.; Chen, J.; Chang, R.; Shieh, J.-S. Development of a Human-Centric Autonomous Heating, Ventilation, and Air Conditioning Control System Enhanced for Industry 5.0 Chemical Fiber Manufacturing. Machines 2025, 13, 421. [Google Scholar] [CrossRef] [Scilit]
  13. Ohueri, C.C.; Seghier, T.E.; Jing, K.T.; Esa, M. AI-powered adaptive exoskeletons for long-term musculoskeletal disorder prevention in dynamic construction environments. Adv. Eng. Inform. 2025, 69, 104042. [Google Scholar] [CrossRef] [Scilit]
  14. Li, Y.; Li, Z.; Duan, Y.; Spulber, A.-B. Physical artificial intelligence (PAI): The next-generation artificial intelligence. Front. Inf. Technol. Electron. Eng. 2023, 24, 1231–1238. [Google Scholar] [CrossRef] [Scilit]
  15. Ray, P.P. Physical AI: Bridging the sim-to-real divide toward embodied, ethical, and autonomous intelligence. Mach. Learn. Comput. Sci. Eng. 2026, 2, 1. [Google Scholar] [CrossRef] [Scilit]
  16. Wu, J.; You, H.; Du, J. AI generations: From AI 1.0 to AI 4.0. Front. Artif. Intell. 2025, 8, 1585629. [Google Scholar] [CrossRef] [Scilit]
  17. Agarwal, N.; Ali, A.; Bala, M.; Balaji, Y.; Barker, E.; Cai, T.; Chattopadhyay, P.; Chen, Y.; Cui, Y.; Ding, Y.; et al. Cosmos world foundation model platform for physical ai. arXiv 2025, arXiv:2501.03575. Available online: https://arxiv.org/abs/2501.03575 (accessed on 13 September 2026).
  18. Booth, A.; Sutton, A.; Clowes, M.; Martyn-St James, M. Systematic Approaches to a Successful Literature Review, 3rd ed.; SAGE Publications: London, UK, 2021. [Google Scholar]
  19. Lee, C.-H.; Liu, C.-L.; Trappey, A.J.; Mo, J.P.T.; Desouza, K.C. Understanding digital transformation in advanced manufacturing and engineering: A bibliometric analysis, topic modeling and research trend discovery. Adv. Eng. Inform. 2021, 50, 101428. [Google Scholar] [CrossRef] [Scilit]
  20. Pranckutė, R. Web of Science (WoS) and Scopus: The titans of bibliographic information in today’s academic world. Publications 2021, 9, 12. [Google Scholar] [CrossRef] [Scilit]
  21. Donthu, N.; Kumar, S.; Mukherjee, D.; Pandey, N.; Lim, W.M. How to conduct a bibliometric analysis: An overview and guidelines. J. Bus. Res. 2021, 133, 285–296. [Google Scholar] [CrossRef] [Scilit]
  22. Archambault, É.; Campbell, D.; Gingras, Y.; Larivière, V. Comparing bibliometric statistics obtained from the Web of Science and Scopus. J. Am. Soc. Inf. Sci. Technol. 2009, 60, 1320–1326. [Google Scholar] [CrossRef] [Scilit]
  23. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. Br. Med. J. 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit]
  24. Wang, B.; Liu, S.; Ding, K.; Liu, Z.; Xu, J. Identifying technological topics and institution-topic distribution probability for patent competitive intelligence analysis: A case study in LTE technology. Scientometrics 2014, 101, 685–704. [Google Scholar] [CrossRef] [Scilit]
  25. Penning de Vries, B.B.L.; van Smeden, M.; Rosendaal, F.R.; Groenwold, R.H.H. Title, abstract, and keyword searching resulted in poor recovery of articles in systematic reviews of epidemiologic practice. J. Clin. Epidemiol. 2020, 121, 55–61. [Google Scholar] [CrossRef] [Scilit]
  26. Callon, M.; Courtial, J.-P.; Turner, W.A.; Bauin, S. From translations to problematic networks: An introduction to co-word analysis. Soc. Sci. Inf. 1983, 22, 191–235. [Google Scholar] [CrossRef] [Scilit]
  27. Blei, D.; Ng, A.; Jordan, M. Latent Dirichlet Allocation. J. Mach. Learn. Res. 2003, 3, 993–1022. [Google Scholar]
  28. Nikolenko, S.I.; Koltcov, S.; Koltsova, O. Topic modelling for qualitative studies. J. Inf. Sci. 2016, 43, 88–102. [Google Scholar] [CrossRef] [Scilit]
  29. Jacobi, C.; van Atteveldt, W.; Welbers, K. Quantitative analysis of large amounts of journalistic texts using topic modelling. In Rethinking Research Methods in an Age of Digital Journalism; Karlsson, M., Sjøvaag, H., Eds.; Routledge: London, UK, 2018; pp. 89–106. [Google Scholar] [CrossRef] [Scilit]
  30. Griffiths, T.L.; Steyvers, M. Finding scientific topics. Proc. Natl. Acad. Sci. USA 2004, 101, 5228–5235. [Google Scholar] [CrossRef] [Scilit]
  31. Bianchi, F.; Terragni, S.; Hovy, D. Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 759–766. [Google Scholar] [CrossRef] [Scilit]
  32. Angelov, D. Top2vec: Distributed representations of topics. arXiv 2020, arXiv:2008.09470. Available online: https://arxiv.org/abs/2008.09470 (accessed on 13 September 2026).
  33. Hoyle, A.; Goel, P.; Hian-Cheong, A.; Peskov, D.; Boyd-Graber, J.; Resnik, P. Is Automated Topic Model Evaluation Broken? The Incoherence of Coherence. Adv. Neural Inf. Process. Syst. 2021, 34, 2018–2033. [Google Scholar]
  34. Antons, D.; Breidbach, C.F. Big Data, Big Insights? Advancing Service Innovation and Design with Machine Learning. J. Serv. Res. 2017, 21, 17–39. [Google Scholar] [CrossRef] [Scilit]
  35. Zupic, I.; Čater, T. Bibliometric Methods in Management and Organization. Organ. Res. Methods 2015, 18, 429–472. [Google Scholar] [CrossRef] [Scilit]
  36. Savva, M.; Kadian, A.; Maksymets, O.; Zhao, Y.; Wijmans, E.; Jain, B.; Straub, J.; Liu, J.; Koltun, V.; Malik, J.; et al. Habitat: A platform for embodied ai research. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 9339–9347. [Google Scholar]
  37. Chaplot, D.S.; Gandhi, D.P.; Gupta, A.; Salakhutdinov, R.R. Object goal navigation using goal-oriented semantic exploration. Adv. Neural Inf. Process. Syst. 2020, 33, 4247–4258. [Google Scholar]
  38. Wijmans, E.; Kadian, A.; Morcos, A.; Lee, S.; Essa, I.; Parikh, D.; Savva, M.; Batra, D. Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames. arXiv 2019, arXiv:1911.00357. Available online: https://arxiv.org/abs/1911.00357 (accessed on 13 September 2026).
  39. Gupta, S.; Davidson, J.; Levine, S.; Sukthankar, R.; Malik, J. Cognitive mapping and planning for visual navigation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2616–2625. [Google Scholar]
  40. Zhu, Y.; Mottaghi, R.; Kolve, E.; Lim, J.J.; Gupta, A.; Fei-Fei, L.; Farhadi, A. Target-driven visual navigation in indoor scenes using deep reinforcement learning. In Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore, 29 May–3 June 2017; IEEE: New York, NY, USA, 2017; pp. 3357–3364. [Google Scholar]
  41. Kolve, E.; Mottaghi, R.; Han, W.; VanderBilt, E.; Weihs, L.; Herrasti, A.; Deitke, M.; Ehsani, K.; Gordon, D.; Zhu, Y.; et al. Ai2-thor: An interactive 3d environment for visual ai. arXiv 2017, arXiv:1712.05474. Available online: https://arxiv.org/abs/1712.05474 (accessed on 13 September 2026).
  42. Xia, F.; Zamir, A.R.; He, Z.; Sax, A.; Malik, J.; Savarese, S. Gibson env: Real-world perception for embodied agents. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 9068–9079. [Google Scholar]
  43. Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft COCO: Common Objects in Context. In Proceedings of the Computer Vision–ECCV 2014; Springer: Cham, Switzerland, 2014; Volume 8693, pp. 740–755. [Google Scholar] [CrossRef] [Scilit]
  44. Szot, A.; Clegg, A.; Undersander, E.; Wijmans, E.; Zhao, Y.; Turner, J.; Maestre, N.; Mukadam, M.; Chaplot, D.S.; Maksymets, O.; et al. Habitat 2.0: Training Home Assistants to Rearrange their Habitat. Adv. Neural Inf. Process. Syst. 2021, 34, 251–266. [Google Scholar]
  45. Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. Available online: https://arxiv.org/abs/1707.06347 (accessed on 13 September 2026).
  46. Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning, Virtual, 18–24 July 2021; Volume 139, pp. 8748–8763. Available online: https://proceedings.mlr.press/v139/radford21a (accessed on 13 September 2026).
  47. Graves, A. Long Short-Term Memory. In Supervised Sequence Labelling with Recurrent Neural Networks; Springer: Berlin/Heidelberg, Germany, 2012; pp. 37–45. [Google Scholar] [CrossRef] [Scilit]
  48. Dai, A.; Chang, A.X.; Savva, M.; Halber, M.; Funkhouser, T.; Nießner, M. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 5828–5839. [Google Scholar]
  49. Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, 2–7 June 2019; pp. 4171–4186. [Google Scholar] [CrossRef] [Scilit]
  50. Fried, D.; Hu, R.; Cirik, V.; Rohrbach, A.; Andreas, J.; Morency, L.P.; Berg-Kirkpatrick, T.; Saenko, K.; Klein, D.; Darrell, T. Speaker-follower models for vision-and-language navigation. Adv. Neural Inf. Process. Syst. 2018, 31, 3314–3325. [Google Scholar]
  51. Qi, Y.; Wu, Q.; Anderson, P.; Wang, X.; Wang, W.Y.; Shen, C.; Hengel, A. Reverie: Remote embodied visual referring expression in real indoor environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 14–19 June 2020; pp. 9982–9991. [Google Scholar]
  52. Tan, H.; Bansal, M. LXMERT: Learning Cross-Modality Encoder Representations from Transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, Hong Kong, China, 3–7 November 2019; pp. 5100–5111. [Google Scholar] [CrossRef] [Scilit]
  53. Anderson, P.; Wu, Q.; Teney, D.; Bruce, J.; Johnson, M.; Sünderhauf, N.; Reid, I.; Gould, S.; van den Hengel, A. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 3674–3683. [Google Scholar]
  54. Röder, M.; Both, A.; Hinneburg, A. Exploring the Space of Topic Coherence Measures. In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining, Shanghai, China, 2–6 February 2015; pp. 399–408. [Google Scholar] [CrossRef] [Scilit]
  55. Wan, Z.; Du, Y.; Ibrahim, M.; Zhao, Y.; Krishna, T.; Raychowdhury, A. Thinking and moving: An efficient computing approach for integrated task and motion planning in cooperative embodied ai systems. In Proceedings of the 43rd IEEE/ACM International Conference on Computer-Aided Design, New York, NY, USA, 27–31 October 2024; pp. 1–7. [Google Scholar]
  56. Hu, D.; Lan, D.; Liu, Y.; Ning, J.; Wang, J.; Yang, Y. Embodied AI Through Cloud-Fog Computing: A Framework for Everywhere Intelligence. In Proceedings of the 2024 IEEE 33rd International Symposium on Industrial Electronics (ISIE), Ulsan, Republic of Korea, 18–21 June 2024; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
  57. Foglia, L.; Wilson, R. Embodied cognition. Wiley Interdiscip. Rev. Cogn. Sci. 2013, 4, 319–325. [Google Scholar] [CrossRef] [Scilit]
  58. Pfeifer, R.; Bongard, J. How the Body Shapes the Way We Think: A New View of Intelligence; MIT Press: Cambridge, MA, USA, 2006. [Google Scholar]
  59. Kwon, W.; Baek, S.; Baek, J.; Shin, W.; Gwak, M.; Park, P.; Lee, S. Reinforced Intelligence Through Active Interaction in Real World: A Survey on Embodied AI. Int. J. Control Autom. Syst. 2025, 23, 1597–1612. [Google Scholar] [CrossRef] [Scilit]
  60. Zhang, J.; Chen, K.; Chen, S.; Zheng, Y.; Huang, T.; Yu, Z. Spikegs: 3d gaussian splatting from spike streams with high-speed camera motion. In Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, VIC, Australia, 28 October–1 November 2024; pp. 9194–9203. [Google Scholar]
  61. Zitkovich, B.; Yu, T.; Xu, S.; Xu, P.; Xiao, T.; Xia, F.; Wu, J.; Wohlhart, P.; Welker, S.; Wahid, A.; et al. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. In Proceedings of the 7th Conference on Robot Learning, Atlanta, GA, USA, 6–9 November 2023; Volume 229, pp. 2165–2183. [Google Scholar]
  62. Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. Available online: https://arxiv.org/abs/2312.00752 (accessed on 13 September 2026).
  63. Hamburg, S.; Jimenez Rodriguez, A.; Htet, A.; Di Nuovo, A. Active Inference for Learning and Development in Embodied Neuromorphic Agents. Entropy 2024, 26, 582. [Google Scholar] [CrossRef] [Scilit]
  64. Friston, K. The free-energy principle: A unified brain theory? Nat. Rev. Neurosci. 2010, 11, 127–138. [Google Scholar] [CrossRef] [Scilit]
  65. You, H.; Zhou, T.; Zhu, Q.; Ye, Y.; Du, E.J. Embodied AI for dexterity-capable construction Robots: DEXBOT framework. Adv. Eng. Inform. 2024, 62, 102572. [Google Scholar] [CrossRef] [Scilit]
  66. Zhao, M.; Xia, J.; Hou, K.; Liu, Y.; Xia, S.; Jiang, X. FlexiFly: Interfacing the Physical World with Foundation Models Empowered by Reconfigurable Drone Systems. In Proceedings of the 23rd ACM Conference on Embedded Networked Sensor Systems, Irvine, CA, USA, 6–9 May 2025; pp. 463–476. [Google Scholar] [CrossRef] [Scilit]
  67. De Haro, L. Using Embodied Artificial Intelligence Agents to Automate Biorisk Management Tasks in High-Containment Laboratories. Appl. Biosaf. 2025, 30, 314–325. [Google Scholar] [CrossRef] [Scilit]
  68. Dennett, D.C. The Intentional Stance; Mit Press: Cambridge, MA, USA, 1987. [Google Scholar]
  69. Sini, R. Does Saudi robot citizen have more rights than women? BBC News, 26 October 2017. Available online: https://www.bbc.com/news/blogs-trending-41761856 (accessed on 13 September 2026).
  70. Tiku, N. The Google engineer who thinks the company’s AI has come to life. The Washington Post, 11 June 2022. Available online: https://www.washingtonpost.com/technology/2022/06/11/google-ai-lamda-blake-lemoine/ (accessed on 17 July 2026).
  71. Gray, H.; Gray, K.; Wegner, D. Dimensions of Mind Perception. Science 2007, 315, 619. [Google Scholar] [CrossRef] [Scilit]
  72. Zhou, X.; Zhou, Y.; Gong, Y.; Cai, Z.; Qiu, A.; Xiao, Q. Bee and I need diversity! Break Filter Bubbles in Recommendation Systems through Embodied AI Learning. In Proceedings of the 23rd Annual ACM Interaction Design and Children Conference, Delft, The Netherlands, 17–20 June 2024; pp. 44–61. [Google Scholar] [CrossRef] [Scilit]
  73. Balazadeh, K.; Kajonius, P. Exploring Intimacy with Artificial Intelligence: Validation of Robot Intimacy Receptivity Scale (RIRS). Int. J. Soc. Robot. 2025, 17, 1453–1465. [Google Scholar] [CrossRef] [Scilit]
  74. Cheetham, M.; Pedroni, A.F.; Antley, A.; Slater, M.; Jäncke, L. Virtual milgram: Empathic concern or personal distress? Evidence from functional MRI and dispositional measures. Front. Hum. Neurosci. 2009, 3, 29. [Google Scholar] [CrossRef] [Scilit]
  75. Cheetham, M.; Suter, P.; Jancke, L. Perceptual discrimination difficulty and familiarity in the Uncanny Valley: More like a “Happy Valley”. Front. Psychol. 2014, 5, 1219. [Google Scholar] [CrossRef] [Scilit]
  76. Marquardt, M.; Graf, P.; Jansen, E.; Hillmann, S.; Voigt-Antons, J.N. Situativität, Funktionalität und Vertrauen: Ergebnisse einer szenariobasierten Interviewstudie zur Erklärbarkeit von KI in der Medizin. TATuP-Z. Tech. Theor. Prax. 2024, 33, 41–47. [Google Scholar] [CrossRef] [Scilit]
  77. Cheon, E.; Zaga, C.; Lee, H.; Lupetti, M.; Dombrowski, L.; Jung, M. Human-Machine Partnerships in the Future of Work: Exploring the Role of Emerging Technologies in Future Workplaces. In Companion Publication of the 2021 Conference on Computer Supported Cooperative Work and Social Computing; Association for Computing Machinery: New York, NY, USA, 2021; pp. 323–326. [Google Scholar] [CrossRef] [Scilit]
  78. Zubala, A.; Pease, A.; Lyszkiewicz, K.; Hackett, S. Art psychotherapy meets creative AI: An integrative review positioning the role of creative AI in art therapy process. Front. Psychol. 2025, 16, 154839. [Google Scholar] [CrossRef] [Scilit]
  79. Song, X.; Liu, C.; Xu, L.; Gao, B.; Lu, Z.; Zhang, Y. Affective computing methods for multimodal embodied AI human–computer interaction. Aslib J. Inf. Manag. 2025, 77, 1–25. [Google Scholar] [CrossRef] [Scilit]
  80. Batra, D.; Gokaslan, A.; Kembhavi, A.; Maksymets, O.; Mottaghi, R.; Savva, M.; Toshev, A.; Wijmans, E. Objectnav revisited: On evaluation of embodied agents navigating to objects. arXiv 2020, arXiv:2006.13171. Available online: https://arxiv.org/abs/2006.13171 (accessed on 13 September 2026).
  81. Anderson, P.; Chang, A.; Chaplot, D.S.; Dosovitskiy, A.; Gupta, S.; Koltun, V.; Kosecka, J.; Malik, J.; Mottaghi, R.; Savva, M.; et al. On evaluation of embodied navigation agents. arXiv 2018, arXiv:1807.06757. Available online: https://arxiv.org/abs/1807.06757 (accessed on 13 September 2026).
  82. Armeni, I.; He, Z.Y.; Gwak, J.; Zamir, A.R.; Fischer, M.; Malik, J.; Fischer, M.; Savarese, S. 3D scene graph: A structure for unified semantics, 3d space, and camera. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 5664–5673. [Google Scholar]
  83. Locatello, F.; Weissenborn, D.; Unterthiner, T.; Mahendran, A.; Heigold, G.; Uszkoreit, J.; Dosovitskiy, A.; Kipf, T. Object-centric learning with slot attention. Adv. Neural Inf. Process. Syst. 2020, 33, 11525–11538. [Google Scholar]
  84. Pal, A.; Qiu, Y.; Christensen, H. Learning hierarchical relationships for object-goal navigation. In Proceedings of the 2020 Conference on Robot Learning, Virtual, 16–18 November 2020; Volume 155, pp. 517–528. [Google Scholar]
  85. Seymour, Z.; Thopalli, K.; Mithun, N.; Chiu, H.P.; Samarasekera, S.; Kumar, R. Maast: Map attention with semantic transformers for efficient visual navigation. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi’an, China, 30 May–5 June 2021; IEEE: New York, NY, USA, 2021; pp. 13223–13230. [Google Scholar]
  86. Li, L.; Chu, W.; Langford, J.; Schapire, R.E. A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th International Conference on World Wide Web, Raleigh, NC, USA, 26–30 April 2010; pp. 661–670. [Google Scholar]
  87. Rudra, S.; Goel, S.; Santara, A.; Gentile, C.; Perron, L.; Xia, F.; Sindhwani, V.; Parada, C.; Aggarwal, G. A contextual bandit approach for learning to plan in environments with probabilistic goal configurations. arXiv 2022, arXiv:2211.16309. Available online: https://arxiv.org/abs/2211.16309 (accessed on 13 September 2026).
  88. Zeng, H.; Song, X.; Jiang, S. Multi-object navigation using potential target position policy function. IEEE Trans. Image Process. 2023, 32, 2608–2619. [Google Scholar] [CrossRef] [Scilit]
  89. Blum, A.; Chalasani, P.; Coppersmith, D.; Pulleyblank, B.; Raghavan, P.; Sudan, M. The minimum latency problem. In Proceedings of the Twenty-Sixth Annual ACM Symposium on Theory of Computing, Montréal, QC, Canada, 23–25 May 1994; pp. 163–171. [Google Scholar]
  90. Höfer, S.; Bekris, K.; Handa, A.; Gamboa, J.C.; Mozifian, M.; Golemo, F.; Atkeson, C.; Fox, D.; Goldberg, K.; Leonard, J.; et al. Sim2real in robotics and automation: Applications and challenges. IEEE Trans. Autom. Sci. Eng. 2021, 18, 398–400. [Google Scholar] [CrossRef] [Scilit]
  91. Wu, S.C.; Wald, J.; Tateno, K.; Navab, N.; Tombari, F. Scenegraphfusion: Incremental 3d scene graph prediction from rgb-d sequences. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtually, 19–25 June 2021; pp. 7515–7525. [Google Scholar]
  92. Wani, S.; Patel, S.; Jain, U.; Chang, A.; Savva, M. Multion: Benchmarking semantic map memory using multi-object navigation. Adv. Neural Inf. Process. Syst. 2020, 33, 9700–9712. [Google Scholar]
  93. Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level Control through Deep Reinforcement Learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit]
  94. Petrenko, A.; Wijmans, E.; Shacklett, B.; Koltun, V. Megaverse: Simulating embodied agents at one million experiences per second. In Proceedings of the 38th International Conference on Machine Learning, Virtual, 18–24 July 2021; Volume 139, pp. 8556–8566. [Google Scholar]
  95. Chang, A.; Dai, A.; Funkhouser, T.; Halber, M.; Niessner, M.; Savva, M.; Song, S.; Zeng, A.; Zhang, Y. Matterport3d: Learning from rgb-d data in indoor environments. arXiv 2017, arXiv:1709.06158. Available online: https://arxiv.org/abs/1709.06158 (accessed on 13 September 2026).
  96. Kadian, A.; Truong, J.; Gokaslan, A.; Clegg, A.; Wijmans, E.; Lee, S.; Savva, M.; Chernova, S.; Batra, D. Sim2Real Predictivity: Does Evaluation in Simulation Predict Real-World Performance? IEEE Robot. Autom. Lett. 2020, 5, 6670–6677. [Google Scholar] [CrossRef] [Scilit]
  97. Jain, U.; Liu, I.J.; Lazebnik, S.; Kembhavi, A.; Weihs, L.; Schwing, A.G. Gridtopix: Training embodied agents with minimal supervision. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Virtually, 10–17 October 2021; pp. 15141–15151. [Google Scholar]
  98. Ramakrishnan, S.K.; Gokaslan, A.; Wijmans, E.; Maksymets, O.; Clegg, A.; Turner, J.; Undersander, E.; Galuba, W.; Westbury, A.; Chang, A.X.; et al. Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai. arXiv 2021, arXiv:2109.08238. Available online: https://arxiv.org/abs/2109.08238 (accessed on 13 September 2026).
  99. Chao, Y.W.; Paxton, C.; Xiang, Y.; Yang, W.; Sundaralingam, B.; Chen, T.; Murali, A.; Cakmak, M.; Fox, D. Handoversim: A simulation framework and benchmark for human-to-robot object handovers. In Proceedings of the 2022 International Conference on Robotics and Automation (ICRA), Philadelphia, PA, USA, 23–27 May 2022; IEEE: New York, NY, USA, 2022; pp. 6941–6947. [Google Scholar]
  100. Christen, S.; Yang, W.; Pérez-D’Arpino, C.; Hilliges, O.; Fox, D.; Chao, Y.W. Learning human-to-robot handovers from point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 9654–9664. [Google Scholar]
  101. Kumar, A.; Fu, Z.; Pathak, D.; Malik, J. Rma: Rapid motor adaptation for legged robots. arXiv 2021, arXiv:2107.04034. Available online: https://arxiv.org/abs/2107.04034 (accessed on 13 September 2026).
  102. Liang, Y.; Ellis, K.; Henriques, J. Rapid motor adaptation for robotic manipulator arms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 16404–16413. [Google Scholar]
  103. Zhu, L.; Jia, K.; Zhao, Y.; Qi, Y.; Wang, L.; Huang, H. Spikenerf: Learning neural radiance fields from continuous spike stream. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 6285–6295. [Google Scholar]
  104. Jiang, Y.; Guo, M.; Li, J.; Exarchos, I.; Wu, J.; Liu, C.K. Dash: Modularized human manipulation simulation with vision and language for embodied ai. In Proceedings of the ACM SIGGRAPH/Eurographics Symposium on Computer Animation, Virtual, 6–9 September 2021; pp. 1–12. [Google Scholar]
  105. Lu, H.; Tan, X.; Chen, M.; Zhang, Z.; Zhang, X.; Chen, J.; Wei, X.; Zhao, T. Cross-Modal Haptic Compression Inspired by Embodied AI for Haptic Communications. IEEE Trans. Multimed. 2025, 27, 4996–5008. [Google Scholar] [CrossRef] [Scilit]
  106. Puig, X.; Shu, T.; Tenenbaum, J.B.; Torralba, A. Nopa: Neurally-guided online probabilistic assistance for building socially intelligent home assistants. arXiv 2023, arXiv:2301.05223. Available online: https://arxiv.org/abs/2301.05223 (accessed on 13 September 2026).
  107. Jain, V.; Magalhaes, G.; Ku, A.; Vaswani, A.; Ie, E.; Baldridge, J. Stay on the path: Instruction fidelity in vision-and-language navigation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 1862–1872. [Google Scholar]
  108. Lin, B.; Nie, Y.; Wei, Z.; Chen, J.; Ma, S.; Han, J.; Xu, H.; Chang, X.; Liang, X. NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 5945–5957. [Google Scholar] [CrossRef] [Scilit]
  109. Min, S.Y.; Puig, X.; Chaplot, D.S.; Yang, T.Y.; Rai, A.; Parashar, P.; Salakhutdinov, R.; Bisk, Y.; Mottaghi, R. Situated instruction following. In European Conference on Computer Vision; Springer Nature: Cham, Switzerland, 2024; pp. 202–228. [Google Scholar]
  110. Gupta, S.K. Embodied ai for smart robotic cells in manufacturing applications. Proc. AAAI Conf. Artif. Intell. 2025, 39, 28630–28636. [Google Scholar] [CrossRef] [Scilit]
  111. Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q.V.; Zhou, D. Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inf. Process. Syst. 2022, 35, 24824–24837. [Google Scholar] [CrossRef] [Scilit]
  112. Shridhar, M.; Thomason, J.; Gordon, D.; Bisk, Y.; Han, W.; Mottaghi, R.; Zettlemoyer, L.; Fox, D. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 10740–10749. [Google Scholar]
  113. Chen, C.; Cong, Y.; Kan, Z. Worldafford: Affordance grounding based on natural language instructions. In Proceedings of the 2024 IEEE 36th International Conference on Tools with Artificial Intelligence (ICTAI), Herndon, VA, USA, 28–30 October 2024; IEEE: New York, NY, USA, 2024; pp. 822–828. [Google Scholar]
  114. Xu, W.; Wang, M.; Zhou, W.; Li, H. P-RAG: Progressive retrieval augmented generation for planning on embodied everyday task. In Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, VIC, Australia, 28 October–1 November 2024; pp. 6969–6978. [Google Scholar]
  115. Gao, F.; Shi, L.; Tang, J.; Wang, J.; Li, S.; Ma, S.; Yu, J. Visual and Textual Commonsense-Enhanced Layout Learning for Vision-and-Language Navigation. IEEE Trans. Autom. Sci. Eng. 2025, 22, 21311–21324. [Google Scholar] [CrossRef] [Scilit]
  116. Rutar, D.; Markelius, A.; Schellaert, W.; Hernández-Orallo, J.; Cheke, L. General interaction battery: Simple object navigation and affordances (GIBSONA). Cogn. Syst. Res. 2025, 94, 101411. [Google Scholar] [CrossRef] [Scilit]
  117. Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.-T.; Rocktäschel, T.; et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Adv. Neural Inf. Process. Syst. 2020, 33, 9459–9474. [Google Scholar]
  118. Johnson-Laird, P.N. Mental models and human reasoning. Proc. Natl. Acad. Sci. USA 2010, 107, 18243–18250. [Google Scholar] [CrossRef] [Scilit]
  119. Wang, R.; Xu, P.; Shi, H.; Schumann, E.; Liu, C.K. FürElise: Capturing and physically synthesizing hand motion of piano performance. In SIGGRAPH Asia 2024 Conference Papers; Association for Computing Machinery: New York, NY, USA, 2024; Article 77; pp. 1–11. [Google Scholar] [CrossRef] [Scilit]
  120. Alayrac, J.B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. Flamingo: A visual language model for few-shot learning. Adv. Neural Inf. Process. Syst. 2022, 35, 23716–23736. [Google Scholar] [CrossRef] [Scilit]
  121. Hong, Y.; Zhen, H.; Chen, P.; Zheng, S.; Du, Y.; Chen, Z.; Gan, C. 3d-llm: Injecting the 3d world into large language models. Adv. Neural Inf. Process. Syst. 2023, 36, 20482–20494. [Google Scholar] [CrossRef] [Scilit]
  122. Ma, X.; Yong, S.; Zheng, Z.; Li, Q.; Liang, Y.; Zhu, S.C.; Huang, S. Sqa3d: Situated question answering in 3d scenes. arXiv 2022, arXiv:2210.07474. Available online: https://arxiv.org/abs/2210.07474 (accessed on 13 September 2026).
  123. Zhang, R.; Zhao, C.; Du, H.; Niyato, D.; Wang, J.; Sawadsitang, S.; Shen, X.; Kim, D.I. Embodied AI-enhanced vehicular networks: An integrated vision language models and reinforcement learning method. IEEE Trans. Mob. Comput. 2025, 24, 11494–11510. [Google Scholar] [CrossRef] [Scilit]
  124. Zhou, J.; Wang, J.; Zhang, J.; Sun, W.; Zhang, J.; Birchfield, S.; Guo, D.; Kong, L.; Wang, M.; Zhong, Y. Audio–visual segmentation. In European Conference on Computer Vision; Springer Nature: Cham, Switzerland, 2022; pp. 386–403. [Google Scholar]
  125. Wang, S.; Huang, X.; Chen, C.; Wu, L.; Li, J. Reform: Error-aware few-shot knowledge graph completion. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, Virtually, 1–5 November 2021; pp. 1979–1988. [Google Scholar]
  126. Wen, C.; Liang, J.; Yuan, S.; Huang, H.; Bethala, G.C.R.; Liu, Y.S.; Wang, M.; Tzes, A.; Fang, Y. How secure are large language models (llms) for navigation in urban environments? arXiv 2024, arXiv:2402.09546. Available online: https://arxiv.org/abs/2402.09546 (accessed on 13 September 2026).
  127. Li, K.; Geng, Q.; Wan, M.; Cao, X.; Zhou, Z. Context and Spatial Feature Calibration for Real-Time Semantic Segmentation. IEEE Trans. Image Process. 2023, 32, 5465–5477. [Google Scholar] [CrossRef] [Scilit]
  128. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
Figure 1. PRISMA flow diagram.
Figure 1. PRISMA flow diagram.
Applsci 16 09306 g001
Figure 2. Methodological Framework.
Figure 2. Methodological Framework.
Applsci 16 09306 g002
Figure 3. Keyword co-occurrence network of Physical AI.
Figure 3. Keyword co-occurrence network of Physical AI.
Applsci 16 09306 g003
Figure 4. Document co-citation network of cited references.
Figure 4. Document co-citation network of cited references.
Applsci 16 09306 g004
Figure 5. Topic coherence scores for K = 3–30.
Figure 5. Topic coherence scores for K = 3–30.
Applsci 16 09306 g005
Figure 6. Annual topic prevalence in PAI research from 2021 to 2025. (Note: Value represents mean document-level topic proportions; 2025 covers publications indexed through October).
Figure 6. Annual topic prevalence in PAI research from 2021 to 2025. (Note: Value represents mean document-level topic proportions; 2025 covers publications indexed through October).
Applsci 16 09306 g006
Figure 7. Hierarchical architecture of physical artificial intelligence.
Figure 7. Hierarchical architecture of physical artificial intelligence.
Applsci 16 09306 g007
Table 1. Definitions of physical artificial intelligence in recent research.
Table 1. Definitions of physical artificial intelligence in recent research.
ReferencesDescriptionDomainsScope and Limitations
Miriyev and Kovač [6]PAI concerns both the conceptual foundations and practical development of physical systems that can perform tasks normally linked to intelligent living beings.RoboticsEmphasizes physical embodiment and morphology, but provides relatively limited consideration of distributed intelligence and higher-level cognition.
Li et al. [14]PAI can be understood as a multidisciplinary field focused on nature-inspired intelligent robots, highlighting the integration of software-based intelligence with hardware elements such as materials and mechanics.Computer scienceBroadens PAI through the integration of software intelligence with hardware, materials, and mechanics, but provides limited detail on higher-level cognition and action mechanisms.
Balasubramani et al. [12]PAI enables systems to learn from both real-time data and simulated environments, supporting adaptive control and ongoing improvement to maintain operational stability in complex industrial contexts.ManufacturingEmphasizes adaptive control and continuous learning in industrial environments, but gives relatively limited attention to material and morphological aspects of embodiment.
Wu et al. [16]AI 3.0, or PAI, expands intelligence into the physical world by combining robotics, autonomous vehicles, and sensor-integrated control systems to operate under uncertain real-world conditions.Intelligent constructionProvides a clear emphasis on sensing, actuation, and operation in uncertain physical environments, but discusses reasoning and learning mechanisms in less detail.
Bousetouane [5]PAI agents are embodied intelligent systems developed to engage directly with the physical environment.Healthcare, logistics, autonomous vehiclesEmphasizes embodiment and direct interaction with the physical environment, but provides limited detail on adaptive learning and higher-level cognitive reasoning.
Agarwal et al. [17]PAI refers to AI systems equipped with sensors and actuators, where sensors enable environmental perception and actuators enable physical interaction and modification of the environment.Robotic manipulation, Virtual worldsClearly defines PAI through sensing and actuation, but places less emphasis on planning, reasoning, cognitive architecture, and material embodiment.
Table 2. The 20 most prolific publication venues for PAI research.
Table 2. The 20 most prolific publication venues for PAI research.
RankSourceFrequencyRelative
Frequency (%)
1IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)268.20
2CHI Conference on Human Factors in Computing Systems (CHI)206.31
3Advances in Neural Information Processing Systems (NeurIPS)175.36
4European Conference on Computer Vision (ECCV)123.79
5IEEE Robotics and Automation Letters (RA-L)113.47
6IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)92.84
7IEEE International Conference on Robotics and Automation (ICRA)92.84
8ACM International Conference on Multimedia (MM)82.52
9IEEE/CVF International Conference on Computer Vision (ICCV)72.21
10Conference on Robot Learning61.89
11International Conference on Information Fusion (FUSION)61.89
12International Joint Conference on Artificial Intelligence 41.26
13AAAI Conference on Artificial Intelligence 41.26
14Artificial Life41.26
15Sensors30.95
16Advanced Engineering Informatics30.95
17IEEE Transactions on Emerging Topics in Computational Intelligence30.95
18Frontiers in Neurorobotics30.95
19ACM/IEEE International Conference on Human-Robot Interaction30.95
20Journal of Field Robotics30.95
Table 3. Major bridging references in the co-citation network.
Table 3. Major bridging references in the co-citation network.
Cited ReferenceClusterTLSCross-Cluster StrengthCross-Cluster Ratio
Savva et al. [36]197257859.5%
Anderson et al. [53]286154763.5%
Radford et al. [46]359045176.4%
Xia et al. [42]452139876.4%
Kolve et al. [41]439231179.3%
Table 4. Research topics on physical artificial intelligence.
Table 4. Research topics on physical artificial intelligence.
LayerTopicKeywordsProportion
Physical Interaction & Infrastructure1. Embodied AIsystem, robot, physical, intelligence, interaction, framework, sensor, control, autonomous, platform19.1%
2. Human–Robot Interactionhuman, interaction, user, social, trust, experience, communication, perception, assistant, emotional13.4%
Policy Learning
& Control
3. Semantic Navigationnavigation, object, scene, goal, map, representation, spatial, obstacle, semantic, exploration13.6%
4. Simulation-Based Learningsimulation, simulator, training, policy, state, benchmark, performance, visual, data, control20.0%
Cognitive Integration & Multimodal
Reasoning
5. Vision-Language Navigationvln, vision, language, instruction, egocentric, communication, goal, dataset, behavior, action5.0%
6. Language-Grounded Actionlanguage, action, instruction, reasoning, planning, llm, knowledge, navigation, performance, prediction21.8%
7. Multimodal Learningvisual, language, llm, knowledge, representation, feature, perception, decision, multi, cross7.1%
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Maeng, K.; Jin, H.; Kim, M. Mapping the Knowledge Structure of Physical Artificial Intelligence: A Data-Driven Systematic Review. Appl. Sci. 2026, 16, 9306. https://doi.org/10.3390/app16189306

AMA Style

Maeng K, Jin H, Kim M. Mapping the Knowledge Structure of Physical Artificial Intelligence: A Data-Driven Systematic Review. Applied Sciences. 2026; 16(18):9306. https://doi.org/10.3390/app16189306

Chicago/Turabian Style

Maeng, Kyuho, Hyeonjun Jin, and Minjun Kim. 2026. "Mapping the Knowledge Structure of Physical Artificial Intelligence: A Data-Driven Systematic Review" Applied Sciences 16, no. 18: 9306. https://doi.org/10.3390/app16189306

APA Style

Maeng, K., Jin, H., & Kim, M. (2026). Mapping the Knowledge Structure of Physical Artificial Intelligence: A Data-Driven Systematic Review. Applied Sciences, 16(18), 9306. https://doi.org/10.3390/app16189306

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop