From Pre-Rendered to Autonomous: A Systematic Review of AI-Driven Character Animation and Embodiment in Virtual Reality
Abstract
1. Introduction
- RQ1: How are traditional animation pipelines being replaced or augmented by Generative AI, diffusion models, and neural rendering (e.g., NeRFs, 3DGS) to create autonomous, real-time avatars in immersive VR?
- RQ2: What are the primary computational bottlenecks of high-fidelity AI animation in standalone VR, and how are emerging hardware–software co-designs (e.g., ASICs, edge accelerators, foveated rendering) resolving these latency constraints?
- RQ3: How do AI-driven behavioral realism, physics-aware kinematics, and multisensory interactions influence user embodiment, agency, and social immersion in Virtual Worlds?
- RQ4: What are the emerging ethical considerations, inclusivity challenges (e.g., accessibility, diverse representation), and societal impacts of deploying autonomous GenAI avatars and “digital afterlives” in the Metaverse?
2. Theoretical Background: Animation and VR
2.1. From Pre-Rendered to Autonomous: The Evolution of VR Animation
2.2. The Pillars of VR-Specific Animation Technologies
- Motion capture (MoCap) remains the historical baseline for high-fidelity motion [1,2]. Inverse Kinematics (IK) complements it by calculating full-body poses from sparse input data [15]. Bridging raw tracking data with neural generation, modern pipelines draw on 3D Morphable Models (3DMMs), parametric mathematical representations of human geometry such as SMPL and FLAME, which serve as geometric priors that guide AI models toward anatomically correct expressions and poses.
- Where IK is fast, it is often physically implausible, while physics-based animation addresses artifacts such as foot sliding by imposing physical constraints, gravity, collisions, and contact forces [3]. Rigorous physical simulation, however, is computationally expensive, and this cost constitutes a critical processing bottleneck for standalone VR headsets [9].
- To circumvent the bottlenecks inherent in traditional physics and rendering pipelines, the industry has turned to GenAI. Diffusion models and Vision–Language–Action (VLA) agents now synthesize complex motion directly from text or audio [8,10]. The rendering paradigm is shifting in tandem, moving from implicit to explicit representations. Early Neural Radiance Fields (NeRFs) relied on implicit Multi-Layer Perceptrons (MLPs), which are computationally heavy and difficult to animate. The field has consequently pivoted toward 3D Gaussian Splatting (3DGS), an explicit volumetric representation that permits direct physics-based deformation and real-time rendering of 4D avatars without the substantial overhead of MLP queries [6,7].
2.3. Computational Paradigms and Rendering Constraints in VR
- Software-level optimization (foveated rendering): to preserve compute resources, systems increasingly rely on gaze-tracked foveated rendering. Exploiting the spatial selectivity of human visual perception, this technique renders ultra-high-resolution graphics only at the center of the user’s gaze (the fovea) while dynamically reducing resolution in the periphery.
- Hardware–software co-design: since software optimization alone is insufficient; mitigating latency requires hardware-level intervention. Explicit representations such as 3DGS eliminate MLP computational overhead but introduce a different bottleneck: severe memory bandwidth pressure from the large volumes of splat data that must be sorted and rasterized per frame. The literature therefore emphasizes edge AI approaches, in which processing is offloaded to ASICs (Application-Specific Integrated Circuits) purpose-built for generative workloads [8,11]. Complementing this, Computing-in-Memory (CIM) architectures perform calculations directly at the memory site, drastically curtailing the energy-intensive transfer of splat data between memory and the GPU.
2.4. User Experience, Embodiment, and Perception in VR
3. Methodology
3.1. Literature Seach
- Discovered main terms by identifying core concepts from the Research Questions (RQs).
- Identified alternative synonyms (e.g., “Spatial Computing” for VR, “neural rendering” or “diffusion” for animation techniques).
- Checked the keywords in prominent state-of-the-art studies.
- Added alternative spellings and synonyms using the Boolean operator OR.
- Linked the main thematic categories using the Boolean operator AND.
3.2. Inclusion and Exclusion Criteria
- Round 1 (Initial Screening): The initial database search yielded a total of 289 records. After removing 8 duplicates, a pool of 281 unique records was established. Titles and abstracts were screened against the criteria, resulting in the exclusion of 89 out-of-scope articles.
- Round 2 (Thorough Assessment): The remaining 192 publications underwent a detailed full-text assessment and quality evaluation. During this phase, 149 articles were excluded based on the predefined exclusion criteria (EC1–EC5), yielding 43 highly relevant studies from the databases.
- Round 3 (Snowballing and Final Selection): The snowballing procedure identified 5 additional critical papers. This resulted in a final corpus of 48 primary studies included in the synthesis, which directly answer the four RQs.
3.3. Quality Assessment
3.4. Data Extraction and Coding Scheme
3.5. Research Assistance Tools and Data Management
4. Results and State-of-the-Art Synthesis
4.1. The Evolution of Avatar Generation: Generative AI and Neural Rendering (RQ1)
4.1.1. Diffusion Models and Masked Motion Synthesis
4.1.2. Neural Rendering and Volumetric Avatars
4.1.3. Automation of Rigging and Dynamic Cloth Simulation
4.1.4. Transformers and End-to-End Generation Frameworks
4.1.5. Methodological Limitations and Gap Analysis
4.2. Overcoming Computational Constraints: Hardware and Performance Optimization (RQ2)
4.2.1. Domain-Specific Accelerators and ASICs
4.2.2. Computing-in-Memory (CIM) Architectures and Bandwidth Reduction
4.2.3. Perception-Aware and Foveated Optimization
4.2.4. The Integration Gap: Interfacing Custom Silicon with XR Ecosystems
4.2.5. Transition to User Centricity
4.3. User Experience, Embodiment, and Interaction (RQ3)
4.3.1. The AI–Agency Paradox and Body Ownership
4.3.2. Kinematic Realism and the Dynamic Uncanny Valley
4.3.3. Social Presence, Affect, and Multimodal Interaction
4.3.4. Transitioning to Sociotechnical Implications
4.4. Sociotechnical and Ethical Implications: Identity, Privacy, and Inclusion (RQ4)
4.4.1. Identity, Consent, and “Generative Ghosts”
4.4.2. Spatial Privacy and Context-Aware Agents
4.4.3. Accessibility, Inclusion, and Kinematic Bias
4.5. Summary of Key Findings
5. Discussion
5.1. Interpreting the State of the Art
5.2. The Causal Chain of Interdisciplinary Interaction
- How algorithmic technology (RQ1) drives hardware demands (RQ2): The transition to autonomous avatars relies heavily on explicit neural rendering (e.g., 3DGS) and continuous diffusion pipelines. However, these algorithms dictate massive memory bandwidth and continuous tensor operations that directly saturate the strict thermal and latency limits of standalone mobile SoCs. Thus, the specific algorithmic choices of RQ1 inherently force the hardware crises of RQ2, necessitating bespoke collaborative solutions like edge ASICs and CIM architectures.
- How the hardware solution (RQ2) becomes the premise of user experience (RQ3): Without the edge-level hardware acceleration and foveated rendering identified in RQ2, the system inevitably suffers from rendering lag and temporal jitter. As established, even milliseconds of latency between a user’s physical movement and the AI’s predicted kinematic response will rupture the sensorimotor loop. Therefore, the hardware–software co-design (RQ2) is not merely a technical optimization; it is the absolute prerequisite for maintaining the sense of agency and preventing the Kinematic Uncanny Valley (RQ3).
- How deep immersion (RQ3) gives rise to sociotechnical risks (RQ4): When RQ1 and RQ2 successfully align, the result is profound psychological embodiment and high social presence (RQ3). However, it is precisely this deep psychological immersion, facilitated by hyper-realistic, context-aware avatars that clone human kinematics and appearance, that triggers the ethical vulnerabilities of RQ4. The psychological power of a personalized “Virtual Twin” or an affective AI agent is precisely what makes the threats of “Generative Ghosts,” deepfakes, and spatial privacy violations so severe.
5.3. Summary of Key Findings—Future Research Directions
- Standardizing Embodiment Metrics for Autonomous Agents: Current evaluations of the sense of embodiment (SoE) rely heavily on subjective self-reporting instruments, such as the Virtual Embodiment Questionnaire, designed for traditional 1:1 tracked avatars. These instruments are poorly suited to AI-driven systems, where the avatar’s behavior is partially autonomous. Future work must develop standardized, multimodal evaluation frameworks that correlate subjective psychological reporting with real-time physiological signals (e.g., EEG, Galvanic Skin Response), enabling objective measurement of the cognitive dissonance that arises when a generative model mispredicts a user’s intended movement, the “Puppeteer Problem.”
- Edge AI and On-Device Neural Rendering: The reliance on cloud computing for generative motion synthesis introduces latency and privacy risks that are structurally incompatible with immersive VR. Architectural research must prioritize TinyML and edge-computing paradigms, focusing specifically on Computing-in-Memory (CIM) designs and dynamic sparsity algorithms capable of executing complex 4D neural rendering entirely on mobile Systems-on-Chip (SoCs) within a sub-30 W power envelope.
- Ethical Frameworks and “Privacy-by-Design”: The emergence of context-aware agents and “Generative Ghosts” requires regulatory and design intervention that goes beyond risk identification. Future studies must focus on engineering Privacy-by-Design solutions: localized, zero-knowledge processing for spatial data that ensures headset camera feeds never reach the cloud, and cryptographic watermarking for 3D assets that allows human and AI-generated kinematic behaviors to be distinguished and verified.
5.4. Design Guidelines for XR Developers
- Balance visual fidelity with kinematic latency bounds: When hardware constraints force a compromise, visual fidelity (e.g., texture resolution, ray-traced lighting) should not exceed the system’s capacity to maintain real-time kinematics. The penalty of sensorimotor collapse is more detrimental to the sense of agency than a slight reduction in visual realism.
- Implement IK as a deterministic fallback: Systems using predictive AI for full-body tracking from sparse inputs must include a deterministic Inverse Kinematics (IK) layer as a seamless fallback. When the AI model’s confidence drops below threshold, the system should smoothly interpolate back to standard IK, preventing the unintended, agency-breaking movements that collapse immersion.
- Mandate inclusive training data: Generative models must be audited for kinematic bias before deployment. Underlying datasets should explicitly include diverse body types, motor abilities, and assistive devices—including wheelchairs—moving away from the “one-size-fits-all” default rig that many current systems assume.
- Adopt gaze-contingent generative inference: Foveation should not be confined to the graphics rasterization pipeline. On standalone HMDs operating under strict thermal and computational budgets, eye-tracking hardware should govern the AI generation threshold directly. Compute-heavy generative models, such as high-density 3D Gaussian Splatting, should be reserved for the foveal region, with the periphery handled by lightweight, pre-scripted animations. This principle of “semantic foveation” ensures that the system’s most expensive inference is expended only where the user’s visual acuity can resolve the kinematic detail it produces.
5.5. Limitations of This Review
6. Conclusions
Supplementary Materials
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Vogel, D.; Lubos, P.; Steinicke, F. Animationvr-interactive controller-based animating in virtual reality. In Proceedings of the 2018 IEEE 1st Workshop on Animation in Virtual and Augmented Environments (ANIVAE), Reutlingen, Germany, 19 March 2018; IEEE: New York, NY, USA, 2018; pp. 1–6. [Google Scholar]
- Berford, B.; Diaz-Padron, C.; Kaleas, T.; Oz, I.; Penney, D. Building an animation pipeline for VR stories. In ACM SIGGRAPH 2017 Talks; Association for Computing Machinery: New York, NY, USA, 2017; pp. 1–2. [Google Scholar]
- Wu, J.; Lu, E.; Kohli, P.; Freeman, B.; Tenenbaum, J. Learning to see physics via visual de-animation. Adv. Neural Inf. Process. Syst. 2017, 30. [Google Scholar]
- Chirico, A.; Gaggioli, A. When virtual feels real: Comparing emotional responses and presence in virtual and natural environments. Cyberpsychol. Behav. Soc. Netw. 2019, 22, 220–226. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Zhong, W. Evolution and innovations in animation: A comprehensive review and future directions. Concurr. Comput. Pract. Exp. 2024, 36, e7904. [Google Scholar] [CrossRef] [Scilit]
- Jiang, Y.; Yu, C.; Xie, T.; Li, X.; Feng, Y.; Wang, H.; Li, M.; Lau, H.; Gao, F.; Yang, Y.; et al. Vr-gs: A physical dynamics-aware interactive gaussian splatting system in virtual reality. In Proceedings of the ACM SIGGRAPH 2024 Conference Papers, Denver, CO, USA, 27 July–1 August 2024; p. 1. [Google Scholar]
- Song, L.; Chen, L.; Liu, C.; Liu, P.; Xu, C. Texttoon: Real-time text toonify head avatar from single video. In Proceedings of the SIGGRAPH Asia 2024 Conference Papers, Tokyo Japan, 3–6 December 2024; pp. 1–11. [Google Scholar]
- Yoon, J.; Yune, S.; Lee, H.; Putra, A.; Heo, J.; Kim, J.-Y. MoDiff: A 11.0 TOPS/W Diffusion Accelerator with Temporal Data Reuse for Real-Time Text-to-Motion Generation. In Proceedings of the 2025 IEEE Asian Solid-State Circuits Conference (A-SSCC), Daejeon, Republic of Korea, 2–5 November 2025; IEEE: New York, NY, USA, 2025; pp. 88–90. [Google Scholar]
- Tessler, C.; Guo, Y.; Nabati, O.; Chechik, G.; Peng, X.B. Maskedmimic: Unified physics-based character control through masked motion inpainting. ACM Trans. Graph. (TOG) 2024, 43, 209. [Google Scholar] [CrossRef] [Scilit]
- Jiang, J.; Xiao, W.; Lin, Z.; Zhang, H.; Ren, T.; Gao, Y.; Lin, Z.; Cai, Z.; Yang, L.; Liu, Z. Solami: Social vision-language-action modeling for immersive interaction with 3d autonomous characters. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 26887–26898. [Google Scholar]
- Hirzle, T.; Müller, F.; Draxler, F.; Schmitz, M.; Knierim, P.; Hornbæk, K. When XR and AI Meet—A Scoping Review on Extended Reality and Artificial Intelligence—Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems; Association for Computing Machinery: New York, NY, USA, 2023. [Google Scholar] [CrossRef] [Scilit]
- Heo, J.; Putra, A.; Yune, S.; Yoon, J.; Lee, H.; Kim, J.; Kim, J.-Y. 23.10 HuMoniX: A 57.3 fps 12.8 TFLOPS/W Text-to-Motion Processor with Inter-Iteration Output Sparsity and Inter-Frame Joint Similarity. In Proceedings of the 2025 IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA, 15–19 February 2025; IEEE: New York, NY, USA, 2025; pp. 424–426. [Google Scholar]
- Peixoto, B.; Melo, M.; Cabral, L.; Bessa, M. Evaluation of animation and lip-sync of avatars, and user interaction in immersive virtual reality learning environments. In Proceedings of the 2021 International Conference on Graphics and Interaction (ICGI), Porto, Portugal, 4–5 November 2021; IEEE: New York, NY, USA, 2021; pp. 1–7. [Google Scholar]
- Waltemate, T.; Gall, D.; Roth, D.; Botsch, M.; Latoschik, M.E. The impact of avatar personalization and immersion on virtual body ownership, presence, and emotional response. IEEE Trans. Vis. Comput. Graph. 2018, 24, 1643–1652. [Google Scholar] [CrossRef] [Scilit]
- Debarba, H.G.; Chague, S.; Charbonnier, C. On the plausibility of virtual body animation features in virtual reality. IEEE Trans. Vis. Comput. Graph. 2020, 28, 1880–1893. [Google Scholar] [CrossRef] [Scilit]
- Yun, H.; Ponton, J.L.; Andujar, C.; Pelechano, N. Animation fidelity in self-avatars: Impact on user performance and sense of agency. In Proceedings of the 2023 IEEE Conference Virtual Reality and 3D User Interfaces (VR), Shanghai, China, 25–29 March 2023; IEEE: New York, NY, USA, 2023; pp. 286–296. [Google Scholar]
- Zibrek, K.; Martin, S.; McDonnell, R. Is photorealism important for perception of expressive virtual humans in virtual reality? ACM Trans. Appl. Percept. (TAP) 2019, 16, 14. [Google Scholar] [CrossRef] [Scilit]
- Volonte, M.; Robb, A.; Duchowski, A.T.; Babu, S.V. Empirical evaluation of virtual human conversational and affective animations on visual attention in inter-personal simulations. In Proceedings of the 2018 IEEE Conference on Virtual Reality and 3D User Interfaces (VR), Reutlingen, Germany, 18–22 March 2018; IEEE: New York, NY, USA, 2018; pp. 25–32. [Google Scholar]
- Kokkinara, E.; McDonnell, R. The Effect of Animation Realism on Face Ownership and Engagement. In Proceedings of the Facial Analysis and Animation, Vienna, Austria, 11 September 2015; pp. 1–2. [Google Scholar]
- Zibrek, K.; Kokkinara, E.; Mcdonnell, R. The effect of realistic appearance of virtual characters in immersive environments-does the character’s personality play a role? IEEE Trans. Vis. Comput. Graph. 2018, 24, 1681–1690. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zibrek, K.; McDonnell, R. Social presence and place illusion are affected by photorealism in embodied VR. In Proceedings of the 12th ACM SIGGRAPH Conference on Motion, Interaction and Games, Newcastle upon Tyne, UK, 28–30 October 2019; pp. 1–7. [Google Scholar]
- Hashim, M.E.A.; Albakry, N.S.; Mustafa, W.A.; Grahita, B.; Ghani, M.M.; Hanafi, H.F.; Nasir, S.M.; Ugap, C.A. Understanding the impact of animation technology in virtual reality: A systematic literature review. Int. J. Comput. Think. Data Sci. 2024, 1, 53–65. [Google Scholar] [CrossRef] [Scilit]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, 71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kitchenham, B.A.; Budgen, D.; Brereton, P. Evidence-Based Software Engineering and Systematic Reviews; CRC Press: Boca Raton, FL, USA, 2015; Volume 4. [Google Scholar]
- Wohlin, C. Guidelines for snowballing in systematic literature studies and a replication in software engineering. In Proceedings of the 18th International Conference on Evaluation and Assessment in Software Engineering, London, UK, 13–14 May 2014; pp. 1–10. [Google Scholar]
- Starke, S.; Starke, P.; He, N.; Komura, T.; Ye, Y. Categorical Codebook Matching for Embodied Character Controllers. ACM Trans. Graph. 2024, 43, 142. [Google Scholar] [CrossRef] [Scilit]
- Zou, S.; Xu, Y.; Sarafianos, N.; Bogo, F.; Tung, T.; Si, W.; Cheng, L. Generating high-fidelity clothed human dynamics with temporal diffusion. ACM Trans. Multimed. Comput. Commun. Appl. 2025, 21, 88. [Google Scholar] [CrossRef] [Scilit]
- Xiang, D.; Bagautdinov, T.; Stuyck, T.; Prada, F.; Romero, J.; Xu, W.; Saito, S.; Guo, J.; Smith, B.; Shiratori, T.; et al. Dressing Avatars: Deep Photorealistic Appearance for Physically Simulated Clothing. ACM Trans. Graph. 2022, 41, 222. [Google Scholar] [CrossRef] [Scilit]
- Wang, R.; Xu, P.; Shi, H.; Schumann, E.; Liu, C.K. Fürelise: Capturing and physically synthesizing hand motion of piano performance. In Proceedings of the SIGGRAPH Asia 2024 Conference Papers, Tokyo, Japan, 3–6 December 2024; pp. 1–11. [Google Scholar]
- Tang, J.; Wu, Y.; Li, M.; Wang, Z. Talking Face Generation Based on Information Bottleneck and Complementary Representations. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, Queensland, Austral, 1–5 November 2021; Association for Computing Machinery: New York, NY, USA, 2021; pp. 3443–3447. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Wang, T.; Hancke, G.; Liu, Z.; Lau, R.W. ThemeStation: Generating Theme-Aware 3D Assets from Few Exemplars—ACM SIGGRAPH 2024 Conference Papers; Association for Computing Machinery: New York, NY, USA, 2024. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Wu, J.; Liu, C.K. Object Motion Guided Human Motion Synthesis. ACM Trans. Graph. 2023, 42, 197. [Google Scholar] [CrossRef] [Scilit]
- Bertiche, H.; Madadi, M.; Escalera, S. Neural Cloth Simulation. ACM Trans. Graph. 2022, 41, 220. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.-P.; Pu, C.-F.; Guo, M.-H.; Cao, Y.-P.; Hu, S.-M. One Model to Rig Them All: Diverse Skeleton Rigging with UniRig. ACM Trans. Graph. 2025, 44, 123. [Google Scholar] [CrossRef] [Scilit]
- Ma, H.; Zhang, T.; Sun, S.; Yan, X.; Han, K.; Xie, X. CVTHead: One-shot Controllable Head Avatar with Vertex-feature Transformer. In Proceedings of the 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 3–8 January 2024; pp. 6119–6129. [Google Scholar]
- Taubner, F.; Zhang, R.; Tuli, M.; Lindell, D.B. CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models. In Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025; pp. 5318–5330. [Google Scholar]
- Pinyoanuntapong, E.; Wang, P.; Lee, M.; Chen, C. MMM: Generative Masked Motion Model; IEEE: New York, NY, USA, 2024; pp. 1546–1555. [Google Scholar] [CrossRef] [Scilit]
- Fu, D.; Sun, T.; Fang, P.; Cai, X.; Kim, H. MOGO: Residual Quantized Hierarchical Causal Transformer for High-Quality and Real-Time 3D Human Motion Generation. arXiv 2025. [Google Scholar] [CrossRef] [Scilit]
- Xin, S.; Wang, H.; Zhang, S.Q. A3FR: Agile 3D Gaussian Splatting with Incremental Gaze Tracked Foveated Rendering in Virtual Reality—Proceedings of the 39th ACM International Conference on Supercomputing; Association for Computing Machinery: New York, NY, USA, 2025; pp. 279–292. [Google Scholar] [CrossRef] [Scilit]
- Feng, Y.; Lin, W.; Cheng, Y.; Liu, Z.; Leng, J.; Guo, M.; Chen, C.; Sun, S.; Zhu, Y. Lumina: Real-Time Neural Rendering by Exploiting Computational Redundancy—Proceedings of the 52nd Annual International Symposium on Computer Architecture; Association for Computing Machinery: New York, NY, USA, 2025; pp. 1925–1939. [Google Scholar] [CrossRef] [Scilit]
- Liu, F.; Li, H.; Zhu, B.; Wang, Z.; Song, Z.; Guan, H.; Jiang, L. ASDR: Exploiting Adaptive Sampling and Data Reuse for CIM-based Instant Neural Rendering—Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3; Association for Computing Machinery: New York, NY, USA, 2025; pp. 18–33. [Google Scholar] [CrossRef] [Scilit]
- Song, X.; Wen, Y.; Hu, X.; Liu, T.; Zhou, H.; Han, H.; Zhi, T.; Du, Z.; Li, W.; Zhang, R.; et al. Cambricon-R: A Fully Fused Accelerator for Real-Time Learning of Neural Scene Representation—Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture; Association for Computing Machinery: New York, NY, USA, 2023; pp. 1305–1318. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Liang, J.; Peng, J.; Xu, J.; Zhang, W. SpNeRF: Memory Efficient Sparse Volumetric Neural Rendering Accelerator for Edge Devices. In Proceedings of the 2025 Design, Automation & Test in Europe Conference (DATE), Lyon, France, 31 March–2 April 2025; pp. 1–7. [Google Scholar]
- Chen, J.; Wang, J.; Zhang, Y.; Pandey, R.; Beeler, T.; Habermann, M.; Theobalt, C. EgoAvatar: Egocentric View-Driven and Photorealistic Full-body Avatars—SIGGRAPH Asia 2024 Conference Papers; Association for Computing Machinery: New York, NY, USA, 2024. [Google Scholar] [CrossRef] [Scilit]
- Medin, S.C.; Li, G.; Du, R.; Garbin, S.; Davidson, P.; Wornell, G.W.; Beeler, T.; Meka, A. FaceFolds: Meshed Radiance Manifolds for Efficient Volumetric Rendering of Dynamic Faces. Proc. ACM Comput. Graph. Interact. Tech. 2024, 7, 23. [Google Scholar] [CrossRef] [Scilit]
- Hung, H.-H.; Do, H.-P.; Li, Y.-H.; Huang, C.-C. TimeNeRF: Building Generalizable Neural Radiance Fields across Time from Few-Shot Input Views—Proceedings of the 32nd ACM International Conference on Multimedia; Association for Computing Machinery: New York, NY, USA, 2024; pp. 253–262. [Google Scholar] [CrossRef] [Scilit]
- Yu, A.; Li, R.; Tancik, M.; Li, H.; Ng, R.; Kanazawa, A. PlenOctrees for Real-time Rendering of Neural Radiance Fields. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, Canada, 10–17 October 2021; pp. 5732–5741. [Google Scholar]
- Qin, D.; Lin, H.; Zhang, Q.; Qiao, K.; Zhang, L.; Saito, J.; Zhao, Z.; Yu, J.; Xu, L.; Komura, T. Instant Gaussian Splatting Generation for High-Quality and Real-Time Facial Asset Rendering. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 1–15. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, J.; Wang, Y.; Wang, Y.; Wang, Y.; Cui, S.; Wang, F. Mobile Volumetric Video Streaming System through Implicit Neural Representation—Proceedings of the 2023 Workshop on Emerging Multimedia Systems; Association for Computing Machinery: New York, NY, USA, 2023; pp. 1–7. [Google Scholar] [CrossRef] [Scilit]
- Rao, C.; Yu, H.; Wan, H.; Zhou, J.; Zheng, Y.; Wu, M.; Ma, Y.; Chen, A.; Yuan, B.; Zhou, P.; et al. ICARUS: A Specialized Architecture for Neural Radiance Fields Rendering. ACM Trans. Graph. 2022, 41, 234. [Google Scholar] [CrossRef] [Scilit]
- Edwards, D.; Rawat, D.B. SleepWalker: Constrastive Fine-tuning Technique for Text to Kinematics Models for Human Computer Interaction. In Proceedings of the 2024 33rd International Conference on Computer Communications and Networks (ICCCN), Kailua-Kona, HI, USA, 29–31 July 2024; pp. 1–7. [Google Scholar]
- Klar, M.; Fischer, F.; Fleig, A.; Bachinski, M.; Müller, J. Simulating interaction movements via model predictive control. ACM Trans. Comput.-Hum. Interact. 2023, 30, 44. [Google Scholar] [CrossRef] [Scilit]
- Xu, Y.; Wang, S.; Hasegawa, S. Realistic Dexterous Manipulation of Virtual Objects with Physics-Based Haptic Rendering—ACM SIGGRAPH 2023 Emerging Technologies; Association for Computing Machinery: New York, NY, USA, 2023. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Jones, B.; Rintel, S.; Neustaedter, C. XRmas: Extended Reality Multi-Agency Spaces for a Magical Remote Christmas—Companion Publication of the 2021 Conference on Computer Supported Cooperative Work and Social Computing; Association for Computing Machinery: New York, NY, USA, 2021; pp. 203–207. [Google Scholar] [CrossRef] [Scilit]
- Jiao, C.; Wang, Y.; Zhang, G.; Bâce, M.; Hu, Z.; Bulling, A. DiffGaze: A Diffusion Model for Modelling Fine-grained Human Gaze Behaviour on 360° Images. ACM Trans. Interact. Intell. Syst. 2026, 16, 5. [Google Scholar] [CrossRef] [Scilit]
- Yang, S.; Tsui, Y.H.; Wang, X.; Alhilal, A.; Mogavi, R.H.; Wang, X.; Hui, P. From Prompt to Metaverse: User Perceptions of Personalized Spaces Crafted by Generative AI—Companion Publication of the 2024 Conference on Computer-Supported Cooperative Work and Social Computing; Association for Computing Machinery: New York, NY, USA, 2024; pp. 497–504. [Google Scholar] [CrossRef] [Scilit]
- Tao, Y.; Wang, C.Y.; Wilson, A.D.; Ofek, E.; Gonzalez-Franco, M. Embodying Physics-Aware Avatars in Virtual Reality. Int. Conf. Hum. Factors Comput. Syst. 2023. [Google Scholar] [CrossRef] [Scilit]
- Boucaud, F.; Pelachaud, C.; Thouvenin, I. “‘\’It patted my arm\”: Investigating Social Touch from a Virtual Agent—Proceedings of the 11th International Conference on Human-Agent Interaction; Association for Computing Machinery: New York, NY, USA, 2023; pp. 72–80. [Google Scholar] [CrossRef] [Scilit]
- Atkare, A.; Bhoyar, A.; Akre, O.; Bhosle, O.; Tembhurne, T.; Bhanuse, S. Development of an Interactive AI Mentor: A Full-Body Digital Human for Real-Time Conversational Learning. In Proceedings of the 2025 IEEE Pune Section International Conference (PuneCon), Pune, India, 12–15 December 2025. [Google Scholar]
- Ng, E.; Zhang, S.; Chen, Z.; Zollhoefer, M.; Richard, A. SARAH: Spatially Aware Real-time Agentic Humans. arXiv 2026, arXiv:2602.18432. [Google Scholar] [CrossRef] [Scilit]
- Guzov, V.; Jiang, Y.; Hong, F.; Pons-Moll, G.; Newcombe, R.; Liu, C.K.; Ye, Y.; Ma, L. Hmd 2: Environment-aware motion generation from single egocentric head-mounted device. In Proceedings of the 2025 International Conference on 3D Vision (3DV), Singapore, 25–28 March 2025; IEEE: New York, NY, USA, 2025; pp. 1394–1405. [Google Scholar]
- Choi, B.; Jang, D.-K.; Yang, D.; Jang, D.-Y. MOVIN TRACIN’: Move Outside the Box—ACM SIGGRAPH 2024 Real-Time Live! Association for Computing Machinery: New York, NY, USA, 2024. [Google Scholar] [CrossRef] [Scilit]
- Morris, M.R.; Brubaker, J.R. Generative Ghosts: Anticipating Benefits and Risks of AI Afterlives—Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems; Association for Computing Machinery: New York, NY, USA, 2025. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Mollyn, V.; Yuan, K.; Carrington, P. WheelPoser: Sparse-IMU Based Body Pose Estimation for Wheelchair Users—Proceedings of the 26th International ACM SIGACCESS Conference on Computers and Accessibility; Association for Computing Machinery: New York, NY, USA, 2024. [Google Scholar] [CrossRef] [Scilit]
- Kim, J.-S. Virtual Performance Augmentation for Senior Users on Universal XR Metaverses. In Proceedings of the 2024 15th International Conference on Information and Communication Technology Convergence (ICTC), Jeju Island, Republic of Korea, 16–18 October 2024; pp. 1334–1337. [Google Scholar]
- Mengoni, P.; Jiandong, D.S.; Zixin, L.; Yun, P.W.P. GenAI avatars in VR: Role of presence, health, and technological factors. Comput. Educ. X Real. 2026, 8, 100141. [Google Scholar] [CrossRef] [Scilit]












| Database | Search Terms | Search Criteria |
|---|---|---|
| IEEE Xplore | (“Virtual Reality” OR “VR” OR “Spatial Computing” OR “Metaverse”) AND (“character animation” OR “procedural animation” OR “physics-based” OR “generative AI” OR “text-to-motion” OR “neural rendering”) AND (“embodiment” OR “immersion” OR “latency” OR “performance” OR “agency”) | Field: abstract; dates: 2021–2026; 52 initial results |
| Scopus/ScienceDirect | TITLE-ABS-KEY ((“Virtual Reality” OR “VR” OR “Spatial Computing” OR “Virtual Worlds”) AND (“animation” OR “procedural” OR “physics-based” OR “generative AI” OR “text-to-motion” OR “avatar”) AND (“embodiment” OR “immersion” OR “presence” OR “real-time performance” OR “latency”)) | Fields: title, abstract or author-specified keywords; dates: 2021–2026; 30 initial results |
| ACM DL | [[Abstract: “Virtual Reality”] OR [Abstract: “Spatial Computing”]] AND [[Abstract: “animation”] OR [Abstract: “physics-based”] OR [Abstract: “generative AI”] OR [Abstract: “text-to-motion”]] AND [[Abstract: “embodiment”] OR [Abstract: “immersion”] OR [Abstract: “performance”]] | All fields; search for articles, books, chapters, conference papers; dates: 2021–2026; 157 initial results |
| Elicit (AI Search) | Query A: “How do generative AI and text-to-motion impact real-time character animation and latency in virtual reality and spatial computing?” Query B: “What is the effect of physics-based animation and AI-driven realistic avatars on user embodiment and immersion in VR?” | AI Semantic Search; dates: 2021–2026; 50 initial results |
| Type | Criteria |
|---|---|
| Basic aspects | Title, authors, publication year, article type, database source, number of citations. |
| Inclusion and exclusion criteria | IC1: Studies focusing on immersive virtual reality (VR), Spatial Computing, or Metaverse environments. |
| IC2: Research proposing or evaluating character animation techniques (e.g., Generative AI, physics-based, procedural, text-to-motion). | |
| IC3: Empirical studies assessing user experience (embodiment, immersion, agency) OR technical performance (latency, frame rates). | |
| IC4: Peer-reviewed journal articles and full conference papers. | |
| IC5: Published within the last 5 years (2021–2026). | |
| IC6: Written in English. | |
| EC1: Studies exclusively focused on augmented reality (AR) without applicability to fully immersive VR. | |
| EC2: Animation techniques strictly for traditional 2D screens/movies without interactive or real-time elements. | |
| EC3: Papers lacking clear empirical evaluation, technical validation, or user data. | |
| EC4: Short abstracts, posters, opinions, or non-peer-reviewed pre-prints (unless highly cited foundational AI models). | |
| EC5: Publications in languages other than English. | |
| Review questions | Extent to which the study addresses RQ1, RQ2, RQ3, or RQ4. |
| Feature | Explanation/Extracted Variables |
|---|---|
| Basic Metadata | Title, authors, year, publisher, article type (journal/conference). |
| Animation Technology (RQ1) | Primary AI model (e.g., Diffusion, NeRF, 3DGS, LLM), representation type (implicit vs. explicit), avatar scope (full-body, face, cloth, hands), input modality (text-to-motion, audio-driven, sparse tracking). |
| Hardware and Performance (RQ2) | Target platform (standalone HMD vs. PC/Cloud), optimization techniques (ASICs, CIM, foveated rendering), latency/FPS, memory footprint and bandwidth reduction (MB/GB), energy/power efficiency (TOPS/W, mW). |
| UX and Embodiment (RQ3) | Psychological metrics (body ownership, agency, social presence, uncanny valley), study design and sample size (user study N = X, method), task/interaction type (object manipulation, social touch, locomotion). |
| Ethics and Inclusion (RQ4) | Mentions of accessibility, racial representation, data privacy, digital afterlife risks, and proposed design guidelines/frameworks. |
| Id Ref | Primary Focus (RQs) | Keywords Thematic Labels | Key Technology AI Model | Target Platform | Main Contribution |
|---|---|---|---|---|---|
| 1—[9] | RQ1, RQ3 | Animation_tech, Procedural_Techniques, Scene_Awareness | Masked Motion Inpainting | VR/3D Engines | Unified physics-based character controller allowing for dynamic adaptation. |
| 2—[26] | RQ1, RQ2 | AI_or_Generative, Animation_tech, Embodiment_Interaction | Codebook Matching/RL | Immersive VR | End-to-end full-body motion synthesis from sparse sensors. |
| 3—[27] | RQ1, RQ2 | AI_or_Generative, Animation_tech, Procedural_Techniques | Temporal Diffusion | VR/Gaming | Generates humans with high-fidelity realistic clothing details. |
| 4—[28] | RQ1, RQ3 | Animation_tech, AI_or_Generative, Embodiment_Interaction | Deep Photorealistic Appearance | VR | Physically simulated clothing integrated with photorealistic rendering. |
| 5—[29] | RQ1, RQ3 | AI_or_Generative, Animation_tech, VR_Core | Diffusion + RL Hybrid | VR/AR | Physically plausible, highly dexterous hand motion synthesis. |
| 6—[30] | RQ1, RQ3 | AI_or_Generative, Animation_tech, VR_Core | Information Bottleneck | VR | Synthesizes high-fidelity, lip-synchronized avatars from audio. |
| 7—[31] | RQ1, RQ2 | AI_or_Generative, Procedural_Techniques | Dual Score Distillation | Spatial Computing | Theme-aware 3D asset generation from few exemplars. |
| 8—[32] | RQ1, RQ3 | AI_or_Generative, Animation_tech, Scene_Awareness | Diffusion Models | VR/AR | Inverse interaction synthesis for human–object manipulation. |
| 9—[33] | RQ1, RQ2 | AI_or_Generative, Animation_tech, Procedural_Techniques | Unsupervised Deep Learning | VR | Real-time cloth dynamics avoiding heavy deterministic solvers. |
| 10—[34] | RQ1, RQ2 | AI_or_Generative, Animation_tech, Technical_Performance | Autoregressive Transformers | 3D Platforms | Automated diverse skeleton rigging bypassing manual modeling. |
| 11—[35] | RQ1, RQ2 | AI_or_Generative, Animation_tech, Technical_Performance | Vertex-Feature Transformer | VR | One-shot controllable neural head avatars driven by minimal data. |
| 12—[36] | RQ1, RQ3 | AI_or_Generative, Animation_tech, Embodiment_Interaction | Multi-View Diffusion | VR | Creates animatable 4D portrait avatars utilizing multi-view diffusion. |
| 13—[37] | RQ1, RQ2 | AI_or_Generative, Animation_tech, Technical_Performance | Masked Modeling | VR/AR | High-speed text-to-motion generation faster than standard diffusion. |
| 14—[38] | RQ1, RQ2 | AI_or_Generative, Animation_tech, Technical_Performance | Hierarchical Causal Transformer | VR/AR | One-pass, real-time continuous motion generation. |
| 15—[7] | RQ1, RQ2 | AI_or_Generative, Animation_tech, Technical_Performance | Neural Rendering | Mobile VR | Real-time text-driven stylization of 3D head avatars. |
| 16—[6] | RQ1, RQ2 | AI_or_Generative, Procedural_Techniques, Technical_Performance | 3D Gaussian Splatting + Physics | VR | Real-time physical dynamics-aware interactive Gaussian Splatting. |
| 17—[39] | RQ2, RQ3 | Technical_Performance, Eye_Tracking_Gaze, VR_Core | 3DGS + Foveated Rendering | Standalone VR | 2x rendering latency reduction exploiting user gaze tracking. |
| 18—[40] | RQ1, RQ2 | Technical_Performance, AI_or_Generative, VR_Core | 3DGS/Neural Rendering | Mobile SoCs/GPU | 4.5x speedup and energy reduction by pruning redundancy. |
| 19—[41] | RQ2 | Technical_Performance, AI_or_Generative, VR_Core | Computing-in-Memory (CIM) | Edge VR | 9.55x speedup via adaptive sampling and memory data reuse. |
| 20—[42] | RQ1, RQ2 | Technical_Performance, AI_or_Generative, VR_Core | Neural Scene Accelerator | Edge Devices | Fully fused ASIC accelerator enabling real-time neural scene learning. |
| 21—[43] | RQ1, RQ2 | Technical_Performance, AI_or_Generative, VR_Core | Sparse Volumetric Rendering | Edge Devices | Memory efficient architecture cutting memory size for edge NeRFs. |
| 22—[8] | RQ1, RQ2 | Technical_Performance, AI_or_Generative, Animation_tech | Diffusion Accelerator (ASIC) | Standalone VR | Achieves 11.0 TOPS/W energy efficiency for text-to-motion. |
| 23—[44] | RQ1, RQ2, RQ3 | AI_or_Generative, VR_Core, Embodiment_Interaction | Egocentric View-Driven Model | Standalone VR | Generates photorealistic full-body avatars using only HMD cameras. |
| 24—[45] | RQ1, RQ2 | AI_or_Generative, Technical_Performance, Animation_tech | Meshed Radiance Manifolds | VR/AR | Efficient volumetric rendering optimized for real-time dynamic faces. |
| 25—[46] | RQ1, RQ2 | AI_or_Generative, Scene_Awareness, Technical_Performance | Few-Shot Temporal NeRF | VR | Generalizable neural rendering across arbitrary times from sparse inputs. |
| 26—[47] | RQ1, RQ2 | Technical_Performance, AI_or_Generative, Procedural_Techniques | Spherical Harmonics + Octrees | VR | Delivers over 150 FPS real-time rendering speed for NeRFs. |
| 27—[48] | RQ1, RQ2 | AI_or_Generative, Technical_Performance, VR_Core | 3DGS | Mobile/VR | High-quality architecture for scalable facial asset rendering. |
| 28—[12] | RQ2 | Technical_Performance, AI_or_Generative, VR_Core | 14 nm ASIC | Edge/VR | Operates at 57.3 fps as a low-power text-to-motion processor. |
| 29—[49] | RQ1, RQ2 | AI_or_Generative, Technical_Performance, VR_Core | Implicit Neural Representations | Mobile VR | Severe bandwidth reduction enabling volumetric video streaming. |
| 30—[50] | RQ2 | Technical_Performance, AI_or_Generative, VR_Core | Specialized NeRF Architecture | Edge/VR | Hardware–software co-designed pipeline accelerating NeRF inferences. |
| 31—[51] | RQ1, RQ3 | AI_or_Generative, Animation_tech, Embodiment_Interaction | Contrastive Fine-Tuning | AR/VR | Accessible few-shot fine-tuning for personalized human motion. |
| 32—[52] | RQ1, RQ2, RQ3 | Procedural_Techniques, VR_Core, User_Study | Model Predictive Control (MPC) | AR/VR | Computes biologically realistic mid-air pointing mimicking constraints. |
| 33—[53] | RQ1, RQ3 | Haptics_Interaction, Procedural_Techniques, Embodiment_Interaction | Physics-Based Haptic Rendering | VR | Precise virtual object manipulation through physics-driven haptics. |
| 34—[54] | RQ1, RQ3 | Embodiment_Interaction, VR_Core, Animation_tech | Extended Reality Multi-Agency | Social VR | Facilitates asymmetric remote communication preserving social presence. |
| 35—[55] | RQ1, RQ3 | AI_or_Generative, Eye_Tracking_Gaze, Scene_Awareness | Diffusion Model | 360/VR | Predicts natural, fine-grained saccadic human gaze behavior. |
| 36—[10] | RQ1, RQ3 | AI_or_Generative, Embodiment_Interaction, VR_Core | Vision–Language–Action (VLA) | VR | Social VLA modeling for immersive multi-modal interaction. |
| 37—[56] | RQ1, RQ2, RQ3 | AI_or_Generative, User_Study, Technical_Performance | AIGC/Generative AI | VR/Metaverse | Evaluates user immersion in personalized spaces crafted by AI. |
| 38—[57] | RQ2, RQ3 | Embodiment_Interaction, Technical_Performance, VR_Core | Physics-Aware Tracking | VR | Demonstrates physically plausible reactions enhance body ownership. |
| 39—[58] | RQ3, RQ4 | Embodiment_Interaction, User_Study, Haptics_Interaction | Virtual Agents + Haptics | VR | Investigates how social touch from virtual agents improves bonding. |
| 40—[59] | RQ1, RQ3 | AI_or_Generative, User_Behavior_Analysis, VR_Core | LLM + Speech Technologies | VR | Implements a full-body digital human reducing cognitive load. |
| 41—[60] | RQ1, RQ2, RQ3 | AI_or_Generative, Technical_Performance, Animation_tech | Spatially Aware Agentic Humans | VR | Provides agents with spatial context to deliver real-time interactive behaviors. |
| 42—[61] | RQ1, RQ2 | AI_or_Generative, VR_Core, Technical_Performance | Diffusion + SLAM | VR | Environment-aware motion generation ensuring accurate foot placement. |
| 43—[62] | RQ1, RQ2 | AI_or_Generative, Animation_tech, Technical_Performance | Deep Learning/GenAI | VR | Markerless self-occlusion solving for full-body avatar reconstruction. |
| 44—[63] | RQ3, RQ4 | Framework_Paper, VR_Core, User_Behavior_Analysis | Personalized Generative Agents | Spatial Computing | Anticipates risks and benefits surrounding AI models mimicking personas. |
| 45—[64] | RQ2, RQ3, RQ4 | Technical_Performance, Embodiment_Interaction, VR_Core | Sparse-IMU Pose Tracking | Inclusive VR | Provides robust and accessible body tracking for wheelchair users. |
| 46—[65] | RQ1, RQ3, RQ4 | Embodiment_Interaction, Animation_tech, User_Behavior_Analysis | IK + Transformers | Inclusive VR | Translates limited physical movements into full avatar actions for seniors. |
| 47—[66] | RQ3, RQ4 | AI_or_Generative, User_Behavior_Analysis, Technical_Performance | GPT-based Avatars | VR | Evaluates functional usability versus realism on reducing motion sickness. |
| 48—[11] | RQ1, RQ2, RQ3, RQ4 | Survey_Paper, VR_Core, Framework_Paper | Scoping Review | XR | Establishes the foundational baseline for the convergence of AI and XR. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Theodoropoulos, A. From Pre-Rendered to Autonomous: A Systematic Review of AI-Driven Character Animation and Embodiment in Virtual Reality. Virtual Worlds 2026, 5, 20. https://doi.org/10.3390/virtualworlds5020020
Theodoropoulos A. From Pre-Rendered to Autonomous: A Systematic Review of AI-Driven Character Animation and Embodiment in Virtual Reality. Virtual Worlds. 2026; 5(2):20. https://doi.org/10.3390/virtualworlds5020020
Chicago/Turabian StyleTheodoropoulos, Anastasios. 2026. "From Pre-Rendered to Autonomous: A Systematic Review of AI-Driven Character Animation and Embodiment in Virtual Reality" Virtual Worlds 5, no. 2: 20. https://doi.org/10.3390/virtualworlds5020020
APA StyleTheodoropoulos, A. (2026). From Pre-Rendered to Autonomous: A Systematic Review of AI-Driven Character Animation and Embodiment in Virtual Reality. Virtual Worlds, 5(2), 20. https://doi.org/10.3390/virtualworlds5020020
