A Lightweight 3DMM-CNN Pipeline for Real-Time Single-Image 3D Face Reconstruction: Prototyping Personalised Avatars for Extended Reality Applications
Abstract
1. Introduction
- RQ1. Can a four-channel (RGB + landmark-mask) 3DMM-CNN fusion pipeline reconstruct an animation-ready 3D face from a single unconstrained image with sufficient geometric fidelity and pose robustness for XR avatar prototyping?
- RQ2. How does the proposed pipeline perform against representative state-of-the-art baselines under stratified yaw-angle conditions that reflect the real-world capture variability of XR end users?
- RQ3. Can the reconstructed mesh be operationalised through an interactive prototype that exports XR-ready assets while preserving the topology requirements of mainstream XR engines, such as Unity and Unreal Engine?
2. Related Work
2.1. Avatars in Extended Reality
2.2. 3D Morphable Models and Their Successors
2.3. Single-Image 3D Face Reconstruction with Deep Learning
2.4. Real-Time and XR-Specific Considerations
2.5. Lightweight and Real-Time 3D Face Reconstruction
3. Materials and Methods
3.1. Overall Pipeline
3.2. Dataset and Pre-Processing
3.3. Network Architecture
3.4. Multi-Task Loss
3.5. Training Configuration
3.6. Interactive Prototype and XR Asset Export
Exported Asset Specifications
3.7. Evaluation Protocol
4. Results
4.1. Convolutional Network Construction and Hyperparameter Selection
4.1.1. Architecture and Multi-Modal Adaptation
4.1.2. Batch Size Ablation
4.1.3. Robustness Across Yaw Strata
4.1.4. Overall Test-Set Metrics
4.1.5. Comparison with State-of-the-Art Baselines
4.1.6. Ablation Studies on the Landmark-Mask Channel
4.2. Interactive Prototype Walk-Through
4.3. XR Asset Export and Engine Round-Trip
End-to-End Time-Cost Breakdown
5. Discussion
5.1. Implications for Real-World XR Applications
5.1.1. Immersive Learning Avatars
5.1.2. AR-Mediated Telepresence
5.1.3. MR-Based Collaborative Design
5.1.4. Information-Systems Productivity Considerations
5.2. Threats to Validity
5.3. Comparison with the Original Conference Version
5.4. Limitations and Future Work
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| 3DMM | 3D Morphable Model |
| AR | Augmented Reality |
| BFM | Basel Face Model |
| CNN | Convolutional Neural Network |
| FACS | Facial Action Coding System |
| FBX | Filmbox (Autodesk 3D asset format) |
| HCI | Human–Computer Interaction |
| HMD | Head-Mounted Display |
| LFW | Labeled Faces in the Wild |
| MAE | Mean Absolute Error |
| MAPE | Mean Absolute Percentage Error |
| MR | Mixed Reality |
| MSE | Mean Squared Error |
| NME | Normalised Mean Error |
| OBJ | Wavefront 3D Object File Format |
| R2 | Coefficient of Determination |
| ResNet | Residual Network |
| RMSE | Root Mean Squared Error |
| ROI | Region of Interest |
| SME | Small and Medium-sized Enterprise |
| UV | UV Texture Coordinate System |
| VR | Virtual Reality |
| XR | Extended Reality |
References
- Schroeder, R. Being There Together: Social Interaction in Shared Virtual Environments; Oxford University Press: New York, NY, USA, 2010. [Google Scholar] [CrossRef] [Scilit]
- Slater, M.; Sanchez-Vives, M.V. Enhancing our lives with immersive virtual reality. Front. Robot. AI 2016, 3, 74. [Google Scholar] [CrossRef] [Scilit]
- Kilteni, K.; Groten, R.; Slater, M. The sense of embodiment in virtual reality. Presence Teleoperators Virtual Environ. 2012, 21, 373–387. [Google Scholar] [CrossRef] [Scilit]
- International Data Corporation (IDC). Worldwide Augmented and Virtual Reality Spending Guide; IDC Report; IDC: Framingham, MA, USA, 2024. [Google Scholar]
- Motion Picture Association. 2023 THEME Report; Motion Picture Association: Washington, DC, USA, 2023. [Google Scholar]
- 3D Scan Store. Industrial 3D Scanning Cost Benchmark; Technical White Paper; 3D Scan Store: London, UK, 2023. [Google Scholar]
- Blanz, V.; Vetter, T. A morphable model for the synthesis of 3D faces. In Proceedings of the 26th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH ’99), Los Angeles, CA, USA, 8–13 August 1999; pp. 187–194. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar] [CrossRef] [Scilit]
- Deng, Y.; Yang, J.; Xu, S.; Chen, D.; Jia, Y.; Tong, X. Accurate 3D face reconstruction with weakly-supervised learning: From single image to image set. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Long Beach, CA, USA, 16–17 June 2019. [Google Scholar] [CrossRef] [Scilit]
- Jackson, A.S.; Bulat, A.; Argyriou, V.; Tzimiropoulos, G. Large pose 3D face reconstruction from a single image via direct volumetric CNN regression. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 1031–1039. [Google Scholar] [CrossRef] [Scilit]
- Zhu, X.; Liu, X.; Lei, Z.; Li, S.Z. Face alignment in full pose range: A 3D total solution. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 41, 78–92. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, H.; Zhu, H.; Wang, Y.; Huang, M.; Shen, Q.; Yang, R.; Cao, X. FaceScape: A large-scale high-quality 3D face dataset and detailed riggable 3D face prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 601–610. [Google Scholar] [CrossRef] [Scilit]
- Feng, Y.; Feng, H.; Black, M.J.; Bolkart, T. Learning an animatable detailed 3D face model from in-the-wild images. ACM Trans. Graph. 2021, 40, 88. [Google Scholar] [CrossRef] [Scilit]
- Sanyal, S.; Bolkart, T.; Feng, H.; Black, M.J. Learning to regress 3D face shape and expression from an image without 3D supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 7763–7772. [Google Scholar] [CrossRef] [Scilit]
- Shang, J.; Shen, T.; Li, Z.; Kang, B.; Ding, J.; Zhu, Z.; Quan, L. Self-supervised monocular 3D face reconstruction by occlusion-aware multi-view geometry consistency. In Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020; pp. 1–16. [Google Scholar] [CrossRef] [Scilit]
- Daněček, R.; Black, M.J.; Bolkart, T. EMOCA: Emotion driven monocular face capture and animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 20311–20322. [Google Scholar] [CrossRef] [Scilit]
- Dwivedi, Y.K.; Hughes, L.; Baabdullah, A.M.; Ribeiro-Navarrete, S.; Giannakis, M.; Al-Debei, M.M.; Wamba, S.F. Metaverse beyond the hype: Multidisciplinary perspectives on emerging challenges, opportunities, and agenda for research, practice and policy. Int. J. Inf. Manag. 2022, 66, 102542. [Google Scholar] [CrossRef] [Scilit]
- Mystakidis, S. Metaverse. Encyclopedia 2022, 2, 486–497. [Google Scholar] [CrossRef] [Scilit]
- Maloney, D.; Freeman, G.; Wohn, D.Y. “Talking without a voice”: Understanding non-verbal communication in social virtual reality. Proc. ACM Hum.-Comput. Interact. 2020, 4, 175. [Google Scholar] [CrossRef] [Scilit]
- Garau, M.; Slater, M.; Vinayagamoorthy, V.; Brogni, A.; Steed, A.; Sasse, M.A. The impact of avatar realism and eye gaze control on perceived quality of communication in a shared immersive virtual environment. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, Fort Lauderdale, FL, USA, 5–10 April 2003; pp. 529–536. [Google Scholar] [CrossRef] [Scilit]
- Banakou, D.; Hanumanthu, P.D.; Slater, M. Virtual embodiment of white people in a black virtual body leads to a sustained reduction in their implicit racial bias. Front. Hum. Neurosci. 2016, 10, 601. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Radianti, J.; Majchrzak, T.A.; Fromm, J.; Wohlgenannt, I. A systematic review of immersive virtual reality applications for higher education: Design elements, lessons learned, and research agenda. Comput. Educ. 2020, 147, 103778. [Google Scholar] [CrossRef] [Scilit]
- Paysan, P.; Knothe, R.; Amberg, B.; Romdhani, S.; Vetter, T. A 3D face model for pose and illumination invariant face recognition. In Proceedings of the 6th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), Genoa, Italy, 2–4 September 2009; pp. 296–301. [Google Scholar] [CrossRef] [Scilit]
- Li, T.; Bolkart, T.; Black, M.J.; Li, H.; Romero, J. Learning a model of facial shape and expression from 4D scans. ACM Trans. Graph. 2017, 36, 194. [Google Scholar] [CrossRef] [Scilit]
- Feng, Y.; Wu, F.; Shao, X.; Wang, Y.; Zhou, X. Joint 3D face reconstruction and dense alignment with position map regression network. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 534–551. [Google Scholar] [CrossRef] [Scilit]
- Ruan, Z.; Zou, C.; Wu, L.; Wu, G.; Wang, L. SADRNet: Self-aligned dual face regression networks for robust 3D dense face alignment and reconstruction. IEEE Trans. Image Process. 2021, 30, 5793–5806. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Filntisis, P.P.; Roussos, A.; Maragos, P. SPECTRE: Visual speech-informed perceptual 3D facial expression reconstruction from videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vancouver, BC, Canada, 18–22 June 2023; pp. 5712–5721. [Google Scholar] [CrossRef] [Scilit]
- Wu, C.Y.; Yen, H.S.; Lai, S.H. Synergy between 3DMM and 3D landmarks for accurate 3D facial geometry. In Proceedings of the 2021 International Conference on 3D Vision (3DV), London, UK, 1–3 December 2021; pp. 453–463. [Google Scholar] [CrossRef] [Scilit]
- Chen, S.; Liu, Y.; Gao, X.; Han, Z. MobileFaceNets: Efficient CNNs for accurate real-time face verification on mobile devices. In Proceedings of the 13th Chinese Conference on Biometric Recognition (CCBR), Urumqi, China, 11–12 August 2018; pp. 428–438. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Liu, Y.; Hu, Y.; Shi, H.; Mei, T. FaceX-Zoo: A PyTorch toolbox for face recognition. In Proceedings of the 29th ACM International Conference on Multimedia (MM ’21), Chengdu, China, 20–24 October 2021; pp. 3779–3782. [Google Scholar] [CrossRef] [Scilit]
- Zielonka, W.; Bolkart, T.; Thies, J. Towards metrical reconstruction of human faces. In Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022; pp. 250–269. [Google Scholar] [CrossRef] [Scilit]
- Dib, A.; Thébault, C.; Ahn, J.; Gosselin, P.-H.; Theobalt, C.; Chevallier, L. Practical face reconstruction via differentiable ray tracing. Comput. Graph. Forum 2021, 40, 153–164. [Google Scholar] [CrossRef] [Scilit]
- Liu, G.; Zhang, X.; Zhang, F.; Wei, Y. Persona: Real-Time Neural 3D Face Reconstruction for Visual Effects on Mobile Devices. In Proceedings of the ACM SIGGRAPH 2021 Talks (SIGGRAPH ’21), Virtual Event, USA, 9–13 August 2021; Association for Computing Machinery: New York, NY, USA, 2021; pp. 1–2. [Google Scholar] [CrossRef] [Scilit]
- Kendall, A.; Gal, Y.; Cipolla, R. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 7482–7491. [Google Scholar] [CrossRef] [Scilit]
- Buolamwini, J.; Gebru, T. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Proceedings of the 1st Conference on Fairness, Accountability and Transparency (PMLR), New York, NY, USA, 23–24 February 2018; pp. 77–91. Available online: https://proceedings.mlr.press/v81/buolamwini18a.html (accessed on 13 June 2026).
- Makransky, G.; Petersen, G.B. The cognitive affective model of immersive learning (CAMIL): A theoretical research-based model of learning in immersive virtual reality. Educ. Psychol. Rev. 2021, 33, 937–958. [Google Scholar] [CrossRef] [Scilit]
- Wu, B.; Yu, X.; Gu, X. Effectiveness of immersive virtual reality using head-mounted displays on learning performance: A meta-analysis. Br. J. Educ. Technol. 2020, 51, 1991–2005. [Google Scholar] [CrossRef] [Scilit]
- Lincoln, P.; Welch, G.; Nashel, A.; Ilie, A.; State, A.; Fuchs, H. Animatronic shader lamps avatars. Virtual Real. 2011, 15, 225–238. [Google Scholar] [CrossRef] [Scilit]
- Bekele, M.K.; Pierdicca, R.; Frontoni, E.; Malinverni, E.S.; Gain, J. A survey of augmented, virtual, and mixed reality for cultural heritage. ACM J. Comput. Cult. Herit. 2018, 11, 7. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Wang, R. AI-augmented creative pipelines in small studios: A cost-benefit analysis. J. Digit. Media Manag. 2023, 11, 215–230. [Google Scholar]
- He, Q.; Nguyen, L.T. AI-based multi-modal fusion for 3D face modeling from single-image recognition. In Proceedings of the 20th International Conference on Humanities and Social Sciences, Khon Kaen, Thailand, 7–8 January 2026; pp. 227–237. Available online: https://ichuso.kku.ac.th/paper/8 (accessed on 21 July 2026).







| Property | Value | Notes |
|---|---|---|
| Vertex count | 53,215 | BFM 2009 topology |
| Triangle count | 105,840 | Manifold; no non-planar polygons |
| OBJ file size | 4.09 MB | ASCII, single mesh, no texture |
| FBX file size (Blender export) | ≈4.7 MB | Binary FBX 7.4 with blend-shape stub |
| FACS blend-shape compatibility | Yes | 68 canonical landmarks mapped to standard AU set |
| Rig-ready | Yes | Vertex groups exported for Unity Humanoid rig |
| Decimation-safe minimum | ≈15,000 vertices | For mobile-VR targets (Quest 3, Pico 4) |
| Batch Size | RMSE | Final Training Loss | Loss Curve Stability |
|---|---|---|---|
| 16 | 2.684 | 0.041 | High oscillation |
| 32 | 2.456 | 0.028 | Stable, moderate noise |
| 64 | 2.583 | 0.031 | Over-smoothed |
| 128 | 2.611 | 0.034 | Over-smoothed; early plateau |
| Yaw Interval | Single-Modality RMSE | Proposed 3DMM-CNN RMSE | Improvement (%) |
|---|---|---|---|
| Frontal | 2.512 | 2.456 | 2.2 |
| 0–30° | 2.845 | 2.580 | 9.3 |
| 30–60° | 3.420 | 2.712 | 20.7 |
| 60–90° | 4.125 | 2.890 | 29.9 |
| Metric | Value | Interpretation |
|---|---|---|
| R2 | 0.854 | Explains 85.4% of variance in 3DMM parameters |
| MSE | 0.022 | Low parameter-level error |
| MAE | 0.110 | Low mean absolute deviation |
| RMSE | 0.148 | Low root-mean-square error |
| MAPE | 10.6% | Within the acceptable industrial range |
| Pearson r | 0.92 | Strong correlation with ground truth |
| Inference latency (BS = 1) | 35 ms | Real-time-compatible on RTX-class GPU |
| Throughput (BS = 8) | 142 FPS | Suitable for batch avatar pre-generation |
| Yaw Stratum | n | NME of 3D Landmarks (%) | Full Vertex Error (%) |
|---|---|---|---|
| Frontal (<15°) | 186 | 2.30 | 1.42 |
| Small (15–30°) | 79 | 2.68 | 1.62 |
| Moderate (30–60°) | 33 | 2.78 | 1.61 |
| Profile (60–90°) | 8 | 2.93 | 1.60 |
| Overall (weighted mean) | 306 | 2.47 | 1.50 |
| Model | R2 | MAPE (%) | Inference Latency (ms) | XR-Ready Export |
|---|---|---|---|---|
| RingNet [14] | ~0.80 | — | ~95 | Manual conversion |
| MGCNet [15] | ~0.80 | — | ~140 | Manual conversion |
| SynergyNet [28] | — | 11.5 | ~60 | Manual conversion |
| DECA [13] | — | — | ~110 | Custom texture tools |
| EMOCA [16] | — | — | ~180 | Custom texture tools |
| SPECTRE [27] | — | — | ~250 | Custom texture tools |
| Proposed | 0.854 | 10.6 | 35 | One-click OBJ/FBX |
| Mask Configuration | Relative EP-MAE (×Baseline) | Effect |
|---|---|---|
| Delta impulses (as trained) | 1.00 | Baseline |
| Landmark channel ablated (−1) | 1.06 | +6% degradation |
| Gaussian σ = 2 px | 1.51 | +51% degradation |
| Gaussian σ = 4 px | 2.29 | +129% degradation |
| Gaussian σ = 8 px | 3.37 | +237% degradation |
| Gaussian σ = 16 px | 4.57 | +357% degradation |
| Stage | Latency (ms) | Hardware |
|---|---|---|
| Image decode + resize | 0.2 | CPU |
| Face detection (dlib HOG) | 155 | CPU |
| 68-point landmark alignment | 5.6 | CPU |
| Four-channel preprocessing + tensor build | 0.2 | CPU |
| CNN forward pass (proposed) | 35 | RTX 4090 GPU |
| CNN forward pass (Mac mini M4, CPU) | 8 | Apple CPU, verified |
| BFM decoding (numpy matmul) | 0.5 | CPU |
| OBJ file write (ASCII, 53k vertices) | 90 | CPU |
| Blender FBX import + rig bind | ≈250 | Blender 3.6 LTS, CPU |
| Unity 2022 LTS ingest (FBX importer) | ≈350 | Unity vendor spec |
| Total end-to-end (typical) | ≈900 | single image to Unity-ready avatar |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
He, Q.; Chansanam, W.; Nguyen, L.T.; Intawong, K.; Puritat, K. A Lightweight 3DMM-CNN Pipeline for Real-Time Single-Image 3D Face Reconstruction: Prototyping Personalised Avatars for Extended Reality Applications. Informatics 2026, 13, 122. https://doi.org/10.3390/informatics13080122
He Q, Chansanam W, Nguyen LT, Intawong K, Puritat K. A Lightweight 3DMM-CNN Pipeline for Real-Time Single-Image 3D Face Reconstruction: Prototyping Personalised Avatars for Extended Reality Applications. Informatics. 2026; 13(8):122. https://doi.org/10.3390/informatics13080122
Chicago/Turabian StyleHe, Qianqian, Wirapong Chansanam, Lan Thi Nguyen, Kannikar Intawong, and Kitti Puritat. 2026. "A Lightweight 3DMM-CNN Pipeline for Real-Time Single-Image 3D Face Reconstruction: Prototyping Personalised Avatars for Extended Reality Applications" Informatics 13, no. 8: 122. https://doi.org/10.3390/informatics13080122
APA StyleHe, Q., Chansanam, W., Nguyen, L. T., Intawong, K., & Puritat, K. (2026). A Lightweight 3DMM-CNN Pipeline for Real-Time Single-Image 3D Face Reconstruction: Prototyping Personalised Avatars for Extended Reality Applications. Informatics, 13(8), 122. https://doi.org/10.3390/informatics13080122

