Reliable Object Pose Alignment in Mixed-Reality Environments Using Background-Referenced 3D Reconstruction
Abstract
1. Introduction
2. Related Work
2.1. Object Pose Estimation
2.2. The MASt3R Network
3. Proposed Method
3.1. Problem Statement
3.2. Overview of the Proposed Method
3.3. Initial 3D Object Reconstruction Using Multi-View Images (MASt3R)
3.4. Estimating Scene Transformation Between Observation Days
3.5. 3D Localization of the Moved Object in the New View
3.6. Computing the Object Displacement
| Algorithm 1 Proposed Method for Object Pose Re-Estimation |
| Require: Images from day and viewpoint |
| Ensure: Object displacement parameters |
|
3.7. Discussion of Limitations of the Proposed Method
3.7.1. Dependence on Foreground-Background Segmentation Accuracy
3.7.2. Challenges in Complex Background Environments
3.7.3. Robustness to Large Difference in the Viewpoints and Lighting Variations
3.7.4. Limitations Under Partial Occlusion Conditions
4. Experimental Results
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Nomenclature
| MASt3R | Matching and Stereo 3D Reconstruction |
| DUSt3R | Dense Unconstrained Stereo 3D Reconstruction |
| AR | Augmented Reality |
| VR | Virtual Reality |
| MR | Mixed Reality |
| SLAM | Simultaneous Localization and Mapping |
| ICP | Iterative Closest Point |
| GPU | Graphics Processing Unit |
| SAM | Segment Anything Model |
| IoU | Intersection over Union |
References
- Leroy, V.; Cabon, Y.; Revaud, J. Grounding Image Matching in 3D with MASt3R. In Proceedings of the European Conference on Computer Vision, ECCV 2024, Milan, Italy, 29 November–4 October 2024; pp. 71–91. [Google Scholar]
- Martinez, J.; Hossain, R.; Romero, J.; Little, J.J. A simple yet effective baseline for 3D human pose estimation. In Proceedings of the IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, 22–29 October 2017; pp. 2659–2668. [Google Scholar]
- Pavllo, D.; Feichtenhofer, C.; Grangier, D.; Auli, M. 3D human pose estimation in video with temporal convolutions and semi-supervised training. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, 16–20 June 2019; pp. 7745–7754. [Google Scholar]
- Ionescu, C.; Papava, D.; Olaru, V.; Sminchisescu, C. Human3.6M Large scale datasets and predictive methods for 3D human sensing. IEEE Trans. Pattern Anal. Mach. Intell. 2014, 36, 1325–1339. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zheng, C.; Zhu, S.; Mendieta, M.; Yang, T.; Chen, C.; Ding, Z. 3D Human Pose Estimation with Spatial and Temporal Transformers. In Proceedings of the IEEE International Conference on Computer Vision, ICCV 2021, Virtual, 11–17 October 2021; pp. 11636–11645. [Google Scholar]
- Mao, W.; Liu, M.; Salzmann, M. History repeats itself: Human motion prediction via motion attention. In Proceedings of the European Conference on Computer Vision, ECCV 2020, Virtual, 23–28 August 2021; pp. 474–489. [Google Scholar]
- Mehraban, S.; Adeli, V.; Taati, B. MotionAGFormer: Enhancing 3D human pose estimation with a transformer-GCNFormer network. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2024, Waikoloa, HI, USA, 4–8 January 2024; pp. 6905–6915. [Google Scholar]
- Du, S.; Yuan, Z.; Lai, P.; Ikenaga, T. JoyPose: Jointly learning evolutionary data augmentation and anatomy-aware global–local representation for 3D human pose estimation. Pattern Recognit. 2024, 147, 110116. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Yang, X.; Li, B.; Gou, W.; Yan, D.; Zeng, A.; Gao, Y.; Wang, J.; Jing, Y.; Zhang, R.; et al. FreeMan: Towards Benchmarking 3D Human Pose Estimation under Real-World Conditions. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, 17–21 June 2024; pp. 21978–21988. [Google Scholar]
- Neupane, B.; Stankovic, V.; Stankovic, L. A Survey of the State of the Art in Monocular 3D Human Pose Estimation: Methods, Benchmarks, and Challenges. Sensors 2025, 25, 2409. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nogueira, A.F.R.; Oliveira, H.P.; Teixeira, L.F. Markerless multi-view 3D human pose estimation: A survey. Image Vis. Comput. 2025, 155, 105437. [Google Scholar] [CrossRef] [Scilit]
- De, S. Efficient Mesh Reconstruction and Texturing of Oracle Bones. Sensors 2026, 26, 2270. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lu, Y.; Wang, H.; Ning, Q.; Liu, Z.; Zang, Y.; Liao, Z.; Yan, Z. In-Orbit MapAnything: An Enhanced Feed-Forward Metric Framework for 3D Reconstruction of Non-Cooperative Space Targets Under Complex Lighting. Sensors 2026, 26, 2026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Murai, R.; Dexheimer, E.; Davison, A.J. MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition, CVPR 2025, Nashville, TN, USA, 11–15 June 2025; pp. 16695–16705. [Google Scholar]
- Li, J.; Lee, Y.; Yadav, A.K.; Peng, C.; Chellappa, R.; Fan, D. Speedy MASt3R. In Proceedings of the IEEE International Conference on Omni-layer Intelligent Systems, COINS 2025, Madison, WI, USA, 4–6 August 2025; pp. 1–6. [Google Scholar]
- Chen, X.; Xia, T.; Xu, S.; Yang, J.; Chai, J.; Cheng, Z. SAB3R: Semantic-Augmented Backbone in 3D Reconstruction. arXiv 2025, arXiv:2506.02112. [Google Scholar]
- Rojas, S.; Armando, M.; Ghamen, B.; Weinzaepfel, P.; Leroy, V.; Rogez, G. HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction. In Proceedings of the IEEE International Conference on Computer Vision, ICCV 2025, Honolulu, HI, USA, 19–23 October 2025; pp. 5027–5037. [Google Scholar]
- Cai, Q.; Hu, X.; Hou, S.; Yao, L.; Huang, Y. Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal Denoiser. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Vancouver, BC, Canada, 20–24 February 2024; pp. 882–890. [Google Scholar]
- Besl, P.J.; McKay, N.D. A Method for Registration of 3-D Shapes. IEEE Trans. Pattern Anal. Mach. Intell. 1992, 14, 239–256. [Google Scholar] [CrossRef] [Scilit]
- Umeyama, S. Least-Squares Estimation of Transformation Parameters Between Two Point Patterns. IEEE Trans. Pattern Anal. Mach. Intell. 1991, 13, 376–380. [Google Scholar] [CrossRef] [Scilit]
- Arun, K.S.; Huang, T.S.; Blostein, S.D. Least-Squares Fitting of Two 3-D Point Sets. IEEE Trans. Pattern Anal. Mach. Intell. 1987, 9, 698–700. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, J.; Li, H.; Jia, Y. Go-ICP: A Globally Optimal Solution to 3D ICP Point-Set Registration. IEEE Trans. Pattern Anal. Mach. Intell. 2016, 38, 2241–2254. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment Anything. In Proceedings of the IEEE International Conference on Computer Vision, ICCV 2023, Paris, France, 2–6 October 2023; pp. 4015–4026. [Google Scholar]











| Image Number | Original Orientation Error (°) | Aligned Orientation Error (°) | Unaligned IoU | Aligned IoU |
|---|---|---|---|---|
| Image1 | 54.02 | 5.05 | 0.503 | 0.953 |
| Image2 | 30.44 | 7.21 | 0.123 | 0.926 |
| Image3 | 65.54 | 11.62 | 0.451 | 0.815 |
| Image4 | 73.21 | 8.24 | 0.391 | 0.921 |
| Image5 | 54.33 | 5.66 | 0.125 | 0.876 |
| Image6 | 44.82 | 6.43 | 0.305 | 0.954 |
| Image7 | 26.37 | 7.27 | 0.528 | 0.963 |
| Image8 | 31.23 | 3.42 | 0.439 | 0.994 |
| Image9 | 52.81 | 9.71 | 0.463 | 0.912 |
| Image10 | 28.73 | 7.58 | 0.217 | 0.925 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Shin, G.-B.; Song, B.-D.; Iordanov, V.B.; Park, S.; Lee, S.; Lee, S.-H. Reliable Object Pose Alignment in Mixed-Reality Environments Using Background-Referenced 3D Reconstruction. Sensors 2026, 26, 2453. https://doi.org/10.3390/s26082453
Shin G-B, Song B-D, Iordanov VB, Park S, Lee S, Lee S-H. Reliable Object Pose Alignment in Mixed-Reality Environments Using Background-Referenced 3D Reconstruction. Sensors. 2026; 26(8):2453. https://doi.org/10.3390/s26082453
Chicago/Turabian StyleShin, Gyu-Bin, Bok-Deuk Song, Vladimirov Blagovest Iordanov, Sangjoon Park, Soyeon Lee, and Suk-Ho Lee. 2026. "Reliable Object Pose Alignment in Mixed-Reality Environments Using Background-Referenced 3D Reconstruction" Sensors 26, no. 8: 2453. https://doi.org/10.3390/s26082453
APA StyleShin, G.-B., Song, B.-D., Iordanov, V. B., Park, S., Lee, S., & Lee, S.-H. (2026). Reliable Object Pose Alignment in Mixed-Reality Environments Using Background-Referenced 3D Reconstruction. Sensors, 26(8), 2453. https://doi.org/10.3390/s26082453

