Review: Techniques in Egocentric Multi-View Image Analysis: Advances, Challenges, and Future Directions
Abstract
1. Introduction
1.1. The Growth of Egocentric Vision
- Reduced occlusion and consistent viewpoints for object manipulation: Hands and manipulated objects usually appear near the center of the frame, making it easier to analyze human–object interactions (HOI), recognize actions, and learn skills. These tasks are often more difficult in third-person views.
- Natural embodiment and contextual awareness: By capturing the wearer’s vision and actions in real time, egocentric systems allow AI models to infer intentions, anticipate actions, and provide context-aware assistance. This is especially important for augmented reality (AR) and virtual reality (VR), where digital content must respond smoothly to gaze direction, hand movement, and environmental changes.
- Support for embodied intelligence and human–AI collaboration: Egocentric data supports applications in robotics (e.g., imitation learning from human demonstrations), assistive technologies, healthcare monitoring, and social behavior analysis. Wearable systems can act as memory aids, provide step-by-step guidance, or improve productivity in training and manufacturing environments.
- Foundation for hybrid and multi-view systems: Although early single-camera setups faced challenges such as motion blur and limited field of view, these limitations are increasingly addressed through wearable multi-camera systems. Such systems enable more reliable 3D reconstruction, novel-view synthesis, and cross-view feature fusion in dynamic, real-world scenarios.
1.2. Limitations of Single-View Egocentric Approaches
1.3. Emergence of Egocentric Multi-View Systems
1.4. Relation to Broader Multi-View Image Analysis Techniques
1.5. Scope and Contributions of This Review
2. Background and Technical Challenges
2.1. Unique Properties of Egocentric Multi-View Data
2.2. Core Technical Challenges
2.3. Comparison with Fixed Multi-Camera and Ego–Exo Setups
3. Data Acquisition Systems and Hardware
3.1. Head-Mounted Multi-View Rigs
3.2. Body-Worn and Multi-User Systems
3.3. Sensor Modalities and Synchronization
3.4. Ground-Truth Acquisition (Motion-Capture, 3D Scanning)
4. Core Techniques in Egocentric Multi-View Image Analysis
4.1. Cross-View Feature Fusion and Geometric Learning
Multi-View Stereo and Deformable Attention
4.2. Egocentric Multi-View Open-World Object Detection
4.3. Egocentric Multi-View Human-Object Interaction and Action Recognition
4.4. 3D Reconstruction, Novel-View Synthesis, and Tracking
5. Datasets and Benchmarks
5.1. Representative Multi-View and Egocentric Datasets
5.1.1. Geometrically Grounded Multi-View Datasets
5.1.2. Large-Scale Egocentric Datasets
5.2. Performance Trends: Multi-View vs. Single-View
5.2.1. 3D Hand Pose Estimation
5.2.2. Hand–Object Interaction
5.2.3. Action Recognition and Anticipation
| Model | Input | Metric | Result |
|---|---|---|---|
| Long-term Action Anticipation (LTA) | |||
| Bertasius et al. [127] | Ego RGB | 0.7169/0.7359/0.9253 | |
| Mittal et al. [125] | Ego + VLM | 0.679/0.681/— | |
| Kim et al. [126] | Ego + VLM/LLM | 0.6471/0.6117/0.8503 | |
| Short-term Object Interaction Anticipation (STA) | |||
| Bertasius et al. [127] | Ego RGB | 26.15/9.45/8.69/3.61 | |
| Ragusa et al. [131] | Ego RGB | 25.06/13.29/9.14/5.12 | |
| Pasca et al. [128] | Ego + Action Context | 30.43/13.45/10.38/5.18 | |
| Model | Setting | Metric | Single | Multi | Gain |
|---|---|---|---|---|---|
| Keystep Recognition | |||||
| Bertasius et al. [127] | Ego vs. Ego + Exo train | Acc. ↑ | 35.24 | 29.84 | −5.40% |
| Pramanick et al. [129] | Ego vs. Ego + Exo | Acc. ↑ | 37.85 | 38.69 | 2.2% |
| Romero et al. [130] | Ego graph vs. Ego + Exo graph | Acc. ↑ | 52.36 | 53.08 | +0.72% |
| Proficiency Estimation | |||||
| Bertasius et al. [127] | Ego vs. Ego + Exos | Acc. ↑ | 42.3 | 40.8 | −1.5% |
| Bianchi et al. [132] | Ego vs. Ego + Exos | Acc. ↑ | 45.9 | 47.5 | +1.6% |
| Bianchi et al. [133] | Ego vs. Ego + Exos | Acc. ↑ | 47.3 | 48.0 | +0.7% |
| Bianchi et al. [134] | Ego vs. Ego + Exos | Acc. ↑ | 44.2 | 48.2 | +4.0% |
| Tanoue et al. [135] | Ego vs. Ego + Exos | Acc. ↑ | 44.3 | 47.8 | +3.5% |
| Braun et al. [136] | Ego vs. Ego + Exo + HR | Acc. ↑ | 39.69 | 43.94 | +4.25% |
5.2.4. Consolidated Meta-Analysis of Multi-View Performance Gains
5.3. Limitations and Research Gaps
6. Applications and Broader Impact
7. Open Challenges and Future Directions
8. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Bolanos, M.; Dimiccoli, M.; Radeva, P. Toward storytelling from visual lifelogging: An overview. IEEE Trans. Hum.-Mach. Syst. 2016, 47, 77–90. [Google Scholar] [CrossRef] [Scilit]
- Miao, Q.; Price, J.; Tassiopoulos, A.; Samaras, D. Behavior-Based Skill Assessment for Open Surgery from Multi-View and Egocentric Videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 8854–8864. [Google Scholar]
- Plizzari, C.; Goletto, G.; Furnari, A.; Bansal, S.; Ragusa, F.; Farinella, G.M.; Damen, D.; Tommasi, T. An Outlook into the Future of Egocentric Vision: C. Plizzari et al. Int. J. Comput. Vis. 2024, 132, 4880–4936. [Google Scholar] [CrossRef] [Scilit]
- Grauman, K.; Westbury, A.; Byrne, E.; Chavis, Z.; Furnari, A.; Girdhar, R.; Hamburger, J.; Jiang, H.; Liu, M.; Liu, X.; et al. Ego4d: Around the world in 3000 h of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 18995–19012. [Google Scholar]
- Grauman, K.; Westbury, A.; Torresani, L.; Kitani, K.; Malik, J.; Afouras, T.; Ashutosh, K.; Baiyya, V.; Bansal, S.; Boote, B.; et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 19383–19400. [Google Scholar]
- Banerjee, P.; Shkodrani, S.; Moulon, P.; Hampali, S.; Han, S.; Zhang, F.; Zhang, L.; Fountain, J.; Miller, E.; Basol, S.; et al. Hot3d: Hand and object tracking in 3d from egocentric multi-view videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025; pp. 7061–7071. [Google Scholar]
- Engel, J.; Somasundaram, K.; Goesele, M.; Sun, A.; Gamino, A.; Turner, A.; Talattof, A.; Yuan, A.; Souti, B.; Meredith, B.; et al. Project aria: A new tool for egocentric multi-modal ai research. arXiv 2023, arXiv:2308.13561. [Google Scholar]
- Li, X.; Qiu, H.; Wang, L.; Zhang, H.; Qi, C.; Han, L.; Xiong, H.; Li, H. Challenges and trends in egocentric vision: A survey. Mach. Intell. Res. 2026, 23, 1–33. [Google Scholar] [CrossRef] [Scilit]
- Jiang, H.; Ramakrishnan, S.K.; Grauman, K. Single-stage visual query localization in egocentric videos. Adv. Neural Inf. Process. Syst. 2023, 36, 24143–24157. [Google Scholar] [CrossRef] [Scilit]
- Hoshen, Y.; Ben-Artzi, G.; Peleg, S. Wisdom of the crowd in egocentric video curation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Columbus, OH, USA, 23–28 June 2014; pp. 573–579. [Google Scholar]
- Bandini, A.; Zariffa, J. Analysis of the hands in egocentric vision: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 45, 6846–6866. [Google Scholar] [CrossRef] [Scilit]
- Thatipelli, A.; Lo, S.Y.; Roy-Chowdhury, A.K. Egocentric and exocentric methods: A short survey. Comput. Vis. Image Underst. 2025, 257, 104371. [Google Scholar] [CrossRef] [Scilit]
- Ardeshir, S.; Borji, A. Egocentric meets top-view. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 41, 1353–1366. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, H.; Singh, M.K.; Torresani, L. Ego-only: Egocentric action detection without exocentric transferring. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 5250–5261. [Google Scholar]
- Huang, J.; Hao, S.; Hu, B.C.; Wang, H.; Wang, G. Understanding dynamic scenes in ego centric 4d point clouds. In Proceedings of the AAAI Conference on Artificial Intelligence, Singapore, 20–27 January 2026; Volume 40, pp. 5031–5039. [Google Scholar]
- Hollidt, D.; Streli, P.; Jiang, J.; Haghighi, Y.; Qian, C.; Liu, X.; Holz, C. Egosim: An egocentric multi-view simulator and real dataset for body-worn cameras during motion and activity. Adv. Neural Inf. Process. Syst. 2024, 37, 106607–106627. [Google Scholar] [CrossRef] [Scilit]
- Bano, S.; Suveges, T.; Zhang, J.; Mckenna, S.J. Multimodal egocentric analysis of focused interactions. IEEE Access 2018, 6, 37493–37505. [Google Scholar] [CrossRef] [Scilit]
- Lee, J.Y.; Scharstein, D.; Bapat, A.; Hu, H.; Fu, A.; Zhao, H.; Sammut, P.; Li, X.; Jeapes, S.; Gupta, A.; et al. Ego-1K-A Large-Scale Multiview Video Dataset for Egocentric Vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 19854–19863. [Google Scholar]
- Li, B.; Zhong, H.; Cheng, Z.; Hu, Q.; Wang, Q.; Song, L.; Zhang, W. MultiEgo: A Multi-View Egocentric Video Dataset for 4D Scene Reconstruction. In Proceedings of the 33rd ACM International Conference on Multimedia, Dublin, Ireland, 27–31 October 2025; pp. 12882–12889. [Google Scholar]
- Rahem, R.; Suleiman, W. Beyond Egocentric Limits: Multi-View Depth-Based Learning for Robust Quadrupedal Locomotion. arXiv 2025, arXiv:2511.22744. [Google Scholar]
- Long, X.; Liu, L.; Li, W.; Theobalt, C.; Wang, W. Multi-view depth estimation using epipolar spatio-temporal networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 8258–8267. [Google Scholar]
- Liu, Y.; Yang, J.; Gu, X.; Chen, Y.; Guo, Y.; Yang, G.Z. Egofish3d: Egocentric 3d pose estimation from a fisheye camera via self-supervised learning. IEEE Trans. Multimed. 2023, 25, 8880–8891. [Google Scholar] [CrossRef] [Scilit]
- Xie, L.; Xu, G.; Cai, D.; He, X. X-view: Non-egocentric multi-view 3D object detector. IEEE Trans. Image Process. 2023, 32, 1488–1497. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Su, W.; Tao, W. Efficient edge-preserving multi-view stereo network for depth estimation. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; Volume 37, pp. 2348–2356. [Google Scholar]
- Wu, M.; Wang, Y.; Hu, Q.; Yu, J. Multi-view neural human rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 1682–1691. [Google Scholar]
- Vogiatzis, G.; Hernández, C. Video-based, real-time multi-view stereo. Image Vis. Comput. 2011, 29, 434–441. [Google Scholar] [CrossRef] [Scilit]
- Furukawa, Y.; Hernández, C. Multi-view stereo: A tutorial. Found. Trends Comput. Graph. Vis. 2015, 9, 1–148. [Google Scholar] [CrossRef] [Scilit]
- Kim, C.; Hornung, A.; Heinzle, S.; Matusik, W.; Gross, M. Multi-perspective stereoscopy from light fields. ACM Trans. Graph. (TOG) 2011, 30, 1–10. [Google Scholar] [CrossRef] [Scilit]
- Son, J.Y.; Lee, H.; Lee, B.R.; Lee, K.H. Holographic and light-field imaging as future 3-D displays. Proc. IEEE 2017, 105, 789–804. [Google Scholar] [CrossRef] [Scilit]
- Han, J.; Chen, H.; Liu, N.; Yan, C.; Li, X. CNNs-based RGB-D saliency detection via cross-view transfer and multiview fusion. IEEE Trans. Cybern. 2017, 48, 3171–3183. [Google Scholar]
- Qiu, H.; Wang, C.; Wang, J.; Wang, N.; Zeng, W. Cross view fusion for 3d human pose estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 4342–4351. [Google Scholar]
- Romanoni, A.; Matteucci, M. Tapa-mvs: Textureless-aware patchmatch multi-view stereo. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 10413–10422. [Google Scholar]
- Wang, F.; Galliani, S.; Vogel, C.; Speciale, P.; Pollefeys, M. Patchmatchnet: Learned multi-view patchmatch stereo. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 14194–14203. [Google Scholar]
- Yao, Y.; Luo, Z.; Li, S.; Fang, T.; Quan, L. Mvsnet: Depth inference for unstructured multi-view stereo. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 767–783. [Google Scholar]
- Yao, Y.; Luo, Z.; Li, S.; Shen, T.; Fang, T.; Quan, L. Recurrent mvsnet for high-resolution multi-view stereo depth inference. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 5525–5534. [Google Scholar]
- Zhang, J.; Li, S.; Luo, Z.; Fang, T.; Yao, Y. Vis-mvsnet: Visibility-aware multi-view stereo network. Int. J. Comput. Vis. 2023, 131, 199–214. [Google Scholar]
- Aanæs, H.; Jensen, R.R.; Vogiatzis, G.; Tola, E.; Dahl, A.B. Large-scale data for multiple-view stereopsis. Int. J. Comput. Vis. 2016, 120, 153–168. [Google Scholar] [CrossRef] [Scilit]
- Knapitsch, A.; Park, J.; Zhou, Q.Y.; Koltun, V. Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Trans. Graph. (ToG) 2017, 36, 1–13. [Google Scholar] [CrossRef] [Scilit]
- Scharstein, D.; Hirschmüller, H.; Kitajima, Y.; Krathwohl, G.; Nešić, N.; Wang, X.; Westling, P. High-resolution stereo datasets with subpixel-accurate ground truth. In Proceedings of the German Conference on Pattern Recognition, Münster, Germany, 2–5 September 2014; Springer: Berlin/Heidelberg, Germany, 2014; pp. 31–42. [Google Scholar]
- Wang, G.; Xiang, W.; Pickering, M.; Chen, C.W. Light field multi-view video coding with two-directional parallel inter-view prediction. IEEE Trans. Image Process. 2016, 25, 5104–5117. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhu, H.; Wang, Q.; Yu, J. Light field imaging: Models, calibrations, reconstructions, and applications. Front. Inf. Technol. Electron. Eng. 2017, 18, 1236–1249. [Google Scholar] [CrossRef] [Scilit]
- Gu, Q.; Lv, Z.; Frost, D.; Green, S.; Straub, J.; Sweeney, C. Egolifter: Open-world 3d segmentation for egocentric perception. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; Springer: Berlin/Heidelberg, Germany, 2024; pp. 382–400. [Google Scholar]
- Choi, C.; Kim, S.M.; Kim, Y.M. Balanced spherical grid for egocentric view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 16590–16599. [Google Scholar]
- Yoo, J.H.; Kim, Y.; Kim, J.; Choi, J.W. 3d-cvf: Generating joint camera and lidar features using cross-view spatial feature fusion for 3d object detection. In Proceedings of the European Conference on Computer Vision, Glasgow, UK, 23–28 August 2020; Springer: Berlin/Heidelberg, Germany, 2020; pp. 720–736. [Google Scholar]
- Gu, Z.; Peng, G.; Yun, Y.; Liu, Y.; Wu, Z.; Zhang, J.; Li, Z.; Suo, X.; Wang, D. Cross-view detection of crowded objects based on multi-sensor fusion. In Proceedings of the 2024 18th International Conference on Control, Automation, Robotics and Vision (ICARCV), Dubai, United Arab Emirates, 12–15 December 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 363–368. [Google Scholar]
- Ma, J.; Xiong, G.; Xu, J.; Chen, X. CVTNet: A cross-view transformer network for LiDAR-based place recognition in autonomous driving environments. IEEE Trans. Ind. Inform. 2023, 20, 4039–4048. [Google Scholar]
- Zhong, H.; Xiang, Z.; Xu, R.; Fu, J.; Xu, P.; Wang, S.; Yang, Z.; Pu, T.; Liu, E. CVFusion: Cross-view fusion of 4D radar and camera for 3D object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–20 October 2025; pp. 28188–28197. [Google Scholar]
- Kim, S.; Ahn, D.; Ko, B.C. Cross-modal learning with 3D deformable attention for action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 10265–10275. [Google Scholar]
- Luo, K.; Guan, T.; Ju, L.; Wang, Y.; Chen, Z.; Luo, Y. Attention-aware multi-view stereo. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 1590–1599. [Google Scholar]
- Wang, X.; Zhu, Z.; Huang, G.; Qin, F.; Ye, Y.; He, Y.; Chi, X.; Wang, X. Mvster: Epipolar transformer for efficient multi-view stereo. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; Springer: Berlin/Heidelberg, Germany, 2022; pp. 573–591. [Google Scholar]
- Liu, T.; Ye, X.; Zhao, W.; Pan, Z.; Shi, M.; Cao, Z. When epipolar constraint meets non-local operators in multi-view stereo. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 18088–18097. [Google Scholar]
- Campbell, N.D.; Vogiatzis, G.; Hernández, C.; Cipolla, R. Using multiple hypotheses to improve depth-maps for multi-view stereo. In Proceedings of the European Conference on Computer Vision, Marseille, France, 12–18 October 2008; Springer: Berlin/Heidelberg, Germany, 2008; pp. 766–779. [Google Scholar]
- Tola, E.; Strecha, C.; Fua, P. Efficient large-scale multi-view stereo for ultra high-resolution image sets. Mach. Vis. Appl. 2012, 23, 903–920. [Google Scholar]
- Galliani, S.; Lasinger, K.; Schindler, K. Gipuma: Massively parallel multi-view stereo reconstruction. Publ. Dtsch. Ges. Photogramm. Fernerkund. Geoinf. e. V 2016, 25, 2. [Google Scholar]
- He, Y.; Huang, Y.; Chen, G.; Lu, L.; Pei, B.; Xu, J.; Lu, T.; Sato, Y. Bridging perspectives: A survey on cross-view collaborative intelligence with egocentric-exocentric vision. Int. J. Comput. Vis. 2026, 134, 62. [Google Scholar] [CrossRef] [Scilit]
- Cao, Y.; Liu, Y.; Wang, G.; Liu, Z.; Wang, K.; Zhang, X.; Yu, J.; Tu, X. EAGLE: Episodic Appearance-and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision. In Proceedings of the AAAI Conference on Artificial Intelligence, Singapore, 20–27 January 2026; Volume 40, pp. 2634–2642. [Google Scholar]
- Tan, Y.; Cheng, X.; Qin, Y.; Li, Z.; Zhang, J. Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 10545–10555. [Google Scholar]
- Fu, Y.; Dai, P.; Zhang, Y.; Yiqiang, F.; Zhang, Y.; Wang, H. SAME: Spatial-Aware Multimodal Egocentric Human Pose Estimation. In Proceedings of the AAAI Conference on Artificial Intelligence, Singapore, 20–27 January 2026; Volume 40, pp. 4067–4075. [Google Scholar]
- Cheng, H.; Ong, S.J.H.; Cai, S.; Koh, A.T.Y.; Ouyang, F.; Khoo, E. EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual Reality. IEEE Trans. Vis. Comput. Graph. 2026, 32, 3211–3221. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jammot, M.; Braun, B.; Streli, P.; Wampfler, R.; Holz, C. egoEMOTION: Egocentric vision and physiological signals for emotion and personality recognition in real-world tasks. Adv. Neural Inf. Process. Syst. 2026, 38, 1–22. [Google Scholar]
- Ballester, I.; Hermosilla, P.; Lin, W.; Glass, J.R.; Mirza, M.J.; Kampel, M. AViON4D: Audio-Visual Open-Vocabulary 4D Egocentric Scene Understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 8355–8365. [Google Scholar]
- Luo, H.; Yue, Z.; Zhang, W.; Feng, Y.; Zheng, S.; Ye, D.; Lu, Z. OpenMMEgo: Enhancing egocentric understanding for LMMs with open weights and data. Adv. Neural Inf. Process. Syst. 2026, 38, 25749–25781. [Google Scholar]
- Özsoy, E.; Mamur, A.; Tristram, F.; Pellegrini, C.; Wysocki, M.; Busam, B.; Navab, N. Egoexor: An ego-exo-centric operating room dataset for surgical activity understanding. Adv. Neural Inf. Process. Syst. 2026, 38, 1–18. [Google Scholar] [CrossRef] [Scilit]
- Ragusa, F.; Leonardi, R.; Mazzamuto, M.; Di Mauro, D.; Quattrocchi, C.; Passanisi, A.; D’Ambra, I.; Furnari, A.; Farinella, G.M. ENIGMA-360: An Ego-Exo Dataset for Human Behavior Understanding in Industrial Scenarios. arXiv 2026, arXiv:2603.09741. [Google Scholar]
- Kang, T.; Kim, K.; Kim, D.; Park, M.; Hyung, J.; Choo, J. EgoX: Egocentric Video Generation from a Single Exocentric Video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 11116–11126. [Google Scholar]
- Jiang, X.; Xu, X.; Han, Z.; Wang, Z.; Song, J.; Shen, H.T. Egocentric Online Action Segmentation via Evidential Temporal Contextualization. IEEE Trans. Multimed. 2026, 9, 1–11. [Google Scholar] [CrossRef] [Scilit]
- Qiu, H.; Wang, L.; Zhao, T.; Shi, Z.; Li, X.; Xu, L.; Li, H. Ego-PMOVE: Prompt-aware Mixture of View Experts Network for Egocentric Gaze Prediction. In Proceedings of the AAAI Conference on Artificial Intelligence, Singapore, 20–27 January 2026; Volume 40, pp. 8565–8573. [Google Scholar]
- Punnakkal, A.R.; Chandrasekaran, A.; Athanasiou, N.; Quiros-Ramirez, A.; Black, M.J. BABEL: Bodies, action and behavior with english labels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 722–731. [Google Scholar]
- Mahmood, N.; Ghorbani, N.; Troje, N.F.; Pons-Moll, G.; Black, M.J. AMASS: Archive of motion capture as surface shapes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 5442–5451. [Google Scholar]
- Yu, L.; Murat, E.; Wang, B.; Zeng, Y.; Luo, T.; Zhou, H.; Li, S.; Feng, H.; Zhao, Z.; Yang, N.; et al. EgoKit: Towards Unified Low-Cost Egocentric Data Collection with Heterogeneous Devices. arXiv 2026, arXiv:2605.16797. [Google Scholar]
- Thapar, D.; Nigam, A.; Arora, C. Anonymizing egocentric videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 2320–2329. [Google Scholar]
- Li, L.; Zhu, F.; Eicher-Miller, H.; Thomas, J.G.; Huang, Y.; Sazonov, E. Extra-lightweight AI-based privacy preserving framework for egocentric wearable cameras. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 401–410. [Google Scholar]
- Zhu, Z.; Sato, Y. Cross-View Correspondence Modeling for Joint Representation Learning Between Egocentric and Exocentric Videos. IEEE Access 2025, 13, 140733–140741. [Google Scholar] [CrossRef] [Scilit]
- Liu, G.; Tang, H.; Latapie, H.M.; Corso, J.J.; Yan, Y. Cross-view exocentric to egocentric video synthesis. In Proceedings of the 29th ACM International Conference on Multimedia, Chengdu, China, 20–24 October 2021; pp. 974–982. [Google Scholar]
- Gao, Y.; Zhang, B.; Tang, Z.; Liao, J.; Wu, W.; Liu, S. VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 21690–21700. [Google Scholar]
- Fonder, M.; Ernst, D.; Van Droogenbroeck, M. Parallax inference for robust temporal monocular depth estimation in unstructured environments. Sensors 2022, 22, 9374. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, S.; Yang, J.; Hao, T.; Wu, S.; Li, M. Temporal feature fusion with deformable attention for multi-view 3D object detection. Digit. Signal Process. 2025, 168, 105518. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Zeng, Z.; Guan, T.; Yang, W.; Chen, Z.; Liu, W.; Xu, L.; Luo, Y. Adaptive patch deformation for textureless-resilient multi-view stereo. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 1621–1630. [Google Scholar]
- Shen, H.; Zhuang, P.; Kou, J.; Zeng, Y.; Xu, H.; Li, J. MGD-SAM2: Multi-view guided detail-enhanced segment anything model 2 for high-resolution class-agnostic segmentation. IEEE Trans. Circuits Syst. Video Technol. 2026. [Google Scholar]
- Hsu, P.H.; Zhang, K.; Wang, F.E.; Tu, T.; Li, M.F.; Liu, Y.L.; Chen, A.Y.; Sun, M.; Kuo, C.H. Openm3d: Open vocabulary multi-view indoor 3d object detection without human annotations. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–23 October 2025; pp. 8688–8698. [Google Scholar]
- Zhang, Q.; Liu, Y.; Zhu, B.; Han, X.; Zhang, R.; Xiao, J.; Wang, Z. Deep multi-modal fusion transformer for emotion recognition. Eng. Appl. Artif. Intell. 2026, 168, 113967. [Google Scholar] [CrossRef] [Scilit]
- Zhou, S.; Wang, H.; Wu, Q.; Meng, F.; Xu, L.; Zhang, W.; Li, H. Adversarially Regularized Tri-Transformer Fusion for continual multimodal egocentric activity recognition. Displays 2025, 88, 102992. [Google Scholar] [CrossRef] [Scilit]
- Shen, Q.; Zhao, Y.; Kwon, N.; Kim, J.; Li, Y.; Kong, S. Solving instance detection from an open-world perspective. In Proceedings of the Computer Vision and Pattern Recognition Conference, Shanghai, China, 15–18 October 2025; pp. 9901–9910. [Google Scholar]
- Geng, Z.; Wang, N.; Xu, S.; Ye, C.; Li, B.; Chen, Z.; Peng, S.; Zhao, H. One view, many worlds: Single-image to 3d object meets generative domain randomization for one-shot 6d pose estimation. arXiv 2025, arXiv:2509.07978. [Google Scholar]
- Liu, Z.; Song, R.; Chuangqi, D.; Li, J.; Ferstl, D.; Hu, Y. Exploring 6D Object Pose Estimation with Deformation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 33078–33087. [Google Scholar]
- Kang, J.Y.; Cho, H.; Lee, T.; Kang, M.; Wen, B.; Kim, Y.; Yoon, K.J. Event6D: Event-based Novel Object 6D Pose Tracking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 15091–15104. [Google Scholar]
- Moroz, A.; Zeman, V.; Mikšík, M.; Isianova, E.; David, M.; Burget, P.; Burde, V. OPFormer: Object Pose Estimation leveraging foundation model with geometric encoding. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Tucson, AR, USA, 6–10 March 2026; pp. 6621–6632. [Google Scholar]
- Cai, Y.; Li, Z.; Lu, T.; Zhu, Y.; Wu, Y.S.; Zhang, Q.; Xu, X.; Jin, Z.; Gowda, M.; Jin, Y. Toward Scalable ASL Education: Egocentric Stereo Sensing with LLM Feedback for Error-Aware Learning. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Barcelona, Spain, 13–17 April 2026; pp. 1–24. [Google Scholar]
- Zhang, Z.; Shi, Y.; Yang, L.; Ni, S.; Ye, Q.; Wang, J. Openhoi: Open-world hand-object interaction synthesis with multimodal large language model. Adv. Neural Inf. Process. Syst. 2026, 38, 166582–166612. [Google Scholar]
- Qi, Z.; Zhang, Z.; Fang, Y.; Wang, J.; Zhao, H. Gpt4scene: Understand 3d scenes from videos with vision-language models. arXiv 2025, arXiv:2501.01428. [Google Scholar]
- Ye, H.; Zhang, H.; Daxberger, E.; Chen, L.; Lin, Z.; Li, Y.; Zhang, B.; You, H.; Xu, D.; Gan, Z.; et al. MMEgo: Towards building egocentric multimodal LLMs for video QA. In Proceedings of the International Conference on Learning Representations, Singapore, 24–28 April 2025; Volume 2025, pp. 71705–71723. [Google Scholar]
- Yang, J.; Liu, S.; Guo, H.; Dong, Y.; Zhang, X.; Zhang, S.; Wang, P.; Zhou, Z.; Xie, B.; Wang, Z.; et al. Egolife: Towards egocentric life assistant. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 10–17 June 2025; pp. 28885–28900. [Google Scholar]
- Park, W.; Lee, I.; Kim, S.; Jang, J.; Noh, M.; Shim, K.; Shim, B. Revealing Multi-View Hallucination in Large Vision-Language Models. arXiv 2026, arXiv:2603.23934. [Google Scholar]
- Ye, Y.; Li, J.; Rong, R.; Liu, C.K. Whole: World-grounded hand-object lifted from egocentric videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 3481–3491. [Google Scholar]
- Luo, H.; Feng, Y.; Zhang, W.; Zheng, S.; Wang, Y.; Yuan, H.; Liu, J.; Xu, C.; Jin, Q.; Lu, Z. Being-h0: Vision-language-action pretraining from large-scale human videos. arXiv 2025, arXiv:2507.15597. [Google Scholar]
- Xu, B.; Zheng, S.; Jin, Q. Pov: Prompt-oriented view-agnostic learning for egocentric hand-object interaction in the multi-view world. In Proceedings of the 31st ACM International Conference on Multimedia, Ottawa, ON, Canada, 29 October–3 November 2023; pp. 2807–2816. [Google Scholar]
- Leonardi, R.; Ragusa, F.; Furnari, A.; Farinella, G.M. Exploiting multimodal synthetic data for egocentric human-object interaction detection in an industrial scenario. Comput. Vis. Image Underst. 2024, 242, 103984. [Google Scholar] [CrossRef] [Scilit]
- Xu, L.; Yang, C.; Lin, Z.; Xu, F.; Liu, Y.; Xu, C.; Zhang, Y.; Qin, J.; Sheng, X.; Liu, Y.; et al. Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–20 October 2025; pp. 12535–12548. [Google Scholar]
- Hamid, D.; Haq, M.E.U.; Yasin, A.; Murtaza, F.; Azam, M.A. Enhancing Recognition of Human–Object Interaction from Visual Data Using Egocentric Wearable Camera. Future Internet 2024, 16, 269. [Google Scholar] [CrossRef] [Scilit]
- Li, G.; Chen, Y.; Wu, Y.; Zhao, K.; Pollefeys, M.; Tang, S. Egom2p: Egocentric multimodal multitask pretraining. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–20 October 2025; pp. 10830–10843. [Google Scholar]
- He, Y.; Huang, Y.; Chen, G.; Pei, B.; Xu, J.; Lu, T.; Pang, J. Egoexobench: A benchmark for first-and third-person view video understanding in mllms. Adv. Neural Inf. Process. Syst. 2026, 38, 1–17. [Google Scholar]
- Akada, H.; Wang, J.; Golyanik, V.; Theobalt, C. Bring your rear cameras for egocentric 3d human pose estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–20 October 2025; pp. 9497–9507. [Google Scholar]
- Zhang, W.; Jia, E.Y.t.; Zhou, J.; Ma, B.; Shi, K.; Liu, Y.S.; Han, Z. NeRFPrior: Learning neural radiance field as a prior for indoor scene reconstruction. In Proceedings of the Computer Vision and Pattern Recognition Conference, Shanghai, China, 15–18 October 2025; pp. 11317–11327. [Google Scholar]
- Zhang, X.; Yu, R.; Ren, S. Neural implicit representations for multi-view surface reconstruction: A survey. IEEE Trans. Vis. Comput. Graph. 2025, 31, 9444–9463. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kawaharazuka, K.; Oh, J.; Yamada, J.; Posner, I.; Zhu, Y. Vision-language-action models for robotics: A review towards real-world applications. IEEE Access 2025, 13, 162467–162504. [Google Scholar] [CrossRef] [Scilit]
- Yoshida, T.; Kurita, S.; Nishimura, T.; Mori, S. Generating 6dof object manipulation trajectories from action description in egocentric vision. In Proceedings of the Computer Vision and Pattern Recognition Conference, Shanghai, China, 15–18 October 2025; pp. 17370–17382. [Google Scholar]
- Saroha, A.; Zeng, H.; Zuo, X.; Cremers, D.; Wang, X. EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 4332–4342. [Google Scholar]
- Xi, Z.; Ao, Z.; Wang, Y.; Gao, M.; Zhang, W.; Feng, J.; Zhou, J. WristPP: A Wrist-Worn System for Hand Pose and Pressure Estimation. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Barcelona, Spain, 13–17 April 2026; pp. 1–30. [Google Scholar]
- Li, Z.; Yang, Q.; Zhuang, Y.; Guo, C.; Zuo, X.; Long, X.; Yao, Y.; Cao, X.; Shen, Q.; Zhu, H. Pressure2Motion: Hierarchical Human Motion Reconstruction from Ground Pressure with Text Guidance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 23495–23505. [Google Scholar]
- Joo, H.; Liu, H.; Tan, L.; Gui, L.; Nabbe, B.; Matthews, I.; Kanade, T.; Nobuhara, S.; Sheikh, Y. Panoptic studio: A massively multiview system for social motion capture. In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 7–13 December 2015; pp. 3334–3342. [Google Scholar]
- Moon, G.; Yu, S.I.; Wen, H.; Shiratori, T.; Lee, K.M. Interhand2. 6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image. In Proceedings of the European Conference on Computer Vision, Glasgow, UK, 23–28 August 2020; Springer: Berlin/Heidelberg, Germany, 2020; pp. 548–564. [Google Scholar]
- Kwon, T.; Tekin, B.; Stühmer, J.; Bogo, F.; Pollefeys, M. H2o: Two hands manipulating objects for first person interaction recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 10138–10148. [Google Scholar]
- Damen, D.; Doughty, H.; Farinella, G.M.; Furnari, A.; Kazakos, E.; Ma, J.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; et al. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. Int. J. Comput. Vis. 2022, 130, 33–55. [Google Scholar]
- Damen, D.; Doughty, H.; Farinella, G.M.; Fidler, S.; Furnari, A.; Kazakos, E.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; et al. Scaling egocentric vision: The epic-kitchens dataset. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 720–736. [Google Scholar]
- Feng, R.; Chen, F.; Xu, D.; Zhong, L. ER: Extract-regress network for precise 3D reconstruction of interacting hands from monocular images. Vis. Comput. 2026, 42, 107. [Google Scholar] [CrossRef] [Scilit]
- Han, G.; Ye, Q.; Chen, A.; Chen, J. Caminterhand: Cooperative attention for multi-view interactive hand pose and mesh reconstruction. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 1041–1047. [Google Scholar]
- Ren, P.; Wang, J.; Sun, H.; Qi, Q.; Liu, X.; Zhang, M.; Zhang, L.; Wang, J.; Liao, J. Prior-aware Dynamic Temporal Modeling Framework for Sequential 3D Hand Pose Estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–20 October 2025; pp. 6476–6487. [Google Scholar]
- Yu, Z.; Zafeiriou, S.; Birdal, T. Dyn-hamr: Recovering 4d interacting hand motion from a dynamic camera. In Proceedings of the Computer Vision and Pattern Recognition Conference, Shanghai, China, 15–18 October 2025; pp. 27716–27726. [Google Scholar]
- Ren, P.; Wen, C.; Zheng, X.; Xue, Z.; Sun, H.; Qi, Q.; Wang, J.; Liao, J. Decoupled iterative refinement framework for interacting hands reconstruction from a single rgb image. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 8014–8025. [Google Scholar]
- Pavlakos, G.; Shan, D.; Radosavovic, I.; Kanazawa, A.; Fouhey, D.; Malik, J. Reconstructing hands in 3d with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 9826–9836. [Google Scholar]
- Pan, H.; Cai, Y.; Yang, J.; Niu, S.; Gao, Q.; Wang, X. HandFI: Multilevel Interacting Hand Reconstruction Based on Multilevel Feature Fusion in RGB Images. Sensors 2024, 25, 88. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Moon, G. Enhancing Hands in 3D Whole-Body Pose Estimation with Conditional Hands Modulator. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 8891–8900. [Google Scholar]
- Han, S.; Wu, P.c.; Zhang, Y.; Liu, B.; Zhang, L.; Wang, Z.; Si, W.; Zhang, P.; Cai, Y.; Hodan, T.; et al. UmeTrack: Unified multi-view end-to-end hand tracking for VR. In Proceedings of the SIGGRAPH Asia 2022 Conference Papers, Daegu, Republic of Korea, 6–9 December 2022; pp. 1–9. [Google Scholar]
- Örnek, E.P.; Labbé, Y.; Tekin, B.; Ma, L.; Keskin, C.; Forster, C.; Hodan, T. Foundpose: Unseen object pose estimation with foundation features. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; Springer: Berlin/Heidelberg, Germany, 2024; pp. 163–182. [Google Scholar]
- Mittal, H.; Agarwal, N.; Lo, S.Y.; Lee, K. Can’t make an omelette without breaking some eggs: Plausible action anticipation using large video-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 18580–18590. [Google Scholar]
- Kim, S.; Huang, D.; Xian, Y.; Hilliges, O.; Van Gool, L.; Wang, X. Palm: Predicting actions through language models. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; Springer: Berlin/Heidelberg, Germany, 2024; pp. 140–158. [Google Scholar]
- Bertasius, G.; Wang, H.; Torresani, L. Is space-time attention all you need for video understanding? In Proceedings of the ICML, Virtual, 18–24 July 2021; Volume 2, p. 4. [Google Scholar]
- Pasca, R.G.; Gavryushin, A.; Hamza, M.; Kuo, Y.L.; Mo, K.; Van Gool, L.; Hilliges, O.; Wang, X. Summarize the past to predict the future: Natural language descriptions of context boost multimodal object interaction anticipation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 18286–18296. [Google Scholar]
- Pramanick, S.; Song, Y.; Nag, S.; Lin, K.Q.; Shah, H.; Shou, M.Z.; Chellappa, R.; Zhang, P. Egovlpv2: Egocentric video-language pre-training with fusion in the backbone. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 5285–5297. [Google Scholar]
- Romero, J.L.; Min, K.; Tripathi, S.; Karimzadeh, M. Keystep Recognition using Graph Neural Networks. arXiv 2025, arXiv:2506.01102. [Google Scholar]
- Ragusa, F.; Farinella, G.M.; Furnari, A. Stillfast: An end-to-end approach for short-term object interaction anticipation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 3636–3645. [Google Scholar]
- Bianchi, E.; Liotta, A. SkillFormer: Unified multiview video understanding for proficiency estimation. In Proceedings of the Eighteenth International Conference on Machine Vision (ICMV 2025), Paris, France, 19–22 October 2025; SPIE: New York, NY, USA, 2026; Volume 14114, pp. 685–692. [Google Scholar]
- Bianchi, E.; Liotta, A. PATS: Proficiency-Aware Temporal Sampling for Multi-View Sports Skill Assessment. In Proceedings of the 2025 IEEE International Workshop on Sport, Technology and Research (STAR), Trento, Italy, 29–31 October 2025; IEEE: Piscataway, NJ, USA, 2025; pp. 1–6. [Google Scholar]
- Bianchi, E.; Staiano, J.; Liotta, A. ProfVLM: A Lightweight Video-Language Model for Multi-View Proficiency Estimation. arXiv 2025, arXiv:2509.26278. [Google Scholar]
- Tanoue, H.; Nishihara, H.; Suzuki, Y.; Hori, T.; Takushima, H.; Manojkumar, A.; Shibata, Y.; Takeda, M.; Beppu, F.; Hengwei, Z.; et al. CuriosAI Submission to the EgoExo4D Proficiency Estimation Challenge 2025. arXiv 2025, arXiv:2507.08022. [Google Scholar]
- Braun, B.; Armani, R.; Meier, M.; Moebus, M.; Holz, C. egoppg: Heart rate estimation from eye-tracking cameras in egocentric systems to benefit downstream vision tasks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–20 October 2025; pp. 5579–5590. [Google Scholar]
- John, R.; Kesari, A.; DiMatteo, V.; Dana, K. EgoCampus: Egocentric Pedestrian Eye Gaze Model and Dataset. arXiv 2025, arXiv:2512.07668. [Google Scholar]
- Pan, X.; Charron, N.; Yang, Y.; Peters, S.; Whelan, T.; Kong, C.; Parkhi, O.; Newcombe, R.; Ren, Y.C. Aria digital twin: A new benchmark dataset for egocentric 3d machine perception. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 20133–20143. [Google Scholar]
- Ma, L.; Ye, Y.; Hong, F.; Guzov, V.; Jiang, Y.; Postyeni, R.; Pesqueira, L.; Gamino, A.; Baiyya, V.; Kim, H.J.; et al. Nymeria: A massive collection of multimodal egocentric daily motion in the wild. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; Springer: Berlin/Heidelberg, Germany, 2024; pp. 445–465. [Google Scholar]
- Musabini, A.; Novikov, I.; Soula, S.; Leonet, C.; Wang, L.; Benmokhtar, R.; Burger, F.; Boulay, T.; Perrotton, X. Enhanced Parking Perception by Multi-Task Fisheye Cross-view Transformers. IET Conf. Proc. CP887 2024, 2024, 31–38. [Google Scholar] [CrossRef] [Scilit]
- Xue, J.; Smirnov, P.; Li, Z.; Shi, Y.; Chen, S.; Yin, X.; Yue, X.; Wang, L.; Wang, Y.; Lin, F.; et al. RePose: A Real-Time 3D Human Pose Estimation and Biomechanical Analysis Framework for Rehabilitation. arXiv 2026, arXiv:2601.00625. [Google Scholar]
- An, S.; Li, Y.; Ogras, U. mri: Multi-modal 3d human pose estimation dataset using mmwave, rgb-d, and inertial sensors. Adv. Neural Inf. Process. Syst. 2022, 35, 27414–27426. [Google Scholar]
- Solbach, M.D.; Tsotsos, J.K. Vision-based fallen person detection for the elderly. In Proceedings of the IEEE International Conference on Computer Vision Workshops, Venice, Italy, 22–29 October 2017; pp. 1433–1442. [Google Scholar]
- Jang, J.; Kim, D.; Park, C.; Jang, M.; Lee, J.; Kim, J. ETRI-activity3D: A large-scale RGB-D dataset for robots to recognize daily activities of the elderly. In Proceedings of the 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 24 October 2020–24 January 2021; IEEE: Piscataway, NJ, USA, 2020; pp. 10990–10997. [Google Scholar]
- Nguyen, T.T.; Kawanishi, Y.; John, V.; Komamizu, T.; Ide, I. MultiSensor-Home: A wide-area multi-modal multi-view dataset for action recognition and Transformer-based sensor fusion. In Proceedings of the 2025 IEEE 19th International Conference on Automatic Face and Gesture Recognition (FG), Clearwater, FL, USA, 26–30 May 2025; IEEE: Piscataway, NJ, USA, 2025; pp. 1–10. [Google Scholar]
- Nguyen, H.D.; Phan, D.T. CMA-ECG: Cross-modal attention for enhanced ECG quality assessment and denoising. Physiol. Meas. 2025, 46, 115001. [Google Scholar] [CrossRef] [Scilit]
- Jun, H.; Shaik, H.; DeVeaux, C.; Lewek, M.; Fuchs, H.; Bailenson, J. An evaluation study of 2D and 3D teleconferencing for remote physical therapy. PRESENCE Virtual Augment. Real. 2022, 31, 47–67. [Google Scholar] [CrossRef] [Scilit]
- Boldo, M.; De Marchi, M.; Martini, E.; Aldegheri, S.; Quaglia, D.; Fummi, F.; Bombieri, N. Real-time multi-camera 3D human pose estimation at the edge for industrial applications. Expert Syst. Appl. 2024, 252, 124089. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Zhu, W.; Gao, B.B.; Gan, Z.; Zhang, J.; Gu, Z.; Qian, S.; Chen, M.; Ma, L. Real-iad: A real-world multi-view dataset for benchmarking versatile industrial anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 22883–22892. [Google Scholar]
- Nazeri, A.; Mishra, S.; Wagner, A.; Ruskowski, M.; Stricker, D.; Rambach, J. A Multi-Camera Vision-Based Approach for Fine-Grained Assembly Quality Control. In Proceedings of the 2025 33rd European Signal Processing Conference (EUSIPCO), Palermo, Italy, 8–12 September 2025; IEEE: Piscataway, NJ, USA, 2025; pp. 1287–1291. [Google Scholar]
- Ragusa, F.; Furnari, A.; Farinella, G.M. Meccano: A multimodal egocentric dataset for humans behavior understanding in the industrial-like domain. Comput. Vis. Image Underst. 2023, 235, 103764. [Google Scholar] [CrossRef] [Scilit]
- Tao, W.; Leu, M.C.; Yin, Z. Multi-modal recognition of worker activity for human-centered intelligent manufacturing. Eng. Appl. Artif. Intell. 2020, 95, 103868. [Google Scholar] [CrossRef] [Scilit]
- González-Alonso, J.; Martín-Tapia, P.; González-Ortega, D.; Antón-Rodríguez, M.; Díaz-Pernas, F.J.; Martínez-Zarzuela, M. ME-WARD: A multimodal ergonomic analysis tool for musculoskeletal risk assessment from inertial and video data in working places. Expert Syst. Appl. 2025, 278, 127212. [Google Scholar] [CrossRef] [Scilit]
- Rahman, F.; Mim, M.S.; Baishakhi, F.B.; Hasan, M.; Morol, M.K. A systematic review on interactive virtual reality laboratory. In Proceedings of the 2nd International Conference on Computing Advancements, Dhaka, Bangladesh, 10–12 March 2022; pp. 491–500. [Google Scholar]
- Bozkir, E.; Kosel, C.; Seidel, T.; Kasneci, E. Automated visual attention detection using mobile eye tracking in behavioral classroom studies. arXiv 2025, arXiv:2505.07552. [Google Scholar]
- Yang, Z.; Wang, S.; Pan, S.; Li, H.; Wang, H.; Li, L.; Li, G.; Wen, Z.; Lin, B.; Tao, J.; et al. Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement. arXiv 2026, arXiv:2604.09473. [Google Scholar]
- Ragusa, F.; Furnari, A.; Battiato, S.; Signorello, G.; Farinella, G.M. EGO-CH: Dataset and fundamental tasks for visitors behavioral understanding using egocentric vision. Pattern Recognit. Lett. 2020, 131, 150–157. [Google Scholar] [CrossRef] [Scilit]
- Leonardi, R.; Furnari, A.; Ragusa, F.; Farinella, G.M. Leveraging Synthetic Data for Enhancing Egocentric Hand-Object Interaction Detection. Int. J. Comput. Vis. 2026, 134, 279. [Google Scholar] [CrossRef] [Scilit]
- Luo, M.; Xue, Z.; Dimakis, A.; Grauman, K. Put myself in your shoes: Lifting the egocentric perspective from exocentric videos. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; Springer: Berlin/Heidelberg, Germany, 2024; pp. 407–425. [Google Scholar]







| Dataset | Year | Views | Scale | Modalities | Primary Tasks |
|---|---|---|---|---|---|
| Panoptic Studio [110] | 2015 | 500+ cams | 700 sequences | RGB, depth | 3D pose, social interaction |
| InterHand2.6M [111] | 2020 | 10+ cams | 2.6M frames | RGB | 3D hand pose |
| H2O [112] | 2021 | Multi-view | 100K frames | RGB-D | HOI |
| HOT3D [6] | 2025 | Multi-view | 833 min | RGB, IR, depth | 3D hand–object tracking |
| EPIC-KITCHENS-100 [113] | 2022 | Ego | 100+ h | RGB, audio | Action recognition |
| Ego4D [4] | 2022 | Ego | 3000+ h | RGB, audio, IMU | Recognition, anticipation |
| Ego–Exo4D [5] | 2024 | Ego + Exo | 100+ h | RGB, audio | Cross-view understanding |
| MultiEgo [19] | 2025 | Ego + Exo | Multi-session | RGB, IMU | 4D reconstruction |
| Model | View | Architecture | MPJPE (mm) ↓ | MPVPE (mm) ↓ | Accel. (mm/s2) ↓ |
|---|---|---|---|---|---|
| Ren et al. [119] | Single | Decoupled Iterative Refinement | 10.49 | 10.26 | 6.28 |
| Pavlakos et al. [120] | Single | Transformers | 9.84 | 10.13 | 5.13 |
| Pan et al. [121] | Single | Feature Fusion | 9.38 | 9.61 | — |
| Moon et al. [122] | Single | Conditional Hand Modulator | — | 9.40 | — |
| Ren et al. [117] | Single (Video) | Temporal Convolution | 7.21 | 7.39 | 4.54 |
| Yu et al. [118] | Single (Video) | Generative Infilling + SLAM | 7.94 | 8.15 | 2.76 |
| Feng et al. [115] | Multi | Extract-Regress Network | 6.65 | 7.00 | — |
| Han et al. [116] | Multi | Cooperative Attention | 5.65 | 5.87 | — |
| Device | Single-View | Multi-View | Metric | Single | Multi | Gain |
|---|---|---|---|---|---|---|
| 3D hand tracking | ||||||
| Q3 | UmeTrack [123] | UmeTrack-2V | MKPE ↓ | 18.0 | 13.1 | 27.2% |
| Q3 | UmeTrack [123] + HOT3D [6] | UmeTrack + HOT3D-2V | MKPE ↓ | 15.4 | 10.9 | 29.2% |
| 6DoF object pose | ||||||
| Aria | FoundPose-1V [124] | FoundPose-3V | R@10 ↑ | 41.7 | 52.9 | 26.9% |
| Q3 | FoundPose-1V [124] | FoundPose-2V | R@10 ↑ | 46.6 | 55.9 | 20.0% |
| In-hand object lifting | ||||||
| Aria | MonoDepth | StereoMatch-3V, GT mask | R@10 ↑ | 30.2 | 86.2 | 185.4% |
| Aria | MonoDepth | StereoMatch-3V, pred. mask | R@10 ↑ | 23.3 | 56.4 | 142.1% |
| Q3 | N/R | StereoMatch-2V, GT mask | R@10 ↑ | – | 96.8 | – |
| Q3 | N/R | StereoMatch-2V, pred. mask | R@10 ↑ | – | 75.3 | – |
| Task | Dataset | Metric | Reported Gain (Multi vs. Single) |
|---|---|---|---|
| 3D hand pose & reconstruction | InterHand2.6M (Table 2) | MPJPE/MPVPE ↓ | ≈30% |
| 3D hand tracking | HOT3D (Table 3) | MKPE ↓ | 27.2–29.2% |
| 6DoF object pose | HOT3D (Table 3) | R@10 ↑ | 20.0–26.9% |
| In-hand object lifting | HOT3D (Table 3) | R@10 ↑ | 142.1–185.4% |
| Multimodal action segmentation | MultiEgo-style benchmarks (Section 4.3) | Segmentation metrics ↑ | 10–25% |
| Keystep recognition (naive fusion) | Ego–Exo4D (Table 5) | Accuracy ↑ | −5.40% to +2.2% |
| Keystep recognition (view-aware fusion) | Ego–Exo4D (Table 5) | Accuracy ↑ | +0.72 pt (52.36→53.08) |
| Proficiency estimation | Ego–Exo4D (Table 5) | Accuracy ↑ | +0.7 to +4.25 pt |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Phan, D.T.; Nguyen, H.D. Review: Techniques in Egocentric Multi-View Image Analysis: Advances, Challenges, and Future Directions. J. Imaging 2026, 12, 324. https://doi.org/10.3390/jimaging12070324
Phan DT, Nguyen HD. Review: Techniques in Egocentric Multi-View Image Analysis: Advances, Challenges, and Future Directions. Journal of Imaging. 2026; 12(7):324. https://doi.org/10.3390/jimaging12070324
Chicago/Turabian StylePhan, Duc Tri, and Hong Duc Nguyen. 2026. "Review: Techniques in Egocentric Multi-View Image Analysis: Advances, Challenges, and Future Directions" Journal of Imaging 12, no. 7: 324. https://doi.org/10.3390/jimaging12070324
APA StylePhan, D. T., & Nguyen, H. D. (2026). Review: Techniques in Egocentric Multi-View Image Analysis: Advances, Challenges, and Future Directions. Journal of Imaging, 12(7), 324. https://doi.org/10.3390/jimaging12070324

