HOSG-Nav: Hierarchical Open-Vocabulary Semantic Graph Navigation for Language-Guided Global Planning in 3D Gaussian Scenes
Abstract
1. Introduction
- We propose HOSG-Nav, a unified framework for natural-language-driven global navigation in complex indoor environments. The framework integrates open-vocabulary 3D scene representation, hierarchical semantic scene graph construction, and LLM-driven planning within a single pipeline, bridging continuous 3D scene modeling and structured navigation reasoning. Unlike prior methods that address isolated stages—such as open-vocabulary 3D mapping without structured planning, or scene-graph-based reasoning without continuous 3D semantic representation—HOSG-Nav tightly couples all three levels into an end-to-end pipeline.
- We develop an open-vocabulary 3D Gaussian scene representation and a navigation-oriented hierarchical scene graph construction method. Compressed CLIP features are lifted into a continuous 3D Gaussian field, while depth supervision is introduced to improve geometric stability and metric consistency. Based on the optimized field, Gaussian primitives are further abstracted into a region–object hierarchical scene graph through spatial–semantic clustering, semantic naming, functional region formation, and traversable topology modeling. This direct abstraction from a jointly optimized continuous Gaussian field to a discrete hierarchical graph differs from prior scene graph methods that build on discrete segmentation or online observation aggregation.
- We introduce an LLM-driven hierarchical planning strategy for natural language navigation. The proposed method decomposes free-form instructions into region-level, object-level, and attribute-level constraints, and combines hierarchical cross-modal retrieval with graph-search-based planning to generate executable global paths. The planning module is explicitly aligned with the region–object graph structure, enabling tighter coupling between language reasoning and structured scene abstraction than prior approaches that plan over flat object graphs or implicit policy representations. Extensive experiments demonstrate the effectiveness of HOSG-Nav in open-vocabulary scene representation, semantic target grounding, and global navigation.
2. Related Work
2.1. Semantic 3D Mapping for Navigation
2.2. Semantic 3D Scene Graphs for Navigation
2.3. Language-Guided Planning and Navigation
2.4. Representations in Advanced Autonomous Navigation Systems
3. Methods
3.1. Overview
3.2. Open-Vocabulary 3D Gaussian Scene Representation
3.2.1. 3DGS Preliminaries
3.2.2. 2D Semantic Feature Extraction and Compression
3.2.3. Joint Optimization with Semantic and Depth Supervision
3.3. Hierarchical Semantic Scene Graph Construction
3.3.1. Spatial–Semantic Joint Clustering
3.3.2. Object-Level Node Construction
3.3.3. Region-Level Node Construction
3.3.4. Topology Relation Generation
3.4. LLM-Driven Global Navigation Planning
3.4.1. Instruction Parsing and Hierarchical Query Generation
3.4.2. Hierarchical Region–Object Retrieval
3.4.3. Hierarchical Graph-Based Path Planning
3.4.4. Online Replanning and Pose Correction
4. Experiments
4.1. Datasets
4.2. Implementation Details
4.3. Evaluation of Open-Vocabulary Scene Representation
4.3.1. Novel-View Rendering Quality
4.3.2. Open-Vocabulary Semantic Understanding
4.4. Evaluation of Hierarchical Semantic Scene Graph
4.5. Global Planning and Navigation Performance
4.6. Ablation Study
4.6.1. Ablation on Open-Vocabulary 3D Scene Representation
4.6.2. Ablation on Hierarchical Scene Graph Construction
4.6.3. Ablation on Hierarchical Planning Strategy
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Elfes, A. Using occupancy grids for mobile robot perception and navigation. Computer 1989, 22, 46–57. [Google Scholar] [CrossRef]
- Choset, H.; Nagatani, K. Topological simultaneous localization and mapping (SLAM): Toward exact localization without explicit localization. IEEE Trans. Robot. Autom. 2001, 17, 125–137. [Google Scholar] [CrossRef]
- Khairuddin, A.R.; Talib, M.S.; Haron, H. Review on simultaneous localization and mapping (SLAM). In Proceedings of the 2015 IEEE International Conference on Control System, Computing and Engineering (ICCSCE), Penang, Malaysia, 27–29 November 2015; IEEE: New York, NY, USA, 2015; pp. 85–90. [Google Scholar]
- Elghazaly, G.; Frank, R.; Harvey, S.; Safko, S. High-Definition Maps: Comprehensive Survey, Challenges, and Future Perspectives. IEEE Open J. Intell. Transp. Syst. 2023, 4, 527–550. [Google Scholar] [CrossRef]
- Ma, Y.; Wang, T.; Bai, X.; Yang, H.; Hou, Y.; Wang, Y.; Qiao, Y.; Yang, R.; Manocha, D.; Zhu, X. Vision-Centric Bird’s Eye View Perception: A Survey. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 10978–10997. [Google Scholar] [CrossRef] [PubMed]
- Xu, H.; Chen, J.; Meng, S.; Wang, Y.; Chau, L.P. A Survey on Occupancy Perception for Autonomous Driving: The Information Fusion Perspective. arXiv 2024, arXiv:2405.05173. [Google Scholar] [CrossRef]
- Liu, Y.; Yuan, T.; Wang, Y.; Wang, Y.; Zhao, H. VectorMapNet: End-to-end Vectorized HD Map Learning. In Proceedings of the 40th International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July 2023; PMLR: Brookline, MA, USA, 2023; Volume 202, pp. 22352–22369. [Google Scholar]
- Zhou, Y.; Yan, L.; Han, Y.; Xie, H.; Zhao, Y. A Survey on the Key Technologies of UAV Motion Planning. Drones 2025, 9, 194. [Google Scholar] [CrossRef]
- Hornung, A.; Wurm, K.M.; Bennewitz, M.; Stachniss, C.; Burgard, W. OctoMap: An Efficient Probabilistic 3D Mapping Framework Based on Octrees. Auton. Robot. 2013, 34, 189–206. [Google Scholar] [CrossRef]
- Oleynikova, H.; Taylor, Z.; Fehr, M.; Nieto, J.; Siegwart, R. Voxblox: Incremental 3D Euclidean Signed Distance Fields for On-Board MAV Planning. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, Vancouver, BC, Canada, 24–28 September 2017; IEEE: New York, NY, USA, 2017; pp. 1366–1373. [Google Scholar] [CrossRef]
- Chaplot, D.S.; Salakhutdinov, R.; Gupta, A.; Gupta, S. Neural topological slam for visual navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; IEEE: New York, NY, USA, 2020; pp. 12875–12884. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference On Machine Learning, Virtual, 18–24 July 2021; PMLR: Brookline, MA, USA, 2021; pp. 8748–8763. [Google Scholar]
- Chen, B.; Xia, F.; Ichter, B.; Rao, K.; Gopalakrishnan, K.; Ryoo, M.S.; Stone, A.; Kappler, D. Open-vocabulary queryable scene representations for real world planning. In Proceedings of the 2023 IEEE International Conference on Robotics and Automation (ICRA), London, UK, 29 May–2 June 2023; IEEE: New York, NY, USA, 2023; pp. 11509–11522. [Google Scholar]
- Peng, S.; Genova, K.; Jiang, C.; Tagliasacchi, A.; Pollefeys, M.; Funkhouser, T. Openscene: 3d scene understanding with open vocabularies. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; IEEE: New York, NY, USA, 2023; pp. 815–824. [Google Scholar]
- Kerr, J.; Kim, C.M.; Goldberg, K.; Kanazawa, A.; Tancik, M. Lerf: Language embedded radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; IEEE: New York, NY, USA, 2023; pp. 19729–19739. [Google Scholar]
- Mildenhall, B.; Srinivasan, P.P.; Tancik, M.; Barron, J.T.; Ramamoorthi, R.; Ng, R. Nerf: Representing scenes as neural radiance fields for view synthesis. Commun. ACM 2021, 65, 99–106. [Google Scholar] [CrossRef]
- Kerbl, B.; Kopanas, G.; Leimkühler, T.; Drettakis, G. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph. 2023, 42, 139:1–139:14. [Google Scholar] [CrossRef]
- Deng, T.; Chen, Y.; Yang, J.; Yuan, S.; Liu, J.; Wang, D.; Chen, W. CGS-SLAM: Compact 3D Gaussian Splatting for Dense Visual SLAM. In Proceedings of the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Hangzhou, China, 19–25 October 2025; IEEE: New York, NY, USA, 2025; pp. 1606–1613. [Google Scholar] [CrossRef]
- Keetha, N.; Karhade, J.; Jatavallabhula, K.M.; Yang, G.; Scherer, S.; Ramanan, D.; Luiten, J. Splatam: Splat track & map 3d gaussians for dense rgb-d slam. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; IEEE: New York, NY, USA, 2024; pp. 21357–21366. [Google Scholar]
- Armeni, I.; He, Z.Y.; Gwak, J.; Zamir, A.R.; Fischer, M.; Malik, J.; Savarese, S. 3d scene graph: A structure for unified semantics, 3d space, and camera. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; IEEE: New York, NY, USA, 2019; pp. 5664–5673. [Google Scholar]
- Wu, S.C.; Wald, J.; Tateno, K.; Navab, N.; Tombari, F. Scenegraphfusion: Incremental 3d scene graph prediction from rgb-d sequences. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; IEEE: New York, NY, USA, 2021; pp. 7515–7525. [Google Scholar]
- Gu, Q.; Kuwajerwala, A.; Morin, S.; Jatavallabhula, K.M.; Sen, B.; Agarwal, A.; Rivera, C.; Paul, W.; Ellis, K.; Chellappa, R.; et al. Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2024; pp. 5021–5028. [Google Scholar]
- Werby, A.; Huang, C.; Büchner, M.; Valada, A.; Burgard, W. Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation. In Proceedings of the First Workshop on Vision-Language Models for Navigation and Manipulation at ICRA 2024, Yokohama, Japan, 17 May 2024. [Google Scholar]
- Anderson, P.; Wu, Q.; Teney, D.; Bruce, J.; Johnson, M.; Sünderhauf, N.; Reid, I.; Gould, S.; Van Den Hengel, A. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; IEEE: New York, NY, USA, 2018; pp. 3674–3683. [Google Scholar]
- Hong, Y.; Wu, Q.; Qi, Y.; Rodriguez-Opazo, C.; Gould, S. Vln bert: A recurrent vision-and-language bert for navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; IEEE: New York, NY, USA, 2021; pp. 1643–1653. [Google Scholar]
- Chen, S.; Guhur, P.L.; Schmid, C.; Laptev, I. History aware multimodal transformer for vision-and-language navigation. Adv. Neural Inf. Process. Syst. 2021, 34, 5834–5847. [Google Scholar]
- Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 2020, 33, 1877–1901. [Google Scholar]
- Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q.V.; Zhou, D. Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inf. Process. Syst. 2022, 35, 24824–24837. [Google Scholar]
- Qiao, Y.; Lyu, W.; Wang, H.; Wang, Z.; Li, Z.; Zhang, Y.; Tan, M.; Wu, Q. Open-nav: Exploring zero-shot vision-and-language navigation in continuous environment with open-source llms. In Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2025; pp. 6710–6717. [Google Scholar]
- Lin, B.; Nie, Y.; Wei, Z.; Chen, J.; Ma, S.; Han, J.; Xu, H.; Chang, X.; Liang, X. Navcot: Boosting llm-based vision-and-language navigation via learning disentangled reasoning. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 5945–5957. [Google Scholar] [CrossRef]
- Chang, M.; Gervet, T.; Khanna, M.; Yenamandra, S.; Shah, D.; Min, S.Y.; Shah, K.; Paxton, C.; Gupta, S.; Batra, D.; et al. Goat: Go to any thing. arXiv 2023, arXiv:2311.06430. [Google Scholar] [CrossRef]
- Gadre, S.Y.; Wortsman, M.; Ilharco, G.; Schmidt, L.; Song, S. Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; IEEE: New York, NY, USA, 2023; pp. 23171–23181. [Google Scholar]
- Shah, D.; Osiński, B.; Levine, S. Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action. In Proceedings of the Conference on Robot Learning; PMLR: Brookline, MA, USA, 2023; pp. 492–504. [Google Scholar]
- Mozos, Ó.M.; Stachniss, C.; Rottmann, A.; Burgard, W. Using adaboost for place labeling and topological map building. In Proceedings of the Robotics Research: Results of the 12th International Symposium ISRR; Springer: Berlin/Heidelberg, Germany, 2007; pp. 453–472. [Google Scholar]
- Salas-Moreno, R.F.; Newcombe, R.A.; Strasdat, H.; Kelly, P.H.; Davison, A.J. Slam++: Simultaneous localisation and mapping at the level of objects. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA, 23–28 June 2013; IEEE: New York, NY, USA, 2013; pp. 1352–1359. [Google Scholar]
- Grinvald, M.; Furrer, F.; Novkovic, T.; Chung, J.J.; Cadena, C.; Siegwart, R.; Nieto, J. Volumetric instance-aware semantic mapping and 3D object discovery. IEEE Robot. Autom. Lett. 2019, 4, 3037–3044. [Google Scholar] [CrossRef]
- McCormac, J.; Clark, R.; Bloesch, M.; Davison, A.; Leutenegger, S. Fusion++: Volumetric object-level slam. In Proceedings of the 2018 International Conference on 3D Vision (3DV); IEEE: New York, NY, USA, 2018; pp. 32–41. [Google Scholar]
- Nicholson, L.; Milford, M.; Sünderhauf, N. Quadricslam: Dual quadrics from object detections as landmarks in object-oriented slam. IEEE Robot. Autom. Lett. 2018, 4, 1–8. [Google Scholar] [CrossRef]
- Yang, S.; Scherer, S. Cubeslam: Monocular 3-d object slam. IEEE Trans. Robot. 2019, 35, 925–938. [Google Scholar] [CrossRef]
- Bloesch, M.; Czarnowski, J.; Clark, R.; Leutenegger, S.; Davison, A.J. Codeslam—Learning a compact, optimisable representation for dense visual slam. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; IEEE: New York, NY, USA, 2018; pp. 2560–2568. [Google Scholar]
- Rosinol, A.; Abate, M.; Chang, Y.; Carlone, L. Kimera: An open-source library for real-time metric-semantic localization and mapping. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2020; pp. 1689–1696. [Google Scholar]
- Huang, C.; Mees, O.; Zeng, A.; Burgard, W. Audio visual language maps for robot navigation. In Proceedings of the International Symposium on Experimental Robotics; Springer: Berlin/Heidelberg, Germany, 2023; pp. 105–117. [Google Scholar]
- Shafiullah, N.M.M.; Paxton, C.; Pinto, L.; Chintala, S.; Szlam, A. CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory. In Proceedings of the Robotics: Science and Systems (RSS), Daegu, Republic of Korea, 10–14 July 2023. [Google Scholar] [CrossRef]
- Huang, H.; Li, L.; Cheng, H.; Yeung, S.K. Photo-slam: Real-time simultaneous localization and photorealistic mapping for monocular stereo and rgb-d cameras. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; IEEE: New York, NY, USA, 2024; pp. 21584–21593. [Google Scholar]
- Yan, C.; Qu, D.; Xu, D.; Zhao, B.; Wang, Z.; Wang, D.; Li, X. Gs-slam: Dense visual slam with 3d gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; IEEE: New York, NY, USA, 2024; pp. 19595–19604. [Google Scholar]
- Li, M.; Liu, S.; Zhou, H.; Zhu, G.; Cheng, N.; Deng, T.; Wang, H. Sgs-slam: Semantic gaussian splatting for neural dense slam. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2024; pp. 163–179. [Google Scholar]
- Zhu, S.; Qin, R.; Wang, G.; Liu, J.; Wang, H. Semgauss-slam: Dense semantic gaussian splatting slam. In Proceedings of the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2025; pp. 21174–21181. [Google Scholar]
- Lee, S.; Ha, S.; Kang, K.; Choi, J.; Tak, S.; Yu, H. LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM. arXiv 2025, arXiv:2511.16144. [Google Scholar]
- Ha, S.; Lee, S.; Kang, K.; Choi, J.; Tak, S.; Yu, H. LangGS-SLAM: Real-Time Language-Feature Gaussian Splatting SLAM. arXiv 2026, arXiv:2602.06991. [Google Scholar]
- Chang, X.; Ren, P.; Xu, P.; Li, Z.; Chen, X.; Hauptmann, A. A comprehensive survey of scene graphs: Generation and application. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 45, 1–26. [Google Scholar] [CrossRef] [PubMed]
- Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.J.; Shamma, D.A.; et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vis. 2017, 123, 32–73. [Google Scholar] [CrossRef]
- Qian, T.; Chen, J.; Chen, S.; Wu, B.; Jiang, Y.G. Scene Graph Refinement Network for Visual Question Answering. IEEE Trans. Multimed. 2023, 25, 3950–3961. [Google Scholar] [CrossRef]
- Nguyen, K.; Tripathi, S.; Du, B.; Guha, T.; Nguyen, T.Q. In defense of scene graphs for image captioning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; IEEE: New York, NY, USA, 2021; pp. 1407–1416. [Google Scholar]
- Greve, E.; Büchner, M.; Vödisch, N.; Burgard, W.; Valada, A. Collaborative dynamic 3d scene graphs for automated driving. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2024; pp. 11118–11124. [Google Scholar]
- Hughes, N.; Chang, Y.; Carlone, L. Hydra: A Real-Time Spatial Perception System for 3D Scene Graph Construction and Optimization. In Proceedings of the Robotics: Science and Systems (RSS), New York, NY, USA, 27 June–1 July 2022. [Google Scholar] [CrossRef]
- Wald, J.; Dhamo, H.; Navab, N.; Tombari, F. Learning 3d semantic scene graphs from 3d indoor reconstructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; IEEE: New York, NY, USA, 2020; pp. 3961–3970. [Google Scholar]
- Rana, K.; Haviland, J.; Garg, S.; Abou-Chakra, J.; Reid, I.; Suenderhauf, N. SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning. In Proceedings of the 7th Conference on Robot Learning; Proceedings of Machine Learning Research; PMLR: Brookline, MA, USA, 2023; Volume 229, pp. 23–72. [Google Scholar]
- Rosinol, A.; Gupta, A.; Abate, M.; Shi, J.; Carlone, L. 3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans. In Proceedings of the Robotics: Science and Systems (RSS), Virtual, 12–16 July 2020. [Google Scholar] [CrossRef]
- Bavle, H.; Sanchez-Lopez, J.L.; Shaheer, M.; Civera, J.; Voos, H. Situational graphs for robot navigation in structured indoor environments. IEEE Robot. Autom. Lett. 2022, 7, 9107–9114. [Google Scholar] [CrossRef]
- Kümmerle, R.; Grisetti, G.; Strasdat, H.; Konolige, K.; Burgard, W. g 2 o: A general framework for graph optimization. In Proceedings of the 2011 IEEE International Conference on Robotics and Automation; IEEE: New York, NY, USA, 2011; pp. 3607–3613. [Google Scholar]
- Fernandez-Cortizas, M.; Bavle, H.; Perez-Saura, D.; Sanchez-Lopez, J.L.; Campoy, P.; Voos, H. Multi S-graphs: An efficient distributed semantic-relational collaborative SLAM. IEEE Robot. Autom. Lett. 2024, 9, 6004–6011. [Google Scholar] [CrossRef]
- Gu, J.; Stefani, E.; Wu, Q.; Thomason, J.; Wang, X. Vision-and-language navigation: A survey of tasks, methods, and future directions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, 22–27 May 2022; Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 7606–7623. [Google Scholar]
- Cai, Y.; He, X.; Wang, M.; Guo, H.; Yau, W.Y.; Lv, C. Cl-cotnav: Closed-loop hierarchical chain-of-thought for zero-shot object-goal navigation with vision-language models. arXiv 2025, arXiv:2504.09000. [Google Scholar]
- Wu, P.; Mu, Y.; Wu, B.; Hou, Y.; Ma, J.; Zhang, S.; Liu, C. Voronav: Voronoi-based zero-shot object navigation with large language model. arXiv 2024, arXiv:2401.02695. [Google Scholar]
- Zhong, Z.; He, Y.; Li, P.; Yu, F.; Ma, F. A Language-Driven Navigation Strategy Integrating Semantic Maps and Large Language Models. In Proceedings of the 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2024; pp. 9753–9760. [Google Scholar] [CrossRef]
- Igelbrink, F.; Renz, M.; Günther, M.; Powell, P.; Niecksch, L.; Lima, O.; Atzmueller, M.; Hertzberg, J. Online Knowledge Integration for 3D Semantic Mapping: A Survey. SSRN 2025. [Google Scholar] [CrossRef]
- Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Leoni Aleman, F.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. GPT-4 Technical Report; Technical Report; OpenAI: San Francisco, CA, USA, 2023. [Google Scholar]
- Rajvanshi, A.; Sikka, K.; Lin, X.; Lee, B.; Chiu, H.P.; Velasquez, A. Saynav: Grounding large language models for dynamic planning to navigation in new environments. In Proceedings of the International Conference on Automated Planning and Scheduling; AAAI Press: Washington, DC, USA, 2024; Volume 34, pp. 464–474. [Google Scholar]
- Honerkamp, D.; Büchner, M.; Despinoy, F.; Welschehold, T.; Valada, A. Language-grounded dynamic scene graphs for interactive object search with mobile manipulation. IEEE Robot. Autom. Lett. 2024, 9, 8298–8305. [Google Scholar] [CrossRef]
- Ni, Z.; Deng, X.; Tai, C.; Zhu, X.; Xie, Q.; Huang, W.; Wu, X.; Zeng, L. Grid: Scene-graph-based instruction-driven robotic task planning. In Proceedings of the 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2024; pp. 13765–13772. [Google Scholar]
- Cheng, L.; Qi, Z.; Zhou, Z.; Lu, C.; Xiong, G. LT-Gaussian: Long-Term Map Update Using 3D Gaussian Splatting for Autonomous Driving. arXiv 2025, arXiv:2508.01704. [Google Scholar]
- Schonberger, J.L.; Frahm, J.M. Structure-from-motion revisited. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; IEEE: New York, NY, USA, 2016; pp. 4104–4113. [Google Scholar]
- Schubert, E.; Sander, J.; Ester, M.; Kriegel, H.P.; Xu, X. DBSCAN revisited, revisited: Why and how you should (still) use DBSCAN. ACM Trans. Database Syst. Tods 2017, 42, 1–21. [Google Scholar] [CrossRef]
- Sulaiman, H.A.; Othman, M.A.; Ismail, M.M.; Said, M.A.M.; Ramlee, A.; Misran, M.H.; Bade, A.; Abdullah, M.H. Distance computation using axis aligned bounding box (AABB) parallel distribution of dynamic origin point. In Proceedings of the 2013 Annual International Conference on Emerging Research Areas and 2013 International Conference on Microelectronics, Communications and Renewable Energy, Kanjirapally, India, 4–6 June 2013; IEEE: New York, NY, USA, 2013; pp. 1–6. [Google Scholar]
- Foead, D.; Ghifari, A.; Kusuma, M.B.; Hanafiah, N.; Gunawan, E. A systematic literature review of A* pathfinding. Procedia Comput. Sci. 2021, 179, 507–514. [Google Scholar] [CrossRef]
- Straub, J.; Whelan, T.; Ma, L.; Chen, Y.; Wijmans, E.; Green, S.; Engel, J.J.; Mur-Artal, R.; Ren, C.; Verma, S.; et al. The Replica Dataset: A Digital Replica of Indoor Spaces. arXiv 2019, arXiv:1906.05797. [Google Scholar] [CrossRef]
- Zhi, S.; Laidlow, T.; Leutenegger, S.; Davison, A.J. In-place scene labelling and understanding with implicit scene representation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; IEEE: New York, NY, USA, 2021; pp. 15838–15847. [Google Scholar]









| Novel-View Rendering | Open-Vocabulary Semantics | |||||
|---|---|---|---|---|---|---|
| Method | PSNR | SSIM | LPIPS | mIoU | F-mIoU | mAcc |
| 3DGS | 28.79 ± 0.05 | 0.943 ± 0.003 | 0.065 ± 0.004 | – | – | – |
| HOV-SG | – | – | – | 0.231 ± 0.004 | 0.386 ± 0.006 | 0.304 ± 0.005 |
| HOSG-Nav | 34.66 ± 0.08 | 0.973 ± 0.002 | 0.096 ± 0.005 | 0.244 ± 0.005 | 0.455 ± 0.007 | 0.362 ± 0.006 |
| Method | Retrieval-SR10 [%] | Navigation-SR [%] |
|---|---|---|
| HOV-SG | 31.48 ± 0.5 | 40.41 ± 0.6 |
| HOSG-Nav | 33.17 ± 0.8 | 42.26 ± 0.9 |
| Novel-View Rendering | Open-Vocabulary Semantics | |||||
|---|---|---|---|---|---|---|
| Variant | PSNR | SSIM | LPIPS | mIoU | F-mIoU | mAcc |
| w/o Semantic | 34.12 | 0.970 | 0.091 | 0.214 | 0.401 | 0.318 |
| w/o Depth | 33.41 | 0.967 | 0.103 | 0.236 | 0.432 | 0.347 |
| Full | 34.66 | 0.973 | 0.096 | 0.244 | 0.455 | 0.362 |
| Variant | Retrieval-SR10 [%] | Navigation-SR [%] |
|---|---|---|
| Geo-only Clustering | 30.92 | 39.18 |
| Flat Object Graph | 32.11 | 40.87 |
| Full Hierarchical Graph | 33.17 | 42.26 |
| Variant | Retrieval-SR10 [%] | Navigation-SR [%] |
|---|---|---|
| Flat Retrieval + A* | 31.36 | 39.95 |
| Region → Object Retrieval w/o Attribute | 32.74 | 41.63 |
| Full Planning | 33.17 | 42.26 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, Y.; Qin, K.; Chen, W.; Wu, H. HOSG-Nav: Hierarchical Open-Vocabulary Semantic Graph Navigation for Language-Guided Global Planning in 3D Gaussian Scenes. Electronics 2026, 15, 2179. https://doi.org/10.3390/electronics15102179
Li Y, Qin K, Chen W, Wu H. HOSG-Nav: Hierarchical Open-Vocabulary Semantic Graph Navigation for Language-Guided Global Planning in 3D Gaussian Scenes. Electronics. 2026; 15(10):2179. https://doi.org/10.3390/electronics15102179
Chicago/Turabian StyleLi, Yuchen, Kai Qin, Weiyi Chen, and Haitao Wu. 2026. "HOSG-Nav: Hierarchical Open-Vocabulary Semantic Graph Navigation for Language-Guided Global Planning in 3D Gaussian Scenes" Electronics 15, no. 10: 2179. https://doi.org/10.3390/electronics15102179
APA StyleLi, Y., Qin, K., Chen, W., & Wu, H. (2026). HOSG-Nav: Hierarchical Open-Vocabulary Semantic Graph Navigation for Language-Guided Global Planning in 3D Gaussian Scenes. Electronics, 15(10), 2179. https://doi.org/10.3390/electronics15102179

