On Segment-Aware Monocular Depth Estimation Using Vision Transformers
Abstract
1. Introduction
- Segment-aware depth experts with transformers. We propose a multi-branch monocular depth architecture in which each semantic class is handled by a dedicated, compact ViT encoder–decoder expert that receives a mask-gated RGB input and predicts a dense depth proposal for that class.
- Two composition strategies (learned vs. parameter-free). We study how to compose per-class depth proposals into a single prediction by comparing (i) a cross-attention fusion module that directly predicts depth from the stack of per-class proposals and masks, and (ii) a parameter-free stitched summation that sums the mask-gated proposals.
- Controlled study with accuracy–efficiency trade-offs. Under a controlled training protocol (same data, optimiser, and per-branch backbone) on Virtual KITTI 2, we evaluate 3-, 6-, and 11-class schemes and quantify both accuracy and computational cost. We find that most gains come from the segment-wise decomposition, while simple summation is a strong default that is competitive with, and often outperforms, learned fusion without additional overhead.
2. Related Work
2.1. Monocular Depth Estimation
2.2. Semantic Guidance for Depth Estimation
2.3. Depth Refinement with Masks
3. Segment-Aware Depth Estimation
3.1. Preliminaries
3.2. Per-Class Depth Branches
3.3. Cross-Attention Fusion
3.4. Losses
3.5. No-Fusion Variant
4. Experimental Setup
4.1. Dataset
- 3-class: Vegetation = {Tree, Vegetation}; Ground = {Road, Terrain}; Artificial = {Building, Truck, Car, Van, GuardRail, TrafficSign, TrafficLight, Pole, Misc}.
- 6-class: Road = {Road}; Terrain = {Terrain}; Building = {Building}; Vegetation = {Tree, Vegetation}; Object = {GuardRail, TrafficSign, TrafficLight, Pole, Misc}; Vehicle = {Truck, Van, Car}.
- 11-class: Terrain, Tree, Vegetation, Building, Road, GuardRail, TrafficSign, TrafficLight, Pole, Misc, Vehicle = {Truck, Car, Van}.
4.2. Model Variants
- Baseline: a single ViT encoder-decoder that predicts depth from RGB only, without access to semantic masks.
- Segment-aware (Fusion): per-class experts (one encoder–decoder per semantic class), followed by a cross-attention fusion module. The fusion module takes the stack of per-class depth maps and their masks as input and directly predicts a fused depth map.
- Segment-aware (No-Fusion)—proposed: per-class experts (one encoder–decoder per semantic class); the class-wise predictions are mask-gated and summed to form the final depth, with no additional fusion or refinement.
4.3. Training Details
4.4. Evaluation Metrics
- Absolute Relative Error (AbsRel): .
- Squared Relative Error (SqRel): .
- Root Mean Squared Error (RMSE): .
- Log RMSE (RMSElog): .
- Threshold accuracy : .
5. Results
5.1. Accuracy Results
5.2. Model Complexity and Runtime
6. Discussion
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Lee, S.H.; Mo, S.; Yu, S.X. SHED Light on Segmentation for Depth Estimation. In Proceedings of the Structural Priors for Vision Workshop at ICCV’25, Honolulu, HI, USA, 19 October 2025. [Google Scholar]
- Jiao, J.; Cao, Y.; Song, Y.; Lau, R. Look Deeper into Depth: Monocular Depth Estimation with Semantic Booster and Attention-Driven Loss. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018. [Google Scholar]
- Cabon, Y.; Murray, N.; Humenberger, M. Virtual KITTI 2. arXiv 2020, arXiv:2001.10773. [Google Scholar] [CrossRef] [Scilit]
- Gaidon, A.; Wang, Q.; Cabon, Y.; Vig, E. Virtual worlds as proxy for multi-object tracking analysis. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 26 June–1 July 2016; pp. 4340–4349. [Google Scholar]
- Wang, L.; Zhang, J.; Wang, O.; Lin, Z.; Lu, H. SDC-Depth: Semantic Divide-and-Conquer Network for Monocular Depth Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 14–19 June 2020. [Google Scholar]
- Gao, N.; He, F.; Jia, J.; Shan, Y.; Zhang, H.; Zhao, X.; Huang, K. PanopticDepth: A Unified Framework for Depth-Aware Panoptic Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 1632–1642. [Google Scholar]
- Arampatzakis, V.; Pavlidis, G.; Mitianoudis, N.; Papamarkos, N. Monocular Depth Estimation: A Thorough Review. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 2396–2414. [Google Scholar] [CrossRef] [Scilit]
- Fu, H.; Gong, M.; Wang, C.; Batmanghelich, K.; Tao, D. Deep Ordinal Regression Network for Monocular Depth Estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018. [Google Scholar]
- Ranftl, R.; Lasinger, K.; Hafner, D.; Schindler, K.; Koltun, V. Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-Shot Cross-Dataset Transfer. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 1623–1637. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, P.Y.; Liu, A.H.; Liu, Y.C.; Wang, Y.C.F. Towards Scene Understanding: Unsupervised Monocular Depth Estimation with Semantic-Aware Representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 16–20 June 2019. [Google Scholar]
- Li, R.; Xue, D.; Su, S.; He, X.; Mao, Q.; Zhu, Y.; Sun, J.; Zhang, Y. Learning depth via leveraging semantics: Self-supervised monocular depth estimation with both implicit and explicit semantic guidance. Pattern Recognit. 2023, 137, 109297. [Google Scholar] [CrossRef] [Scilit]
- Ladicky, L.; Shi, J.; Pollefeys, M. Pulling Things out of Perspective. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA, 23–28 June 2014. [Google Scholar]
- Xing, D.; Shen, J.; Ho, C.; Tzes, A. ROIFormer: Semantic-Aware Region of Interest Transformer for Efficient Self-Supervised Monocular Depth Estimation. Proc. AAAI Conf. Artif. Intell. 2023, 37, 2983–2991. [Google Scholar] [CrossRef] [Scilit]
- Kim, S.Y.; Zhang, J.; Niklaus, S.; Fan, Y.; Chen, S.; Lin, Z.; Kim, M. Layered Depth Refinement with Mask Guidance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 3855–3865. [Google Scholar]
- Saeedan, F.; Roth, S. Boosting Monocular Depth with Panoptic Segmentation Maps. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Virtual, 5–9 January 2021; pp. 3853–3862. [Google Scholar]
- Eigen, D.; Puhrsch, C.; Fergus, R. Depth Map Prediction from a Single Image using a Multi-Scale Deep Network. In Proceedings of the 28th International Conference on Neural Information Processing Systems; Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N., Weinberger, K., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2014; Volume 27. [Google Scholar]
- Laina, I.; Rupprecht, C.; Belagiannis, V.; Tombari, F.; Navab, N. Deeper Depth Prediction with Fully Convolutional Residual Networks. In Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV), Stanford, CA, USA, 25–28 October 2016; pp. 239–248. [Google Scholar] [CrossRef] [Scilit]






| Model | Classes | AbsRel | SqRel | RMSE | RMSE_log | |||
|---|---|---|---|---|---|---|---|---|
| Baseline | - | 0.243 | 3.762 | 11.952 | 0.297 | 0.684 | 0.894 | 0.957 |
| Segment-aware (Fusion) | 3 | 0.153 | 3.090 | 11.153 | 0.230 | 0.841 | 0.945 | 0.975 |
| Segment-aware (No-Fusion) | 3 | 0.185 | 2.259 | 9.547 | 0.237 | 0.764 | 0.926 | 0.972 |
| Segment-aware (Fusion) | 6 | 0.155 | 3.146 | 11.177 | 0.227 | 0.844 | 0.946 | 0.975 |
| Segment-aware (No-Fusion) | 6 | 0.156 | 1.908 | 9.101 | 0.207 | 0.809 | 0.944 | 0.980 |
| Segment-aware (Fusion) | 11 | 0.162 | 3.252 | 11.304 | 0.235 | 0.828 | 0.941 | 0.974 |
| Segment-aware (No-Fusion) | 11 | 0.152 | 1.895 | 9.250 | 0.204 | 0.813 | 0.946 | 0.981 |
| Condition | Model | AbsRel | |
|---|---|---|---|
| clean | Baseline | 0.2359 | 0.6889 |
| clean | No-Fusion | 0.1419 | 0.8314 |
| clean | Fusion | 0.1597 | 0.8281 |
| erode (r = 1) | No-Fusion | 0.1658 | 0.8029 |
| erode (r = 1) | Fusion | 0.1649 | 0.8185 |
| dilate (r = 1) | No-Fusion | 0.2159 | 0.7462 |
| dilate (r = 1) | Fusion | 0.1842 | 0.8099 |
| flip (p = 0.03) | No-Fusion | 0.2960 | 0.7522 |
| flip (p = 0.03) | Fusion | 0.1778 | 0.7927 |
| Model | Classes | Params (M) | FLOPs (G) |
|---|---|---|---|
| Baseline | - | 0.91 | 6.88 |
| Segment-aware (Fusion) | 3 | 2.85 | 38.04 |
| Segment-aware (No-Fusion) | 3 | 1.76 | 18.50 |
| Segment-aware (Fusion) | 6 | 4.62 | 56.55 |
| Segment-aware (No-Fusion) | 6 | 3.52 | 37.01 |
| Segment-aware (Fusion) | 11 | 7.55 | 87.39 |
| Segment-aware (No-Fusion) | 11 | 6.46 | 67.85 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Arampatzakis, V.; Pavlidis, G.; Mitianoudis, N.; Papamarkos, N. On Segment-Aware Monocular Depth Estimation Using Vision Transformers. Information 2026, 17, 145. https://doi.org/10.3390/info17020145
Arampatzakis V, Pavlidis G, Mitianoudis N, Papamarkos N. On Segment-Aware Monocular Depth Estimation Using Vision Transformers. Information. 2026; 17(2):145. https://doi.org/10.3390/info17020145
Chicago/Turabian StyleArampatzakis, Vasileios, George Pavlidis, Nikolaos Mitianoudis, and Nikos Papamarkos. 2026. "On Segment-Aware Monocular Depth Estimation Using Vision Transformers" Information 17, no. 2: 145. https://doi.org/10.3390/info17020145
APA StyleArampatzakis, V., Pavlidis, G., Mitianoudis, N., & Papamarkos, N. (2026). On Segment-Aware Monocular Depth Estimation Using Vision Transformers. Information, 17(2), 145. https://doi.org/10.3390/info17020145

