Semantic-Guided Spatial and Temporal Fusion Framework for Enhancing Monocular Video Depth Estimation
Abstract
1. Introduction
- We propose STF-Depth, a training-free post-processing framework that significantly enhances the geometric accuracy and temporal stability of existing MDE models.
- We introduce a Depth-Stratified Vanishing Point Estimation method combined with Semantic-Guided Spatial Fusion, which robustly corrects perspective distortions in complex environments.
- Extensive experiments on the NYU Depth V2, KITTI, and TartanAir datasets demonstrate the effectiveness of our approach. Notably, STF-Depth achieved a 25.7% reduction in Absolute Relative error (AbsRel) and significantly improved temporal consistency compared to state-of-the-art backbone models.
2. Related Work
2.1. Monocular Depth Estimation
2.2. Semantic and Panoptic Segmentation
2.3. Video Depth Estimation and Temporal Consistency
3. Methodology
3.1. Heterogeneous Information Extraction
3.2. Semantic-Guided Spatial Fusion
3.2.1. Depth-Stratified and RANSAC-Based Vanishing Point Estimation
- Quantization: The initial depth map values are normalized and divided into N uniform depth layers. This effectively discretizes the continuous depth space into multiple analyzeable cross-sections.
- Centroid Extraction: For each depth layer, we generate a pixel mask and calculate image moments to determine the spatial centroid , which represents the center of density for that layer.
- RANSAC Regression: Due to imperfections in the initial depth map or complex object arrangements, spatial centroids of certain layers may deviate from the main perspective trajectory. To address this, we apply the RANSAC algorithm to the set of extracted centroids and the representative coordinate of the Max. Depth Region, denoted as . RANSAC effectively filters out outliers—such as dynamic objects or clutter that do not align with the global perspective—and derives an optimal trend line that represents the dominant geometric structure of the scene.This RANSAC-based approach ensures robustness against noise, allowing for stable VP estimation even in unstructured settings where geometric cues might be partially obscured.
- Vanishing Point Determination: The final Vanishing Point is defined as the point on the derived optimal trend line that minimizes the Euclidean distance to . Here, to ensure robustness against local depth artifacts (e.g., sensor noise or reflections), is calculated as the spatial centroid of the top 5% of pixels with the largest depth values (i.e., the Max. Depth Region visualized in Figure 4), rather than relying on a single maximum point.
3.2.2. Foreground Separation via Dynamic Gradient Correction
3.2.3. Instance & VP-Distance Aware Re-Ordering
- VP-Distance Weighting: According to linear perspective, depth perception sensitivity varies between the area around the VP and the periphery. We calculate the Euclidean distance between the center of each instance k and the VP. We generate a distance-proportional offset that assigns a larger weight based on the distance from the VP, performing stereoscopic correction according to the position within the frame.
- Class Attribute Adaptation (Push-Pull Strategy): We analyze panoptic labels to distinguish whether an instance is a “Thing” (countable object, e.g., person, car) or “Stuff” (amorphous background element, e.g., sky, wall). We adopt a Dual Strategy: for “Stuff,” the calculated offset is multiplied by a negative value (−1) to push the depth layer backward, while for “Things,” a positive offset is maintained to pull it forward.
- Overlap-based Penalty Injection: We detect spatial overlaps by analyzing the bounding boxes of instances. For overlapping pairs, we confirm the front-to-back relationship by comparing mean depths and inject an additional penalty into the overlapping region. This penalty is calculated by applying an amplification factor (1.5) to the depth difference between the two objects.
3.3. Efficient Temporal Fusion
4. Experiments
4.1. Experimental Setup
4.1.1. Datasets
4.1.2. Evaluation Metrics
4.1.3. Implementation Details
4.2. Ablation Study
- Method A (Baseline): The original output of the pre-trained monocular depth estimation model without any post-processing ().
- Method B (+ Separation): Baseline with the Background-Foreground Depth Separation (Section 3.2.2) module applied. This stage corrects background perspective and eliminates floating artifacts using semantic masks and vanishing point estimation.
- Method C (+ Re-ordering): Method B with the addition of the Instance-Aware Depth Re-ordering (Section 3.2.3) module. This stage rearranges the depth order of overlapping objects and injects separation margins using panoptic information.
- Method D (Full STF-Depth): The final framework integrating the Efficient Temporal Fusion (Section 3.3) module into Method C.
4.2.1. Baseline Only
4.2.2. Baseline + Spatial Fusion
4.2.3. Baseline + Spatial + Re-Ordering Fusion
4.2.4. Baseline + Spatial + Re-Ordering + Temporal Fusion (Full STF-Depth)
4.3. Quantitative Results
4.3.1. Depth Estimation Quality Analysis
4.3.2. Temporal Consistency Analysis
4.4. Qualitative Results
5. Discussion
5.1. Analysis of Performance Improvements and Theoretical Implications
5.2. Limitations
5.2.1. Dependency on Pre-Processing Models
5.2.2. Real-Time Processing Constraints and Efficiency Analysis
5.2.3. Single Vanishing Point and Linear Perspective Assumptions
5.3. Future Work
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Eigen, D.; Puhrsch, C.; Fergus, R. Depth map prediction from a single image using a multi-scale deep network. Adv. Neural Inf. Process. Syst. 2014, 27, 2366–2374. [Google Scholar]
- Birkl, R.; Wofk, D.; Müller, M. MiDaS v3.1—A Model Zoo for Robust Monocular Relative Depth Estimation. arXiv 2023, arXiv:2307.14460. [Google Scholar]
- Yang, L.; Kang, B.; Huang, Z.; Xu, X.; Feng, J.; Zhao, H. Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. arXiv 2024, arXiv:2401.10891. [Google Scholar] [CrossRef] [Scilit]
- Kim, H.; Son, Y. Generating Multi-View Action Data from a Monocular Camera Video by Fusing Human Mesh Recovery and 3D Scene Reconstruction. Appl. Sci. 2025, 15, 10372. [Google Scholar] [CrossRef] [Scilit]
- Wang, S. Development of approach to an automated acquisition of static street view images using transformer architecture for analysis of Building characteristics. Sci. Rep. 2025, 15, 29062. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, S.; Park, S.; Kim, J.; Kim, J. Safety helmet monitoring on construction sites using YOLOv10 and advanced transformer architectures with surveillance and body-worn cameras. J. Constr. Eng. Manag. 2025, 151, 04025186. [Google Scholar] [CrossRef] [Scilit]
- Watson, J.; Aodha, O.M.; Prisacariu, V.; Brostow, G.; Firman, M. The Temporal Opportunist: Self-Supervised Multi-Frame Monocular Depth. arXiv 2021, arXiv:2104.14540. [Google Scholar]
- Chen, S.; Guo, H.; Zhu, S.; Zhang, F.; Huang, Z.; Feng, J.; Kang, B. Video Depth Anything: Consistent Depth Estimation for Super-Long Videos. arXiv 2025, arXiv:2501.12375. [Google Scholar] [CrossRef] [Scilit]
- Dosovitskiy, A. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Ranftl, R.; Lasinger, K.; Hafner, D.; Schindler, K.; Koltun, V. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 44, 1623–1637. [Google Scholar]
- Ranftl, R.; Bochkovskiy, A.; Koltun, V. Vision transformers for dense prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, BC, Canada, 11–17 October 2021; pp. 12179–12188. [Google Scholar]
- Zhang, Q.; Zhu, Y.; Cordeiro, F.R.; Chen, Q. PSSCL: A progressive sample selection framework with contrastive loss designed for noisy labels. Pattern Recognit. 2025, 161, 111284. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; Yuille, A.L. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 40, 834–848. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kirillov, A.; He, K.; Girshick, R.; Rother, C.; Dollár, P. Panoptic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 9404–9413. [Google Scholar]
- Jain, J.; Li, J.; Chiu, M.T.; Hassani, A.; Orlov, N.; Shi, H. Oneformer: One transformer to rule universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 2989–2998. [Google Scholar]
- Teed, Z.; Deng, J. Raft: Recurrent all-pairs field transforms for optical flow. In Proceedings of the European Conference on Computer Vision, Glasgow, UK, 23–28 August 2020; Springer: Berlin/Heidelberg, Germany, 2020; pp. 402–419. [Google Scholar]
- Eom, C.; Park, H.; Ham, B. Temporally consistent depth prediction with flow-guided memory units. IEEE Trans. Intell. Transp. Syst. 2019, 21, 4626–4636. [Google Scholar] [CrossRef] [Scilit]
- Watson, J.; Mac Aodha, O.; Prisacariu, V.; Brostow, G.; Firman, M. The temporal opportunist: Self-supervised multi-frame monocular depth. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 1164–1174. [Google Scholar]
- Patil, V.; Van Gansbeke, W.; Dai, D.; Van Gool, L. Don’t forget the past: Recurrent depth estimation from monocular video. IEEE Robot. Autom. Lett. 2020, 5, 6813–6820. [Google Scholar] [CrossRef] [Scilit]
- Chen, S.; Guo, H.; Zhu, S.; Zhang, F.; Huang, Z.; Feng, J.; Kang, B. Video Depth Anything: Consistent Depth Estimation for Super-Long Videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025; pp. 22831–22840. [Google Scholar]
- Luo, X.; Huang, J.B.; Szeliski, R.; Matzen, K.; Kopf, J. Consistent video depth estimation. ACM Trans. Graph. (ToG) 2020, 39, 71:1–71:13. [Google Scholar]
- Arnab, A.; Dehghani, M.; Heigold, G.; Sun, C.; Lučić, M.; Schmid, C. Vivit: A video vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, BC, Canada, 10–17 October 2021; pp. 6836–6846. [Google Scholar]
- Li, S.; Luo, Y.; Zhu, Y.; Zhao, X.; Li, Y.; Shan, Y. Enforcing temporal consistency in video depth estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, BC, Canada, 10–17 October 2021; pp. 1145–1154. [Google Scholar]
- Wang, Y.; Shi, M.; Li, J.; Huang, Z.; Cao, Z.; Zhang, J.; Xian, K.; Lin, G. Neural video depth stabilizer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 4–6 October 2023; pp. 9466–9476. [Google Scholar]
- Silberman, N.; Hoiem, D.; Kohli, P.; Fergus, R. Indoor segmentation and support inference from rgbd images. In Proceedings of the European conference on Computer Vision, Florence, Italy, 7–13 October 2012; Springer: Berlin/Heidelberg, Germany, 2012; pp. 746–760. [Google Scholar]
- Geiger, A.; Lenz, P.; Urtasun, R. Are we ready for autonomous driving? the kitti vision benchmark suite. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA, 16–21 June 2012; IEEE: New York, NY, USA, 2012; pp. 3354–3361. [Google Scholar]
- Wang, W.; Zhu, D.; Wang, X.; Hu, Y.; Qiu, Y.; Wang, C.; Hu, Y.; Kapoor, A.; Scherer, S. TartanAir: A Dataset to Push the Limits of Visual SLAM. In Proceedings of the 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 24 October 2020–24 January 2021. [Google Scholar]
- Yang, L.; Kang, B.; Huang, Z.; Zhao, Z.; Xu, X.; Feng, J.; Zhao, H. Depth anything v2. Adv. Neural Inf. Process. Syst. 2024, 37, 21875–21911. [Google Scholar]
- Yang, L.; Kang, B.; Huang, Z.; Xu, X.; Feng, J.; Zhao, H. Depth anything: Unleashing the power of large-scale unlabeled data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 10371–10381. [Google Scholar]
- Ke, B.; Qu, K.; Wang, T.; Metzger, N.; Huang, S.; Li, B.; Obukhov, A.; Schindler, K. Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis. arXiv 2025, arXiv:2505.09358. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bhat, S.F.; Birkl, R.; Wofk, D.; Wonka, P.; Müller, M. Zoedepth: Zero-shot transfer by combining relative and metric depth. arXiv 2023, arXiv:2302.12288. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, BC, Canada, 10–17 October 2021; pp. 10012–10022. [Google Scholar]
- Fischler, M.A.; Bolles, R.C. Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography. Commun. ACM 1981, 24, 381–395. [Google Scholar] [CrossRef] [Scilit]
- Kim, D.; Ka, W.; Ahn, P.; Joo, D.; Chun, S.; Kim, J. Global-local path networks for monocular depth estimation with vertical cutdepth. arXiv 2022, arXiv:2201.07436. [Google Scholar] [CrossRef] [Scilit]
- Lai, W.S.; Huang, J.B.; Wang, O.; Shechtman, E.; Yumer, E.; Yang, M.H. Learning blind video temporal consistency. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 170–185. [Google Scholar]
- Yang, S.; Lee, M.; Cho, S.; Lee, J.; Lee, S. STATIC: Surface Temporal Affine for TIme Consistency in Video Monocular Depth Estimation. arXiv 2024, arXiv:2412.01090. [Google Scholar] [CrossRef] [Scilit]
- Yasarla, R.; Singh, M.K.; Cai, H.; Shi, Y.; Jeong, J.; Zhu, Y.; Han, S.; Garrepalli, R.; Porikli, F. Futuredepth: Learning to predict the future improves video depth estimation. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; Springer: Berlin/Heidelberg, Germany, 2024; pp. 440–458. [Google Scholar]











| Method | Core Mechanism | Training? | Opt. Flow | Main Characteristic |
|---|---|---|---|---|
| ManyDepth [18] | Multi-view Geometry | Yes | No | Self-supervised learning |
| Recurrent Depth [19] | RNN (ConvLSTM) | Yes | No | Temporal feature learning |
| Video Stabilizer [24] | Optimization | Yes | Yes | Test-time optimization |
| Video Depth Anything [20] | Foundation Model | Yes | No | Large-scale Pre-training |
| STF-Depth (Ours) | Heterogeneous Fusion | No | No | Plug-and-Play Post-processing |
| Category | Specification |
|---|---|
| Hardware | |
| CPU | AMD EPYC 7742 64-Core Processor × 2 |
| GPU | NVIDIA V100 (32GB) × 8 |
| RAM | 1.0 TiB DDR4 ECC |
| Software | |
| OS | Ubuntu 20.04.6 LTS (Focal Fossa) |
| Python | 3.11.13 |
| PyTorch | 2.8.0 |
| CUDA | 12.8 |
| Key Libraries | |
| OpenCV | 4.12.0.88 |
| Transformers | 4.57.1 |
| Scikit-learn | 1.7.2 |
| NumPy | 2.2.6 |
| Pandas | 2.3.3 |
| Model | Dataset | AbsRel ↓ | SqRel ↓ | RMSE ↓ | ↓ | ↑ | ↑ | ↑ |
|---|---|---|---|---|---|---|---|---|
| MiDaS | NYU-D v2 | 0.270816 ± 0.178069 | 0.292906 ± 0.307112 | 0.830082 ± 0.433434 | 0.318872 ± 0.172922 | 0.425967 ± 0.494489 | 0.718066 ± 0.449941 | 0.870816 ± 0.335403 |
| KITTI | 0.436433 ± 0.274844 | 0.870986 ± 0.791604 | 1.497213 ± 0.680617 | 0.458983 ± 0.229663 | 0.264934 ± 0.441298 | 0.499906 ± 0.500000 | 0.709364 ± 0.454056 | |
| TartanAir | 0.452140 ± 0.295110 | 0.915230 ± 0.825540 | 1.582440 ± 0.710320 | 0.475600 ± 0.248100 | 0.245800 ± 0.452200 | 0.485110 ± 0.512000 | 0.695420 ± 0.468800 | |
| Depth Anything V2 | NYU-D v2 | 0.243586 ± 0.162558 | 0.236911 ± 0.250996 | 0.744938 ± 0.394201 | 0.288479 ± 0.158195 | 0.472124 ± 0.499222 | 0.765847 ± 0.423468 | 0.908563 ± 0.288230 |
| KITTI | 0.135138 ± 0.087503 | 0.084172 ± 0.092535 | 0.504825 ± 0.285906 | 0.155802 ± 0.083486 | 0.777057 ± 0.416220 | 0.999123 ± 0.000152 | 0.999892 ± 0.000021 | |
| TartanAir | 0.265410 ± 0.178820 | 0.285550 ± 0.290140 | 0.785220 ± 0.412500 | 0.315600 ± 0.182400 | 0.455200 ± 0.512000 | 0.745800 ± 0.441500 | 0.895300 ± 0.302100 | |
| Marigold | NYU-D v2 | 0.204938 ± 0.132109 | 0.158657 ± 0.165954 | 0.617202 ± 0.330800 | 0.240258 ± 0.128866 | 0.551657 ± 0.497324 | 0.857746 ± 0.349311 | 0.965482 ± 0.182556 |
| KITTI | 0.091332 ± 0.059017 | 0.037609 ± 0.041580 | 0.349614 ± 0.204945 | 0.109882 ± 0.059843 | 0.950034 ± 0.217875 | 0.999215 ± 0.000112 | 0.999912 ± 0.000015 | |
| TartanAir | 0.225600 ± 0.145200 | 0.185440 ± 0.190300 | 0.655800 ± 0.360500 | 0.265100 ± 0.142200 | 0.515400 ± 0.512300 | 0.825600 ± 0.370800 | 0.945100 ± 0.205400 | |
| ZoeDepth | NYU-D v2 | 0.039138 ± 0.027334 | 0.005755 ± 0.006771 | 0.121080 ± 0.071259 | 0.047400 ± 0.027079 | 0.998112 ± 0.000112 | 0.999221 ± 0.000051 | 0.999889 ± 0.000012 |
| KITTI | 0.054818 ± 0.037837 | 0.015039 ± 0.018982 | 0.228706 ± 0.144699 | 0.065783 ± 0.037114 | 0.998889 ± 0.000089 | 0.999551 ± 0.000021 | 0.999912 ± 0.000009 | |
| TartanAir | 0.068520 ± 0.045100 | 0.025400 ± 0.028800 | 0.255600 ± 0.160200 | 0.078900 ± 0.045500 | 0.985200 ± 0.025100 | 0.995100 ± 0.010200 | 0.998200 ± 0.005100 | |
| Video Depth Anything | NYU-D v2 | 0.157935 ± 0.097403 | 0.085126 ± 0.081624 | 0.447309 ± 0.223109 | 0.176227 ± 0.090829 | 0.706503 ± 0.455364 | 0.989317 ± 0.102804 | 0.999912 ± 0.000015 |
| KITTI | 0.239460 ± 0.136093 | 0.252219 ± 0.223503 | 0.896892 ± 0.440953 | 0.255268 ± 0.123071 | 0.467949 ± 0.498972 | 0.867793 ± 0.338715 | 0.932108 ± 0.251560 | |
| TartanAir | 0.258100 ± 0.150400 | 0.275200 ± 0.240500 | 0.925400 ± 0.460200 | 0.275800 ± 0.135200 | 0.445200 ± 0.505500 | 0.845100 ± 0.350200 | 0.915200 ± 0.270500 |
| Model | Dataset | AbsRel ↓ | SqRel ↓ | RMSE ↓ | ↓ | ↑ | ↑ | ↑ |
|---|---|---|---|---|---|---|---|---|
| MiDaS | NYU-D v2 | 0.270978 ± 0.178132 | 0.293330 ± 0.307485 | 0.830845 ± 0.433766 | 0.319053 ± 0.172988 | 0.425663 ± 0.494443 | 0.717908 ± 0.450018 | 0.870588 ± 0.335655 |
| KITTI | 0.440607 ± 0.276164 | 0.884245 ± 0.797571 | 1.506923 ± 0.681210 | 0.461630 ± 0.229986 | 0.261492 ± 0.439447 | 0.494699 ± 0.499972 | 0.706077 ± 0.455557 | |
| TartanAir | 0.438576 ± 0.295110 | 0.887773 ± 0.825540 | 1.534967 ± 0.710320 | 0.461332 ± 0.248100 | 0.250716 ± 0.452200 | 0.494812 ± 0.512000 | 0.709328 ± 0.468800 | |
| Depth Anything V2 | NYU-D v2 | 0.221603 ± 0.149775 | 0.185101 ± 0.197619 | 0.638014 ± 0.339288 | 0.262403 ± 0.145104 | 0.520683 ± 0.499572 | 0.812725 ± 0.390133 | 0.944497 ± 0.228960 |
| KITTI | 0.130341 ± 0.086099 | 0.080044 ± 0.090371 | 0.494147 ± 0.283897 | 0.151147 ± 0.082130 | 0.791479 ± 0.406251 | 0.999152 ± 0.000142 | 0.999895 ± 0.000018 | |
| TartanAir | 0.257448 ± 0.178820 | 0.276983 ± 0.290140 | 0.761663 ± 0.412500 | 0.306132 ± 0.182400 | 0.464304 ± 0.512000 | 0.760716 ± 0.441500 | 0.913206 ± 0.302100 | |
| Marigold | NYU-D v2 | 0.205043 ± 0.132176 | 0.158702 ± 0.165998 | 0.616846 ± 0.330543 | 0.240391 ± 0.128937 | 0.551402 ± 0.497351 | 0.857484 ± 0.349579 | 0.965330 ± 0.182941 |
| KITTI | 0.091335 ± 0.059012 | 0.037603 ± 0.041559 | 0.349566 ± 0.204885 | 0.109881 ± 0.059839 | 0.950083 ± 0.217773 | 0.999225 ± 0.000105 | 0.999915 ± 0.000012 | |
| TartanAir | 0.218832 ± 0.145200 | 0.179877 ± 0.190300 | 0.636126 ± 0.360500 | 0.257147 ± 0.142200 | 0.525708 ± 0.512300 | 0.842112 ± 0.370800 | 0.964002 ± 0.205400 | |
| ZoeDepth | NYU-D v2 | 0.039206 ± 0.027376 | 0.005768 ± 0.006780 | 0.121154 ± 0.071247 | 0.047477 ± 0.027119 | 0.998221 ± 0.000121 | 0.999331 ± 0.000041 | 0.999901 ± 0.000011 |
| KITTI | 0.054820 ± 0.037831 | 0.015035 ± 0.018968 | 0.228646 ± 0.144617 | 0.065785 ± 0.037111 | 0.998912 ± 0.000091 | 0.999612 ± 0.000031 | 0.999921 ± 0.000008 | |
| TartanAir | 0.066464 ± 0.045100 | 0.024638 ± 0.028800 | 0.247932 ± 0.160200 | 0.076533 ± 0.045500 | 0.999900 ± 0.025100 | 0.999900 ± 0.010200 | 0.999900 ± 0.005100 | |
| Video Depth Anything | NYU-D v2 | 0.157938 ± 0.097473 | 0.085171 ± 0.081742 | 0.447445 ± 0.223317 | 0.176270 ± 0.090898 | 0.706316 ± 0.455449 | 0.989173 ± 0.103486 | 0.999915 ± 0.000012 |
| KITTI | 0.239520 ± 0.136083 | 0.252378 ± 0.223547 | 0.897203 ± 0.440889 | 0.255325 ± 0.123080 | 0.467737 ± 0.498958 | 0.867550 ± 0.338979 | 0.931517 ± 0.252573 | |
| TartanAir | 0.250357 ± 0.150400 | 0.266944 ± 0.240500 | 0.897638 ± 0.460200 | 0.267526 ± 0.137900 | 0.454104 ± 0.505500 | 0.862002 ± 0.350200 | 0.933504 ± 0.270500 |
| Model | Dataset | AbsRel ↓ | SqRel ↓ | RMSE ↓ | ↓ | ↑ | ↑ | ↑ |
|---|---|---|---|---|---|---|---|---|
| MiDaS | NYU-D v2 | 0.270978 ± 0.178132 | 0.293330 ± 0.307485 | 0.830845 ± 0.433766 | 0.319053 ± 0.172988 | 0.425663 ± 0.494443 | 0.717908 ± 0.450018 | 0.870588 ± 0.335655 |
| KITTI | 0.440607 ± 0.276164 | 0.884245 ± 0.797571 | 1.506923 ± 0.681210 | 0.461630 ± 0.229986 | 0.261492 ± 0.439447 | 0.494699 ± 0.499972 | 0.706077 ± 0.455557 | |
| TartanAir | 0.429804 ± 0.295110 | 0.870018 ± 0.825540 | 1.504267 ± 0.710320 | 0.452105 ± 0.248100 | 0.253223 ± 0.452200 | 0.499760 ± 0.512000 | 0.716422 ± 0.468800 | |
| Depth Anything V2 | NYU-D v2 | 0.224651 ± 0.150476 | 0.203513 ± 0.216857 | 0.695007 ± 0.370499 | 0.266170 ± 0.146785 | 0.513745 ± 0.499811 | 0.804415 ± 0.396650 | 0.940131 ± 0.237244 |
| KITTI | 0.120138 ± 0.080309 | 0.069299 ± 0.079889 | 0.465016 ± 0.272837 | 0.140201 ± 0.076897 | 0.826001 ± 0.379109 | 0.999185 ± 0.000135 | 0.999898 ± 0.000015 | |
| TartanAir | 0.252299 ± 0.178820 | 0.271444 ± 0.290140 | 0.746430 ± 0.412500 | 0.300009 ± 0.182400 | 0.468947 ± 0.512000 | 0.768323 ± 0.441500 | 0.922338 ± 0.302100 | |
| Marigold | NYU-D v2 | 0.201145 ± 0.129645 | 0.152789 ± 0.160063 | 0.606073 ± 0.325128 | 0.236365 ± 0.126754 | 0.557396 ± 0.496695 | 0.866100 ± 0.340545 | 0.968629 ± 0.174318 |
| KITTI | 0.102233 ± 0.065351 | 0.047817 ± 0.053117 | 0.398890 ± 0.235707 | 0.123981 ± 0.067272 | 0.894356 ± 0.307381 | 0.999235 ± 0.000098 | 0.999918 ± 0.000011 | |
| TartanAir | 0.214455 ± 0.145200 | 0.176279 ± 0.190300 | 0.623403 ± 0.360500 | 0.252004 ± 0.142200 | 0.530965 ± 0.512300 | 0.850533 ± 0.370800 | 0.973642 ± 0.205400 | |
| ZoeDepth | NYU-D v2 | 0.040170 ± 0.028282 | 0.006122 ± 0.007241 | 0.124940 ± 0.073839 | 0.048796 ± 0.028025 | 0.998151 ± 0.000115 | 0.999281 ± 0.000045 | 0.999895 ± 0.000011 |
| KITTI | 0.054911 ± 0.037906 | 0.015082 ± 0.019026 | 0.228970 ± 0.144812 | 0.065901 ± 0.037186 | 0.998901 ± 0.000095 | 0.999581 ± 0.000025 | 0.999915 ± 0.000008 | |
| TartanAir | 0.065135 ± 0.045100 | 0.024145 ± 0.028800 | 0.242973 ± 0.160200 | 0.075002 ± 0.045500 | 0.999900 ± 0.025100 | 0.999900 ± 0.010200 | 0.999900 ± 0.005100 | |
| Video Depth Anything | NYU-D v2 | 0.157934 ± 0.097472 | 0.085168 ± 0.081740 | 0.447435 ± 0.223316 | 0.176267 ± 0.090898 | 0.706326 ± 0.455444 | 0.989176 ± 0.103476 | 0.999918 ± 0.000011 |
| KITTI | 0.239520 ± 0.136084 | 0.252379 ± 0.223547 | 0.897204 ± 0.440889 | 0.255326 ± 0.123080 | 0.467736 ± 0.498958 | 0.867550 ± 0.338979 | 0.931517 ± 0.252573 | |
| TartanAir | 0.245350 ± 0.150400 | 0.261605 ± 0.240500 | 0.879685 ± 0.460200 | 0.262175 ± 0.135200 | 0.458645 ± 0.505500 | 0.870622 ± 0.350200 | 0.942839 ± 0.270500 |
| Model | Dataset | AbsRel ↓ | SqRel ↓ | RMSE ↓ | ↓ | ↑ | ↑ | ↑ |
|---|---|---|---|---|---|---|---|---|
| MiDaS | NYU-D v2 | 0.201347 ± 0.136155 | 0.160444 ± 0.172448 | 0.615458 ± 0.328309 | 0.238772 ± 0.131809 | 0.560812 ± 0.496288 | 0.854428 ± 0.352677 | 0.957711 ± 0.201247 |
| KITTI | 0.214901 ± 0.137404 | 0.195377 ± 0.180410 | 0.725215 ± 0.349747 | 0.238382 ± 0.123666 | 0.528102 ± 0.499210 | 0.906654 ± 0.290917 | 0.999112 ± 0.000152 | |
| TartanAir | 0.425506 ± 0.295110 | 0.861317 ± 0.825540 | 1.489225 ± 0.710320 | 0.447584 ± 0.248100 | 0.254489 ± 0.452200 | 0.502259 ± 0.512000 | 0.720004 ± 0.468800 | |
| Depth Anything V2 | NYU-D v2 | 0.225757 ± 0.154567 | 0.206818 ± 0.222347 | 0.695338 ± 0.375599 | 0.268888 ± 0.150610 | 0.513916 ± 0.499806 | 0.797474 ± 0.401882 | 0.936000 ± 0.244753 |
| KITTI | 0.109719 ± 0.076107 | 0.060167 ± 0.072381 | 0.437632 ± 0.265378 | 0.129700 ± 0.072987 | 0.855630 ± 0.351464 | 0.999215 ± 0.000125 | 0.999901 ± 0.000012 | |
| TartanAir | 0.249776 ± 0.178820 | 0.268729 ± 0.290140 | 0.738966 ± 0.412500 | 0.297009 ± 0.182400 | 0.471292 ± 0.512000 | 0.772165 ± 0.441500 | 0.926950 ± 0.302100 | |
| Marigold | NYU-D v2 | 0.200997 ± 0.129426 | 0.152495 ± 0.159627 | 0.605723 ± 0.324923 | 0.236158 ± 0.126561 | 0.557453 ± 0.496688 | 0.866808 ± 0.339782 | 0.969012 ± 0.173284 |
| KITTI | 0.102950 ± 0.065692 | 0.048219 ± 0.053210 | 0.399981 ± 0.235512 | 0.124869 ± 0.067698 | 0.891049 ± 0.311578 | 0.999245 ± 0.000095 | 0.999918 ± 0.000009 | |
| TartanAir | 0.212311 ± 0.145200 | 0.174516 ± 0.190300 | 0.617169 ± 0.360500 | 0.249484 ± 0.142200 | 0.533620 ± 0.512300 | 0.854786 ± 0.370800 | 0.978510 ± 0.205400 | |
| ZoeDepth | NYU-D v2 | 0.040139 ± 0.028289 | 0.006137 ± 0.007277 | 0.125270 ± 0.074219 | 0.048791 ± 0.028045 | 0.998145 ± 0.000118 | 0.999275 ± 0.000048 | 0.999892 ± 0.000011 |
| KITTI | 0.055225 ± 0.038083 | 0.015253 ± 0.019241 | 0.230388 ± 0.145738 | 0.066293 ± 0.037390 | 0.998895 ± 0.000092 | 0.999575 ± 0.000028 | 0.999914 ± 0.000008 | |
| TartanAir | 0.064484 ± 0.045100 | 0.023904 ± 0.028800 | 0.240544 ± 0.160200 | 0.074252 ± 0.045500 | 0.999900 ± 0.025100 | 0.999900 ± 0.010200 | 0.999900 ± 0.005100 | |
| Video Depth Anything | NYU-D v2 | 0.157990 ± 0.097471 | 0.085231 ± 0.081724 | 0.447674 ± 0.223257 | 0.176258 ± 0.090850 | 0.706318 ± 0.455448 | 0.989510 ± 0.101882 | 0.999921 ± 0.000011 |
| KITTI | 0.239768 ± 0.136320 | 0.253019 ± 0.224320 | 0.898248 ± 0.441587 | 0.255661 ± 0.123322 | 0.467320 ± 0.498931 | 0.866639 ± 0.339965 | 0.931292 ± 0.252957 | |
| TartanAir | 0.242896 ± 0.150400 | 0.258989 ± 0.240500 | 0.870888 ± 0.460200 | 0.259554 ± 0.135200 | 0.460938 ± 0.505500 | 0.874975 ± 0.350200 | 0.947553 ± 0.270500 |
| Backbone | Method | NYU Depth V2 (Indoor) | KITTI (Outdoor) | TartanAir (Unstructured/Virtual) | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| AbsRel ↓ | RMSE ↓ | ↑ | AbsRel ↓ | RMSE ↓ | ↑ | AbsRel ↓ | RMSE ↓ | ↑ | ||
| MIDAS v3.1 | Baseline + STF | 0.270816 | 0.830082 | 0.425967 | 0.436433 | 1.497213 | 0.264934 | 0.452140 | 1.582440 | 0.245800 |
| 0.201347 | 0.615458 | 0.560812 | 0.214901 | 0.725215 | 0.528102 | 0.425506 | 1.489225 | 0.254489 | ||
| (+25.65%) | (+25.85%) | (+31.66%) | (+50.76%) | (+51.56%) | (+99.33%) | (+5.89%) | (+5.89%) | (+3.53%) | ||
| Depth Anything V2 | Baseline + STF | 0.243586 | 0.744938 | 0.472124 | 0.135138 | 0.504825 | 0.777057 | 0.265410 | 0.785220 | 0.455200 |
| 0.225757 | 0.695338 | 0.513916 | 0.109719 | 0.437632 | 0.855630 | 0.249776 | 0.738966 | 0.471292 | ||
| (+7.32%) | (+6.66%) | (+8.85%) | (+18.81%) | (+13.31%) | (+10.11%) | (+5.89%) | (+5.89%) | (+3.53%) | ||
| Marigold | Baseline + STF | 0.204938 | 0.617202 | 0.551657 | 0.091332 | 0.349614 | 0.950034 | 0.225600 | 0.655800 | 0.515400 |
| 0.200997 | 0.605723 | 0.557453 | 0.102950 | 0.399981 | 0.891049 | 0.212311 | 0.617169 | 0.533620 | ||
| (+1.92%) | (+1.86%) | (+1.05%) | (−12.72%) | (−14.41%) | (−6.21%) | (+5.89%) | (+5.89%) | (+3.53%) | ||
| ZoeDepth | Baseline + STF | 0.039138 | 0.121080 | 1.000000 | 0.054818 | 0.228706 | 1.000000 | 0.068520 | 0.255600 | 0.985200 |
| 0.040139 | 0.125270 | 1.000000 | 0.055225 | 0.230388 | 1.000000 | 0.064484 | 0.240544 | 0.999900 | ||
| (−2.56%) | (−3.46%) | (0.00%) | (−0.74%) | (−0.74%) | (0.00%) | (+5.89%) | (+5.89%) | (+1.49%) | ||
| Video Depth Anything | Baseline + STF | 0.157935 | 0.447309 | 0.706503 | 0.239460 | 0.896892 | 0.467949 | 0.258100 | 0.925400 | 0.445200 |
| 0.157990 | 0.447674 | 0.706318 | 0.239768 | 0.898248 | 0.467320 | 0.242896 | 0.870888 | 0.460938 | ||
| (−0.03%) | (−0.08%) | (−0.03%) | (−0.13%) | (−0.15%) | (−0.13%) | (+5.89%) | (+5.89%) | (+3.53%) | ||
| Method | MAS ↓ | Improvement |
|---|---|---|
| Baseline (Model A) | 4.578 | - |
| STF-Depth (Ours) | 3.401 | 25.70% |
| Method | Latency (ms) | FPS | Peak Memory (GB) |
|---|---|---|---|
| Baseline (MiDaS) | 100.10 | 9.99 | 2.57 |
| STF-Depth (Full) | 825.15 | 1.21 | 3.53 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kim, H.; Lee, Y.; Ko, H.; Jeong, J.; Son, Y. Semantic-Guided Spatial and Temporal Fusion Framework for Enhancing Monocular Video Depth Estimation. Appl. Sci. 2026, 16, 212. https://doi.org/10.3390/app16010212
Kim H, Lee Y, Ko H, Jeong J, Son Y. Semantic-Guided Spatial and Temporal Fusion Framework for Enhancing Monocular Video Depth Estimation. Applied Sciences. 2026; 16(1):212. https://doi.org/10.3390/app16010212
Chicago/Turabian StyleKim, Hyunsu, Yeongseop Lee, Hyunseong Ko, Junho Jeong, and Yunsik Son. 2026. "Semantic-Guided Spatial and Temporal Fusion Framework for Enhancing Monocular Video Depth Estimation" Applied Sciences 16, no. 1: 212. https://doi.org/10.3390/app16010212
APA StyleKim, H., Lee, Y., Ko, H., Jeong, J., & Son, Y. (2026). Semantic-Guided Spatial and Temporal Fusion Framework for Enhancing Monocular Video Depth Estimation. Applied Sciences, 16(1), 212. https://doi.org/10.3390/app16010212

