Improved DPT-Hybrid for Monocular Depth Estimation with Geometry-Enhanced Encoding and Structure-Aware Gated Fusion
Abstract
1. Introduction
- We present a coordinated structure enhancement framework for DPT-Hybrid monocular depth estimation. The framework is motivated by the observation that local geometric cues may be weakened during shallow encoding, diluted during encoder–decoder fusion, and insufficiently constrained by value-oriented regression losses.
- We implement this framework through three complementary mechanisms: residual depth-wise separable refinement in shallow high-resolution encoder stages, gradient-assisted adaptive gating for encoder–decoder fusion, and joint supervision of depth values, gradient transitions, and boundaries. These mechanisms align structural preservation across feature representation, feature fusion, and optimization.
- We conduct controlled component-level ablations on NYUv2 to quantify the individual and complementary effects of the proposed design. The complete framework obtains lower error metrics than the controlled DPT-Hybrid baseline on both NYUv2 and KITTI, with a moderate increase in the number of parameters and computational cost.
2. Related Work
3. Method
3.1. Overall Network Architecture
3.2. Geometry-Enhanced DPT Encoder
3.3. Structure-Aware Cross-Scale Gated Attention Fusion (S-GAF)
3.4. Joint Structure–Geometric Consistency Loss
4. Experiments
4.1. Datasets
4.2. Implementation Details
4.3. Evaluation Metrics
4.4. Quantitative Comparison
4.5. Ablation Studies
4.5.1. Overall Component Ablation
4.5.2. Ablation on the S-GAF Module
4.5.3. Ablation on the Geometry-Enhanced Encoder
4.5.4. Ablation on the Joint Loss
4.5.5. Sensitivity to Representative Loss Weight Settings
4.6. Efficiency Analysis
5. Discussion
5.1. Mechanistic Analysis and Comparison with Published Methods
5.2. Limitations
5.3. Future Research Goals and Methodological Extensions
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| AbsRel | Absolute Relative Error |
| DPT | Dense Prediction Transformer |
| DWConv | Depth-Wise Separable Convolution |
| FLOPs | Floating-Point Operations |
| FPS | Frames Per Second |
| GE | Geometry-Enhanced |
| KITTI | Karlsruhe Institute of Technology and Toyota Technological Institute Dataset |
| NYUv2 | New York University Depth Dataset v2 |
| RMSE | Root Mean Square Error |
| S-GAF | Structure-Aware Cross-Scale Gated Attention Fusion |
| SI-Log | Scale-Invariant Logarithmic |
| SqRel | Squared Relative Error |
References
- Eigen, D.; Puhrsch, C.; Fergus, R. Depth Map Prediction from a Single Image using a Multi-Scale Deep Network. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS 2014), Montreal, QC, Canada, 8–13 December 2014; pp. 2366–2374. [Google Scholar] [CrossRef]
- Scharstein, D.; Szeliski, R. A Taxonomy and Evaluation of Dense Two-Frame Stereo Correspondence Algorithms. Int. J. Comput. Vis. 2002, 47, 7–42. [Google Scholar] [CrossRef]
- Ranftl, R.; Bochkovskiy, A.; Koltun, V. Vision Transformers for Dense Prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2021), Montreal, QC, Canada, 11–17 October 2021; pp. 12179–12188. [Google Scholar]
- Azuma, R.T. A Survey of Augmented Reality. Presence Teleoperators Virtual Environ. 1997, 6, 355–385. [Google Scholar] [CrossRef]
- Qi, C.R.; Liu, W.; Wu, C.; Su, H.; Guibas, L.J. Frustum PointNets for 3D Object Detection from RGB-D Data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018), Salt Lake City, UT, USA, 18–22 June 2018; pp. 918–927. [Google Scholar] [CrossRef]
- Mur-Artal, R.; Montiel, J.M.M.; Tardos, J.D. ORB-SLAM: A Versatile and Accurate Monocular SLAM System. IEEE Trans. Robot. 2015, 31, 1147–1163. [Google Scholar] [CrossRef]
- Saxena, A.; Sun, M.; Ng, A.Y. Make3D: Learning 3D Scene Structure from a Single Still Image. IEEE Trans. Pattern Anal. Mach. Intell. 2009, 31, 824–840. [Google Scholar] [CrossRef] [PubMed]
- Liu, F.; Shen, C.; Lin, G. Deep Convolutional Neural Fields for Depth Estimation from a Single Image. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2015), Boston, MA, USA, 7–12 June 2015; pp. 5162–5170. [Google Scholar] [CrossRef]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the International Conference on Learning Representations (ICLR 2021), Virtual Event, 3–7 May 2021. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2021), Montreal, QC, Canada, 11–17 October 2021; pp. 10012–10022. [Google Scholar] [CrossRef]
- Silberman, N.; Hoiem, D.; Kohli, P.; Fergus, R. Indoor Segmentation and Support Inference from RGBD Images. In Proceedings of the European Conference on Computer Vision (ECCV 2012), Florence, Italy, 7–13 October 2012; pp. 746–760. [Google Scholar] [CrossRef]
- Geiger, A.; Lenz, P.; Urtasun, R. Are We Ready for Autonomous Driving? The KITTI Vision Benchmark Suite. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2012), Providence, RI, USA, 16–21 June 2012; pp. 3354–3361. [Google Scholar] [CrossRef]
- Laina, I.; Rupprecht, C.; Belagiannis, V.; Tombari, F.; Navab, N. Deeper Depth Prediction with Fully Convolutional Residual Networks. In Proceedings of the International Conference on 3D Vision (3DV 2016), Stanford, CA, USA, 25–28 October 2016; pp. 239–248. [Google Scholar] [CrossRef]
- Fu, H.; Gong, M.; Wang, C.; Batmanghelich, K.; Tao, D. Deep Ordinal Regression Network for Monocular Depth Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018), Salt Lake City, UT, USA, 18–22 June 2018; pp. 2002–2011. [Google Scholar] [CrossRef] [PubMed]
- Lin, T.Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature Pyramid Networks for Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Honolulu, HI, USA, 21–26 July 2017; pp. 2117–2125. [Google Scholar] [CrossRef]
- Lin, G.; Milan, A.; Shen, C.; Reid, I. RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Honolulu, HI, USA, 21–26 July 2017; pp. 1925–1934. [Google Scholar]
- Yin, W.; Liu, Y.; Shen, C.; Yan, Y. Enforcing Geometric Constraints of Virtual Normal for Depth Prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2019), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 5684–5693. [Google Scholar] [CrossRef]
- Hu, J.; Ozay, M.; Zhang, Y.; Okatani, T. Revisiting Single Image Depth Estimation: Toward Higher Resolution Maps with Accurate Object Boundaries. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV 2019), Waikoloa Village, HI, USA, 7–11 January 2019; pp. 1043–1051. [Google Scholar] [CrossRef]
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision (ECCV 2018), Munich, Germany, 8–14 September 2018; pp. 3–19. [Google Scholar] [CrossRef]
- Wang, Q.; Wu, B.; Zhu, P.; Li, P.; Zuo, W.; Hu, Q. ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2020), Virtual Event, 14–19 June 2020; pp. 11534–11542. [Google Scholar] [CrossRef]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018), Salt Lake City, UT, USA, 18–22 June 2018; pp. 7132–7141. [Google Scholar] [CrossRef]
- Godard, C.; Mac Aodha, O.; Brostow, G.J. Unsupervised Monocular Depth Estimation with Left-Right Consistency. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Honolulu, HI, USA, 21–26 July 2017; pp. 270–279. [Google Scholar]
- Zhou, L.; Shen, C.; van den Hengel, A. Edge-Guided Depth Estimation Network. In Proceedings of the Asian Conference on Computer Vision (ACCV 2018), Perth, Australia, 2–6 December 2018; pp. 253–268. [Google Scholar]
- Chen, L.C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; Yuille, A.L. Attention to Scale: Scale-Aware Semantic Image Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), Las Vegas, NV, USA, 27–30 June 2016; pp. 3640–3649. [Google Scholar] [CrossRef]
- Bhat, S.F.; Alhashim, I.; Wonka, P. AdaBins: Depth Estimation using Adaptive Bins. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2021), Virtual Event, 19–25 June 2021; pp. 4009–4018. [Google Scholar]
- Chen, Y.; Yin, Q.; Zhao, L.; Wang, J.; Zhou, S.; Tang, J. Enhancing long-range depth estimation via heterogeneous CNN-transformer encoding and cross-dimensional semantic fusion. Sci. Rep. 2026, 16, 9396. [Google Scholar] [CrossRef] [PubMed]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention (MICCAI 2015), Munich, Germany, 5–9 October 2015; pp. 234–241. [Google Scholar] [CrossRef]
- Xian, K.; Zhang, J.; Wang, O.; Mai, L.; Xu, Z.; Cao, Z. Structure-Guided Ranking Loss for Single Image Depth Prediction. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 611–620. [Google Scholar]
- Yuan, W.; Gu, X.; Dai, Z.; Zhu, S.; Tan, P. NeW CRFs: Neural Window Fully-Connected CRFs for Monocular Depth Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022), New Orleans, LA, USA, 18–24 June 2022; pp. 3916–3925. [Google Scholar]
- El Ogri, O.; El-Mekkaoui, J.; Hjouji, A. A computer-assisted medical diagnosis system for cancer diseases based on quaternion orthogonal Rademacher-Fourier moments and deep learning. Biomed. Signal Process. Control 2026, 89, 108744. [Google Scholar] [CrossRef]
- Chan, K.H.; Im, S.K. Sentiment analysis by using Naïve-Bayes classifier with stacked CARU. Electron. Lett. 2022, 58, 411–413. [Google Scholar] [CrossRef]
- Chollet, F. Xception: Deep Learning with Depthwise Separable Convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Honolulu, HI, USA, 21–26 July 2017; pp. 1251–1258. [Google Scholar]
- Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.C. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018), Salt Lake City, UT, USA, 18–22 June 2018; pp. 4510–4520. [Google Scholar] [CrossRef]
- Wang, P.; Shen, X.; Lin, Z.; Cohen, S.; Price, B.; Yuille, A.L. Towards Unified Depth and Semantic Prediction from a Single Image. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2015), Boston, MA, USA, 7–12 June 2015; pp. 2800–2809. [Google Scholar] [CrossRef]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS 2019), Vancouver, BC, Canada, 8–14 December 2019; pp. 8024–8035. [Google Scholar]
- Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. In Proceedings of the International Conference on Learning Representations (ICLR 2019), New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Loshchilov, I.; Hutter, F. SGDR: Stochastic Gradient Descent with Warm Restarts. In Proceedings of the International Conference on Learning Representations (ICLR 2017), Toulon, France, 24–26 April 2017. [Google Scholar]
- Kim, D.; Sohn, C. Global-Local Path Networks for Monocular Depth Estimation with Vertical CutDepth. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022), New Orleans, LA, USA, 18–24 June 2022; pp. 2858–2868. [Google Scholar]
- Bochkovskii, A.; Delaunoy, A.; Germain, H.; Santos, M.; Zhou, Y.; Richter, S.R.; Koltun, V. Depth Pro: Sharp Monocular Metric Depth In Less Than A Second. In Proceedings of the International Conference on Learning Representations (ICLR), Singapore, 24–28 April 2025. [Google Scholar]











| Method Category | Representative Works | Core Mechanism | Primary Limitations Drawbacks | Our Proposed |
|---|---|---|---|---|
| Multi-Scale CNNs | Eigen et al. [1], Laina et al. [13], DORN [14] | Multi-scale convolutional feature extraction and ordinal regression. | Limited receptive field; struggles to capture long-range global scene semantics. | Adopt DPT-Hybrid backbone to capture robust long-range global contextual dependencies. |
| Dense Transformers | DPT [3], Swin [10], NeW CRFs [29] | Self-attention mechanisms for dense prediction and windowed CRFs. | Shallow stages lose high-frequency local geometric details due to patchification. | Geometry-enhanced (GE) encoder: Inject residual DWConv branches into shallow Stages 1 and 2. |
| Attention Fusion | FPN [15], CBAM [19], ECA [20], AdaBins [25] | Top-down feature pyramids and intra-feature channel–spatial recalibration. | Lacks explicit cross-scale gating between encoder details and decoder semantics. | S-GAF module: Gradient-assisted cross-scale channel-spatial soft-selection gating. |
| Geometric Losses | Virtual Normal [17], Ranking Loss [28] | Surface-normal constraints and pair-wise ordinal ranking supervision. | High sampling complexity; sensitive to annotation noise and normal estimation errors. | Joint consistency loss: Direct spatial gradient consistency and Sobel edge-focused supervision. |
| Method | AbsRel ↓ | SqRel ↓ | RMSE ↓ | RMSE-log ↓ | ↑ | ↑ | ↑ |
|---|---|---|---|---|---|---|---|
| AdaBins [25] | 0.105 | 0.070 | 0.350 | 0.142 | 0.908 | 0.989 | 0.998 |
| NeW CRFs [29] | 0.103 | 0.068 | 0.347 | 0.140 | 0.910 | 0.990 | 0.998 |
| GLPN [38] | 0.101 | 0.066 | 0.343 | 0.137 | 0.913 | 0.991 | 0.999 |
| Depth Pro [39] | 0.099 | 0.065 | 0.335 | 0.133 | 0.919 | 0.992 | 0.999 |
| DPT-Hybrid | 0.107 | 0.072 | 0.357 | 0.145 | 0.904 | 0.988 | 0.998 |
| DPT + CBAM | 0.102 | 0.069 | 0.345 | 0.138 | 0.912 | 0.990 | 0.998 |
| DPT + ECA | 0.100 | 0.067 | 0.340 | 0.135 | 0.915 | 0.991 | 0.999 |
| Ours | 0.099 | 0.066 | 0.334 | 0.133 | 0.918 | 0.991 | 0.999 |
| Method | AbsRel ↓ | SqRel ↓ | RMSE ↓ | RMSE-log ↓ | ↑ | ↑ | ↑ |
|---|---|---|---|---|---|---|---|
| AdaBins [25] | 0.061 | 0.360 | 2.550 | 0.091 | 0.960 | 0.995 | 0.999 |
| NeW CRFs [29] | 0.060 | 0.355 | 2.520 | 0.090 | 0.961 | 0.995 | 0.999 |
| GLPN [38] | 0.059 | 0.350 | 2.495 | 0.089 | 0.962 | 0.996 | 0.999 |
| Depth Pro [39] | 0.057 | 0.345 | 2.440 | 0.088 | 0.965 | 0.996 | 0.999 |
| DPT-Hybrid | 0.062 | 0.365 | 2.573 | 0.092 | 0.959 | 0.995 | 0.998 |
| DPT-Large | 0.058 | 0.352 | 2.487 | 0.089 | 0.963 | 0.996 | 0.999 |
| Ours | 0.058 | 0.348 | 2.455 | 0.088 | 0.964 | 0.996 | 0.999 |
| Model | Components | Metrics | ||||
|---|---|---|---|---|---|---|
| GE | S-GAF | JL | AbsRel ↓ | RMSE ↓ | ↑ | |
| DPT-Hybrid | × | × | × | 0.107 | 0.357 | 0.904 |
| DPT-Hybrid + S-GAF | × | ✓ | × | 0.104 | 0.351 | 0.907 |
| DPT-Hybrid + JL | × | × | ✓ | 0.105 | 0.353 | 0.908 |
| DPT-Hybrid + GE | ✓ | × | × | 0.103 | 0.348 | 0.910 |
| DPT-Hybrid + GE + S-GAF | ✓ | ✓ | × | 0.100 | 0.340 | 0.916 |
| Ours | ✓ | ✓ | ✓ | 0.099 | 0.334 | 0.918 |
| Variant/ Fusion Strategy | Channel Gate | Spatial Gate | Cross-Scale Gating | Auxiliary RGB-Gradient Input | AbsRel ↓ | RMSE ↓ |
|---|---|---|---|---|---|---|
| Simple Concatenation | × | × | × (Direct Concat) | × | 0.107 | 0.357 |
| CBAM-Style Recalibration | ✓ | ✓ | × (Self-Attention on Concat) | × | 0.106 | 0.354 |
| S-GAF (Channel Only) | ✓ | × | ✓ | × | 0.106 | 0.355 |
| S-GAF (Spatial Only) | × | ✓ | ✓ | × | 0.106 | 0.355 |
| S-GAF (Channel + Spatial) | ✓ | ✓ | ✓ | × | 0.105 | 0.353 |
| Full S-GAF (Ours) | ✓ | ✓ | ✓ | ✓ | 0.104 | 0.351 |
| Variant | GE Stage 1 | GE Stage 2 | Conv. Type | AbsRel ↓ | RMSE ↓ |
|---|---|---|---|---|---|
| No GE | × | × | - | 0.104 | 0.351 |
| GE only Stage 1 | ✓ | × | DW Conv | 0.102 | 0.344 |
| GE only Stage 2 | × | ✓ | DW Conv | 0.103 | 0.345 |
| GE on Stage 1 + Stage 2 | ✓ | ✓ | DW Conv | 0.100 | 0.340 |
| GE (standard conv) | ✓ | ✓ | Standard Conv | 0.102 | 0.343 |
| Variant | SI-Log | Grad Loss | Edge Loss | AbsRel ↓ | RMSE ↓ |
|---|---|---|---|---|---|
| SI-Log only | ✓ | × | × | 0.100 | 0.340 |
| SI-Log + grad | ✓ | ✓ | × | 0.099 | 0.335 |
| SI-Log + edge | ✓ | × | ✓ | 0.099 | 0.336 |
| SI-Log + grad + edge (Ours) | ✓ | ✓ | ✓ | 0.099 | 0.334 |
| Configuration | (Grad) | (Edge) | AbsRel ↓ | RMSE ↓ | RMSE-log ↓ | ↑ |
|---|---|---|---|---|---|---|
| Config A (Under-weighted) | 0.2 | 0.2 | 0.101 | 0.341 | 0.136 | 0.914 |
| Config B (Ours—Optimal) | 0.5 | 0.5 | 0.099 | 0.334 | 0.133 | 0.918 |
| Config C (Gradient-heavy) | 0.8 | 0.2 | 0.100 | 0.336 | 0.134 | 0.916 |
| Config D (Edge-heavy) | 0.2 | 0.8 | 0.100 | 0.337 | 0.135 | 0.916 |
| Config E (Over-weighted) | 1.0 | 1.0 | 0.102 | 0.343 | 0.137 | 0.913 |
| Model | Params (M) | FLOPs (G) | FPS ↑ | Latency (ms/Frame) ↓ | NYUv2 AbsRel ↓ | KITTI AbsRel ↓ |
|---|---|---|---|---|---|---|
| DPT-Hybrid | 123.15 | 109.96 | 13.04 | 76.68 | 0.107 | 0.062 |
| Full model (Ours) | 129.75 | 115.56 | 12.11 | 82.57 | 0.099 | 0.058 |
| Component/Variant | Params (M) | Added Params | FLOPs (G) | Added FLOPs | Latency (ms) | Added Latency |
|---|---|---|---|---|---|---|
| DPT-Hybrid (Baseline) | 123.15 | – | 109.96 | – | 76.68 | – |
| +GE Encoder Only | 123.55 | +0.40 M (+0.32%) | 111.12 | +1.16 G (+1.05%) | 77.82 | +1.14 ms |
| +S-GAF Module Only | 129.35 | +6.20 M (+5.03%) | 114.40 | +4.44 G (+4.04%) | 81.45 | +4.77 ms |
| Full Model (GE + S-GAF) | 129.75 | +6.60 M (+5.36%) | 115.56 | +5.60 G (+5.09%) | 82.57 | +5.89 ms |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Liu, W.; Hu, S.; Qin, Y.; Hong, S.; Zhang, D. Improved DPT-Hybrid for Monocular Depth Estimation with Geometry-Enhanced Encoding and Structure-Aware Gated Fusion. Electronics 2026, 15, 3465. https://doi.org/10.3390/electronics15153465
Liu W, Hu S, Qin Y, Hong S, Zhang D. Improved DPT-Hybrid for Monocular Depth Estimation with Geometry-Enhanced Encoding and Structure-Aware Gated Fusion. Electronics. 2026; 15(15):3465. https://doi.org/10.3390/electronics15153465
Chicago/Turabian StyleLiu, Wei, Shilei Hu, Yi Qin, Shengkai Hong, and Dehua Zhang. 2026. "Improved DPT-Hybrid for Monocular Depth Estimation with Geometry-Enhanced Encoding and Structure-Aware Gated Fusion" Electronics 15, no. 15: 3465. https://doi.org/10.3390/electronics15153465
APA StyleLiu, W., Hu, S., Qin, Y., Hong, S., & Zhang, D. (2026). Improved DPT-Hybrid for Monocular Depth Estimation with Geometry-Enhanced Encoding and Structure-Aware Gated Fusion. Electronics, 15(15), 3465. https://doi.org/10.3390/electronics15153465

