Next Article in Journal
Innovative Multi-View Strategies for AI-Assisted Breast Cancer Detection in Mammography
Previous Article in Journal
Three-Dimensional Ultraviolet Fluorescence Imaging in Cultural Heritage: A Review of Applications in Multi-Material Artworks
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

DP-AMF: Depth-Prior–Guided Adaptive Multi-Modal and Global–Local Fusion for Single-View 3D Reconstruction

1
Doctoral Program in Empowerment Informatics, University of Tsukuba, 1-1-1 Tennodai, Tsukuba 3058577, Japan
2
Center for Computational Science, University of Tsukuba, 1-1-1 Tennodai, Tsukuba 3058577, Japan
*
Authors to whom correspondence should be addressed.
J. Imaging 2025, 11(7), 246; https://doi.org/10.3390/jimaging11070246
Submission received: 15 June 2025 / Revised: 13 July 2025 / Accepted: 16 July 2025 / Published: 21 July 2025
(This article belongs to the Section AI in Imaging)

Abstract

Single-view 3D reconstruction remains fundamentally ill-posed, as a single RGB image lacks scale and depth cues, often yielding ambiguous results under occlusion or in texture-poor regions. We propose DP-AMF, a novel Depth-Prior–Guided Adaptive Multi-Modal and Global–Local Fusion framework that integrates high-fidelity depth priors—generated offline by the MARIGOLD diffusion-based estimator and cached to avoid extra training cost—with hierarchical local features from ResNet-32/ResNet-18 and semantic global features from DINO-ViT. A learnable fusion module dynamically adjusts per-channel weights to balance these modalities according to local texture and occlusion, and an implicit signed-distance field decoder reconstructs the final mesh. Extensive experiments on 3D-FRONT and Pix3D demonstrate that DP-AMF reduces Chamfer Distance by 7.64%, increases F-Score by 2.81%, and boosts Normal Consistency by 5.88% compared to strong baselines, while qualitative results show sharper edges and more complete geometry in challenging scenes. DP-AMF achieves these gains without substantially increasing model size or inference time, offering a robust and effective solution for complex single-view reconstruction tasks.
Keywords: single-view reconstruction; multi-model; 3D vision; visual feature fusion; indoor scene single-view reconstruction; multi-model; 3D vision; visual feature fusion; indoor scene

Share and Cite

MDPI and ACS Style

Zhang, L.; Xie, C.; Kitahara, I. DP-AMF: Depth-Prior–Guided Adaptive Multi-Modal and Global–Local Fusion for Single-View 3D Reconstruction. J. Imaging 2025, 11, 246. https://doi.org/10.3390/jimaging11070246

AMA Style

Zhang L, Xie C, Kitahara I. DP-AMF: Depth-Prior–Guided Adaptive Multi-Modal and Global–Local Fusion for Single-View 3D Reconstruction. Journal of Imaging. 2025; 11(7):246. https://doi.org/10.3390/jimaging11070246

Chicago/Turabian Style

Zhang, Luoxi, Chun Xie, and Itaru Kitahara. 2025. "DP-AMF: Depth-Prior–Guided Adaptive Multi-Modal and Global–Local Fusion for Single-View 3D Reconstruction" Journal of Imaging 11, no. 7: 246. https://doi.org/10.3390/jimaging11070246

APA Style

Zhang, L., Xie, C., & Kitahara, I. (2025). DP-AMF: Depth-Prior–Guided Adaptive Multi-Modal and Global–Local Fusion for Single-View 3D Reconstruction. Journal of Imaging, 11(7), 246. https://doi.org/10.3390/jimaging11070246

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop