remotesensing-logo

Journal Browser

Journal Browser

3D City Modeling and Observation Using Remote Sensing and Artificial Intelligence

A Special Issue of Remote Sensing (ISSN 2072-4292) belonging to the section "AI Remote Sensing".

Deadline for manuscript submissions: 31 October 2026 | Viewed by 11474

Editors


E-Mail Website
Guest Editor
School of Remote Sensing and Information Engineering, Wuhan University, Wuhan 430079, China
Interests: 3D reconstruction; UAV intelligent remote sensing; 3D measurement; 3D real scene modelling; VR-AR application; deep learning algorithm; stereo visual system; laser scanning

E-Mail Website
Guest Editor

Special Issue Information

Dear Colleagues,

3D City modeling stands as a fundamental pillar for the burgeoning field of digital twin applications in modern cities. The evolution of cutting-edge remote sensing equipment, platforms, and methodologies has facilitated the acquisition of three-dimensional spatial data, making it an increasingly accessible process. Leveraging urban maps and three-dimensional datasets, the observation and analysis of critical urban objects in both static and dynamic states have been integrated into a variety of sectors, including urban planning, management, safety monitoring, and the burgeoning domain of smart cities.

While remote sensing offers the tools for modeling and observation, artificial intelligence (AI) has elevated these processes to new heights. AI's broad spectrum of object recognition and execution of complex tasks is characterized by its efficiency and reliability. The proliferation of mobile remote sensing platforms, such as drones and laser scanners, has democratized low-cost, automated urban remote sensing monitoring.

The evolution of 3D modeling techniques from various optimized models of multi-view geometry reconstruction is now progressing towards methods based on neural networks, such as deep learning for stereo reconstruction and neural radiance fields.

The quest to integrate more sophisticated AI algorithms into the realms of modeling and observational data analysis has emerged as a pivotal area of interest. The ongoing advancements in SOTA methods have not only validated this trend but have also set the stage for a new era of innovation in the field of 3D modeling.

This Special Issue aims at studies covering remote sensing platforms, systems,new algorithms, and the latest experimental datasets. Relevant research can contribute to enhancing the quality, efficiency, and application level of 3D city modeling and observational data analysis. This Special Issue includes, but is not limited to, the following topics:

  • Intelligent mobile measurement platform
  • 3D building modelling
  • 3D modelling with multi-source space data
  • Object recognition using AI algorithm
  • City space plane & management
  • UAV smart task plan
  • Realtime data processing
  • Neural network 3D algorithm 
  • Laser scanning data processing
  • City objects datasets

Dr. Zheng Ji
Dr. Brian Alan Johnson
Guest Editors

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Remote Sensing is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2700 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • 3D city modlling
  • city observation
  • AI algorithm
  • mobile measurement platform
  • laser scanning data processing
  • object recognition and positioning
  • neural network for 3D reconstruction
  • multi-source space data processing

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (7 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

Jump to: Review

21 pages, 6813 KB  
Article
GSANet: Geometric Structure-Aware Siamese Network for 3D Change Detection
by Jiakang Chen, Rongfang Wang, Libin Sun and Changzhe Jiao
Remote Sens. 2026, 18(16), 2647; https://doi.org/10.3390/rs18162647 - 7 Aug 2026
Viewed by 319
Abstract
Three-dimensional (3D) point cloud change detection is essential for urban monitoring and environmental analysis, yet existing methods mainly rely on point-wise semantic differences and overlook change–unchange boundaries and object edges—key geometric cues for precise localization, especially for subtle or gradual changes. To address [...] Read more.
Three-dimensional (3D) point cloud change detection is essential for urban monitoring and environmental analysis, yet existing methods mainly rely on point-wise semantic differences and overlook change–unchange boundaries and object edges—key geometric cues for precise localization, especially for subtle or gradual changes. To address this, we propose GSANet, a Geometric Structure-Aware Siamese Network that explicitly integrates boundary and edge priors into sampling and feature learning. First, a Boundary-Aware Subsampling (BAS) strategy preserves key points near change boundaries while reducing redundancy, and a Boundary-Aware Binary Cross-Entropy (BA-BCE) loss assigns higher supervision weights to boundary points, enhancing learning in ambiguous regions. Second, an Edge-Aware Siamese Network captures robust local shapes by embedding edge priors into feature extraction, incorporating Edge-Aware Adaptive Graph Convolution, Edge-Aware Downsampling, and Cross-Attention Upsampling to maintain structural consistency across temporal branches. Additionally, a Difference Enhancement Module (DEM) amplifies feature discrepancies between bitemporal point clouds, improving sensitivity to subtle changes. Extensive experiments on a street-level dataset and urban dataset show our method outperforms state-of-the-art approaches. Full article
Show Figures

Figure 1

33 pages, 48783 KB  
Article
VRPF: A Fine-Grained 3D Radar Power-Density Computation Framework Based on Photogrammetric City Models for Urban Observation
by Linhui Jiao, Anran Yang, Qingren Jia, Mengyu Ma, Yifan Zhang, Linyue Wang and Jun Li
Remote Sens. 2026, 18(12), 1936; https://doi.org/10.3390/rs18121936 - 11 Jun 2026
Viewed by 415
Abstract
Radar is critical for urban security against Unmanned Aerial Vehicles (UAVs), yet signal occlusion caused by dense buildings and complex urban structures remains a major challenge for coverage assessment. Existing approaches commonly rely on 2D maps or 2.5D Digital Surface Models (DSMs), which [...] Read more.
Radar is critical for urban security against Unmanned Aerial Vehicles (UAVs), yet signal occlusion caused by dense buildings and complex urban structures remains a major challenge for coverage assessment. Existing approaches commonly rely on 2D maps or 2.5D Digital Surface Models (DSMs), which have difficulty representing vertical facades, vegetation, bridges, overhanging structures, and void spaces. These geometric limitations can introduce errors in radar occlusion determination and direct-path power-density estimation. Full 3D ray-tracing methods offer high fidelity, but their multi-path modeling and material-parameter requirements can be costly for large oblique photogrammetric city meshes. To address this problem, this paper proposes the Visible Radar Power-Density Field (VRPF), a 3D radar power-density field computation framework based on high-resolution oblique photogrammetric models. The method constructs a reusable spatial index for large numbers of triangular facets and performs two-stage occlusion queries: rapid Axis-Aligned Bounding Box (AABB) pruning followed by ray-triangle intersection tests. Together, these components enable efficient direct-path power-density calculation while accounting for line-of-sight occlusion in complex urban scenes. Qualitative and quantitative experiments show that VRPF better preserves occlusion boundaries around building edges, vegetation, and elevated structures than DSM-based baselines. VRPF also requires less time for repeated occlusion queries than a conventional 3D BVH ray-casting baseline while maintaining highly consistent radar-signal occlusion determinations. With 32 threads, VRPF computes power density for 108 target points in 5.92 s, about 2.66× faster than the 1 m DSM method. These results indicate that VRPF provides a practical balance between geometric fidelity and computational efficiency for direct-path radar power-density assessment with urban geometric occlusion. Full article
Show Figures

Figure 1

25 pages, 260979 KB  
Article
RDAH-Net: Bridging Relative Depth and Absolute Height for Monocular Height Estimation in Remote Sensing
by Liting Jiang, Feng Wang, Niangang Jiao, Jingxing Zhu, Yuming Xiang and Hongjian You
Remote Sens. 2026, 18(7), 1024; https://doi.org/10.3390/rs18071024 - 29 Mar 2026
Cited by 1 | Viewed by 1074
Abstract
Generating high-precision normalized digital surface models (nDSMs) from a single remote sensing image remains a challenging and ill-posed problem due to the absence of reliable geometric constraints. In this work, we show that monocular depth provides structurally stable cues of local geometry but [...] Read more.
Generating high-precision normalized digital surface models (nDSMs) from a single remote sensing image remains a challenging and ill-posed problem due to the absence of reliable geometric constraints. In this work, we show that monocular depth provides structurally stable cues of local geometry but lacks the global scale and vertical reference required for absolute height recovery. This intrinsic mismatch limits direct depth-to-height regression, particularly when transferring across heterogeneous terrains, land-cover compositions, and imaging conditions. Building on this idea, we propose the Relative Depth–Absolute Height Prediction Network (RDAH-Net), a framework that exploits relative depth as a geometry-aware prior while learning terrain-dependent height mappings from image appearance to absolute height. As the backbone, we employ a lightweight MobileNetV2 enhanced with a Convolutional Block Attention Module (CBAM), and further incorporate a cross-modal bidirectional attention fusion scheme with positional encoding to achieve a deep and effective fusion of image appearance and depth prior cues. Finally, a PixelShuffle-based upsampling strategy is used to sharpen prediction details and mitigate typical upsampling artifacts. Extensive experiments across diverse regions demonstrate that RDAH-Net achieves robust and generalizable height estimation, providing a practical alternative for large-scale mapping and rapid update scenarios. Full article
Show Figures

Figure 1

19 pages, 3672 KB  
Article
DualRecon: Building 3D Reconstruction from Dual-View Remote Sensing Images
by Ruizhe Shao, Hao Chen, Jun Li, Mengyu Ma and Chun Du
Remote Sens. 2025, 17(23), 3793; https://doi.org/10.3390/rs17233793 - 22 Nov 2025
Cited by 1 | Viewed by 1730
Abstract
Large-scale and rapid 3D reconstruction of urban areas holds significant practical value. Recently, methods that reconstruct buildings from off-nadir imagery have gained attention for their potential to meet the demand for large-scale, time-sensitive reconstruction applications. These methods typically estimate the building height and [...] Read more.
Large-scale and rapid 3D reconstruction of urban areas holds significant practical value. Recently, methods that reconstruct buildings from off-nadir imagery have gained attention for their potential to meet the demand for large-scale, time-sensitive reconstruction applications. These methods typically estimate the building height and footprint position by extracting building roof and the roof-to-footprint offset within a single off-nadir image. However, the reconstruction accuracy of these methods is primarily constrained by two issues: first, errors in the single-view building detection, and second, the inaccurate extraction of offsets, which is often a consequence of these detection errors as well as interference from shadow occlusion. To address these challenges, we propose DualRecon, a method for 3D building reconstruction from heterogeneous dual-view remote sensing imagery. In contrast to single-image detection methods, DualRecon achieves more accurate 3D information extraction for reconstruction by fusing and correlating building information across different views. This success can be attributed to three key advantages of DualRecon. First, DualRecon fuses the two input views and extracts building objects based on the fused image features, thereby improving the accuracy of building detection and localization. Second, compared to the roof-to-footprint offset, the disparity offset of the same rooftop between different views is less affected by interference from shadows and occlusions. Our method leverages this disparity offset to determine building height, which enhances the accuracy of height estimation. Third, we designed DualRecon with a three-branch architecture to be optimally tailored for the dual-view 3D information extraction task. Third, we designed DualRecon with a three-branch architecture to be optimally tailored for the dual-view 3D information extraction task. Moreover, this paper introduces BuildingDual—the first large-scale dual-view 3D building reconstruction dataset. It comprises 3789 image pairs containing 288,787 building instances, where each instance is annotated with its respective roofs in both views, roof-to-footprint offset, footprint, and the disparity offset of the roof. Experiments on this dataset demonstrate that DualRecon achieves more accurate reconstruction results than existing methods when performing 3D building reconstruction from dual-view remote sensing imagery. Our data and code will be made publicly available. Full article
Show Figures

Figure 1

19 pages, 43835 KB  
Article
A Stereo Disparity Map Refinement Method Without Training Based on Monocular Segmentation and Surface Normal
by Haoxuan Sun and Taoyang Wang
Remote Sens. 2025, 17(9), 1587; https://doi.org/10.3390/rs17091587 - 30 Apr 2025
Cited by 1 | Viewed by 3025
Abstract
Stereo disparity estimation is an essential component in computer vision and photogrammetry with many applications. However, there is a lack of real-world large datasets and large-scale models in the domain. Inspired by recent advances in the foundation model for image segmentation, we explore [...] Read more.
Stereo disparity estimation is an essential component in computer vision and photogrammetry with many applications. However, there is a lack of real-world large datasets and large-scale models in the domain. Inspired by recent advances in the foundation model for image segmentation, we explore the RANSAC disparity refinement based on zero-shot monocular surface normal prediction and SAM segmentation masks, which combine stereo matching models and advanced monocular large-scale vision models. The disparity refinement problem is formulated as follows: extracting geometric structures based on SAM masks and surface normal prediction, building disparity map hypotheses of the geometric structures, and selecting the hypotheses-based weighted RANSAC method. We believe that after obtaining geometry structures, even if there is only a part of the correct disparity in the geometry structure, the entire correct geometry structure can be reconstructed based on the prior geometry structure. Our method can best optimize the results of traditional models such as SGM or deep learning models such as MC-CNN. The model obtains 15.48% D1-error without training on the US3D dataset and obtains 6.09% bad 2.0 error and 3.65% bad 4.0 error on the Middlebury dataset. The research helps to promote the development of scene and geometric structure understanding in stereo disparity estimation and the application of combining advanced large-scale monocular vision models with stereo matching methods. Full article
Show Figures

Figure 1

18 pages, 22424 KB  
Article
Class-Incremental Semantic Segmentation for Mobile Laser Scanning Point Clouds Using Feature Representation Preservation and Loss Cross-Coupling
by Xucheng Chen, Haifeng Luo, Tianqiang Huang, Hanxian He and Wenyan Hu
Remote Sens. 2025, 17(3), 541; https://doi.org/10.3390/rs17030541 - 5 Feb 2025
Cited by 5 | Viewed by 2933
Abstract
Significant progress has been made in the semantic segmentation of mobile laser scanning (MLS) point clouds based on deep learning. However, the segmentation classes of deep learning models depend on the label classes of the source point clouds used for training, which makes [...] Read more.
Significant progress has been made in the semantic segmentation of mobile laser scanning (MLS) point clouds based on deep learning. However, the segmentation classes of deep learning models depend on the label classes of the source point clouds used for training, which makes it difficult to generalize the models to target point clouds with novel classes. In addition, retraining models using complete class label datasets is time-consuming, and the source point clouds are often unavailable or occupy a large amount of storage space. In this paper, we propose a new class-incremental semantic segmentation framework for MLS point clouds. Firstly, to prevent catastrophic forgetting of original class knowledge when the model learns novel classes, we design a feature representation preservation-based knowledge distillation module to maintain the encoding ability of the target models for original classes. Then, to further separate novel classes from the original background classes, we introduce a background shift mechanism based on loss cross-coupling and pseudo-label collaborative training, which adaptively balances the model plasticity when learning novel class knowledge. Finally, we conducted extensive experiments on two benchmark datasets (Paris-Lille-3D and Toronto-3D), and our proposed method achieved impressive results, which indicate that the proposed framework could effectively achieve class-incremental semantic segmentation for MLS point clouds. Full article
Show Figures

Figure 1

Review

Jump to: Research

33 pages, 58211 KB  
Review
Binocular Stereo Vision in Remote Sensing: A Review
by Xing Li, Hongwei Zhou, Mingyu Sun, Bangshu Xiong, Yuchao Dai, Renjie He, Zhihua Chen and Zhibo Rao
Remote Sens. 2026, 18(10), 1480; https://doi.org/10.3390/rs18101480 - 9 May 2026
Viewed by 592
Abstract
Stereo vision leverages binocular imagery to emulate the human visual system in perceiving three-dimensional (3D) structures by estimating disparity from rectified image pairs and converting it to depth via geometric triangulation. In recent years, deep learning-based stereo matching has significantly advanced in accuracy, [...] Read more.
Stereo vision leverages binocular imagery to emulate the human visual system in perceiving three-dimensional (3D) structures by estimating disparity from rectified image pairs and converting it to depth via geometric triangulation. In recent years, deep learning-based stereo matching has significantly advanced in accuracy, efficiency, and generalization, surpassing traditional methods and demonstrating great potential in remote sensing applications. However, stereo matching in remote sensing faces unique challenges not commonly seen in terrestrial datasets. These include limited access to satellite imagery, seasonal differences between image pairs, difficulty in identifying small objects, and widespread regions with repetitive textures, such as lakes and forests. Unlike prior surveys that primarily address ground-level scenes, this paper presents a comprehensive review of stereo matching techniques tailored for remote sensing. It synthesizes the progress and limitations of representative models, analyzes the characteristics and domain-specific constraints of remote sensing stereo datasets, and outlines future research directions and application prospects in this field. Full article
Show Figures

Figure 1

Back to TopTop