In-Orbit MapAnything: An Enhanced Feed-Forward Metric Framework for 3D Reconstruction of Non-Cooperative Space Targets Under Complex Lighting
Abstract
1. Introduction
- Construction of a Space-Specific Dataset: We established a multi-modal dataset featuring specular and weak-texture characteristics, effectively alleviating the data scarcity in space scenarios.
- Hierarchical Sampling and Fusion-Aware Mechanism: We design a fusion module that incorporates hierarchical cascade sampling and lightweight highlight-aware features. Through the dynamic integration of multi-level features, this module enhances the geometric reconstruction completeness and accuracy in regions with strong reflections and weak textures.
- Realization of Parameter-Efficient Domain Adaptation: By leveraging DoRA for lightweight fine-tuning of MapAnything, we transformed the general vision baseline into an expert model for spatial 3D reconstruction at a minimal computational cost.
2. Related Work
2.1. Space-Specific Reconstruction Algorithms
2.2. General Feed-Forward 3D Reconstruction Models
2.3. Public Datasets for Space Targets
3. Method
3.1. Overview of the Baseline Framework
- (1)
- Input Encoding and Alignment:
- (2)
- Transformer Backbone:
- (3)
- Factored Decoding:
3.2. SatMap-Adapter Perception Fusion Module
- (1)
- Hierarchical Feature Extraction
- Shallow features : Focus on edges, texture gradients, and high-frequency illumination variations, capable of retaining specular reflection information;
- Middle features : Correspond to part-level representations, enabling differentiation between structures such as solar panels and the main body, and supporting local geometric smoothing constraints;
- Deep features : Contain high-level semantic information before excessive smoothing occurs;
- Final features : The final-layer features originally used by MapAnything, providing global semantic consistency and ensuring topologically plausible reconstruction.
- (2)
- Feature Alignment and Position Encoding
- (3)
- Illumination-Adaptive Feature Fusion Module
3.3. DoRA-Based Decoder Adaptation
- (1)
- Principle of DoRA
- (2)
- Targeted Optimization Strategy
- Fully Frozen Modules: The heavy DINOv2 encoder backbone (approximately 300 million parameters) and the mask/confidence computation heads in the MapAnything decoder are completely frozen.
- DoRA-Adapted Modules: We selectively insert DoRA adapters into the attention computation modules, the newly proposed SatMap-Adapter, and the specific decoding branches responsible for pose, depth, and scale estimation.
- (3)
- Train Loss
- Scale-Independent Losses: and constrain the predicted dense ray directions and camera quaternions, respectively. Since rotation and ray direction are independent of scene scale, these are evaluated directly against the ground truth using angular and distance metrics.
- Up-to-Scale Geometry Losses: To ensure robust convergence of the spatial structure, scale-invariant losses are applied to the translation vectors (), ray depths (), local pointmaps (), and global world-frame pointmaps (). Following standard practices, these regressions are computed in log-space, with acting as a confidence-weighted loss.
- Global Metric Scale Loss: penalizes the deviation between the predicted global scaling factor and the ground-truth metric scale, forcing the network to recover the absolute physical dimensions of the non-cooperative targets.
- Detail and Mask Constraints: Leveraging our high-fidelity synthetic dataset, we apply a normal loss () and a multi-scale gradient matching loss () to explicitly capture the high-frequency structural details of the spacecraft (e.g., antennas and solar panels). Lastly, utilizes binary cross-entropy to supervise the valid foreground masks.
4. Datasets Construction for Space Targets
4.1. Synthetic Image Generation
4.2. Real Image Acquisition
5. Experiment and Analysis
5.1. Experimental Settings
- AbsRel (Absolute Relative Error): Measures the absolute relative deviation between predicted and ground truth depths; lower values indicate better performance.
- RMSE (Root Mean Squared Error): Quantifies the global deviation of depth values; lower is better.
- δ (Threshold Accuracy): The percentage of pixels where the ratio between the predicted and ground truth; higher is better.
- CD (Chamfer Distance): Evaluates the geometric structural alignment between the reconstructed point cloud and the Ground Truth (GT); lower values denote higher geometric fidelity.
- ATE (Absolute Trajectory Error): Directly measures the translational drift of camera poses; lower is better.
- Inference Time: The duration of each reconstruction inference is documented to directly reflect the model’s real-time processing capability.
5.2. Comparison with Baselines
5.3. Ablation Studies
5.4. Generalization Assessment Experiment
6. Conclusions and Discussion
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Lowe, D.G. Distinctive image features from scale-invariant keypoints. Int. J. Comput. Vis. 2004, 60, 91–110. [Google Scholar] [CrossRef]
- Snavely, N.; Seitz, S.M.; Szeliski, R. Photo tourism: Exploring photo collections in 3D. ACM Trans. Graph. 2006, 25, 835–846. [Google Scholar] [CrossRef]
- Schönberger, J.L.; Frahm, J.-M. Structure-from-motion revisited. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 4104–4113. [Google Scholar]
- Schönberger, J.L.; Zheng, E.; Frahm, J.-M.; Pollefeys, M. Pixelwise view selection for unstructured multi-view stereo. In Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands, 11–14 October 2016; pp. 501–518. [Google Scholar]
- Sharma, S.; Beierle, C.; D’Amico, S. Pose estimation for non-cooperative spacecraft rendezvous using neural networks. IEEE Trans. Aerosp. Electron. Syst. 2020, 56, 4638–4658. [Google Scholar] [CrossRef]
- Xu, H.; Zhang, G.; Cai, J.; Liu, X.; Di, J.M. Kinematics-informed neural implicit representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; pp. 3368–3378. [Google Scholar]
- Sitzmann, V.; Thies, J.; Heide, F.; Nießner, M.; Wetzstein, G.; Zollhöfer, M. DeepVoxels: Learning persistent 3D feature embeddings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 2437–2446. [Google Scholar]
- Mildenhall, B.; Srinivasan, P.P.; Tancik, M.; Barron, J.T.; Ramamoorthi, R.; Ng, R. NeRF: Representing scenes as neural radiance fields for view synthesis. In Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020; pp. 405–421. [Google Scholar]
- Gong, Z.; Marroquim, R.; He, Y. Sat-NeRF: Learning multi-view satellite photogrammetry with transient objects and shadow modeling using RPC cameras. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 21087–21096. [Google Scholar]
- Liu, Y.; Sun, Z.; Zhang, L.; Wang, D. Spacecraft-NeRF: High-fidelity reconstruction of spacecraft by neural radiance field based implicit representation. IEEE Trans. Aerosp. Electron. Syst. 2025, 61, 15182–15194. [Google Scholar] [CrossRef]
- Yang, D.; Zhang, Y.; Yu, G.; Jiao, J.; Wang, B.; Huang, P. NeRF-based simultaneous pose estimation and 3D reconstruction for non-cooperative space target. Aerosp. Sci. Technol. 2026, 157, 110167. [Google Scholar] [CrossRef]
- Fan, Z.; Zhai, D.; Li, H.; Tian, Y.; Ni, G. Compact single-photon LiDAR for satellite laser ranging. Opt. Express 2025, 33, 40876–40889. [Google Scholar]
- Tian, Y.; Zhang, J.; Li, H.; Fan, Z. High accuracy ranging for space debris with spaceborne single photon Lidar. Opt. Express 2024, 32, 12318–12339. [Google Scholar] [CrossRef] [PubMed]
- Li, S.; Wang, Y.; Chen, Z. Three-dimensional quantum imaging of dynamic targets using quantum compressed sensing. Opt. Express 2024, 32, 6025–6040. [Google Scholar] [CrossRef] [PubMed]
- Zhang, J.; Liu, S.; Wang, H. LocNet: Deep learning-based localization on a rotating point spread function with applications to telescope imaging. Opt. Express 2023, 31, 39341–39356. [Google Scholar]
- Wang, G.; Chen, Z.; Li, J.; Nie, Z.; Liu, C.; Deng, S.; Liang, X. SparseNeRF: Distilling depth ranking for few-shot novel view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; pp. 17192–17203. [Google Scholar]
- Han, X.; Li, W.; Zhang, Y.; Hu, Q. SparseRecon: Neural implicit surface reconstruction from sparse views with feature and depth consistencies. In IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2025. [Google Scholar]
- Yu, A.; Ye, V.; Tancik, M.; Kanazawa, A. pixelNeRF: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 4578–4587. [Google Scholar]
- Hong, Y.; Zhang, K.; Gu, J.; Bi, S.; Zhou, Y.; Liu, D.; Liu, F.; Sunkavalli, K.; Bui, T.; Tan, H. LRM: Large reconstruction model for single image to 3D. In Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Fan, Z.; Wang, N.; Zhang, Y.; Wang, H. InstantSplat: Sparse-view SfM-free Gaussian splatting in seconds. arXiv 2024, arXiv:2403.20230. [Google Scholar]
- Wang, S.; Leroy, V.; Cabon, Y.; Chidlovskii, B.; Revaud, J. DUSt3R: Geometric 3D vision made easy. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 20697–20709. [Google Scholar]
- Cabon, Y.; Stoffl, L.; Antsfeld, L.; Csurka, G.; Chidlovskii, B.; Revaud, J.; Leroy, V. MUSt3R: Multi-view network for stereo 3D reconstruction. arXiv 2024, arXiv:2410.12652. [Google Scholar]
- Wang, J.; Chen, M.; Karaev, N.; Vedaldi, A.; Rupprecht, C.; Novotny, D. VGGT: Visual geometry grounded transformer. arXiv 2025, arXiv:2503.11651. [Google Scholar] [CrossRef]
- Chen, C.; Chen, Z.; Wang, Y. MapAnything: Universal feed-forward metric 3D reconstruction. arXiv 2024, arXiv:2406.02314. [Google Scholar]
- Kisantal, M.; Sharma, S.; Park, T.H.; Izzo, D.; Märtens, S.; D’Amico, S. Satellite pose estimation challenge: Dataset, competition design, and results. IEEE Trans. Aerosp. Electron. Syst. 2020, 56, 4083–4098. [Google Scholar] [CrossRef]
- Park, T.H.; Märtens, M.; Lecuyer, G.; Izzo, D.; D’Amico, S. SPEED+: Next-generation dataset for spacecraft pose estimation across domain gap. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 21087–21096. [Google Scholar]
- Hu, Y.; Speierer, S.; Jakob, W.; Fua, P.; Salzmann, M. Wide-depth-range 6D object pose estimation in space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 15870–15879. [Google Scholar]
- Afara, M.; Aly, H.A.; Lab, C.V.I. SPADES: A realistic spacecraft pose estimation dataset using event sensing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vancouver, BC, Canada, 18–22 June 2023. [Google Scholar]
- Proença, P.F.; Gao, Y. Deep learning for spacecraft pose estimation from photorealistic rendering. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May–31 August 2020; pp. 6007–6013. [Google Scholar]
- Ranftl, R.; Bochkovskiy, A.; Koltun, V. Vision transformers for dense prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 12179–12188. [Google Scholar]
- Rumelhart, D.E.; Hinton, G.E.; Williams, R.J. Learning representations by back-propagating errors. Nature 1986, 323, 533–536. [Google Scholar] [CrossRef]
- Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafranski, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. DINOv2: Learning robust visual features without supervision. arXiv 2024, arXiv:2304.07193. [Google Scholar] [CrossRef]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 7132–7141. [Google Scholar]
- Xia, X.; Zhang, D.; Song, W.; Huang, W.; Hurni, L. MapSAM: Adapting segment anything model for automated feature detection in historical maps. GISci. Remote Sens. 2025, 62, 2494883. [Google Scholar] [CrossRef]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual Event, 25–29 April 2022. [Google Scholar]
- Liu, S.-Y.; Wang, C.-Y.; Hong, H.; Ma, X.; Li, Y.; Wu, P.-H.; Venkataraman, S. DoRA: Weight-decomposed low-rank adaptation. In Proceedings of the International Conference on Machine Learning (ICML), Vienna, Austria, 21–27 July 2024. [Google Scholar]
- National Aeronautics and Space Administration. NASA 3D Resources Database. Available online: https://nasa3d.arc.nasa.gov/ (accessed on 21 January 2026).
- European Space Agency. ESA Science Satellite Fleet (Scifleet) Database. Available online: https://www.cosmos.esa.int/ (accessed on 21 January 2026).
- Lu, Y.; Wang, H.; Zang, Y.; Liao, Z.; Ning, Q.; Yan, Z. Elaborate imaging simulation and validation of the space target based on surface fractal features. Opt. Express 2025, 33, 24813–24830. [Google Scholar] [CrossRef] [PubMed]












| Dataset | Modality | Targets | Lighting Condition | Data Supplementation | Material Complexity |
|---|---|---|---|---|---|
| SwissCube | Syn | 1 | Variable | Pose + Mask | Medium |
| SPEED+ | Syn + Real | 1 | Variable | Pose Only | Low |
| URSO | Syn | 2 | Variable | Pose Only | Low |
| SPARK | Syn + Real | 1 | Simple | Pose + Model | Medium |
| Ours | Syn + Real | 50+ | Variable | Mask + Depth + Model | High (MLI/Wrinkles) |
| Parameters | Value |
|---|---|
| Illumination (mm2) | 305 × 305 |
| Maximum angle of incidence (°) | <±0.5° |
| Typical power output (mW/cm2) | 100 (1SUN), ±20% Adjustable |
| Uniformity | <±2% |
| Spectral match | 9.7–16.1% (800–900 nm) |
| Parameter Category | Parameter Symbol | Calibrated Value |
|---|---|---|
| Image Resolution | Width × Height | 1920 × 1080 |
| Focal Length | 1450.25 | |
| 1451.10 | ||
| Principal Point | 965.50 | |
| 542.30 | ||
| Radial Distortion | −0.1254 | |
| 0.0842 | ||
| −0.0021 | ||
| Tangential Distortion | 0.0015 | |
| −0.0008 | ||
| Reprojection Error | RMSE | 5.6 pixels |
| Method | AbsRel ↓ | RMSE ↓ | δ ↑ | ATE ↓ | CD ↓ | Time ↓ |
|---|---|---|---|---|---|---|
| DUSt3R | 0.112 ± 0.005 | 0.456 ± 0.012 | 0.854 ± 0.008 | 0.045 ± 0.003 | 1.23 ± 0.04 | 9.7 s |
| VGGT | 0.135 ± 0.006 | 0.512 ± 0.015 | 0.810 ± 0.009 | 0.052 ± 0.004 | 1.45 ± 0.05 | 3.2 s |
| MapAnything | 0.105 ± 0.004 | 0.420 ± 0.010 | 0.885 ± 0.006 | 0.038 ± 0.002 | 1.10 ± 0.03 | 2.1 s |
| Ours | 0.089 ± 0.002 | 0.350 ± 0.008 | 0.920 ± 0.004 | 0.032 ± 0.001 | 0.88 ± 0.02 | 2.2 s |
| Model Setting | SatMap-Adapter | DoRA | AbsRel ↓ | CD ↓ | Trainable Params | Peak VRAM | Training Time |
|---|---|---|---|---|---|---|---|
| Baseline (MapAnything) | × | × | 0.115 ± 0.005 | 1.12 ± 0.04 | 390 M | 22.3 G | 77.6 h |
| Baseline + SatMap-Adapter | √ | × | 0.091 ± 0.003 | 0.95 ± 0.03 | 394 M | 22.6 G | 79.5 h |
| Ours (Full) | √ | √ | 0.089 ± 0.002 | 0.88 ± 0.02 | 85 M | 5.9 G | 58.4 h |
| Parameters | Value |
|---|---|
| Satellite body size (mm) | 45 × 45 × 75 |
| Satellite panel size (mm) | 160 × 52 × 3 |
| Satellite body material | Gold polyimide film |
| Satellite panel material | Monocrystalline silicon cell |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Lu, Y.; Wang, H.; Ning, Q.; Liu, Z.; Zang, Y.; Liao, Z.; Yan, Z. In-Orbit MapAnything: An Enhanced Feed-Forward Metric Framework for 3D Reconstruction of Non-Cooperative Space Targets Under Complex Lighting. Sensors 2026, 26, 2026. https://doi.org/10.3390/s26072026
Lu Y, Wang H, Ning Q, Liu Z, Zang Y, Liao Z, Yan Z. In-Orbit MapAnything: An Enhanced Feed-Forward Metric Framework for 3D Reconstruction of Non-Cooperative Space Targets Under Complex Lighting. Sensors. 2026; 26(7):2026. https://doi.org/10.3390/s26072026
Chicago/Turabian StyleLu, Yinxi, Hongyuan Wang, Qianhao Ning, Ziyang Liu, Yunzhao Zang, Zhen Liao, and Zhiqiang Yan. 2026. "In-Orbit MapAnything: An Enhanced Feed-Forward Metric Framework for 3D Reconstruction of Non-Cooperative Space Targets Under Complex Lighting" Sensors 26, no. 7: 2026. https://doi.org/10.3390/s26072026
APA StyleLu, Y., Wang, H., Ning, Q., Liu, Z., Zang, Y., Liao, Z., & Yan, Z. (2026). In-Orbit MapAnything: An Enhanced Feed-Forward Metric Framework for 3D Reconstruction of Non-Cooperative Space Targets Under Complex Lighting. Sensors, 26(7), 2026. https://doi.org/10.3390/s26072026

