Dynamic-Parameterized Reconstruction Model for Resource-Aware Spatial Intelligence
Abstract
1. Introduction
- We propose a multi-exit monocular vehicle 3D reconstruction framework that maintains valid pose-and-mesh outputs at all exits and provides multiple predefined accuracy–latency operating points within a single model, making it suitable for resource-aware spatial intelligence.
- We introduce an exit-specific dynamic parameterization strategy by binding each exit to a predefined mesh specification, thereby coupling branch depth with geometric representation scale and enabling coarse-to-fine structured reconstruction across branches.
- We design branch-dependent keypoint heads for different exits, with lightweight coordinate-classification decoding adopted for low-latency early exits and heatmap regression retained in the Main branch to preserve the accuracy ceiling, effectively reducing RoI-head latency without sacrificing the strong performance of the full branch.
2. Related Work
2.1. Spatial Intelligence and Spatial Representations
2.2. Monocular 3D Reconstruction for Traffic Scenarios
2.3. Elastic Inference Under Resource Constraints
3. Methods
3.1. Model Overview
3.2. Dynamic-Parameterized Multi-Resolution Mesh Reconstruction
3.3. Attention-Guided Modeling Mechanism
3.4. Keypoint Localization and Visibility Prediction
3.5. Loss Functions
4. Experiments
4.1. ApolloCar3D Dataset
4.2. Implementation Details
4.3. Main Results
4.4. Ablation Study
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Rosinol, A.; Abate, M.; Chang, Y.; Carlone, L. Kimera: An open-source library for real-time metric-semantic localization and mapping. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2020; pp. 1689–1696. [Google Scholar]
- Wang, Z. 3D representation methods: A survey. arXiv 2024, arXiv:2410.06475. [Google Scholar] [CrossRef]
- Mehta, V.; Sharma, C.; Thiyagarajan, K. Large language models and 3D vision for intelligent robotic perception and autonomy. Sensors 2025, 25, 6394. [Google Scholar] [CrossRef] [PubMed]
- Wu, X.; Jiang, L.; Wang, P.S.; Liu, Z.; Liu, X.; Qiao, Y.; Ouyang, W.; He, T.; Zhao, H. Point transformer v3: Simpler faster stronger. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 4840–4851. [Google Scholar]
- Chen, Y.; Liu, J.; Zhang, X.; Qi, X.; Jia, J. Voxelnext: Fully sparse voxelnet for 3d object detection and tracking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 21674–21683. [Google Scholar]
- Li, Z.; Müller, T.; Evans, A.; Taylor, R.H.; Unberath, M.; Liu, M.Y.; Lin, C.H. Neuralangelo: High-fidelity neural surface reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 8456–8465. [Google Scholar]
- Zhang, C.; Delitzas, A.; Wang, F.; Zhang, R.; Ji, X.; Pollefeys, M.; Engelmann, F. Open-vocabulary functional 3d scene graphs for real-world indoor spaces. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 19401–19413. [Google Scholar]
- Renz, K.; Chitta, K.; Mercea, O.B.; Koepke, A.; Akata, Z.; Geiger, A. Plant: Explainable planning transformers via object-level representations. arXiv 2022, arXiv:2210.14222. [Google Scholar]
- Jiang, B.; Chen, S.; Xu, Q.; Liao, B.; Chen, J.; Zhou, H.; Zhang, Q.; Liu, W.; Huang, C.; Wang, X. Vad: Vectorized scene representation for efficient autonomous driving. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 8340–8350. [Google Scholar]
- Tong, H.; Chu, L.; Chen, Z.; Liu, Y.; Zhang, Y.; Hu, J. Multi-objective autonomous eco-driving strategy: A pathway to future green mobility. Green Energy Intell. Transp. 2025, 4, 100279. [Google Scholar] [CrossRef]
- Irshayyid, A.; Chen, J.; Xiong, G. A review on reinforcement learning-based highway autonomous vehicle control. Green Energy Intell. Transp. 2024, 3, 100156. [Google Scholar] [CrossRef]
- Li, T.; Ruan, J.; Zhang, K. The investigation of reinforcement learning-based end-to-end decision-making algorithms for autonomous driving on the road with consecutive sharp turns. Green Energy Intell. Transp. 2025, 4, 100288. [Google Scholar] [CrossRef]
- Jia, D.; Ruan, X.; Xia, K.; Zou, Z.; Wang, L.; Tang, W. Analysis-by-synthesis transformer for single-view 3d reconstruction. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2024; pp. 259–277. [Google Scholar]
- Yang, X.; Lin, G.; Zhou, L. Single-view 3D mesh reconstruction for seen and unseen categories. IEEE Trans. Image Process. 2023, 32, 3746–3758. [Google Scholar] [CrossRef] [PubMed]
- Long, X.; Guo, Y.C.; Lin, C.; Liu, Y.; Dou, Z.; Liu, L.; Ma, Y.; Zhang, S.H.; Habermann, M.; Theobalt, C.; et al. Wonder3d: Single image to 3d using cross-domain diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 9970–9980. [Google Scholar]
- Chabot, F.; Chaouch, M.; Rabarisoa, J.; Teuliere, C.; Chateau, T. Deep manta: A coarse-to-fine many-task network for joint 2d and 3d vehicle analysis from monocular image. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2040–2049. [Google Scholar]
- Ke, L.; Li, S.; Sun, Y.; Tai, Y.W.; Tang, C.K. Gsnet: Joint vehicle pose and shape reconstruction with geometrical and scene-aware supervision. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2020; pp. 515–532. [Google Scholar]
- Lee, H.J.; Kim, H.; Choi, S.M.; Jeong, S.G.; Koh, Y.J. BAAM: Monocular 3D pose and shape reconstruction with bi-contextual attention module and attention-guided modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 9011–9020. [Google Scholar]
- Hu, H.; Dey, D.; Hebert, M.; Bagnell, J.A. Learning anytime predictions in neural networks via adaptive loss balancing. In Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA, 27 January–1 February 2019; Association for the Advancement of Artificial Intelligence: Palo Alto, CA, USA, 2019; Volume 33, pp. 3812–3821. [Google Scholar]
- Lee, H.; Shin, J. Anytime neural prediction via slicing networks vertically. arXiv 2018, arXiv:1807.02609. [Google Scholar] [CrossRef]
- Iuzzolino, M.; Mozer, M.C.; Bengio, S. Improving anytime prediction with parallel cascaded networks and a temporal-difference loss. Adv. Neural Inf. Process. Syst. 2021, 34, 27631–27644. [Google Scholar]
- Huang, G.; Chen, D.; Li, T.; Wu, F.; Van Der Maaten, L.; Weinberger, K.Q. Multi-scale dense networks for resource efficient image classification. arXiv 2017, arXiv:1703.09844. [Google Scholar]
- Zhou, W.; Xu, C.; Ge, T.; McAuley, J.; Xu, K.; Wei, F. Bert loses patience: Fast and robust inference with early exit. Adv. Neural Inf. Process. Syst. 2020, 33, 18330–18341. [Google Scholar]
- Panda, P.; Sengupta, A.; Roy, K. Conditional deep learning for energy-efficient and enhanced pattern recognition. In Proceedings of the 2016 Design, Automation & Test in Europe Conference & Exhibition (DATE); IEEE: New York, NY, USA, 2016; pp. 475–480. [Google Scholar]
- Teerapittayanon, S.; McDanel, B.; Kung, H.T. Branchynet: Fast inference via early exiting from deep neural networks. In Proceedings of the 2016 23rd International Conference on Pattern Recognition (ICPR); IEEE: New York, NY, USA, 2016; pp. 2464–2469. [Google Scholar]
- LeCun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-based learning applied to document recognition. Proc. IEEE 2002, 86, 2278–2324. [Google Scholar] [CrossRef]
- Krizhevsky, A.; Sutskever, I.; Hinton, G.E. Imagenet classification with deep convolutional neural networks. Adv. Neural Inf. Process. Syst. 2012, 25, 1097–1105. [Google Scholar] [CrossRef]
- Liu, D.; Kan, M.; Shan, S.; Chen, X. A simple romance between multi-exit vision transformer and token reduction. In Proceedings of the Twelfth International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Yoon, J.W.; Woo, B.J.; Kim, N.S. Hubert-ee: Early exiting hubert for efficient speech recognition. arXiv 2022, arXiv:2204.06328. [Google Scholar]
- Lasbordes, M.; Falavigna, D.; Brutti, A. Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices. arXiv 2025, arXiv:2506.18035. [Google Scholar] [CrossRef]
- Di Francesco, A.G.; Bucarelli, M.S.; Nardini, F.M.; Perego, R.; Tonellotto, N.; Silvestri, F. Early-Exit Graph Neural Networks. arXiv 2025, arXiv:2505.18088. [Google Scholar] [CrossRef]
- Yu, J.; Yang, L.; Xu, N.; Yang, J.; Huang, T. Slimmable neural networks. arXiv 2018, arXiv:1812.08928. [Google Scholar] [CrossRef]
- Li, C.; Wang, G.; Wang, B.; Liang, X.; Li, Z.; Chang, X. Dynamic slimmable network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 19–25 June 2021; pp. 8607–8617. [Google Scholar]
- Chen, Y.; Dai, X.; Liu, M.; Chen, D.; Yuan, L.; Liu, Z. Dynamic convolution: Attention over convolution kernels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 14–19 June 2020; pp. 11030–11039. [Google Scholar]
- Pan, Z.; Cai, J.; Zhuang, B. Stitchable neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 16102–16112. [Google Scholar]
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 2961–2969. [Google Scholar]
- Gao, S.H.; Cheng, M.M.; Zhao, K.; Zhang, X.Y.; Yang, M.H.; Torr, P. Res2net: A new multi-scale backbone architecture. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 43, 652–662. [Google Scholar] [CrossRef] [PubMed]
- Tan, M.; Pang, R.; Le, Q.V. Efficientdet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 14–19 June 2020; pp. 10781–10790. [Google Scholar]
- Song, X.; Wang, P.; Zhou, D.; Zhu, R.; Guan, C.; Dai, Y.; Su, H.; Li, H.; Yang, R. Apollocar3d: A large 3d car instance understanding benchmark for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; pp. 5452–5462. [Google Scholar]
- Liu, S.; Li, T.; Chen, W.; Li, H. Soft rasterizer: A differentiable renderer for image-based 3d reasoning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 7708–7717. [Google Scholar]
- Li, Y.; Yang, S.; Liu, P.; Zhang, S.; Wang, Y.; Wang, Z.; Yang, W.; Xia, S.T. Simcc: A simple coordinate classification perspective for human pose estimation. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2022; pp. 89–106. [Google Scholar]
- Chen, Y.; Tai, L.; Sun, K.; Li, M. Monopair: Monocular 3d object detection using pairwise spatial relationships. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 14–19 June 2020; pp. 12093–12102. [Google Scholar]
- Lu, Y.; Ma, X.; Yang, L.; Zhang, T.; Liu, Y.; Chu, Q.; Yan, J.; Ouyang, W. Geometry uncertainty projection network for monocular 3d object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 3111–3121. [Google Scholar]
- Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft coco: Common objects in context. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2014; pp. 740–755. [Google Scholar]



| Equipment | Computer Configuration Parameters |
|---|---|
| Operating system | Linux |
| RAM | 32 G |
| Type of operating system | Ubuntu20.04 |
| CPU | Intel Core i7-12700K |
| GPU | RTX 4090 (24 GB) × 1 |
| Development language | Python 3.9 |
| Deep learning framework | PyTorch 1.10.1 |
| Methods | A3DP-Abs Mean ↑ | A3DP-Abs c-l ↑ | A3DP-Abs c-s ↑ | A3DP-Rel Mean ↑ | A3DP-Rel c-l ↑ | A3DP-Rel c-s ↑ | Inf. Time (ms) ↓ | Vertices |
|---|---|---|---|---|---|---|---|---|
| DyPRSI-EE1 | 21.31 | 42.89 | 19.27 | 17.08 | 38.89 | 13.06 | 235.48 | 488 |
| DyPRSI-EE2 | 22.57 | 46.42 | 19.49 | 19.63 | 44.16 | 15.30 | 243.80 | 828 |
| DyPRSI-Main | 23.39 | 47.68 | 20.09 | 21.83 | 46.35 | 17.88 | 407.52 | 1352 |
| GSNet | 18.91 | 37.42 | 18.35 | 20.19 | 40.38 | 19.47 | 180.26 | 1352 |
| BAAM | 23.40 | 44.07 | 22.01 | 19.39 | 40.60 | 15.85 | 409.01 | 1352 |
| Model | A3DP-Abs-Mean ↑ | A3DP-Rel-Mean ↑ | Inf. Time (ms) ↓ | Vertices | |||
|---|---|---|---|---|---|---|---|
| Backbone + BiFPN | RoI Heads | Total | |||||
| DyPRSI (Default) | EE1-CC | 21.31 | 17.08 | 63.39 | 172.64 | 236.03 | 488 |
| EE2-CC | 22.57 | 19.63 | 71.75 | 172.05 | 243.80 | 828 | |
| Main-HR | 23.39 | 21.83 | 79.64 | 327.87 | 407.52 | 1352 | |
| DyPRSI (All-HR) | EE1-HR | 21.56 | 17.20 | 63.62 | 332.40 | 396.02 | 488 |
| EE2-HR | 22.66 | 19.73 | 71.29 | 327.73 | 399.02 | 828 | |
| Main-HR | 23.24 | 20.83 | 79.96 | 330.09 | 410.05 | 1352 | |
| DyPRSI (All-CC) | EE1-CC | 21.54 | 17.29 | 63.21 | 172.03 | 235.24 | 488 |
| EE2-CC | 22.29 | 19.10 | 71.68 | 171.34 | 243.03 | 828 | |
| Main-CC | 22.74 | 20.39 | 79.94 | 174.38 | 254.32 | 1352 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Huang, H.; Zhang, Y.; Song, L.; Zhao, Z.; Yang, X. Dynamic-Parameterized Reconstruction Model for Resource-Aware Spatial Intelligence. Sensors 2026, 26, 2355. https://doi.org/10.3390/s26082355
Huang H, Zhang Y, Song L, Zhao Z, Yang X. Dynamic-Parameterized Reconstruction Model for Resource-Aware Spatial Intelligence. Sensors. 2026; 26(8):2355. https://doi.org/10.3390/s26082355
Chicago/Turabian StyleHuang, Hongyi, Yanni Zhang, Liang Song, Zhen Zhao, and Xiaopeng Yang. 2026. "Dynamic-Parameterized Reconstruction Model for Resource-Aware Spatial Intelligence" Sensors 26, no. 8: 2355. https://doi.org/10.3390/s26082355
APA StyleHuang, H., Zhang, Y., Song, L., Zhao, Z., & Yang, X. (2026). Dynamic-Parameterized Reconstruction Model for Resource-Aware Spatial Intelligence. Sensors, 26(8), 2355. https://doi.org/10.3390/s26082355

