LiDAR-Free 3D Auto-Labeling via Radar–Visual Spatio-Temporal Consistency
Abstract
1. Introduction
- We present a LiDAR-free 3D auto-labeler for roadside camera–radar data, where VFM pseudo-point clouds, 2D mask tracks, and radar trajectories are organized into object-centric associated trajectories for subsequent consistency reasoning.
- We propose an uncertainty-aware heading fusion mechanism that robustly combines motion-derived (radar) and structure-derived (VFM) orientation estimates by calibrating their discrepancies on automatically identified reliable frames, yielding stable object headings for downstream optimization.
- We adapt canonical-space bundle adjustment and Thin-Plate Spline (TPS)-based propagation to VFM-derived pseudo-geometry, using temporally consistent landmarks to refine sparse reliable corrections and propagate them to full object point clouds.
- We evaluate the labeling framework on a diverse dataset in terms of data quality and downstream detection gains, showcasing performance improvement compared to current approaches.
2. Related Work
2.1. Visual Foundation Models for 3D Geometry
2.2. Radar–Camera Fusion for 3D Perception
2.3. LiDAR-Based 3D Auto-Labeling
3. Preliminaries and Problem Formulation
3.1. Geometry Limitation of Vision Foundation Models
3.2. Problem Formulation and Notations
4. Methods
4.1. Associated Trajectory Generation
- 3D pseudo-object trajectory: For each mask , we extract all 3D points from whose pixel coordinates lie within the mask, forming . The sequence constitutes the VFM-derived 3D trajectory.
- Radar trajectory: Radar points are projected onto the image plane using calibrated extrinsics. A radar detection is assigned to object k if it falls within , yielding the radar trajectory .
4.2. Uncertainty-Aware Heading Fusion
4.2.1. Heading Candidate Estimation
- Shape-based heading : We perform discrete L-shape fitting on the object point cloud by evaluating a set of candidate orientations . For each , we rotate , compute the axis-aligned bounding rectangle, and select the orientation that minimizes its area:where and are the width and height of the rotated bounding box. The interval is used only as a coarse orientation proposal grid over the symmetry range of vehicle boxes. This coarse-shape proposal is subsequently fused with radar-motion cues and refined by multi-frame optimization, so the final heading is not limited to the initial grid resolution.
- Motion-based heading : We fit a cubic B-spline to the radar trajectory . The raw motion heading is taken as the tangent direction of the spline at time t. To mitigate outliers from radar ghost returns, we regularize this direction using the Doppler velocity vector projected onto the ground plane. Given the radar trajectory, the cubic B-spline is derived by minimizing the following:where is a smoothing factor selected by generalized cross-validation. The integral term penalizes curvature, suppressing oscillations from noisy radar returns.
4.2.2. Uncertainty Calibration via Trusted Frames
4.2.3. Fused Heading Computation
4.3. Landmark-Guided Global Refinement in Canonical Space
4.3.1. 2D Mask Alignment
4.3.2. Canonical Object-Aware Bundle Adjustment
| Algorithm 1 Adaptive Bidirectional Landmark Tracking |
|
4.4. Propagation
5. Experiments
5.1. Dataset Selection
5.2. Effectiveness of Canonical Bundle Adjustment
5.3. Label Quality Assessment
5.4. Downstream Detection
6. Discussion
6.1. Discussion on Offline Labeling
6.2. Discussion on Adverse Weather Condition
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Dao, M.Q.; Caesar, H.; Berrio, J.S.; Shan, M.; Worrall, S.; Frémont, V.; Malis, E. Label-efficient 3d object detection for road-side units. In Proceedings of the 2024 IEEE Intelligent Vehicles Symposium (IV), Jeju Island, Republic of Korea, 2–5 June 2024; pp. 1572–1579. [Google Scholar]
- Engelmann, F.; Stuckler, J.; Leibe, B. SAMP: Shape and Motion Priors for 4D Vehicle Reconstruction. In Proceedings of the 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), Santa Rosa, CA, USA, 24–31 March 2017; pp. 400–408. [Google Scholar] [CrossRef]
- Liu, C.; Qian, X.; Huang, B.; Qi, X.; Lam, E.; Tan, S.C.; Wong, N. Multimodal transformer for automatic 3d annotation and object detection. In Computer Vision—ECCV 2022, Proceedings of the 17th European Conference, Tel Aviv, Israel, 23–27 October 2022; Springer: Cham, Switzerland, 2022; pp. 657–673. [Google Scholar]
- Shi, S.; Jiang, L.; Deng, J.; Wang, Z.; Guo, C.; Shi, J.; Wang, X.; Li, H. PV-RCNN++: Point-voxel feature set abstraction with local vector representation for 3D object detection. Int. J. Comput. Vis. 2023, 131, 531–551. [Google Scholar] [CrossRef]
- Ma, T.; Yang, X.; Zhou, H.; Li, X.; Shi, B.; Liu, J.; Yang, Y.; Liu, Z.; He, L.; Qiao, Y.; et al. DetZero: Rethinking Offboard 3D Object Detection with Long-term Sequential Point Clouds. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 6713–6724. [Google Scholar] [CrossRef]
- Yang, B.; Bai, M.; Liang, M.; Zeng, W.; Urtasun, R. Auto4d: Learning to label 4d objects from sequential point clouds. arXiv 2021, arXiv:2101.06586. [Google Scholar] [CrossRef]
- Zakharov, S.; Kehl, W.; Bhargava, A.; Gaidon, A. Autolabeling 3D Objects With Differentiable Rendering of SDF Shape Priors. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 12221–12230. [Google Scholar] [CrossRef]
- Wang, R.; Xu, S.; Dai, C.; Xiang, J.; Deng, Y.; Tong, X.; Yang, J. Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025; pp. 5261–5271. [Google Scholar]
- Wang, R.; Xu, S.; Dong, Y.; Deng, Y.; Xiang, J.; Lv, Z.; Sun, G.; Tong, X.; Yang, J. MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details. arXiv 2025, arXiv:2507.02546. [Google Scholar] [CrossRef]
- Yang, L.; Kang, B.; Huang, Z.; Zhao, Z.; Xu, X.; Feng, J.; Zhao, H. Depth anything v2. In Advances in Neural Information Processing Systems, Proceedings of the 38th Conference on Neural Information Processing Systems (NeurIPS 2024), Vancouver, BC, Canada, 10–15 December 2024; Curran Associates, Inc.: Red Hook, NY, USA, 2024; Volume 37, pp. 21875–21911. [Google Scholar]
- Lin, H.; Chen, S.; Liew, J.; Chen, D.Y.; Li, Z.; Shi, G.; Feng, J.; Kang, B. Depth Anything 3: Recovering the Visual Space from Any Views. arXiv 2025, arXiv:2511.10647. [Google Scholar] [CrossRef]
- Yang, L.; Zhang, X.; Li, J.; Wang, L.; Zhang, C.; Ju, L.; Li, Z.; Shen, Y.; Lv, C.; Wang, H. SGV3D: Toward Scenario Generalization for Vision-Based Roadside 3D Object Detection. IEEE Trans. Intell. Transp. Syst. 2025, 26, 11782–11793. [Google Scholar] [CrossRef]
- Lin, H.; Peng, S.; Chen, J.; Peng, S.; Sun, J.; Liu, M.; Bao, H.; Feng, J.; Zhou, X.; Kang, B. Prompting depth anything for 4k resolution accurate metric depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025; pp. 17070–17080. [Google Scholar]
- Zhang, F.; Yu, Z.; Li, C.; Zhang, R.; Bai, X.; Zhou, Z.; Cao, S.Y.; Wang, F.; Shen, H.L. Structure-Aware Radar-Camera Depth Estimation. In Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA), Atlanta, GA, USA, 19–23 May 2025; pp. 13028–13035. [Google Scholar] [CrossRef]
- Ranftl, R.; Lasinger, K.; Hafner, D.; Schindler, K.; Koltun, V. Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-Shot Cross-Dataset Transfer. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 1623–1637. [Google Scholar] [CrossRef] [PubMed]
- Yang, L.; Kang, B.; Huang, Z.; Xu, X.; Feng, J.; Zhao, H. Depth anything: Unleashing the power of large-scale unlabeled data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 10371–10381. [Google Scholar]
- Hu, M.; Yin, W.; Zhang, C.; Cai, Z.; Long, X.; Chen, H.; Wang, K.; Yu, G.; Shen, C.; Shen, S. Metric3D v2: A Versatile Monocular Geometric Foundation Model for Zero-Shot Metric Depth and Surface Normal Estimation. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 10579–10596. [Google Scholar] [CrossRef] [PubMed]
- Piccinelli, L.; Yang, Y.H.; Sakaridis, C.; Segu, M.; Li, S.; Gool, L.V.; Yu, F. UniDepth: Universal Monocular Metric Depth Estimation. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 10106–10116. [Google Scholar] [CrossRef]
- Nabati, R.; Qi, H. CenterFusion: Center-based Radar and Camera Fusion for 3D Object Detection. In Proceedings of the 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 3–8 January 2021; pp. 1526–1535. [Google Scholar] [CrossRef]
- Wu, Z.; Chen, G.; Gan, Y.; Wang, L.; Pu, J. MVFusion: Multi-View 3D Object Detection with Semantic-Aligned Radar and Camera Fusion. In Proceedings of the 2023 IEEE International Conference on Robotics and Automation (ICRA), London, UK, 29 May–2 June 2023. [Google Scholar] [CrossRef]
- Kim, Y.; Shin, J.; Kim, S.; Lee, I.J.; Choi, J.W.; Kum, D. Crn: Camera radar net for accurate, robust, efficient 3d perception. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 17615–17626. [Google Scholar]
- Lin, Z.; Liu, Z.; Xia, Z.; Wang, X.; Wang, Y.; Qi, S.; Dong, Y.; Dong, N.; Zhang, L.; Zhu, C. Rcbevdet: Radar-camera fusion in bird’s eye view for 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 14928–14937. [Google Scholar]
- Musiat, A.; Reichardt, L.; Schulze, M.; Wasenmüller, O. RadarPillars: Efficient Object Detection from 4D Radar Point Clouds. In Proceedings of the 2024 IEEE 27th International Conference on Intelligent Transportation Systems (ITSC), Edmonton, AB, Canada, 24–27 September 2024. [Google Scholar] [CrossRef]
- Long, Y.; Morris, D.; Liu, X.; Castro, M.; Chakravarty, P.; Narayanan, P. Radar-Camera Pixel Depth Association for Depth Completion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 19–25 June 2021; pp. 12507–12516. [Google Scholar]
- Li, H.; Ma, Y.; Gu, Y.; Hu, K.; Liu, Y.; Zuo, X. RadarCam-Depth: Radar-Camera Fusion for Depth Estimation with Learned Metric Scale. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; pp. 10665–10672. [Google Scholar] [CrossRef]
- Cheng, L.; Guo, L.; Zhang, T.; Bang, T.; Harris, A.; Hajij, M.; Sartipi, M.; Cao, S. CalibRefine: Deep Learning-Based Online Automatic Targetless LiDAR-Camera Calibration With Iterative and Attention-Driven Post-Refinement. IEEE Trans. Instrum. Meas. 2026, 75, 9500418. [Google Scholar] [CrossRef]
- Cheng, L.; Cao, S. Online Targetless Radar-Camera Extrinsic Calibration Based on the Common Features of Radar and Camera. In Proceedings of the NAECON 2023—IEEE National Aerospace and Electronics Conference, Dayton, OH, USA, 28–31 August 2023; pp. 294–299. [Google Scholar] [CrossRef]
- Liu, X.; Deng, Z.; Zhang, G. Targetless Radar-Camera Extrinsic Parameter Calibration Using Track-to-Track Association. Sensors 2025, 25, 949. [Google Scholar] [CrossRef] [PubMed]
- Zhu, B.; Hu, Z.; Lu, Z.; Wen, X. Trajectory-Driven Automatic Extrinsic Calibration for Roadside Radar-Camera Fusion. IEEE Sens. J. 2025, 25, 15502–15510. [Google Scholar] [CrossRef]
- Deng, J.; Hu, Z.; Lu, Z.; Wen, X. 3-D Multiple Extended Object Tracking by Fusing Roadside Radar and Camera Sensors. IEEE Sens. J. 2025, 25, 1885–1899. [Google Scholar] [CrossRef]
- Cheng, L.; Cao, S. Radar-Camera Fused Multi-Object Tracking: Online Calibration and Common Feature. IEEE Trans. Intell. Transp. Syst. 2026, 27, 1295–1311. [Google Scholar] [CrossRef]
- Ren, T.; Liu, S.; Zeng, A.; Lin, J.; Li, K.; Cao, H.; Chen, J.; Huang, X.; Chen, Y.; Yan, F.; et al. Grounded sam: Assembling open-world models for diverse visual tasks. arXiv 2024, arXiv:2401.14159. [Google Scholar] [CrossRef]
- Geiger, A.; Lenz, P.; Urtasun, R. Are we ready for autonomous driving? the kitti vision benchmark suite. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA, 16–21 June 2012; pp. 3354–3361. [Google Scholar]
- Bookstein, F.L. Principal warps: Thin-plate splines and the decomposition of deformations. IEEE Trans. Pattern Anal. Mach. Intell. 2002, 11, 567–585. [Google Scholar] [CrossRef]
- Yang, L.; Zhang, X.; Li, J.; Wang, C.; Ma, J.; Song, Z.; Zhao, T.; Song, Z.; Wang, L.; Zhou, M.; et al. V2x-radar: A multi-modal dataset with 4d radar for cooperative perception. arXiv 2024, arXiv:2411.10962. [Google Scholar]
- Yang, L.; Yu, K.; Tang, T.; Li, J.; Yuan, K.; Wang, L.; Zhang, X.; Chen, P. Bevheight: A robust framework for vision-based roadside 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 21611–21620. [Google Scholar]






| Model | Align | 0–20 m | 20–40 m | 40–60 m | 60–80 m | Overall | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| RMSE | RMSE | RMSE | RMSE | RMSE | |||||||
| DA2 | No Align | 0.251 | 10.92 | 0.000 | 19.51 | 0.005 | 26.67 | 0.000 | 35.04 | 0.004 | 20.43 |
| DA2 | Disp-space | 0.583 | 13.51 | 0.990 | 1.56 | 0.976 | 4.71 | 0.155 | 16.78 | 0.977 | 3.11 |
| DA2 | Depth-space | 0.499 | 15.75 | 0.989 | 2.19 | 0.957 | 4.72 | 0.968 | 5.35 | 0.976 | 3.36 |
| DA3 | No Align | 0.583 | 12.30 | 0.629 | 5.17 | 0.532 | 9.46 | 0.509 | 14.37 | 0.508 | 7.70 |
| DA3 | Disp-space | 0.586 | 14.83 | 0.990 | 1.90 | 0.954 | 4.56 | 0.845 | 9.57 | 0.980 | 3.16 |
| DA3 | Depth-space | 0.579 | 14.90 | 0.990 | 1.83 | 0.961 | 4.35 | 0.966 | 6.25 | 0.981 | 3.04 |
| MoGe v1 | No Align | 0.000 | 16.70 | 0.000 | 27.51 | 0.000 | 43.19 | 0.000 | 63.87 | 0.000 | 29.89 |
| MoGe v1 | Disp-space | 0.586 | 13.70 | 0.991 | 1.57 | 0.972 | 3.83 | 0.942 | 9.21 | 0.983 | 2.77 |
| MoGe v1 | Depth-space | 0.578 | 14.46 | 0.992 | 1.76 | 0.969 | 3.85 | 0.967 | 5.14 | 0.983 | 2.88 |
| MoGe v2 | No Align | 0.038 | 10.72 | 0.027 | 8.24 | 0.077 | 12.19 | 0.112 | 17.20 | 0.034 | 8.88 |
| MoGe v2 | Disp-space | 0.585 | 13.77 | 0.990 | 1.59 | 0.965 | 4.09 | 0.910 | 10.13 | 0.981 | 2.85 |
| MoGe v2 | Depth-space | 0.578 | 14.53 | 0.990 | 1.77 | 0.962 | 4.00 | 0.964 | 6.05 | 0.981 | 2.92 |
| Dataset | Methods † | ||||||
|---|---|---|---|---|---|---|---|
| CamRadRoad | Deng et al. [30] | 65.0 | 43.6 | 37.1 | 40.0 | 35.3 | 35.2 |
| Raw VFM [8] output | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | |
| VFM (Depth-aligned) [8] output | 53.2 | 28.7 | 25.3 | 30.7 | 26.0 | 26.5 | |
| VFM (Disparity-aligned) [8] output | 51.5 | 25.3 | 26.2 | 25.8 | 23.2 | 20.0 | |
| Ours (MA) | 92.3 | 30.5 | 28.2 | 33.2 | 29.7 | 28.0 | |
| Ours (MA + HF) | 82.3 | 36.7 | 31.6 | 35.9 | 32.8 | 32.6 | |
| Ours (MA + HF + CBA) | 85.5 | 42.7 | 38.8 | 38.6 | 34.7 | 33.3 | |
| Ours (MA + HF + CBA + SP) | 84.5 | 49.1 | 43.0 | 47.5 | 43.5 | 43.0 | |
| V2X-Radar-I | Deng et al. [30] | 68.5 | 48.8 | 35.9 | 42.8 | 38.4 | 37.9 |
| Raw VFM [8] output | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | |
| VFM (Depth-aligned) [8] output | 58.6 | 35.4 | 24.8 | 33.1 | 29.2 | 28.7 | |
| VFM (Disparity-aligned) [8] output | 56.9 | 37.1 | 25.3 | 34.0 | 29.7 | 29.4 | |
| Ours (MA) | 90.8 | 35.9 | 25.4 | 34.5 | 30.8 | 29.9 | |
| Ours (MA + HF) | 82.1 | 43.5 | 31.0 | 39.4 | 36.3 | 35.1 | |
| Ours (MA + HF + CBA) | 83.4 | 50.6 | 37.6 | 43.2 | 38.7 | 38.4 | |
| Ours (MA + HF + CBA + SP) | 81.7 | 55.2 | 40.9 | 46.4 | 42.2 | 40.6 | |
| Dataset | Detector | Modality | Labels | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Easy | Mod. | Hard | Easy | Mod. | Hard | ||||
| CamRadRoad | PVRCNN++ [4] | Point cloud | GT | 51.3 | 33.1 | 33.1 | 48.4 | 28.9 | 28.9 |
| Ours w/o SE | 48.4 | 28.9 | 28.9 | 46.3 | 26.4 | 26.4 | |||
| Ours w. SE | 54.0 | 35.7 | 35.7 | 49.8 | 29.8 | 29.8 | |||
| CRN [21] Backbone | Radar | GT | 56.3 | 49.8 | 49.7 | 53.3 | 44.5 | 44.5 | |
| Ours w/o SE | 52.9 | 50.0 | 49.9 | 50.0 | 43.4 | 43.3 | |||
| Ours w. SE | 59.2 | 56.6 | 56.3 | 55.5 | 53.2 | 53.2 | |||
| Modified RCBEVDet [22,36] | Camera + Radar | GT | 93.8 | 88.2 | 88.2 | 91.8 | 86.5 | 86.5 | |
| Ours w/o SE | 91.9 | 85.3 | 85.2 | 89.9 | 83.3 | 83.2 | |||
| Ours w. SE | 96.2 | 90.1 | 90.4 | 92.1 | 87.1 | 87.2 | |||
| V2X-Radar-I | PVRCNN++ [4] | Point cloud | GT | 58.6 | 42.4 | 41.8 | 52.1 | 35.6 | 35.0 |
| Ours w/o SE | 55.2 | 39.0 | 38.4 | 49.0 | 32.4 | 31.9 | |||
| Ours w. SE | 61.8 | 45.7 | 44.9 | 53.6 | 37.0 | 36.8 | |||
| CRN [21] Backbone | Radar | GT | 63.5 | 57.8 | 57.0 | 58.4 | 50.2 | 49.5 | |
| Ours w/o SE | 61.2 | 55.4 | 54.7 | 56.0 | 48.0 | 47.2 | |||
| Ours w. SE | 66.4 | 61.1 | 61.3 | 60.7 | 54.6 | 53.5 | |||
| Modified RCBEVDet [22,36] | Camera + Radar | GT | 95.0 | 90.5 | 90.0 | 92.0 | 86.8 | 86.2 | |
| Ours w/o SE | 93.2 | 88.1 | 87.5 | 90.0 | 84.2 | 83.6 | |||
| Ours w. SE | 96.6 | 92.0 | 92.3 | 93.0 | 88.2 | 87.4 | |||
| Split | Label Quality | Aligned VFM [8] Output | 2D Mask Quality | |||
|---|---|---|---|---|---|---|
| 3D IoU (%) | Hard | RMSE (m) | Recall@0.5 | Box IoU | ||
| Sun | 41.0 | 37.6 | 0.912 | 5.96 | 0.8333 | 0.6796 |
| Night | 39.7 | 37.0 | 0.945 | 5.95 | 0.8750 | 0.6682 |
| N–S | +0.033 | +0.0417 | ||||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhu, B.; Hu, Z.; Lu, Z. LiDAR-Free 3D Auto-Labeling via Radar–Visual Spatio-Temporal Consistency. Sensors 2026, 26, 2956. https://doi.org/10.3390/s26102956
Zhu B, Hu Z, Lu Z. LiDAR-Free 3D Auto-Labeling via Radar–Visual Spatio-Temporal Consistency. Sensors. 2026; 26(10):2956. https://doi.org/10.3390/s26102956
Chicago/Turabian StyleZhu, Boning, Zhiqun Hu, and Zhaoming Lu. 2026. "LiDAR-Free 3D Auto-Labeling via Radar–Visual Spatio-Temporal Consistency" Sensors 26, no. 10: 2956. https://doi.org/10.3390/s26102956
APA StyleZhu, B., Hu, Z., & Lu, Z. (2026). LiDAR-Free 3D Auto-Labeling via Radar–Visual Spatio-Temporal Consistency. Sensors, 26(10), 2956. https://doi.org/10.3390/s26102956

