AI-Enhanced Mono-View Geometry for Digital Twin 3D Visualization in Autonomous Driving †
Abstract
1. Introduction
2. AI-Enhanced Mono View Geometry Method
- Camera calibration: Using the OpenCV camera calibration code and workflow [10], we obtain the camera’s distortion coefficients and intrinsic calibration matrix, which are saved to a .pkl file. All subsequent steps can load this file to rectify images from the same lens and then apply the series of operations defined by the 3 × 3 calibration matrix K, where is the focal length of the camera, and is the principal point, i.e., the intersection of the optical axis with the image plane (near the image center).
- 2.
- Reuse the vanishing line under a fixed pose: If the relative pose between the camera and the ground plane remains fixed (i.e., the camera only translates along its Z-axis or rotates about its Y-axis), the vanishing line computed from the ground-plane normal need not be recomputed. In that case, skip step 3 and proceed directly to step 4.
- 3.
- Locate Ground-plane Vanishing Line: Let the ground-plane normal be . The corresponding vanishing line in the image coordinate system is then
- 4.
- YOLOv11 object detection: Feed the rectified 2D image into YOLOv11 to detect object classes and their 2D bounding boxes, and from each box, compute the object’s image height.
- 5.
- Height ratio analysis: For each detected object on the ground plane, obtain its image height and use the vanishing line from Step 3 to compute the object’s real-world height ratio [11].
- 6.
- Real-world height estimation: Using a reference object with a known real-world height that stands perpendicular to the ground plane, together with the height ratios from the previous step, compute each target object’s true height .
- 7.
- 3D positional coordinate calculation: As shown in Figure 3, given the camera focal length f, the object’s image height , and its real height , the object-to-camera distance follows by similar triangles:
- 8.
- License-plate-based depth refinement: As illustrated in Figure 4a,b, in close-range scenes, perspective distortion causes the car bounding box (green box) to include both the roof and the hood, resulting in an observed height (green arrow) that is greater than the ideal height (red arrow), which represents the true vertical span from the roof to the ground in the image. Using this overestimated height to compute depth yields an inaccurate result, often placing somewhere between the front bumper and the actual center of the vehicle . To resolve this, we detect the license plate (yellow box) within the car’s 2D bounding box; if successful, we take the license plate box height as the new image-space reference height , and use the standard physical height of a license plate ( = 16 cm in Taiwan) to recompute depth using a similarity triangle formulation. This yields a refined depth estimate , which reflects the minimum distance from the vehicle to the camera, and replaces the original depth value to improve 3D localization accuracy in close-range scenes.
- 9.
- 2D/3D image output: Finally, we output 2D annotated images that display each detected object’s class and its 3D positional coordinates, and 3D renderings visualizing the driving-scene digital twin model.
3. Results
- Roof-to-ground: Use the 2D bounding box height as the image height , and then apply the height ratio analysis to estimate the car’s actual height, which is the value of for depth computation.
- License plate: Use the license plate height as and the known plate height (16 cm) as to compute depth accordingly.
4. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Yan, L.; Wu, X.; Wei, C.; Zhao, S. Human-Vehicle Shared Steering Control for Obstacle Avoidance: A Reference-Free Approach with Reinforcement Learning. IEEE Trans. Intell. Transp. Syst. 2024, 25, 17888–17901. [Google Scholar] [CrossRef]
- Mao, J.; Shi, S.; Wang, X.; Li, H. 3D Object Detection for Autonomous Driving: A Comprehensive Survey. Int. J. Comput. Vis. 2023, 131, 1909–1963. [Google Scholar] [CrossRef]
- Pan, H.; Jia, Y.; Wang, J.; Sun, W. MonoAMNet: Three-Stage Real-Time Monocular 3D Object Detection with Adaptive Methods. IEEE Trans. Intell. Transp. Syst. 2025, 26, 3574–3587. [Google Scholar] [CrossRef]
- Xia, C.; Zhao, W.; Han, H.; Tao, Z.; Ge, B.; Gao, X.; Li, K.-C.; Zhang, Y. MonoSAID: Monocular 3D Object Detection Based on Scene-Level Adaptive Instance Depth Estimation. J. Intell. Robot. Syst. 2024, 110, 2. [Google Scholar] [CrossRef]
- Gao, Y.; Wang, P.; Li, X.; Sun, M.; Di, R.; Li, L.; Hong, W. MonoDFNet: Monocular 3D Object Detection with Depth Fusion and Adaptive Optimization. Sensors 2025, 25, 760. [Google Scholar] [CrossRef]
- Cheng, Z.; Zhang, Y. A Comparative Analysis of Traditional and CNN-Based Object Recognition Techniques in Robotics. In Proceedings of the 2024 IEEE 7th Information Technology, Networking, Electronic and Automation Control Conference (ITNEC), Chongqing, China, 20–22 September 2024. [Google Scholar]
- Zhang, H.; Ge, S.; Luo, G.; Tian, Y.; Ye, P.; Li, Y. Internet of Vehicular Intelligence: Enhancing Connectivity and Autonomy in Smart Transportation Systems. IEEE Trans. Intell. Veh. 2024, 1–5. [Google Scholar] [CrossRef]
- Barricelli, B.R.; Casiraghi, E.; Fogli, D. A Survey on Digital Twin: Definitions, Characteristics, Applications, and Design Implications. IEEE Access 2019, 7, 167653–167671. [Google Scholar] [CrossRef]
- Khanam, R.; Hussain, M. YOLOv11: An Overview of the Key Architectural Enhancements. arXiv 2024, arXiv:2410.17725. [Google Scholar] [CrossRef]
- OpenCV: Camera Calibration. Available online: https://docs.opencv.org/4.x/dc/dbb/tutorial_py_calibration.html (accessed on 19 July 2025).
- Hartley, R.; Zisserman, A. Multiple View Geometry in Computer Vision, 2nd ed.; Cambridge University Press: New York, NY, USA, 2003. [Google Scholar]
- Ultralytics. Integrations—TensorRT: NVIDIA A100. Available online: https://docs.ultralytics.com/integrations/tensorrt/#nvidia-a100 (accessed on 19 July 2025).
- Agrawal, S. Global License Plate Dataset. arXiv 2024, arXiv:2405.10949. [Google Scholar] [CrossRef]






| Method | AI-Enhanced Mono-View Geometry (This Study) | MonoAMNet | MonoSAID | MonoDFNet |
|---|---|---|---|---|
| Property | ||||
| Low computational complexity | Yes | No | No | No |
| Interpretability | Yes | No | No | No |
| No large dataset required | Yes | No | No | No |
| High generalization in data-scarce and diverse scenarios | Yes | No | No | No |
| Height | Value |
|---|---|
| Motorcycle height (seat → ground) | 80 cm |
| Car height (roof → ground) | 143.5 cm |
| Shot | Depth (Z-Axis) Distances | Value |
|---|---|---|
| 1 | Camera → motorcycle | 375.5 cm |
| Camera → car | 743.8 cm | |
| 2 | camera → motorcycle | 534.8 cm |
| camera → car | 885 cm | |
| 3 | camera → motorcycle | 715 cm |
| camera → car | 1057.6 cm | |
| 4 | camera → motorcycle | 906.3 cm |
| camera → car | 1214.9 cm |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Chang, I.-C.; Chang, Y.-C.; Kuo, C.; Yen, C.-E. AI-Enhanced Mono-View Geometry for Digital Twin 3D Visualization in Autonomous Driving. Eng. Proc. 2025, 120, 6. https://doi.org/10.3390/engproc2025120006
Chang I-C, Chang Y-C, Kuo C, Yen C-E. AI-Enhanced Mono-View Geometry for Digital Twin 3D Visualization in Autonomous Driving. Engineering Proceedings. 2025; 120(1):6. https://doi.org/10.3390/engproc2025120006
Chicago/Turabian StyleChang, Ing-Chau, Yu-Chiao Chang, Chunghui Kuo, and Chin-En Yen. 2025. "AI-Enhanced Mono-View Geometry for Digital Twin 3D Visualization in Autonomous Driving" Engineering Proceedings 120, no. 1: 6. https://doi.org/10.3390/engproc2025120006
APA StyleChang, I.-C., Chang, Y.-C., Kuo, C., & Yen, C.-E. (2025). AI-Enhanced Mono-View Geometry for Digital Twin 3D Visualization in Autonomous Driving. Engineering Proceedings, 120(1), 6. https://doi.org/10.3390/engproc2025120006

