Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (82)

Search Parameters:
Keywords = monocular 3D object detection

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
26 pages, 4175 KB  
Article
Graph Enhanced Multi-Modal Network of 4-D Radar-Camera Fusion for Perception in Autonomous Systems
by Yuanzhi Deng, Cheng Chi, Jianhao Shen, Yu Han, Shanyin He and Shaolong Chen
Sensors 2026, 26(14), 4635; https://doi.org/10.3390/s26144635 - 22 Jul 2026
Abstract
Modern autonomous systems rely on heterogeneous sensing modalities—including vision sensors, millimeter-wave radar, and LiDAR—yet each individual sensor exhibits characteristic failure modes in challenging real-world conditions. While LiDAR-vision co-processing has received extensive attention, the synergistic potential of 4D radar paired with monocular optics remains [...] Read more.
Modern autonomous systems rely on heterogeneous sensing modalities—including vision sensors, millimeter-wave radar, and LiDAR—yet each individual sensor exhibits characteristic failure modes in challenging real-world conditions. While LiDAR-vision co-processing has received extensive attention, the synergistic potential of 4D radar paired with monocular optics remains comparatively unexplored. To fill this gap, we develop a graph-enhanced multi-modal architecture that jointly leverages sparse 4D radar returns and high-resolution camera imagery for scene-level 3D perception. The proposed system is organized around four tightly coupled processing stages: (i) an image-guided point densification scheme (SAA) that augments sparse radar clouds with camera-derived pseudo measurements; (ii) a pose-invariant cross-modal fusion layer that harmonizes enriched radar features with image descriptors and object saliency maps; (iii) a dynamic hypergraph assembly stage that captures higher-order inter-object and cross-sensor dependencies; and (iv) a HyperGCN inference module that regresses 3D bounding parameters and class labels on the resulting relational graph. Integrating the temporal velocity cues native to radar with the rich appearance information from cameras, the system delivers reliable perception under diverse environmental conditions. On the View-of-Delft (VOD) evaluation suite, the proposed model records an mAP of 69.3 and mAOS of 59.8. A systematic ablation further quantifies how geometric invariance—across translation, rotation, and scale transformations—individually affects end-to-end detection fidelity. Full article
(This article belongs to the Topic Advances in Autonomous Vehicles, Automation, and Robotics)
Show Figures

Figure 1

27 pages, 24147 KB  
Article
IntelligentVehicle Security: Real-Time Anomaly Detection and Anti-Theft Surveillance Using Monocular Depth Estimation and Behavioral Analysis
by Umar Adeel, Ammar Rashid, Shafiz Affendi Bin Mohd Yusof and Usman Javed Butt
Information 2026, 17(7), 676; https://doi.org/10.3390/info17070676 - 12 Jul 2026
Viewed by 212
Abstract
Vehicle theft and vandalism remain significant urban security challenges commonly addressed through reactive, post-incident forensic measures. This paper proposes a proactive, real-time computer vision system designed to detect potentially suspicious behavior around parked vehicles, with a specific focus on unauthorized proximity and loitering. [...] Read more.
Vehicle theft and vandalism remain significant urban security challenges commonly addressed through reactive, post-incident forensic measures. This paper proposes a proactive, real-time computer vision system designed to detect potentially suspicious behavior around parked vehicles, with a specific focus on unauthorized proximity and loitering. The proposed architecture integrates state-of-the-art object detection using YOLOv11 (You Only Look Once version 11), multi-object tracking via a lightweight custom association tracker inspired by the ByteTrack/StrongSORT/OC-SORT paradigm, and monocular depth estimation based on the Intel DPT-Large framework.A key contribution is the identification and mitigation of the Perspective Challenge: the two-dimensional (2D) scale ambiguity that causes distant background pedestrians to appear falsely proximate to foreground vehicles in monocular camera feeds. To address this, three spatial analysis strategies are implemented and evaluated: (A) fixed Euclidean thresholding, (B) adaptive perspective thresholding, and (C) three-dimensional (3D) depth injection. Experimental results on real-world urban surveillance footage (27,000 annotated frames across two datasets) demonstrate that Strategy C achieves the highest precision (0.95) with an F1-score of 0.92, while Strategy B provides the best balance between accuracy (precision 0.88, recall 0.91, F1 0.89) and computational efficiency (32.7 frames per second, FPS). Compared to naive 2D thresholding (Strategy A), Strategy B reduces false alarms by approximately 80%, while Strategy C further improves precision to 0.95 through depth-plane verification. The system maintains real-time performance exceeding 30 FPS under Strategy B, making it a strong candidate for practical urban vehicle monitoring, subject to further large-scale validation across diverse environments. Full article
(This article belongs to the Special Issue Generative AI for Data Privacy and Anomaly Detection)
Show Figures

Figure 1

24 pages, 955 KB  
Review
Sensor Fusion and Perception for Autonomous Driving: A Critical Review of Modalities, AI Models, Algorithms, and Industry Configurations
by Esraa Khatab, Fares Fathy, Abdallah AlKholy and Omar Shalash
Mach. Learn. Knowl. Extr. 2026, 8(7), 199; https://doi.org/10.3390/make8070199 - 7 Jul 2026
Viewed by 403
Abstract
Autonomous driving systems rely on a sophisticated pipeline of artificial intelligence models to perceive, predict, and plan in dynamic environments. This review presents a systematic analysis of the machine learning and deep learning models underpinning vehicle autonomy, spanning classical convolutional neural networks (CNNs) [...] Read more.
Autonomous driving systems rely on a sophisticated pipeline of artificial intelligence models to perceive, predict, and plan in dynamic environments. This review presents a systematic analysis of the machine learning and deep learning models underpinning vehicle autonomy, spanning classical convolutional neural networks (CNNs) for object detection and semantic segmentation to recurrent and Transformer-based architectures for trajectory prediction and motion planning. It also provides a critical examination of the autonomous vehicle sensor stack, including cameras, LiDAR, radar, ultrasonics, and GNSS/IMU as data acquisition systems, highlighting modality-specific AI challenges such as monocular depth estimation, 3D point cloud processing, and radar Doppler interpretation. The evolution of perception and decision-making pipelines is reviewed, contrasting modular architectures with end-to-end learning paradigms that directly map raw sensor data to control commands, and discussing their trade-offs in interpretability, safety assurance, and robustness to rare edge cases. We further survey specialized hardware accelerators and heterogeneous automotive SoCs designed to meet stringent real-time and power constraints. Industrial strategies are compared, including multi-modal sensor fusion and vision-centric approaches based on large-scale imitation learning. Finally, we identify open challenges related to robustness under adverse conditions, domain shift, causal ambiguity, and the need for interpretable and certifiable AI in safety-critical autonomous driving systems. Full article
Show Figures

Figure 1

22 pages, 6385 KB  
Article
Targetless Calibration of Wide-Baseline and Wide-Angle Surround-View Fisheye Cameras Using Cylindrical Projection Model
by Gee Hoon Lee and Soon-Yong Park
Sensors 2026, 26(12), 3622; https://doi.org/10.3390/s26123622 - 6 Jun 2026
Viewed by 412
Abstract
We propose a novel targetless extrinsic calibration method for wide-baseline and wide-angle fisheye cameras, which are mounted on a driving vehicle for surround view monitoring. Sequences of image frames from three fisheye cameras are obtained, and the object instance and depth around the [...] Read more.
We propose a novel targetless extrinsic calibration method for wide-baseline and wide-angle fisheye cameras, which are mounted on a driving vehicle for surround view monitoring. Sequences of image frames from three fisheye cameras are obtained, and the object instance and depth around the vehicle are used for calibration. Thus, the proposed method can be applied to online vehicle camera calibration. Fisheye images are first transformed into the cylindrical coordinate system by considering the panoramic formation of the cameras. Then, the state-of-the-art object detection and monocular depth estimation models are applied to the cylindrical images. Vehicle instances matched across different views are reconstructed into 3D point clouds, and their depths are scaled by employing the pose geometry of the front camera. The per-point depths and global scale are then jointly optimized to achieve accurate cross-view alignment and extrinsic calibration. Experiments on both real-world and synthetic video datasets show that the proposed method achieves higher accuracy than COLMAP and DUSt3R under challenging conditions such as wide baselines and low frame rates, without requiring an artificial calibration target. Full article
Show Figures

Figure 1

16 pages, 58544 KB  
Article
D3SSTrack: Center-Focused State-Space Modeling for Monocular 3D Multi-Object Tracking
by Darius-Ovidiu Firan and Călin-Adrian Popa
Mathematics 2026, 14(10), 1737; https://doi.org/10.3390/math14101737 - 18 May 2026
Viewed by 306
Abstract
Monocular 3D multi-object tracking (3D MOT) remains challenging because it is hard to model how objects move over time and to keep correct identities without explicit depth information. In this context, we introduce D3SSTrack, a novel tracking-by-detection framework that integrates Mamba state-space modeling [...] Read more.
Monocular 3D multi-object tracking (3D MOT) remains challenging because it is hard to model how objects move over time and to keep correct identities without explicit depth information. In this context, we introduce D3SSTrack, a novel tracking-by-detection framework that integrates Mamba state-space modeling into the 3D tracking pipeline. At its core is the Solid State Track (SST) block, which extends the original Mamba block with dropout regularization and an additional projection layer to improve feature integration before temporal fusion. This design enables efficient modeling of long-range temporal dependencies while maintaining real-time performance at 38 FPS on a single GPU. The proposed tracker combines structured sequence modeling with effective temporal association, improving robustness against occlusions and abrupt motion changes. On the KITTI benchmark, D3SSTrack achieves the best sAMOTA (97.12%) and AMOTA (49.95%) among recent monocular 3D MOT methods, outperforming the best model S3MOT by 0.16% and 0.22%, respectively. Our results highlight the potential of state space-based architectures for real-world monocular 3D MOT applications. Full article
Show Figures

Figure 1

21 pages, 12844 KB  
Article
Unsupervised Domain Adaptation with Multimodal Fusion for Monocular 3D Object Detection
by Jin Jiang, Jidong Dai, Wei Li, Yuquan Zhou, Maozhang Ye, Jianhuan Zhang and Chentao Zhang
Vehicles 2026, 8(5), 98; https://doi.org/10.3390/vehicles8050098 - 1 May 2026
Viewed by 475
Abstract
This paper presents UM3D, an end-to-end unsupervised domain adaptation framework for monocular 3D object detection. Monocular 3D object detection is appealing due to its low cost, yet it suffers from limited depth cues and poor cross-domain generalization when labeled data are scarce. Existing [...] Read more.
This paper presents UM3D, an end-to-end unsupervised domain adaptation framework for monocular 3D object detection. Monocular 3D object detection is appealing due to its low cost, yet it suffers from limited depth cues and poor cross-domain generalization when labeled data are scarce. Existing Pseudo-LiDAR methods require supervised training and propagate depth estimation errors to downstream detection, while current unsupervised domain adaptation (UDA) approaches exploit only a single modality and lack effective pseudo-label quality control. UM3D addresses these limitations through two key designs: (1) a quality-aware pseudo-label generation strategy with object-level random scaling and a memory bank refinement mechanism; and (2) an end-to-end differentiable pipeline that integrates multimodal fusion of image and Pseudo-LiDAR features with a multi-network consistency loss, which jointly optimizes depth estimation and 3D detection via backpropagation. Notably, the entire pipeline requires only a single monocular camera at inference; the Pseudo-LiDAR representation is generated internally from the same image, and thus the multimodal fusion integrates image and Pseudo-LiDAR features without requiring additional sensors. Extensive experiments across KITTI, nuScenes, Waymo, and Lyft demonstrate that UM3D generally outperforms existing UDA methods. In particular, a 19.30% relative APBEV improvement is achieved under easy conditions through end-to-end joint training compared to independent depth estimation, and up to 76.81% of the domain gap is closed on the WOD → KITTI benchmark. Full article
(This article belongs to the Section Intelligent and Connected Mobility)
Show Figures

Figure 1

19 pages, 30364 KB  
Article
CLIP-Mono3D: End-to-End Open-Vocabulary Monocular 3D Object Detection via Semantic–Geometric Similarity
by Zichong Gu, Shiyi Mu, Hanqi Lyu and Shugong Xu
Sensors 2026, 26(8), 2380; https://doi.org/10.3390/s26082380 - 13 Apr 2026
Viewed by 915
Abstract
Open-vocabulary 3D object detection (OV-3DOD) is crucial for real-world perception, yet existing monocular methods are often limited by predefined categories or heavy reliance on external 2D detectors. In this paper, we propose CLIP-Mono3D, an end-to-end one-stage transformer framework that directly integrates vision–language semantics [...] Read more.
Open-vocabulary 3D object detection (OV-3DOD) is crucial for real-world perception, yet existing monocular methods are often limited by predefined categories or heavy reliance on external 2D detectors. In this paper, we propose CLIP-Mono3D, an end-to-end one-stage transformer framework that directly integrates vision–language semantics into monocular 3D detection. By leveraging CLIP-derived semantic priors and grounding object queries in semantically salient regions, our model achieves robust zero-shot generalization to novel categories without requiring auxiliary 2D detectors. Furthermore, we introduce OV-KITTI, a large-scale benchmark extending KITTI with 40 new categories and over 7000 annotated 3D bounding boxes. Extensive experiments on OV-KITTI, KITTI, and Argoverse demonstrate that CLIP-Mono3D achieves competitive performance in open-vocabulary scenarios. Full article
(This article belongs to the Section Sensing and Imaging)
Show Figures

Figure 1

38 pages, 3132 KB  
Article
Lightweight Semantic-Aware Route Planning on Edge Hardware for Indoor Mobile Robots: Monocular Camera–2D LiDAR Fusion with Penalty-Weighted Nav2 Route Server Replanning
by Bogdan Felician Abaza, Andrei-Alexandru Staicu and Cristian Vasile Doicin
Sensors 2026, 26(7), 2232; https://doi.org/10.3390/s26072232 - 4 Apr 2026
Viewed by 2075
Abstract
The paper introduces a computationally efficient semantic-aware route planning framework for indoor mobile robots, designed for real-time execution on resource-constrained edge hardware (Raspberry Pi 5, CPU-only). The proposed architecture fuses monocular object detection with 2D LiDAR-based range estimation and integrates the resulting semantic [...] Read more.
The paper introduces a computationally efficient semantic-aware route planning framework for indoor mobile robots, designed for real-time execution on resource-constrained edge hardware (Raspberry Pi 5, CPU-only). The proposed architecture fuses monocular object detection with 2D LiDAR-based range estimation and integrates the resulting semantic annotations into the Nav2 Route Server for penalty-weighted route selection. Object localization in the map frame is achieved through the Angular Sector Fusion (ASF) pipeline, a deterministic geometric method requiring no parameter tuning. The ASF projects YOLO bounding boxes onto LiDAR angular sectors and estimates the object range using a 25th-percentile distance statistic, providing robustness to sparse returns and partial occlusions. All intrinsic and extrinsic sensor parameters are resolved at runtime via ROS 2 topic introspection and the URDF transform tree, enabling platform-agnostic deployment. Detected entities are classified according to mobility semantics (dynamic, static, and minor) and persistently encoded in a GeoJSON-based semantic map, with these annotations subsequently propagated to navigation graph edges as additive penalties and velocity constraints. Route computation is performed by the Nav2 Route Server through the minimization of a composite cost functional combining geometric path length with semantic penalties. A reactive replanning module monitors semantic cost updates during execution and triggers route invalidation and re-computation when threshold violations occur. Experimental evaluation over 115 navigation segments (legs) on three heterogeneous robotic platforms (two single-board RPi5 configurations and one dual-board setup with inference offloading) yielded an overall success rate of 97% (baseline: 100%, adaptive: 94%), with 42 replanning events observed in 57% of adaptive trials. Navigation time distributions exhibited statistically significant departures from normality (Shapiro–Wilk, p < 0.005). While central tendency differences between the baseline and adaptive modes were not significant (Mann–Whitney U, p = 0.157), the adaptive planner reduced temporal variance substantially (σ = 11.0 s vs. 31.1 s; Levene’s test W = 3.14, p = 0.082), primarily by mitigating AMCL recovery-induced outliers. On-device YOLO26n inference, executed via the NCNN backend, achieved 5.5 ± 0.7 FPS (167 ± 21 ms latency), and distributed inference reduced the average system CPU load from 85% to 48%. The study further reports deployment-level observations relevant to the Nav2 ecosystem, including GeoJSON metadata persistence constraints, graph discontinuity (“path-gap”) artifacts, and practical Route Server configuration patterns for semantic cost integration. Full article
(This article belongs to the Special Issue Advances in Sensing, Control and Path Planning for Robotic Systems)
Show Figures

Figure 1

30 pages, 135773 KB  
Article
Robust 3D Multi-Object Tracking via 4D mmWave Radar-Camera Fusion and Disparity-Domain Depth Recovery
by Yunfei Xie, Xiaohui Li, Dingheng Wang, Zhuo Wang, Shiliang Li, Jia Wang and Zhenping Sun
Sensors 2026, 26(7), 2096; https://doi.org/10.3390/s26072096 - 27 Mar 2026
Cited by 1 | Viewed by 1092
Abstract
4D millimeter-wave radar provides high-precision ranging capability and exhibits strong robustness under adverse weather and low-visibility conditions, but its point clouds are relatively sparse and suffer from severe elevation-angle measurement noise. Monocular cameras, by contrast, provide rich semantic information and high recall, yet [...] Read more.
4D millimeter-wave radar provides high-precision ranging capability and exhibits strong robustness under adverse weather and low-visibility conditions, but its point clouds are relatively sparse and suffer from severe elevation-angle measurement noise. Monocular cameras, by contrast, provide rich semantic information and high recall, yet are fundamentally limited by scale ambiguity. To exploit the complementary characteristics of these two sensors, this paper proposes a radar-camera fusion 3D multi-object tracking framework that does not rely on complex 3D annotated data. First, on the radar signal-processing side, a Gaussian distribution-based adaptive angle compression method and IMU-based velocity compensation are introduced to effectively suppress measurement noise, and an improved DBSCAN clustering scheme with recursive cluster splitting and historical static-box guidance is employed to generate high-quality radar detections. Second, a disparity-domain metric depth recovery method is proposed. This method uses filtered radar points as sparse metric anchors, performs robust fitting with RANSAC, and applies Kalman filtering for temporal smoothing, thereby converting the relative depth output of the visual foundation model Depth Anything V2 into metric depth. Finally, a hierarchical fusion strategy is designed at both the detection and tracking levels to achieve stable cross-modal state association. Experimental results on a self-collected dataset show that the proposed method achieves an overall MOTA of 77.93%, outperforming single-modality baselines and other comparison methods by 11 to 31 percentage points. This study provides an effective solution for low-cost and robust environment perception in complex dynamic scenarios. Full article
(This article belongs to the Section Vehicular Sensing)
Show Figures

Figure 1

21 pages, 1669 KB  
Article
Robust BEV Perception via Dual 4D Radar–Camera Fusion Under Adverse Conditions with Fog-Aware Enhancement
by Zhengqing Li and Baljit Singh
Electronics 2026, 15(6), 1284; https://doi.org/10.3390/electronics15061284 - 19 Mar 2026
Viewed by 1158
Abstract
Bird’s-eye-view (BEV) perception has emerged as a key representation for unified scene understanding in autonomous driving. However, current BEV methods relying solely on monocular cameras suffer from severe degradation under adverse weather and dynamic scenes due to limited depth cues and illumination dependency. [...] Read more.
Bird’s-eye-view (BEV) perception has emerged as a key representation for unified scene understanding in autonomous driving. However, current BEV methods relying solely on monocular cameras suffer from severe degradation under adverse weather and dynamic scenes due to limited depth cues and illumination dependency. To address these challenges, we propose a robust multi-modal BEV perception framework that integrates dual-source 4D millimeter-wave radar and multi-view camera images. The proposed architecture systematically exploits Doppler velocity and temporal information from 4D radar to model dynamic object motion, while introducing a deformable fusion strategy in the BEV space for accurate semantic alignment across modalities. Our design includes four key modules: a Doppler-Aware Radar Encoder (DARE) that enhances motion-sensitive features via velocity-guided attention; a Fog-Aware Feature Denoising Module (FADM) that suppresses modality inconsistency in low-visibility conditions through cross-modal attention and residual enhancement; a Multi-Modal Temporal Fusion Module (TFM) that encodes radar temporal sequences using a Transformer encoder for motion continuity modeling; and a confidence-aware multi-task loss that jointly supervises semantic segmentation, motion estimation, and object detection. Extensive experiments on the DualRadar dataset and adverse-weather simulations demonstrate that our method achieves significant gains over state-of-the-art baselines in BEV segmentation accuracy, detection robustness, and motion stability. The proposed framework offers a scalable and resilient solution for real-world autonomous perception, especially under challenging environmental conditions. Full article
(This article belongs to the Special Issue Image Processing Based on Convolution Neural Network: 2nd Edition)
Show Figures

Figure 1

25 pages, 8614 KB  
Article
Underwater Image Restoration Integrating Monocular Depth Estimation with a Physical Imaging Model
by Tianchi Zhang, Hongwei Qin, Qiang Liu and Xing Liu
J. Mar. Sci. Eng. 2026, 14(6), 563; https://doi.org/10.3390/jmse14060563 - 18 Mar 2026
Cited by 1 | Viewed by 663
Abstract
Underwater images suffer from quality degradation such as haze, detail blurring, color distortion, and low contrast due to factors like light scattering and wavelength-dependent attenuation in water. This severely hinders the high-quality completion of target detection tasks for Autonomous Underwater Vehicles (AUV) relying [...] Read more.
Underwater images suffer from quality degradation such as haze, detail blurring, color distortion, and low contrast due to factors like light scattering and wavelength-dependent attenuation in water. This severely hinders the high-quality completion of target detection tasks for Autonomous Underwater Vehicles (AUV) relying on image information. Although deep learning-based methods have gained widespread attention, existing approaches still face challenges such as insufficient feature extraction and limited generalization in complex real-world scenes. Methods based on physical models, on the other hand, heavily rely on depth information which is difficult to obtain accurately. To address these issues, this paper proposes a novel underwater image restoration method that integrates depth estimation with the Akkaynak-Treibitz physical imaging model. In the depth estimation stage, efficient and robust feature extraction is achieved through a lightweight encoder–decoder architecture combined with a channel–spatial hybrid attention mechanism. To overcome the inherent scale ambiguity problem in monocular depth estimation, which prevents direct output of absolute depth consistent with the real scene, sparse depth priors are introduced. Subsequently, adaptive depth binning and depth map optimization are realized via m-Vision Transformer and convolutional regression. In the image restoration stage, the acquired high-quality depth map is combined with the Akkaynak-Treibitz physical imaging model for inverse solving, achieving high-quality restoration from degraded to clear images. Experimental results demonstrate that the proposed method outperforms mainstream depth estimation methods (LapDepth, UDepth, etc.) and mainstream image restoration methods (CLAHE, FUnIE-GAN, etc.) in terms of evaluation metrics and visual perceptual quality. When processing the extremely degraded UIEB-S dataset, the proposed method achieves evaluation metrics of SSIM = 0.8954, UCIQE = 0.6107, and PSNR = 23.35 dB. Compared to the CLAHE and FUnIE-GAN methods, SSIM improved by 2.8% and 16.7%, UCIQE improved by 9.6% and 14.3%, and PSNR improved by 22.5% and 13.9%, respectively. Comprehensive subjective and objective evaluation results validate the effectiveness of the proposed method in addressing image quality degradation, particularly demonstrating outstanding capability in severe color cast correction and detail recovery. Full article
(This article belongs to the Section Ocean Engineering)
Show Figures

Figure 1

9 pages, 2357 KB  
Proceeding Paper
AI-Enhanced Mono-View Geometry for Digital Twin 3D Visualization in Autonomous Driving
by Ing-Chau Chang, Yu-Chiao Chang, Chunghui Kuo and Chin-En Yen
Eng. Proc. 2025, 120(1), 6; https://doi.org/10.3390/engproc2025120006 - 25 Dec 2025
Viewed by 846
Abstract
To address the critical problem of 3D object detection in autonomous driving scenarios, we developed a novel digital twin architecture. This architecture combines AI models with geometric optics algorithms of camera systems for autonomous vehicles, characterized by low computational cost and high generalization [...] Read more.
To address the critical problem of 3D object detection in autonomous driving scenarios, we developed a novel digital twin architecture. This architecture combines AI models with geometric optics algorithms of camera systems for autonomous vehicles, characterized by low computational cost and high generalization capability. The architecture leverages monocular images to estimate the real-world heights and 3D positions of objects using vanishing lines and the pinhole camera model. The You Only Look Once (YOLOv11) object detection model is employed for accurate object category identification. These components are seamlessly integrated to construct a digital twin system capable of real-time reconstruction of the surrounding 3D environment. This enables the autonomous driving system to perform real-time monitoring and optimized decision-making. Compared with conventional deep-learning-based 3D object detection models, the architecture offers several notable advantages. Firstly, it mitigates the significant reliance on large-scale labeled datasets typically required by deep learning approaches. Secondly, its decision-making process inherently provides interpretability. Thirdly, it demonstrates robust generalization capabilities across diverse scenes and object types. Finally, its low computational complexity makes it particularly well-suited for resource-constrained in-vehicle edge devices. Preliminary experimental results validate the reliability of the proposed approach, showing a depth prediction error of less than 5% in driving scenarios. Furthermore, the proposed method achieves significantly faster runtime, corresponding to only 42, 27, and 22% of MonoAMNet, MonoSAID, and MonoDFNet, respectively. Full article
(This article belongs to the Proceedings of 8th International Conference on Knowledge Innovation and Invention)
Show Figures

Figure 1

23 pages, 6012 KB  
Article
A Pseudo-Point-Based Adaptive Fusion Network for Multi-Modal 3D Detection
by Chenghong Zhang, Wei Wang, Bo Yu and Hanting Wei
Electronics 2026, 15(1), 59; https://doi.org/10.3390/electronics15010059 - 23 Dec 2025
Viewed by 756
Abstract
A 3D multi-modal detection method using a monocular camera and LiDAR has drawn much attention due to its low cost and strong applicability, making it highly valuable for autonomous driving and unmanned aerial vehicles (UAVs). However, conventional fusion approaches relying on static arithmetic [...] Read more.
A 3D multi-modal detection method using a monocular camera and LiDAR has drawn much attention due to its low cost and strong applicability, making it highly valuable for autonomous driving and unmanned aerial vehicles (UAVs). However, conventional fusion approaches relying on static arithmetic operations often fail to adapt to dynamic, complex scenarios. Furthermore, existing ROI alignment techniques, such as local projection and cross-attention, are inadequate for mitigating the feature misalignment triggered by depth estimation noise in pseudo-point clouds. To address these issues, this paper proposes a pseudo-point-based 3D object detection method that achieves biased fusion of multi-modal data. First, a meta-weight fusion module dynamically generates fusion weights based on global context, adaptively balancing the contributions of point clouds and images. Second, a module combining bidirectional cross-attention and a gating filter mechanism is introduced to eliminate the ROI feature misalignment caused by depth completion noise. Finally, a class-agnostic box fusion strategy is introduced to aggregate highly overlapping detection boxes at the decision level, improving localization accuracy. Experiments on the KITTI dataset show that the proposed method achieves APs of 92.22%, 85.03%, and 82.25% on Easy, Moderate, and Hard difficulty levels, respectively, demonstrating leading performance. Ablation studies further validate the effectiveness and computational efficiency of each module. Full article
Show Figures

Figure 1

20 pages, 10328 KB  
Article
Toward Autonomous Pavement Inspection: An End-to-End Vision-Based Framework for PCI Computation and Robotic Deployment
by Nada El Desouky, Ahmed A. Torky, Mohamed Elbheiri, Mohamed S. Eid and Mohamed Ibrahim
Automation 2025, 6(4), 67; https://doi.org/10.3390/automation6040067 - 4 Nov 2025
Cited by 1 | Viewed by 1854
Abstract
Advancements in robotics and computer vision are transforming how infrastructure is monitored and maintained. This paper presents a novel, fully automated pipeline for pavement condition assessment that integrates real-time image analysis with PCI (Pavement Condition Index) computation, which is specifically designed for deployment [...] Read more.
Advancements in robotics and computer vision are transforming how infrastructure is monitored and maintained. This paper presents a novel, fully automated pipeline for pavement condition assessment that integrates real-time image analysis with PCI (Pavement Condition Index) computation, which is specifically designed for deployment on mobile and robotic platforms. Unlike traditional methods that rely on costly equipment or manual input, the proposed system uses deep learning-based object detection and ensemble segmentation to identify and measure multiple types of road distress directly from 2D imagery, including surface weathering, a key precursor to pothole formation often overlooked in previous studies. Depth estimation is achieved using a monocular diffusion model, enabling volumetric assessment without specialized sensors. Validated on real-world footage captured by a smartphone, the pipeline demonstrated reliable performance across detection, measurement, and scoring stages. Its potential hardware-agnostic design and modular architecture position it as a practical solution for autonomous inspection by drones or ground robots in future smart infrastructure systems. Full article
(This article belongs to the Section Robotics and Autonomous Systems)
Show Figures

Figure 1

14 pages, 13455 KB  
Article
Enhancing 3D Monocular Object Detection with Style Transfer for Nighttime Data Augmentation
by Alexandre Evain, Firas Jendoubi, Redouane Khemmar, Sofiane Ahmedali and Mathieu Orzalesi
Appl. Sci. 2025, 15(20), 11288; https://doi.org/10.3390/app152011288 - 21 Oct 2025
Cited by 2 | Viewed by 1419
Abstract
Monocular 3D object detection (Mono3D) is essential for autonomous driving and augmented reality, yet its performance degrades significantly at night due to the scarcity of annotated nighttime data. In this paper, we investigate the use of style transfer for nighttime data augmentation and [...] Read more.
Monocular 3D object detection (Mono3D) is essential for autonomous driving and augmented reality, yet its performance degrades significantly at night due to the scarcity of annotated nighttime data. In this paper, we investigate the use of style transfer for nighttime data augmentation and evaluate its effect on individual components of 3D detection. Using CycleGAN, we generated synthetic night images from daytime scenes in the nuScenes dataset and trained a modular Mono3D detector under different configurations. Our results show that training solely on style-transferred images improves certain metrics, such as AP@0.95 (from 0.0299 to 0.0778, a 160% increase) and depth error (11% reduction), compared to daytime-only baselines. However, performance on orientation and dimension estimation deteriorates. When real nighttime data is included, style transfer provides complementary benefits: for cars, depth error decreases from 0.0414 to 0.021, and AP@0.95 remains stable at 0.66; for pedestrians, AP@0.95 improves by 13% (0.297 to 0.336) with a 35% reduction in depth error. Cyclist detection remains unreliable due to limited samples. We conclude that style transfer cannot replace authentic nighttime data, but when combined with it, it reduces false positives and improves depth estimation, leading to more robust detection under low-light conditions. This study highlights both the potential and the limitations of style transfer for augmenting Mono3D training, and it points to future research on more advanced generative models and broader object categories. Full article
Show Figures

Figure 1

Back to TopTop