Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (103)

Search Parameters:
Keywords = road scene perception

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
54 pages, 19062 KB  
Article
Research on Multi-Dimensional and Multi-Level Environmental Parameter System for Visual Perception of Urban Tunnel Portal Sections Based on Structure–Light–Traffic (SLT) Coupling
by Mengdie Xu, Bo Liang, Haonan Long and Shuangkai Zhu
Appl. Sci. 2026, 16(18), 8994; https://doi.org/10.3390/app16188994 - 10 Sep 2026
Viewed by 226
Abstract
As a transition zone between open road environments and enclosed tunnel spaces, urban tunnel entrances undergo rapid variations in spatial structure, lighting conditions, and traffic-related semantic information over short distances. These abrupt environmental changes impose considerable visual demands on drivers and may adversely [...] Read more.
As a transition zone between open road environments and enclosed tunnel spaces, urban tunnel entrances undergo rapid variations in spatial structure, lighting conditions, and traffic-related semantic information over short distances. These abrupt environmental changes impose considerable visual demands on drivers and may adversely affect the reliability of autonomous driving perception systems. However, existing studies have primarily investigated individual environmental factors, while lacking a systematic parameterized framework capable of characterizing multi-source environmental features and their coupled effects on visual perception. To address this limitation, this study proposes a multidimensional and multilevel environmental parameter system for urban tunnel entrances based on a Structure–Lighting–Traffic (SLT) coupling framework. First, considering the formation mechanism of visual information, environmental factors influencing perception performance in tunnel entrance zones are categorized into three dimensions: spatial structure, lighting environment, and traffic semantics, thereby establishing a unified representation framework. Subsequently, an SLT coupling model is developed, incorporating parameter gradient intensity, coupling strength, and dispersion characteristics to quantitatively characterize the spatial variation and interaction patterns of environmental parameters. Furthermore, by integrating autonomous driving perception tasks, the relationships between environmental parameter variations, image quality degradation, and perception performance are investigated. The proposed framework is validated using field measurements collected from the entrance zones of 20 urban tunnels in Chongqing, China. The results reveal pronounced spatial heterogeneity and directional asymmetry in tunnel entrance environments. The entrance transition sections exhibit the characteristics of “high variation, strong coupling, and low stability,” whereas exit transition sections demonstrate more complex multi-factor interactions due to the combined effects of structural variations, intense light intrusion, and overlapping traffic information. Compared with structural and traffic-related factors, lighting variations play a dominant role in the degradation of overall visual perception performance. The proposed SLT-based environmental parameter system establishes a unified representation framework linking human visual perception mechanisms with machine vision perception modeling, providing theoretical support for autonomous driving perception optimization, complex scene understanding, and safety risk assessment in urban tunnel environments. Full article
Show Figures

Figure 1

21 pages, 107365 KB  
Article
M3-RGB: An Imaging Sensor System Using Multicore, Multimode Optical Fiber and Neural Networks
by Seigo Ito, Isamu Takai, Akari Kawasaki, Tadashi Ichikawa, Shin Motooka and Minoru Tanaka
Sensors 2026, 26(17), 5582; https://doi.org/10.3390/s26175582 - 2 Sep 2026
Viewed by 389
Abstract
Conventional image acquisition requires an electrically powered image sensor to be placed directly behind the camera lens, constraining camera placement. To overcome this issue, we introduce M3-RGB as an incoherent-light fiber imaging system in which a multicore, multimode optical fiber passively relays lens [...] Read more.
Conventional image acquisition requires an electrically powered image sensor to be placed directly behind the camera lens, constraining camera placement. To overcome this issue, we introduce M3-RGB as an incoherent-light fiber imaging system in which a multicore, multimode optical fiber passively relays lens images to a remotely located image sensor. Unlike conventional approaches, M3-RGB is designed to operate directly on incoherent light and requires no electrical power or active components at the sensing interface. Because propagation through the fiber yields spatially scrambled patterns, a neural network is used to reconstruct the original scene by exploiting the spatial locality preserved by the multicore structure. In a controlled optical bench setup, where a liquid crystal display monitor displays road-scene images, we construct a paired dataset of scrambled and ground-truth images and quantitatively evaluate reconstruction performance across different fiber core counts, fiber lengths, and calibration settings, utilizing the peak signal-to-noise ratio and structural similarity index measure as performance metrics. By decoupling imaging electronics from the sensing point, this passive remote image relay approach may expand sensor placement options for potential applications such as all-around perception for mobile robots and autonomous vehicles, surveillance, and inspection in confined spaces. Evaluations in real outdoor environments constitute future work. Full article
(This article belongs to the Section Industrial Sensors)
Show Figures

Figure 1

18 pages, 17371 KB  
Article
A Lighting-Aware Infrared–Visible Image Fusion Network for Security Surveillance
by Yi Jiang, Xin Nie and Imad Rida
Algorithms 2026, 19(9), 746; https://doi.org/10.3390/a19090746 - 2 Sep 2026
Viewed by 267
Abstract
Infrared and visible image fusion combines the thermal cues captured by infrared sensors with the rich structural and texture information provided by visible images. This technique is particularly valuable for security surveillance, nighttime scene perception, and target recognition under challenging environmental conditions. Existing [...] Read more.
Infrared and visible image fusion combines the thermal cues captured by infrared sensors with the rich structural and texture information provided by visible images. This technique is particularly valuable for security surveillance, nighttime scene perception, and target recognition under challenging environmental conditions. Existing methods generally adopt fixed fusion strategies and neglect dynamic dual-modality changes under low illumination, overexposure and strong light interference, failing to preserve both thermal target saliency and visible structural details in fused images. To address this problem, this paper proposes a Lighting-Aware Spatial–Frequency Fusion Network (LASFNet) for security surveillance. It first estimates modality reliability across low-light, overexposed and infrared-salient regions, incorporating it into the fusion of frequency-domain amplitude and phase. Spatial infrared, visible and frequency-domain compensation features are then jointly fused to generate the output. Experiments on M3FD, MSRS and RoadScene datasets show that LASFNet achieves competitive fusion performance with strong cross-dataset generalization. On M3FD, YOLOv8s with LASFNet-fused inputs achieves 84.825% mAP@0.5 and 57.319% mAP@0.5:0.95, outperforming visible, infrared and YDTR-fused inputs. The proposed method balances target saliency and scene structure, providing more effective visual input for object detection under complex illumination. Full article
Show Figures

Figure 1

23 pages, 37053 KB  
Article
Odometry and Mapping for Complex Environment Perception Under Partial-View Sensing
by Xinye Dai, Dingxi Wang, Jin Xing, Zhibo Zhang, Xiaoxiao Zhang, Xiao Wang, Shiqi Zheng, Yusheng Wang and Lijian Feng
Remote Sens. 2026, 18(17), 2959; https://doi.org/10.3390/rs18172959 - 2 Sep 2026
Viewed by 323
Abstract
Dense partial-view LiDAR observations are attractive for outdoor perception, but limited overlap and viewpoint sensitivity make odometry and mapping less reliable than with spinning LiDARs. Many recent algorithms for this sensing regime are built as LiDAR-inertial odometry frameworks, whose localization and mapping performance [...] Read more.
Dense partial-view LiDAR observations are attractive for outdoor perception, but limited overlap and viewpoint sensitivity make odometry and mapping less reliable than with spinning LiDARs. Many recent algorithms for this sensing regime are built as LiDAR-inertial odometry frameworks, whose localization and mapping performance can degrade or fail when the IMU state estimation becomes unstable. This paper presents a LiDAR-only framework for complex outdoor scenes using a factor-graph back-end. After denoising and motion compensation, the point cloud is projected onto a range image for ground, planar, edge, and line extraction. Pose estimation is strengthened by degeneracy-aware feature selection, while loop closing combines scan-based and path-based cues to handle partial-view revisits. Experiments in tunnels, urban roads, residential areas, and other challenging scenes show reduced drift and improved mapping consistency for dense limited-FoV LiDAR data. Full article
(This article belongs to the Special Issue LiDAR Technology for Autonomous Navigation and Mapping)
Show Figures

Figure 1

24 pages, 4725 KB  
Article
DCA-Net: Dilated Context Attention Network for MLS Point Cloud Semantic Segmentation
by Bingchen Du, Bozhao Li, Zhenkun Zhang, Peng Cheng and Zhongliang Cai
Remote Sens. 2026, 18(16), 2740; https://doi.org/10.3390/rs18162740 - 14 Aug 2026
Viewed by 259
Abstract
Mobile LiDAR systems (MLS) enable rapid acquisition of large-scale 3D point cloud data. Semantic segmentation of the acquired point clouds is an important task in outdoor scene understanding and environmental perception for autonomous driving. However, existing methods tend to suffer from boundary confusion [...] Read more.
Mobile LiDAR systems (MLS) enable rapid acquisition of large-scale 3D point cloud data. Semantic segmentation of the acquired point clouds is an important task in outdoor scene understanding and environmental perception for autonomous driving. However, existing methods tend to suffer from boundary confusion when segmenting MLS point clouds with long-tail categories. To address this problem, we propose the Dilated Context Attention Network (DCA-Net), which consists of a dilated local geometric encoding module, a channel attention pooling module, and a category-boundary sampling strategy. The dilated local geometric encoding module expands point-to-point connections within a fixed neighborhood to strengthen contextual modeling among neighboring points. The channel attention pooling module uses a channel attention mechanism to enhance informative channel responses in neighborhood features, thereby improving local feature representation. The category-boundary sampling strategy increases the sampling probabilities of minority-category points and boundary points, reducing feature information loss during down-sampling. Experimental results on the S3DIS, Toronto3D, and MLS road scene datasets show that DCA-Net achieves mIoU scores of 69.4%, 84.1%, and 96.6%, respectively. These results demonstrate that the proposed method alleviates boundary confusion in point cloud segmentation with long-tail categories, without causing a noticeable degradation in the segmentation performance of majority categories. Full article
Show Figures

Figure 1

32 pages, 9964 KB  
Review
Robust Perception for Autonomous Driving Under Low-Visibility Conditions: A Review of Low-Light Enhancement, Multimodal Fusion, and Task-Oriented Detection
by Jongbae Kim
Appl. Sci. 2026, 16(16), 8037; https://doi.org/10.3390/app16168037 - 12 Aug 2026
Viewed by 372
Abstract
Nighttime driving and adverse weather expose persistent weaknesses in autonomous-driving perception pipelines. Low illumination, fog, rain, snow, glare, wet-road reflections, and motion blur degrade camera, LiDAR, radar, and event-camera inputs in modality-specific ways. This review examines robust perception under low-visibility conditions as a [...] Read more.
Nighttime driving and adverse weather expose persistent weaknesses in autonomous-driving perception pipelines. Low illumination, fog, rain, snow, glare, wet-road reflections, and motion blur degrade camera, LiDAR, radar, and event-camera inputs in modality-specific ways. This review examines robust perception under low-visibility conditions as a pipeline-level problem requiring joint consideration of image enhancement, sensor fusion, detection, and evaluation. Rather than treating enhancement as an isolated restoration task, it analyzes whether recent methods preserve detector-relevant structures, exploit cross-sensor complementarity, and improve downstream 2D and 3D perception. A taxonomy-driven narrative approach compares representative studies along five axes, from input modality and supervision strategy to evaluation protocol and deployment feasibility, supported by a structured verification search of literature published between January 2020 and July 2026, with the search strategy, eligibility criteria, and corpus composition documented. Recent work indicates a shift from image-quality-oriented restoration toward perception-driven optimization, in which enhancement and fusion modules are evaluated by their effect on object detection and 3D perception. Remaining challenges include generalization to compound degradations, cross-sensor misalignment, scene-dependent sensor reliability, latency on in-vehicle edge platforms, and inconsistent benchmark protocols. Future systems should therefore jointly model degradation severity, sensor reliability, downstream task performance, and real-time constraints rather than optimizing restoration, fusion, and detection modules in isolation. Full article
Show Figures

Figure 1

32 pages, 36061 KB  
Article
Residual Conditional Diffusion with Transformer Refinement for Unsupervised Infrared–Visible Image Fusion
by Sirui Huang, Lin Tian and Yao Zhang
Electronics 2026, 15(15), 3449; https://doi.org/10.3390/electronics15153449 - 4 Aug 2026
Viewed by 390
Abstract
Infrared–visible image fusion aims to integrate thermal target information from infrared images and structural texture information from visible images into a single informative image. Existing deep fusion methods still face challenges in preserving fine textures, maintaining structural consistency, and balancing complementary information under [...] Read more.
Infrared–visible image fusion aims to integrate thermal target information from infrared images and structural texture information from visible images into a single informative image. Existing deep fusion methods still face challenges in preserving fine textures, maintaining structural consistency, and balancing complementary information under low-light conditions. To address these issues, this paper proposes MRCDFusion, an unsupervised infrared–visible image fusion network based on residual conditional diffusion and Transformer refinement. Specifically, a shared dense encoder is used to extract modality-specific and cross-modal complementary features from infrared and visible images. A Modality-Level Attention Module (MLAM) is then introduced to aggregate strong responses from infrared and visible features and construct modality-aware condition features for guiding the diffusion process. Instead of generating fused features from scratch, the proposed method adopts a base-plus-residual diffusion strategy, in which base features preserve global structures and residual diffusion enhances local details. A deterministic noise strategy is further introduced to improve inference reproducibility. The diffusion-enhanced features are refined by a window Transformer and depthwise separable convolutions, followed by gated feature fusion and progressive image reconstruction. Experiments are conducted primarily on the low-light LLVIP dataset, while FMB, TNO, and RoadScene are used for zero-shot cross-dataset evaluation without additional fine-tuning. The results show that MRCDFusion achieves particularly strong performance in gradient- and edge-related metrics while remaining competitive in visual information fidelity and cross-modal correlation metrics. Ablation studies verify the effectiveness of the main components, and downstream detection and auxiliary segmentation experiments further demonstrate the potential utility of the fused representations for subsequent visual perception tasks. Full article
Show Figures

Figure 1

36 pages, 653 KB  
Review
Bird’s-Eye-View Road Occupancy Prediction for Autonomous Driving: A Survey of Representations, Methods, and Benchmarks
by Abdelrahman S. Heikal, Mostafa Farouk Senussi, Ahmed Salem and Hyun-Soo Kang
Mathematics 2026, 14(15), 2720; https://doi.org/10.3390/math14152720 - 31 Jul 2026
Viewed by 506
Abstract
Bird’s-eye-view (BEV) perception has become the dominant paradigm for camera-centric scene understanding in autonomous driving, as well as road occupancy prediction, which involves the dense estimation of which regions of space are occupied and by what has emerged as its most expressive form. [...] Read more.
Bird’s-eye-view (BEV) perception has become the dominant paradigm for camera-centric scene understanding in autonomous driving, as well as road occupancy prediction, which involves the dense estimation of which regions of space are occupied and by what has emerged as its most expressive form. Between 2020 and 2026, the field underwent three overlapping transitions: from two-dimensional BEV semantic map segmentation to dense three-dimensional voxel-based 3D semantic occupancy, catalyzed by the 2022 industrial adoption of “occupancy networks,” and, most recently, to efficient, generative, and four-dimensional forecasting formulations. This survey organizes the literature along six orthogonal axes output representation, view-transformation mechanism, input modality, supervision paradigm, temporal scope, and efficiency strategy and uses the representation lineage as a primary spine connecting the 2020 BEV-segmentation works to the 2026 Gaussian and 4D frontier. Alongside the ego-centric mainstream, we review the parallel multi-view and infrastructure-side lineage from multi-view pedestrian occupancy to roadside traffic occupancy, which shares the BEV occupancy-map output and contributes generalization tools the ego-centric thread has yet to absorb. We review the canonical methods at each stage, summarize the standard datasets (CARLA, GMVD, MultiviewX, WildTrack, nuScenes, SemanticKITTI, Occ3D, OpenOccupancy) and evaluation metrics (MODA, mIoU, RayIoU, RayPQ), and consolidate reported results on the Occ3D-nuScenes benchmark into a single comparison. We close by identifying open problems in label efficiency, robustness, temporal forecasting, and deployment. Our intent is to bridge the historically separate BEV-segmentation and 3D-occupancy literatures within a single taxonomy. Full article
(This article belongs to the Special Issue New Advances in Image Processing and Computer Vision)
Show Figures

Figure 1

25 pages, 14950 KB  
Article
TopoGraph-Fusion: Hierarchical Task-Conditioned Topology Reasoning for RGB–Thermal Object Detection
by Pu Yu, Yanshan Ma, Yuheng Li and Chunhao Li
Symmetry 2026, 18(8), 1272; https://doi.org/10.3390/sym18081272 - 27 Jul 2026
Viewed by 390
Abstract
Robust object detection for autonomous driving requires perception models that remain reliable when visible imagery is degraded by darkness, glare, rain, fog, motion blur, or long-range small targets. Visible and thermal infrared cameras provide complementary evidence, yet many RGB–thermal detectors fuse modalities, mainly [...] Read more.
Robust object detection for autonomous driving requires perception models that remain reliable when visible imagery is degraded by darkness, glare, rain, fog, motion blur, or long-range small targets. Visible and thermal infrared cameras provide complementary evidence, yet many RGB–thermal detectors fuse modalities, mainly as aligned tensors, and may underuse relational structure in channel responses, spatial layouts, semantic scales, and modality-specific uncertainty. This paper presents TopoGraph-Fusion, a hierarchical graph-guided dual-modal object detector that formulates fusion as topology-aware reasoning rather than direct feature concatenation. The proposed framework builds a dual-stream backbone for RGB and thermal images, constructs channel-wise topology through a channel-topology graph aggregation module, derives relation-aware spatial and channel global attention from affinity graphs, and replaces fixed feature-pyramid communication with a Graph-Guided Feature-Pyramid Network. A topology-regularized detection objective further encourages stable cross-modal correspondence while suppressing noisy all-to-all connections. Experiments on M3FD, FLIR, RGBTDronePerson, and VEDAI512 cover road scenes, adverse illumination, drone–person perception, and aerial vehicle detection. Within this validation scope, the results and visual analyses indicate that topology-guided fusion improves small-object recall, cross-modal consistency, and robustness under modality imbalance. Full article
Show Figures

Figure 1

35 pages, 15509 KB  
Article
Roadside Monocular Camera Calibration Based on Adaptive Structure Tensor and Hierarchical Vanishing Point Estimation
by Xuecong Liu, Kun Kang and Kangchao Gao
Electronics 2026, 15(14), 3215; https://doi.org/10.3390/electronics15143215 - 21 Jul 2026
Viewed by 490
Abstract
Accurate camera parameter calibration is fundamental to scene geometric reconstruction, target localization, and metric measurement. Existing roadside monocular camera calibration methods usually rely on geometric information such as lane markings. However, in real-world roadside traffic scenarios, such information is easily affected by road [...] Read more.
Accurate camera parameter calibration is fundamental to scene geometric reconstruction, target localization, and metric measurement. Existing roadside monocular camera calibration methods usually rely on geometric information such as lane markings. However, in real-world roadside traffic scenarios, such information is easily affected by road textures, shadows, reflections, and noise responses, resulting in unstable extracted edge directions, scattered line–intersection distributions, and reduced vanishing point localization accuracy, which significantly affects camera parameter estimation and scene scale recovery. To address these issues, this paper proposes an automatic calibration method for roadside monocular cameras based on an adaptive structure tensor and hierarchical vanishing point estimation, and further applies the calibrated parameters to vehicle speed measurement. First, an adaptive structure tensor model based on local gradient extrema is constructed to reduce directional deviations caused by noise, reflections, and local textures, thereby enabling stable extraction of edge-direction features in complex scenarios. Second, a hierarchical vanishing point estimation framework integrating spatial focusing, Mean-Shift clustering, and nonlinear optimization is proposed. By progressively filtering outlier intersections and introducing geometric consistency constraints, the proposed framework achieves robust self-calibration in cluttered traffic environments. On this basis, the geometric prior information of road markings is exploited to establish a road-marking contour optimization model constrained by the two vanishing points. Through joint optimization, scene scale recovery is achieved without the need for calibration targets, thereby completing automatic calibration of the roadside monocular camera. Experimental results on the public BrnoCompSpeed dataset demonstrate that the proposed method achieves a mean relative camera calibration error of 5.31%, indicating higher calibration accuracy and stability than existing automatic calibration methods. In addition, the calibrated camera parameters are applied to vehicle speed estimation, yielding a mean speed estimation error of 1.79 km/h. These results verify the effectiveness of the proposed calibration method in real-world traffic scenarios and provide reliable support for geometric measurement and visual perception applications in intelligent transportation systems. Full article
Show Figures

Figure 1

24 pages, 61324 KB  
Article
Target Detection for Traffic Flow in Low-Altitude Unmanned Aerial Vehicle Scenarios
by Tian Luan, Fan Yang, Huanxia Wei and Weijun Pan
Mathematics 2026, 14(14), 2615; https://doi.org/10.3390/math14142615 - 18 Jul 2026
Viewed by 363
Abstract
Low-altitude unmanned aerial vehicle (UAV)-based traffic object detection is challenged by substantial scale variations from aerial perspectives, the extremely small pixel proportions of distant traffic participants, complex road background interference, unstable illumination, and severe occlusion in dense traffic scenes. To address these problems, [...] Read more.
Low-altitude unmanned aerial vehicle (UAV)-based traffic object detection is challenged by substantial scale variations from aerial perspectives, the extremely small pixel proportions of distant traffic participants, complex road background interference, unstable illumination, and severe occlusion in dense traffic scenes. To address these problems, this paper proposes ACP2-YOLO, an improved YOLO11-based detection framework for low-altitude UAV traffic scenarios, with the goal of enhancing the detection of vehicles, pedestrians, and non-motorized traffic participants. The proposed framework introduces two key improvements. First, a lightweight hybrid ACmix module that integrates convolution and self-attention is embedded into the network, enabling the model to jointly capture local detailed features and global contextual dependencies and thereby strengthen feature representation under complex backgrounds. Second, a P2 small-object detection layer is added to the original three-scale detection structure of YOLO11 to construct a four-scale P2–P5 feature pyramid. By allowing shallow high-resolution features to directly participate in object prediction, this design effectively reduces spatial information loss caused by deep downsampling and improves small-object perception. Experiments on the VisDrone2019 dataset show that the improved model achieves 53.1% Precision, 41.1% Recall, 42.9% mAP@50, and 26.3% mAP@50–95, outperforming the baseline YOLO11 by 4.2, 4.2, 5.0, and 3.6 percentage points, respectively. Comparisons with mainstream YOLO-series detectors further demonstrate its superior overall accuracy, small-object detection capability, and adaptability to complex scenes, indicating its potential for UAV-based traffic monitoring, road safety inspection, and intelligent transportation perception. Full article
Show Figures

Figure 1

33 pages, 10785 KB  
Article
Lightweight Semantic Perception from UAV-Borne Visual Sensors via Conflict-Suppressed Heterogeneous Expert Distillation
by Feng Ouyang, Yongpeng Ding, Miao Qin, Weiting Xie and Chao Zhou
Sensors 2026, 26(14), 4509; https://doi.org/10.3390/s26144509 - 16 Jul 2026
Viewed by 474
Abstract
UAV-borne visual sensors provide high-resolution aerial observations for low-altitude scene understanding, urban monitoring, traffic observation, emergency inspection, and infrastructure assessment. However, semantic perception from UAV visual sensor data remains challenging because aerial images often contain dense small objects, elongated road structures, fragmented boundaries, [...] Read more.
UAV-borne visual sensors provide high-resolution aerial observations for low-altitude scene understanding, urban monitoring, traffic observation, emergency inspection, and infrastructure assessment. However, semantic perception from UAV visual sensor data remains challenging because aerial images often contain dense small objects, elongated road structures, fragmented boundaries, scale variations caused by flight-altitude changes, oblique viewpoints, and strict onboard or edge computational constraints. To address these challenges, this paper proposes MEKD-UAVSeg, a lightweight semantic perception framework based on conflict-suppressed heterogeneous expert distillation. During training, a Transformer-based semantic expert provides global contextual understanding and region-level class consistency, while a Mamba-based spatial expert provides complementary structural guidance for roads, roofs, boundaries, and other continuous aerial structures. Both experts are used only during training, and the final inference model remains a compact CNN-based segmentation network. In addition, UAV-aware density and hard-region priors are designed to emphasize small-object-dense areas, boundary-sensitive regions, rare classes, and uncertain aerial categories. A conflict-suppressed reliability routing strategy is further developed to reduce inconsistent supervision between heterogeneous experts and selectively transfer reliable knowledge to the student model. Experiments on UAVid and UDD6 demonstrate that the proposed framework achieves a favorable accuracy–efficiency trade-off compared with representative CNN-, Transformer-, Mamba-, and hybrid-based UAV segmentation methods, without introducing expert-induced inference complexity. Full article
(This article belongs to the Section Vehicular Sensing)
Show Figures

Figure 1

34 pages, 22783 KB  
Article
An Explainable Multimodal Framework for Cyclist Safety Perception in Mixed Traffic Environments
by Chia-Yen Chiang, Meihui Wang, Yasmin Fathy, Mona Jaber and Ahmed M. Abdelmoniem
Appl. Sci. 2026, 16(13), 6690; https://doi.org/10.3390/app16136690 - 3 Jul 2026
Viewed by 538
Abstract
Despite growing policy support for active travel, the fatality rate of vulnerable road users has remained persistently high in recent years, while the emergence of autonomous vehicles has further increased the complexity of mixed traffic environments. Interactions between cyclists and motorized vehicles are [...] Read more.
Despite growing policy support for active travel, the fatality rate of vulnerable road users has remained persistently high in recent years, while the emergence of autonomous vehicles has further increased the complexity of mixed traffic environments. Interactions between cyclists and motorized vehicles are a major contributor to these fatalities, highlighting the urgent need for effective cyclist protection strategies. As one of the most widely adopted active transport modes, cycling safety cannot be assessed solely through crash statistics; understanding cyclists’ perceived safety is equally critical, as it reflects how infrastructure design and dynamic traffic conditions influence cycling behavior. In this study, we propose a cyclist safety perception framework that combines vision–language models with interpretable machine learning to analyze perceived safety in mixed traffic scenarios. A vision–language model is employed to generate semantic descriptions of traffic scenes, while an Explainable Boosting Machine quantifies both individual and interactive contributions of traffic-related features. By integrating visual information with road attributes extracted from OpenStreetMap, the proposed framework achieves a binary safety classification accuracy of 71% and a mean absolute error of 1.01 on a safety score scale ranging from 1 to 9. The results demonstrate the potential of combining multimodal perception and explainable models to support cyclist-centered safety assessment and inform sustainable and intelligent transportation system design. More specifically, the results show that protected cycling infrastructure is the most significant factor in improving perceived safety, whereas road construction has the opposite effect. Full article
(This article belongs to the Special Issue Advances in Intelligent Transportation and Sustainable Mobility)
Show Figures

Figure 1

54 pages, 2578 KB  
Review
Traversability Driven Perception and Planning Coupling Mechanisms for Autonomous Driving in Unstructured Environments: A Review
by Qingxin Ge, Haobin Jiang, Shidian Ma, Yixiao Chen and Lei Yin
Machines 2026, 14(7), 713; https://doi.org/10.3390/machines14070713 - 23 Jun 2026
Viewed by 498
Abstract
Autonomous driving in unstructured environments faces challenges such as missing road boundaries, terrain variations, random obstacle distributions, and complex vehicle–terrain interactions, making it difficult to achieve safe navigation by relying on lane-level priors from structured roads. To address the problems of the relative [...] Read more.
Autonomous driving in unstructured environments faces challenges such as missing road boundaries, terrain variations, random obstacle distributions, and complex vehicle–terrain interactions, making it difficult to achieve safe navigation by relying on lane-level priors from structured roads. To address the problems of the relative separation between traversability analysis and trajectory planning, the ineffective propagation of perception uncertainty, and the insufficient scene adaptability of coupling mechanisms, this paper takes traversability as the main thread and systematically reviews the research progress of perception–planning coupling mechanisms in unstructured environments. First, traversability analysis methods based on geometric terrain, semantic understanding, and physical dynamics are reviewed, and the representation and propagation mechanisms of uncertainty in the perception–planning chain are analyzed. Second, the role of traversability information in global path search, local trajectory optimization, and data-driven planning is discussed, and the applicable boundaries of different coupling architectures are summarized from the perspectives of representation level and system organization form. Finally, datasets, simulation platforms, and evaluation metric systems are summarized, and a risk-state-oriented adaptive perception–planning coupling framework is proposed to dynamically adjust coupling strength based on risk-state information, thereby improving the safety, interpretability, and environmental adaptability of autonomous driving in unstructured environments. Full article
(This article belongs to the Section Vehicle Engineering)
Show Figures

Figure 1

35 pages, 48685 KB  
Article
Efficient Multitask Onboard Vision Sensing for Open-Pit Mining Advanced Driver Assistance System with Classification-Guided Adaptive Temporal Inference
by Maximiliano Vélez and Claudio Urrea
Sensors 2026, 26(12), 3860; https://doi.org/10.3390/s26123860 - 17 Jun 2026
Viewed by 577
Abstract
Cameras and IMUs on heavy mining trucks supply the visual signal that Advanced Driver Assistance Systems (ADASs) use in open-pit operations. Haul roads in a surface mine are unstructured and unmarked, so a perception model must be both accurate and fast. We address [...] Read more.
Cameras and IMUs on heavy mining trucks supply the visual signal that Advanced Driver Assistance Systems (ADASs) use in open-pit operations. Haul roads in a surface mine are unstructured and unmarked, so a perception model must be both accurate and fast. We address this with a video-based multitask pipeline for a mining Driver Support System (DSS): a single BiSeNetV1 network produces drivable-area segmentation and steering-direction classification in one forward pass. Training used only 100 frames sampled non-sequentially from in-cab recordings of a real open-pit mine; evaluation used two full onboard sequences. To exploit temporal redundancy without annotating video, we propose an Adaptive Clockwork (A-CW) inference scheme: the spatial path runs on every frame, while the context path is refreshed only on keyframes whose cadence is set by the classification output, the same signal shown to the driver as a steering hint. This classification-guided policy increases context updates on curved segments, where the scene changes more rapidly, and reduces them on straight sections, where semantic redundancy is higher. The selected A-CW configuration was evaluated on full temporal test sequences, including one route kept entirely outside the training source. On this unseen route, A-CW achieved 94.70% road-class IoU and 73.68% Top-1 Accuracy. GPU-only throughput increased from about 55 FPS with frame-by-frame inference to 168.01 FPS, and display-excluded end-to-end processing in the simulated ADAS pipeline remained at approximately 37.5 FPS. Full article
(This article belongs to the Section Vehicular Sensing)
Show Figures

Figure 1

Back to TopTop