Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (800)

Search Parameters:
Keywords = accurate vehicle localization

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
27 pages, 687 KB  
Article
Two-Tier Anomaly Detection for V2I Alerting on IoT Vehicle Counts: A Kuwait Corridor Benchmark
by Yousef AlSaqabi
Sensors 2026, 26(17), 5368; https://doi.org/10.3390/s26175368 - 25 Aug 2026
Abstract
Anomaly detection in vehicular networks focuses on cybersecurity, leaving physical traffic-flow anomalies at urban intersections underserved by IoT sensing. This paper benchmarks anomaly detection on five days of hourly vehicle counts from four consecutive signalized intersections in Kuwait City, with a vehicle-to-infrastructure (V2I) [...] Read more.
Anomaly detection in vehicular networks focuses on cybersecurity, leaving physical traffic-flow anomalies at urban intersections underserved by IoT sensing. This paper benchmarks anomaly detection on five days of hourly vehicle counts from four consecutive signalized intersections in Kuwait City, with a vehicle-to-infrastructure (V2I) latency feasibility analysis. Anomalies are synthetically injected because verified incident labels are unavailable; scores reflect detectability under the injection protocol rather than validated incident detection. Across ten injection seeds, CUSUM is the most accurate (mean F1 0.945, perfect precision on every seed), Isolation Forest attains the highest recall (0.955), and the LSTM-AE reaches F1 0.347; on misaligned anomaly classes, the margin narrows, and the LSTM-AE matches CUSUM on gradual drift. A corridor rule localizes detected corridor anomalies (9/9, conditional on detection). Hourly aggregation alone imposes an expected 1800 s detection delay, over 130 times the 13.5 s V2I budget at 80 km/h. A sub-second Tier-1 edge detector, evaluated in traffic-calibrated simulation, detects surges within budget (median 6.8 to 9.1 s, robust to signal-cycle platooning), whereas flow-cutoff detection requires roughly 21 s and overnight hours remain a blind spot. Results support a two-tier edge-cloud design and provide, to our knowledge, the first such benchmark on real Gulf-region corridor count data. Full article
(This article belongs to the Section Internet of Things)
Show Figures

Figure 1

26 pages, 7007 KB  
Article
OVR-GS: Open-Vocabulary 3D Object Removal via Semantic Gaussian Selection and Local Diffusion-Guided Completion
by Yongpeng Ding, Feng Ouyang, Jiawei Fan, Ting Chen and Hongyan Xu
Sensors 2026, 26(16), 5258; https://doi.org/10.3390/s26165258 - 19 Aug 2026
Viewed by 199
Abstract
Camera-reconstructed 3D scenes often require offline visual cleanup before inspection, presentation, or reuse as renderable virtual-scene assets. Representative applications include removing temporary furniture, parked vehicles, equipment, signage, and other distracting or obsolete objects from reconstructed indoor and outdoor environments. Such editing requires not [...] Read more.
Camera-reconstructed 3D scenes often require offline visual cleanup before inspection, presentation, or reuse as renderable virtual-scene assets. Representative applications include removing temporary furniture, parked vehicles, equipment, signage, and other distracting or obsolete objects from reconstructed indoor and outdoor environments. Such editing requires not only accurate target localization across viewpoints but also plausible recovery of the previously occluded background. Existing methods often depend on manually specified masks or category-restricted detectors, while projection-based pipelines independently inpaint multiple views and subsequently refine the 3D representation, potentially introducing cross-view appearance and geometry inconsistencies. We present OVR-GS (Open-Vocabulary Removal in Gaussian Splatting), an instruction-driven object-removal framework for pre-trained 3D Gaussian Splatting (3DGS) scenes. Given a free-form instruction, a language parser generates target-oriented queries and a textual background-completion condition. Grounding DINO and the Segment Anything Model (SAM) produce multi-view candidate masks, which are filtered using Contrastive Language–Image Pre-training (CLIP). The proposed Semantic-Aware Gaussian Selector (SAGS) aggregates rendering-contribution-weighted mask evidence, groups spatially coherent candidates, and identifies the target Gaussian subset through rendered-cluster semantic verification. After removal, new Gaussians are initialized from boundary-adjacent primitives and interior samples and optimized locally using Score Distillation Sampling (SDS), while the original background remains fixed. On IMFine, SPIn-NeRF, and Inpaint360GS, OVR-GS achieves peak signal-to-noise ratio (PSNR) values of 19.78, 17.82, and 24.62 dB and Fréchet inception distance (FID) values of 142.30, 148.60, and 34.80, respectively. The results demonstrate the effectiveness of localized Gaussian optimization for instruction-driven cleanup of reconstructed environments before visual inspection, presentation, or reuse as renderable virtual-scene assets. Full article
(This article belongs to the Section Optical Sensors)
Show Figures

Figure 1

31 pages, 36482 KB  
Article
Geo-Consistent Centralized Multi-UAV Gaussian SLAM for Incremental Orthophoto Generation
by Xiao Zhang, Shuaixin Li, Hongbin Dong, Xiaozhou Zhu, Haoxin Zhang and Baosong Deng
Remote Sens. 2026, 18(16), 2804; https://doi.org/10.3390/rs18162804 - 19 Aug 2026
Viewed by 245
Abstract
Online incremental orthophoto generation with multiple unmanned aerial vehicles (UAVs) remains challenging, as it requires accurate, efficient, and scalable mapping from distributed aerial observations. In this paper, we present a centralized GNSS-assisted multi-UAV 3D Gaussian Splatting SLAM framework for online incremental orthophoto mapping. [...] Read more.
Online incremental orthophoto generation with multiple unmanned aerial vehicles (UAVs) remains challenging, as it requires accurate, efficient, and scalable mapping from distributed aerial observations. In this paper, we present a centralized GNSS-assisted multi-UAV 3D Gaussian Splatting SLAM framework for online incremental orthophoto mapping. Each UAV independently performs visual odometry to build local submaps, which are first aligned into a unified global coordinate system using GNSS constraints and further refined via inter-agent visual loop closures for improved cross-agent consistency. To enable scalable and high-quality mapping, we introduce two complementary Gaussian map maintenance modules: plane-guided grid-based collaborative densification, which improves mapping quality and accelerates convergence under multi-UAV conditions, and visibility-aware adaptive pruning, which effectively controls redundancy and memory usage. These components allow efficient joint optimization within a unified Gaussian representation. Experiments on multiple aerial datasets using video-derived image frames captured by consumer-grade UAV cameras demonstrate that the proposed system provides a favorable trade-off between geo-consistency, visual fidelity, and efficiency compared with existing methods. Quantitatively, the proposed method achieves a GCP RMSE of 2.32 m, completes multi-UAV orthophoto generation within 3.3–5.5 min, and reduces the total mapping time by approximately 35–55% compared with the corresponding single-UAV setting, while supporting online tracking and incremental orthophoto updates with bounded latency and memory consumption. Full article
Show Figures

Figure 1

25 pages, 27003 KB  
Article
RA-SIDO: Robust and Adaptive Sonar–Inertial–Depth Odometry for Consistent Underwater Acoustic 3D Mapping
by Yabei Guo, Huigang Wang, Wei Qiang, Runhe Yao and Zhizhen Xie
J. Mar. Sci. Eng. 2026, 14(16), 1520; https://doi.org/10.3390/jmse14161520 - 17 Aug 2026
Viewed by 183
Abstract
Autonomous acoustic remote sensing of underwater infrastructure is challenging due to the physical characteristics of 3D sonar and the geometric degeneracy commonly encountered in feature-poor underwater environments. Accurate localization is essential for integrating sequential sonar observations into globally consistent 3D maps; however, existing [...] Read more.
Autonomous acoustic remote sensing of underwater infrastructure is challenging due to the physical characteristics of 3D sonar and the geometric degeneracy commonly encountered in feature-poor underwater environments. Accurate localization is essential for integrating sequential sonar observations into globally consistent 3D maps; however, existing odometry methods often rely on isotropic noise assumptions despite the highly directional nature of acoustic sensing. This mismatch may cause unreliable measurements to be over-trusted, leading to severe trajectory drift and distortion in sonar-derived 3D reconstructions. To address these challenges, we propose RA-SIDO, a robust and adaptive tightly coupled 3D sonar–inertial–depth odometry framework based on the Error-State Iterated Kalman Filter (ESIKF), which fuses measurements from a 3D sonar, an inertial measurement unit (IMU), and a depth sensor for reliable underwater acoustic mapping. The proposed method introduces two mechanisms to handle sonar-specific uncertainties: (1) a physics-based anisotropic acoustic measurement model that distinguishes high-resolution radial range measurements from highly uncertain cross-range angular measurements; (2) an online degeneracy-awareness module that continuously evaluates the minimum eigenvalue of the translational information matrix and dynamically adjusts sensor fusion weights to avoid over-trusting ill-conditioned constraints. Real-world experiments were conducted with an unmanned surface vehicle in underwater infrastructure inspection scenarios. RA-SIDO achieved an ATE RMSE of 0.8924m, reducing the error by 16.8% compared with SIDO, the strongest baseline. In addition, the proposed method effectively suppresses longitudinal slip and produces globally consistent 3D acoustic maps of submerged structures. These results validate the potential of RA-SIDO as a robust localization and mapping solution for underwater remote sensing, infrastructure inspection, and acoustic 3D reconstruction in challenging aquatic environments. Full article
(This article belongs to the Section Ocean Engineering)
Show Figures

Figure 1

20 pages, 7984 KB  
Article
Vision-Map Fusion Multi-Object Tracking at Complex Intersections Using HD Map Priors and Nonlinear Filtering
by Dezheng Ma and Lan Tang
Automation 2026, 7(4), 130; https://doi.org/10.3390/automation7040130 - 16 Aug 2026
Viewed by 373
Abstract
Accurate multi-object tracking and metric localization support traffic monitoring and cooperative intelligent transportation at complex intersections. This study presents a fixed-camera vision-map fusion framework that addresses two practical difficulties: axis-aligned boxes poorly represent turning vehicles, and unconstrained image-plane tracking can produce physically implausible [...] Read more.
Accurate multi-object tracking and metric localization support traffic monitoring and cooperative intelligent transportation at complex intersections. This study presents a fixed-camera vision-map fusion framework that addresses two practical difficulties: axis-aligned boxes poorly represent turning vehicles, and unconstrained image-plane tracking can produce physically implausible trajectories. A map-aided frontend first generates candidate detections using improved You Only Look Once version 8 nano (YOLOv8n) horizontal bounding box (HBB) branch and an improved YOLOv8 oriented bounding box (OBB) branch. A high-definition (HD) map selector then retains the candidate geometry consistent with the straight-driving or turning region and converts it into a unified detection record. The selected reference point is projected to the ground plane through an offline-estimated homography, whereas the appearance feature bypasses the homography and is passed directly to the association stage. The tracking backend uses a 12-dimensional joint image/metric state, symmetric central-difference evaluations of the process and measurement functions, appearance-motion association, and a feasible-road projection derived from HD-map lane polygons. On the evaluated public sequences, the complete configuration achieved a multiple object tracking accuracy (MOTA) of 74.5%, an identification F1 score (IDF1) of 82.6%, 614 identity switches, and a throughput of 26.8 frames per second (FPS) on an RTX 4090 workstation. In a descriptive Vehicle-in-the-Loop case study involving one instrumented vehicle at one intersection, the overall localization mean absolute error (MAE) was 0.180 m, compared with 0.208 m for the baseline end-to-end configuration. These results indicate the feasibility of combining branch-specific vehicle geometry with map-constrained tracking; controlled same-detector comparisons, repeated multi-vehicle trials, and embedded-device latency and power profiling remain necessary for broader claims. Full article
(This article belongs to the Section Smart Transportation and Autonomous Vehicles)
Show Figures

Figure 1

37 pages, 4042 KB  
Article
VIWNO: Vehicle–Bridge Interaction Wavelet Neural Operator for Controlled Bridge Simulation and Laboratory Damage Identification
by Zixu Hu, Haitao Li, Wei He and Yongweng Wu
Buildings 2026, 16(16), 3235; https://doi.org/10.3390/buildings16163235 - 14 Aug 2026
Viewed by 264
Abstract
Controlled bridge simulation and laboratory damage identification require models that can simulate structural responses and infer localized stiffness loss from limited measurements. Existing Fourier Neural Operator (FNO)-based vehicle–bridge interaction (VBI) models provide efficient surrogates for these mappings, but the global Fourier representation can [...] Read more.
Controlled bridge simulation and laboratory damage identification require models that can simulate structural responses and infer localized stiffness loss from limited measurements. Existing Fourier Neural Operator (FNO)-based vehicle–bridge interaction (VBI) models provide efficient surrogates for these mappings, but the global Fourier representation can smooth localized damage transitions and introduce boundary-related errors for finite-span bridge responses. This study adapts the Wavelet Neural Operator (WNO) to the VBI setting and develops the Vehicle–Bridge Interaction Wavelet Neural Operator (VIWNO), an application-oriented framework for wavelet-domain operator learning between structural response fields and damage fields. VIWNO is pre-trained on a numerical VBI finite-element dataset (VBI-FE) and fine-tuned using only healthy-state measurements from a scaled VBI experimental dataset (VBI-EXP), before being evaluated on unseen laboratory damage scenarios. Under the controlled VBI-FE setting, where bridge, vehicle, speed, and measured road-profile parameters are fixed and the main variation is the damage field, VIWNO reduces forward response errors by 20–30% and inverse damage-estimation errors by 26–32% relative to the FNO-based Vehicle–Bridge Interaction Neural Operator (VINO) baseline. Additional morphology and operating-condition stress tests show that the error increases under sharper damage fields and perturbed VBI conditions, but VIWNO remains more accurate than VINO and the added convolutional or frequency-domain baselines in the tested cases. On VBI-EXP, projection-only healthy-state fine-tuning reduces intact false-damage levels and yields sharper damage estimates than VINO under both displacement and acceleration inputs. Stability checks over five initializations and repeated vehicle passages show limited variation in the reported inverse metrics. These results support the feasibility of wavelet-domain neural operators for calibrated VBI simulation and scaled laboratory damage identification, while field-scale bridge health monitoring still requires validation under broader traffic, environmental, support, and damage-morphology variability. Full article
(This article belongs to the Special Issue Structural Health Monitoring and Vibration Control)
Show Figures

Figure 1

20 pages, 9593 KB  
Article
Multi-Scale Spatio-Temporal Graph Transformer with xLSTM for Electric Vehicle Charging Demand Prediction
by Yedan Li, Zhuoran Ye, Bo Li and Zijun Chen
Electronics 2026, 15(16), 3618; https://doi.org/10.3390/electronics15163618 - 14 Aug 2026
Viewed by 163
Abstract
Accurate prediction of electric vehicle (EV) charging demand is critical for planning charging infrastructure and allocating resources efficiently. However, existing methods often fail to capture multi-scale spatial dependencies and struggle to model both long-range temporal dependencies and short-term fluctuations. To address these limitations, [...] Read more.
Accurate prediction of electric vehicle (EV) charging demand is critical for planning charging infrastructure and allocating resources efficiently. However, existing methods often fail to capture multi-scale spatial dependencies and struggle to model both long-range temporal dependencies and short-term fluctuations. To address these limitations, a deep learning framework, STGFormer, is developed for citywide EV charging demand forecasting. First, a temporal dilated convolution module (TDConv) is proposed to extract multi-scale local temporal patterns. Second, an adaptive spatial dilated graph attention module (ADGAT) is proposed to mine multi-hop spatial correlations between geographically adjacent and functionally similar regions. Third, a hybrid xLSTM-Transformer encoder captures global temporal dependencies while preserving local continuity. The performance of STGFormer is evaluated on a real-world dataset from Shenzhen. Extensive experiments demonstrate that STGFormer consistently outperforms fifteen representative baseline models. It achieves average improvements of 8.58% in RMSE, 13.95% in RAE, and 14.64% in MAE. Full article
(This article belongs to the Special Issue Electric Vehicle Power Systems: Design, Control and Integration)
Show Figures

Figure 1

15 pages, 44273 KB  
Article
SlotNet: A Lightweight Network with Skeleton-Driven and Adaptive Completion for Robust Detection of Degraded Parking Slot Lines
by Jiaxin Cheng, Yanhong Ning, Yongxing Huang and Shugang Liu
Appl. Sci. 2026, 16(16), 8098; https://doi.org/10.3390/app16168098 - 14 Aug 2026
Viewed by 150
Abstract
In response to degraded parking slot markings caused by wear and tear, water accumulation, or occlusion, which significantly impair the perception accuracy and localization robustness of automated parking systems, this paper proposes SlotNet, a lightweight enhancement network. The proposed method incorporates a skeleton-driven [...] Read more.
In response to degraded parking slot markings caused by wear and tear, water accumulation, or occlusion, which significantly impair the perception accuracy and localization robustness of automated parking systems, this paper proposes SlotNet, a lightweight enhancement network. The proposed method incorporates a skeleton-driven adaptive width completion algorithm to mitigate segmentation errors and restore the topological continuity of fractured parking slot lines. The network integrates three lightweight modules: Lightweight Reparameterized VGG (LightRepVGG) for enhancing the extraction of fine-grained structural features via structural reparameterization, Parallel Perceptual Structured Attention—Light (PASA_Light) for multi-scale feature fusion, and Adaptive Decoupled Detect and Segment (AdaDecDS) for anchor-free decoupled detection and segmentation. The experimental results show that SlotNet achieves an inference speed of 65.75 Frames Per Second (FPS). The mask average precision (mask mAP@0.5) reaches 90.2% under an Intersection over Union (IoU) threshold of 0.5, enabling robust completion and accurate detection of degraded parking slot lines. Compared with existing detection, SlotNet achieves a superior balance among accuracy, robustness, and real-time performance, making it suitable for deployment on embedded in-vehicle platforms. Full article
(This article belongs to the Topic Intelligent Image Processing Technology, 2nd Edition)
Show Figures

Figure 1

29 pages, 4499 KB  
Article
Fog-YOLO11n: A Lightweight Traffic Object Detection Framework for Autonomous Driving Under Foggy Conditions
by Furui Kuang and Zhi Chen
Appl. Sci. 2026, 16(16), 8059; https://doi.org/10.3390/app16168059 - 12 Aug 2026
Viewed by 209
Abstract
Foggy weather degrades road images by reducing visibility, attenuating object contrast, and blurring boundaries, while autonomous ground vehicles require accurate and lightweight environment perception on resource-constrained onboard platforms. This study proposes Fog-YOLO11n, a lightweight traffic object detector based on YOLO11n. A GhostRTA backbone [...] Read more.
Foggy weather degrades road images by reducing visibility, attenuating object contrast, and blurring boundaries, while autonomous ground vehicles require accurate and lightweight environment perception on resource-constrained onboard platforms. This study proposes Fog-YOLO11n, a lightweight traffic object detector based on YOLO11n. A GhostRTA backbone combines RepGhostNet with the proposed C2RTA module: RepGhostNet reduces redundant computation, whereas C2RTA uses channel statistics and multi-scale spatial context to reinforce weak low-contrast responses. FRISA, a dual-branch interactive attention module, separates semantic responses from residual details and uses bidirectional gates to suppress fog- and reflection-related textures while preserving object contours. SD-MPDIoU adaptively modulates corner-distance penalties according to relative box geometry and reweights samples by localization quality, improving regression stability for ambiguous boundaries. Experiments on the RTTS dataset show improvements of 3.80 and 2.52 percentage points in mAP@0.5 and mAP@0.5:0.95, respectively, over YOLO11n, while reducing parameters from 2.59 M to 2.14 M and computation from 6.44 to 5.66 GFLOPs. The resulting accuracy-complexity balance supports real-time foggy-road perception for autonomous unmanned vehicles. Full article
Show Figures

Figure 1

37 pages, 91707 KB  
Article
EdgeNeXt-Attn: A Lightweight Attention-Enhanced Deep Learning Framework for Fire Detection in Remote Sensing Imagery
by Hikmat Yar, Nehad Ali Shah, Weiwei Jiang, Norah Saleh Alghamdi and Heung Soo Kim
Remote Sens. 2026, 18(16), 2706; https://doi.org/10.3390/rs18162706 - 12 Aug 2026
Viewed by 282
Abstract
Wildfires are a major environmental hazard with severe consequences for ecosystems, air quality, infrastructure, and public safety. The rising incidence and severity of wildfire events worldwide have increased the need for reliable early detection and monitoring systems. Remote sensing technologies, such as satellite [...] Read more.
Wildfires are a major environmental hazard with severe consequences for ecosystems, air quality, infrastructure, and public safety. The rising incidence and severity of wildfire events worldwide have increased the need for reliable early detection and monitoring systems. Remote sensing technologies, such as satellite and unmanned aerial vehicle (UAV) imagery, along with ground-based Closed-Circuit Television (CCTV) cameras, provide valuable geospatial data for large-scale wildfire monitoring. Recent advances in deep learning, particularly Convolutional Neural Networks (CNNs) and Transformer-based architectures, have significantly improved the accuracy of wildfire detection systems. Despite these advances, balancing local feature representation with global contextual modeling remains challenging. CNNs effectively capture local spatial features but have limited receptive fields, whereas Vision Transformers (ViTs) model long-range dependencies but often overlook fine-grained local details and require substantial computational resources. Consequently, accurately detecting small, occluded, and visually ambiguous fire regions remains difficult, particularly for real-time deployment on resource-constrained edge devices. To address these challenges, this study proposes EdgeNeXt-Attn, an enhanced EdgeNeXt-based framework that effectively integrates local feature learning and global contextual modeling through channel and spatial attention mechanisms. The proposed model improves the detection of small, occluded, and visually ambiguous fire regions while maintaining the computational efficiency required for real-time edge deployment. The proposed framework is evaluated on four multi-platform benchmarks spanning ground-based CCTV (DFAN, Complex-Fire), aerial drone (FLAME), and mixed drone–satellite (ADSF) imagery, achieving 92.09%, 95.16%, 96.65%, and 87.81% accuracy, respectively, and outperforming recent state-of-the-art baselines. With only 5.3M parameters, the model achieves real-time inference at 85.9, 27.3, and 8.4 FPS on GPU, CPU, and Raspberry Pi, respectively. Furthermore, ablation studies and Grad-CAM analysis validate its effectiveness and accurate fire localization. These results demonstrate an accurate and computationally efficient framework for real-time wildfire monitoring using multi-platform remote sensing and ground-based imaging systems. Full article
(This article belongs to the Special Issue Image Analysis for Forest Environmental Monitoring (2nd Edition))
Show Figures

Figure 1

26 pages, 3594 KB  
Article
Master Mix Localization Algorithm for Autonomous Systems in Indoor Environments
by Zakaryae Ezzouine, Adil Salbi, Mohamed Abouzahir, Ilham Elmourabit, Adil Brouri and Sébastien Roy
Entropy 2026, 28(8), 903; https://doi.org/10.3390/e28080903 - 12 Aug 2026
Viewed by 286
Abstract
Reliable navigation in GPS-denied environments remains a critical challenge for autonomous vehicles (AVs), particularly in complex indoor and urban settings. GPS-based localization systems often fail under these conditions, highlighting the need for resilient multimodal solutions. In this article, we present a radar-assisted tracking [...] Read more.
Reliable navigation in GPS-denied environments remains a critical challenge for autonomous vehicles (AVs), particularly in complex indoor and urban settings. GPS-based localization systems often fail under these conditions, highlighting the need for resilient multimodal solutions. In this article, we present a radar-assisted tracking system that integrates LiDAR and inertial measurements within a sensor-fusion architecture to achieve robust navigation. The principal methodological contribution is a unified tracking and prediction framework that combines Bayesian state estimation with learning-based temporal prediction, enabling accurate tracking while continuously forecasting the slave robot’s short-term future state from mapping observations generated by the master robot, with a typical end-to-end perception-to-action latency of 20–60 ms. The communication and prediction forecasting module operates with an update interval below 35 ms, enabling real-time cooperative robotic operation. Sensor data are fused through a pipeline incorporating Gaussian Mixture Models (GMMs) for post-processing, which helps mitigate the limitations associated with individual sensors during edge processing. Moreover, Kalman filtering is employed to mitigate sensor noise and drift, thereby improving state estimation accuracy through trajectory smoothing. The fused spatiotemporal information is subsequently exploited by a Convolutional Recurrent Neural Network (CRNN) coupled with a Nonlinear Autoregressive model with eXogenous Inputs (NARX) to model the robot’s motion dynamics and provide short-horizon state prediction. Through simulations and real-world indoor experiments conducted in GPS-denied environments, we validate the system’s ability to provide accurate and continuous pose estimation with low localization errors. Experimental results show that the proposed framework achieves root-mean-square errors of 0.12 m, 0.15 m, and 0.28 m along the X, Y, and Z axes, respectively, while maintaining sub-meter maximum position deviations throughout the evaluated trajectories. These results confirm that the proposed framework provides reliable localization and predictive state estimation for cooperative robotic navigation in indoor GPS-denied environments. Future work will investigate outdoor validation and extend the framework to additional data-driven decision-making models for future robotic services. Full article
(This article belongs to the Special Issue Topics from the 2025 Biennial Symposium on Communications)
Show Figures

Figure 1

33 pages, 2920 KB  
Article
Characterizing the Operating Envelope of an Anomaly-Aware Adaptive EKF for GNSS-Denied USV Formation Relative Localization
by Ling Tan, Jianqiang Zhang, Yiping Liu, Pengfei Zhang and Xingda Li
J. Mar. Sci. Eng. 2026, 14(16), 1490; https://doi.org/10.3390/jmse14161490 - 11 Aug 2026
Viewed by 226
Abstract
Unmanned surface vehicle (USV) formations operating under GNSS denial require accurate relative localization using proprioceptive sensors and inter-vehicle ranging. This paper presents an anomaly-aware adaptive extended Kalman filter for four-USV formations using inertial measurements, compass, and ultra-wideband ranging, and systematically characterizes its operating [...] Read more.
Unmanned surface vehicle (USV) formations operating under GNSS denial require accurate relative localization using proprioceptive sensors and inter-vehicle ranging. This paper presents an anomaly-aware adaptive extended Kalman filter for four-USV formations using inertial measurements, compass, and ultra-wideband ranging, and systematically characterizes its operating envelope. Observability analysis establishes that S-curve maneuvering achieves structural rank 24, with only global translation unobservable, while straight-line motion leads to a rank deficiency of exactly seven dimensions All four gyroscope biases remain observable under both trajectories. The proposed filter integrates chi-square testing, cumulative sum (CUSUM) detection, and bias drift rate monitoring to trigger coordinated R adaptation and Q-boost mechanisms. Controlled experiments spanning outlier magnitudes and drift rates reveal three performance regimes, clean conditions with equivalent performance across all variants, moderate outliers [3σd,10σd] where the proposed method achieves 4.8–13.4% improvement, and extreme outliers where all robust methods converge. Critically, pure bias drift experiments expose a structural limitation of single-hypothesis, residual domain robustification within the tested drift range—all variants exhibit equivalent performance across the tested drift rates, analytically attributable to Kalman gain partitioning that distributes innovations between position and bias subspaces. The characterized operating envelope establishes that robust mechanisms provide measurable benefits for transient anomalies but encounter hard boundaries under persistent drift conditions, with all variants converging to equivalent performance across the tested range, necessitating multi-hypothesis or constraint-based approaches. Full article
(This article belongs to the Section Ocean Engineering)
Show Figures

Figure 1

32 pages, 7047 KB  
Article
A Transformer-Based Framework with Multi-Scale Feature Reconstruction for UAV Power Inspection
by Bing Zhang, Mengyao Sun, Haolong Meng and Lei Yang
Mathematics 2026, 14(16), 2901; https://doi.org/10.3390/math14162901 - 11 Aug 2026
Viewed by 220
Abstract
Accurate detection of transmission line components is crucial for the stability and security of power grid operations. However, accurate power line inspection is always affected by complex factors, such as multi-scale objects, complex background interference, and object occlusion, etc. To tackle these complexities, [...] Read more.
Accurate detection of transmission line components is crucial for the stability and security of power grid operations. However, accurate power line inspection is always affected by complex factors, such as multi-scale objects, complex background interference, and object occlusion, etc. To tackle these complexities, this paper leverages the long-range dependency modeling advantages of the Transformer architecture, and an improved real-time end-to-end Detection Transformer (RT-DETR) with multi-scale feature reconstruction, referred to as MFRRT-DETR, for Unmanned Aerial Vehicle (UAV) inspection systems is presented. Specifically, an enhanced attention-based backbone network integrated via an aggregated pixel-focus attention (APFA) module is built which uses a dual-path design with fine-grained and coarse-grained branches to combine pixel-level focus with global perception to enhance the interaction between local and global features, alleviating the limitations of the local receptive field in Convolutional Neural Networks (CNNs). To further overcome the issues of target overlap, occlusion, and foreground–background confusion, a context-guided spatial feature reconstruction feature pyramid network (CGR-FPN) module is proposed which strengthens foreground representation and effectively fuses multi-scale features, improving performance in crowded scenes. Additionally, a Focaler–Shape IoU loss function is introduced to mitigate class imbalance issues and localization errors by focusing on hard samples and optimizing bounding box regression, particularly for long and wide irregular rectangular targets. Experiments show that the proposed MFRRT-DETR significantly outperforms advanced detection models, which effectively validates the detection efficiency and accuracy of the proposed model in complex inspection scenarios, making it a promising solution for UAV-based power line inspection. Full article
Show Figures

Figure 1

20 pages, 6282 KB  
Article
DOU-Pose: Robust Camera-Based Visual Localization for Autonomous Vehicles in Repetitive and Low-Texture Intelligent Transportation Environments
by Xin’an Qiu, Liwen Wang, Zezheng Dong, Xiao Xiao, Zhihao Liu, Lin Zhu, Jingyin Wang and Hao Xu
Sensors 2026, 26(16), 5070; https://doi.org/10.3390/s26165070 - 10 Aug 2026
Viewed by 278
Abstract
Accurate and robust vehicle localization is essential for autonomous driving. However, existing visual pose estimation methods often struggle in scenarios dominated by repetitive structures or sparse textures. These conditions lead to ambiguous predictions of 3D scene coordinates and a high proportion of structured [...] Read more.
Accurate and robust vehicle localization is essential for autonomous driving. However, existing visual pose estimation methods often struggle in scenarios dominated by repetitive structures or sparse textures. These conditions lead to ambiguous predictions of 3D scene coordinates and a high proportion of structured outliers—erroneous predictions forming coherent clusters that deceive standard estimators. To address these limitations, this paper proposes DOU-Pose (Depthwise Over-parameterized U-shaped Pose estimation), a visual pose estimation framework built upon the Differentiable SAmple Consensus (DSAC)* pipeline. The core idea is to enhance the discriminative capability of scene coordinate regression through improved feature extraction. Specifically, we replace standard convolutional layers with Depthwise Over-parameterized Convolution (DO-Conv), which introduces auxiliary learnable depthwise kernels during training to enrich the representational capacity of the network, while allowing their fusion into a single kernel for inference. Furthermore, a U-shaped regression network with transposed convolutions is designed to preserve spatial details and strengthen fine-grained geometric reasoning. The entire pipeline is trained end-to-end by coupling dense scene coordinate prediction with a differentiable robust estimator. Extensive experiments demonstrate that DOU-Pose achieves competitive performance on public benchmarks and clear robustness improvements on the self-collected Campus-AV dataset, especially in repetitive and low-texture outdoor driving scenarios. Full article
Show Figures

Figure 1

27 pages, 12215 KB  
Article
Trajectory Prediction-Aided Deep Reinforcement Learning for Autonomous Vehicle Decision-Making at Unsignalized Intersections
by Shufeng Wang, Yuhang Wang, Yongxin Lei and Lu Jin
Machines 2026, 14(8), 900; https://doi.org/10.3390/machines14080900 - 6 Aug 2026
Viewed by 194
Abstract
Due to the absence of traffic signal control and the difficulty in accurately estimating the future movements of surrounding vehicles, autonomous vehicle decision-making faces challenges at unsignalized intersections. This study proposes a trajectory prediction-aided deep reinforcement learning framework. First, a composite prioritized replay [...] Read more.
Due to the absence of traffic signal control and the difficulty in accurately estimating the future movements of surrounding vehicles, autonomous vehicle decision-making faces challenges at unsignalized intersections. This study proposes a trajectory prediction-aided deep reinforcement learning framework. First, a composite prioritized replay mechanism is introduced into the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm, jointly considering temporal-difference error and reward-based event severity to enhance critical-experience reuse. Second, a convolutional multi-layer long short-term memory (CM-LSTM) model predicts surrounding-vehicle trajectories through convolutional local-motion encoding and stacked LSTM temporal modeling, and the predicted trajectories are incorporated into the deep reinforcement learning state representation. A multi-objective reward function is designed to balance collision avoidance, passing efficiency, lane keeping, and task completion. In CARLA go-straight and left-turn tests, CLS-TD3 achieves success rates of 93.8% and 90.2%, collision rates of 2.5% and 4.2%, and average passing times of 5.18 s and 5.58 s. Compared with TD3, the success rates increase by 6.3 and 8.6 percentage points, while average passing times decrease by 18.8% and 20.5%. These results demonstrate that the proposed framework improves the safety and crossing efficiency of autonomous vehicle decision-making at unsignalized intersections. Full article
(This article belongs to the Section Vehicle Engineering)
Show Figures

Figure 1

Back to TopTop