Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (1,160)

Search Parameters:
Keywords = scene representation

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
27 pages, 3514 KB  
Article
Latent Reorganization by Fixed-Point Regularization in Recurrent Frame Prediction
by Arpan Ghosh, Jinho Kim, Ye-Chan An and Tae-Yong Kuc
Electronics 2026, 15(15), 3288; https://doi.org/10.3390/electronics15153288 - 25 Jul 2026
Abstract
Regularization can improve generalization by constraining how a model uses its internal representation. In this paper, we study whether algebraic fixed-point constraints applied to the final LSTM hidden state during training can reorganize the recurrent latent space and improve held-out frame prediction. Rather [...] Read more.
Regularization can improve generalization by constraining how a model uses its internal representation. In this paper, we study whether algebraic fixed-point constraints applied to the final LSTM hidden state during training can reorganize the recurrent latent space and improve held-out frame prediction. Rather than modifying the inference-time architecture, we introduce four training-time operators, GlobalHouseholder (reflection), GlobalGivens (rotation), Composition, and Lie Algebra, that bias the hidden state toward geometrically structured regions without changing the decoder pathway. Experiments across four datasets (indoor robot sequences, KITTI driving, Flying Shapes 2D, and Moving 3D Shapes) show that lightweight constraints consistently improve prediction on structured scenes, with GlobalGivens achieving up to +1.04 dB PSNR and 11.3% MAE over the unconstrained baseline on held-out Indoor sequences. The latent analysis reveals that the operators that generalize best are not those that compress the representation most aggressively but those that redistribute latent energy while preserving broad dimensional participation. Lie Algebra, despite collapsing activation variance by 81–96%, degrades under latent perturbation and does not match the lighter operators on structured datasets identifying over-constraint as a clear failure mode. These results suggest that geometric regularization of a recurrent bottleneck can act as a useful training-time prior without adding any inference overhead. Full article
Show Figures

Figure 1

23 pages, 4175 KB  
Article
BEV-Nexus: BEV Perception Algorithm Based on Depth Perception Enhancement and Dynamic Adaptive Fusion
by Xiaona Song, Haozhe Zhang, Zhengyi Huang, Jianlin Zhao and Lijun Wang
Sensors 2026, 26(15), 4720; https://doi.org/10.3390/s26154720 - 25 Jul 2026
Viewed by 62
Abstract
This paper proposes an improved multimodal fusion framework for 3D object detection, termed BEV-Nexus, which aims to address the issues of inaccurate depth estimation and inefficient fusion paradigms in existing image-point cloud fusion methods. We introduce a Point-Cloud-Guided Depth Prediction Network (PCGD-Net), which [...] Read more.
This paper proposes an improved multimodal fusion framework for 3D object detection, termed BEV-Nexus, which aims to address the issues of inaccurate depth estimation and inefficient fusion paradigms in existing image-point cloud fusion methods. We introduce a Point-Cloud-Guided Depth Prediction Network (PCGD-Net), which enhances the image branch’s depth prediction capability by embedding point cloud spatial prior, ground-truth loss constraint, and projected point cloud depth filling. Additionally, we design a Dynamic Self-adaptive Feature Fusion Module (DSF-Module), which computes multimodal feature similarity using window attention and performs weighted fusion based on self-adaptive weights, resolving alignment deviations in BEV features. Finally, we propose a Dilated Attention Enhancement Block (DAEB), which expands the receptive field through dilated convolution and integrates parameter-free attention mechanism (SimAM) for feature enhancement, ensuring efficiency while improving overall feature representation. Experimental results on nuScenes validation set show that BEV-Nexus outperforms it baseline (BEVFusion) by 1.8% mAP and 1.5% NDS. On the test set, BEV-Nexus improves mAP and NDS by 1.6% and 1.4%, respectively. Furthermore, the detection FPS remains nearly unchanged, demonstrating significant lightweight advantages. Full article
(This article belongs to the Special Issue AI-Powered Vision Sensing for Autonomous Driving)
Show Figures

Figure 1

18 pages, 48650 KB  
Article
PMS-Net: A Real-Time Instance Segmentation Framework with Large Receptive Fields for Urban Driving Scenes
by Ling Zhang and Zhuang Xiong
Algorithms 2026, 19(8), 620; https://doi.org/10.3390/a19080620 - 24 Jul 2026
Viewed by 114
Abstract
With the rapid development of edge AI chips and autonomous driving algorithms, autonomous driving perception systems are evolving toward higher accuracy, lower latency, and lightweight deployment. This trend places greater demands on real-time instance segmentation algorithms in terms of multi-scale feature representation, spatial [...] Read more.
With the rapid development of edge AI chips and autonomous driving algorithms, autonomous driving perception systems are evolving toward higher accuracy, lower latency, and lightweight deployment. This trend places greater demands on real-time instance segmentation algorithms in terms of multi-scale feature representation, spatial detail modeling, and edge deployment efficiency. To address these challenges, this paper proposes PMS-Net (Progressive Multi-Scale Network), a network designed for real-time instance segmentation. PMS-Net adopts a progressive multi-scale feature modeling mechanism that progressively enlarges the receptive field while integrating semantic and fine-grained spatial information across different scales. This enables efficient collaboration between local features and global contextual information, thereby enhancing scale-awareness and feature representation while maintaining a lightweight architecture. In addition, efficient feature encoding and dynamic feature reconstruction are incorporated to further improve spatial alignment, boundary recovery, and semantic continuity for complex scene modeling. Experimental results show that PMS-Net achieves 36.7% Mask mAP50 and 175 FPS on the Cityscapes dataset, outperforming the baseline by 2.9%. Deployment experiments on the NVIDIA Jetson Orin NX platform further demonstrate that PMS-Net achieves 35.2% Mask mAP50 and 98 FPS, improving the baseline by 3.7% and 14.0%, respectively. These results validate the effectiveness and practicality of PMS-Net for real-time edge-deployed autonomous driving applications. Full article
Show Figures

Figure 1

26 pages, 867 KB  
Article
Physics-Guided Multi-GSO Spectral Filtering for Degradation-Aware Automotive Radar Point-Cloud Detection
by Xiuping Li, Xiyan Sun, Yuanfa Ji, Jingjing Li, Wentao Fu, Songke Zhao, Wenbin Liang, Xizi Jia and Jian Liu
Sensors 2026, 26(15), 4714; https://doi.org/10.3390/s26154714 - 24 Jul 2026
Viewed by 70
Abstract
Automotive millimeter-wave radar produces sparse point clouds with Doppler velocity and radar cross-section (RCS), but graph detectors typically use a shared representation for semantic prediction and box regression despite their different propagation requirements. We propose multi-GSO spectral filtering (MGSF), a residual module that [...] Read more.
Automotive millimeter-wave radar produces sparse point clouds with Doppler velocity and radar cross-section (RCS), but graph detectors typically use a shared representation for semantic prediction and box regression despite their different propagation requirements. We propose multi-GSO spectral filtering (MGSF), a residual module that filters radar features over geometry-, Doppler-, and RCS-defined graph shift operators and fuses diffusion and residual components with a node-adaptive gate. MGSF-TD applies full multi-GSO refinement to semantic prediction and geometry-only refinement to box regression. On the complete RadarScenes validation set, MGSF-TD improves the official RadarGNN checkpoint from 60.19 to 60.59 mAP and from 74.06 to 75.10 mean foreground F1 (FG-F1). Across three MGSF-TD training seeds, the FG-F1 margin under RCS noise increases from +1.15 at 3 dBsm to +2.38 at 20 dBsm; seed-42 full-validation mAP margins are +0.24, +1.07, and +1.82. Controls show that geometry-only diffusion explains part of the gain and the RCS operator contributes most clearly at low-to-moderate noise, whereas a parameter-matched widened baseline matches or exceeds MGSF-TD under severe RCS and Doppler corruption. Cross-sensor diagnostics reproduce the Doppler failure trend but not the severity-dependent RCS gain. MGSF-TD therefore offers a balanced, physically interpretable operating point rather than a universal robustness gain. Full article
24 pages, 3902 KB  
Article
SonarReg-GS SLAM: Sparse Sonar-Guided Depth Regularization for Underwater Gaussian Splatting SLAM
by Wen Yang, Xiaolong Qian, Xulin Liu and Jianxing Leng
Sensors 2026, 26(15), 4713; https://doi.org/10.3390/s26154713 - 24 Jul 2026
Viewed by 98
Abstract
3D Gaussian Splatting (3DGS) SLAM provides an explicit scene representation for dense tracking and mapping, which is useful for underwater robotic perception. However, underwater monocular 3DGS SLAM lacks reliable metric depth cues: monocular depth estimation can provide dense structural priors, but its scale [...] Read more.
3D Gaussian Splatting (3DGS) SLAM provides an explicit scene representation for dense tracking and mapping, which is useful for underwater robotic perception. However, underwater monocular 3DGS SLAM lacks reliable metric depth cues: monocular depth estimation can provide dense structural priors, but its scale and reliability often degrade under underwater appearance changes. Forward-looking sonar (FLS) provides range–azimuth acoustic measurements whose range coordinate is related to physical distance, but raw sonar observations are sparse, noisy, and ambiguous. Our key insight is that FLS returns can serve as sparse metric depth anchors when they are associated with visually detected object regions. Based on this insight, we propose SonarReg-GS SLAM, an underwater visual–acoustic 3DGS SLAM framework with sparse sonar-guided depth regularization. Given synchronized RGB and sonar inputs, SonarReg-GS SLAM uses object masks to constrain the search space for acoustic range association. Filtered sonar responses are selected as sparse metric anchors through object-aware sampling, bearing-to-beam gating, and valid-pair checking. These anchors regularize the scale of monocular depth and generate metric depth priors for Gaussian initialization and tracking. An object-aware RGB mask loss further increases supervision on detected object regions while preserving full-scene mapping. Experiments on two public RGB–sonar underwater datasets show that SonarReg-GS SLAM improves tracking accuracy and mapping quality compared with representative classical SLAM and Gaussian Splatting SLAM baselines. Compared with Splat-SLAM, our method reduces the average ATE RMSE from 0.1296 m to 0.1015 m on UXO and from 0.5687 m to 0.4640 m on OPTI, corresponding to relative reductions of 21.7% and 18.4%, respectively. For rendering-based mapping, it increases the average PSNR from 28.05 dB to 29.61 dB on UXO and from 20.91 dB to 28.63 dB on OPTI while reducing the average LPIPS from 0.345 to 0.173 and from 0.450 to 0.303, respectively. Full article
(This article belongs to the Section Sensors and Robotics)
Show Figures

Figure 1

14 pages, 4417 KB  
Article
MFIA-YOLO: A Small-Object Defect Detection Model for Transmission Line Spacers
by Jinlong Du, Xiangyu Wang, Haiyang Lu, Xiaoye Zhang, Quanlei Cui, Tong Zhang, Zhongyu Wei and Yujian Ding
Energies 2026, 19(15), 3481; https://doi.org/10.3390/en19153481 - 24 Jul 2026
Viewed by 154
Abstract
UAV inspection images of transmission line spacers often contain small-scale defect targets, complex backgrounds, and weak fine-grained features. We propose MFIA-YOLO, a small-object defect detection model for spacer defects. Using the lightweight YOLOv8n variant as the baseline, we introduce a Multi-scale Spatial Heterogeneous [...] Read more.
UAV inspection images of transmission line spacers often contain small-scale defect targets, complex backgrounds, and weak fine-grained features. We propose MFIA-YOLO, a small-object defect detection model for spacer defects. Using the lightweight YOLOv8n variant as the baseline, we introduce a Multi-scale Spatial Heterogeneous Convolution (MSHC) into the backbone network. This module enhances the extraction of multi-scale features from spacer defect targets. Before the SPPF module, we further construct a Feature Complementary module (FCM). The FCM embeds shallow spatial location information into deep semantic features, thereby alleviating spatial information degradation during small-object detection. In the detection head, a C2f_IAFF module is adopted to adaptively fuse features at different scales through iterative attention-based feature fusion. This design improves the representation of defect targets in complex scenes. In addition, a transmission line spacer defect dataset is constructed from UAV inspection images collected in Xilingol League, Inner Mongolia. Experimental results showed that MFIA-YOLO achieved an mAP50 of 96.53% and an mAP50–95 of 83.50%. Compared with representative YOLO-series models, MFIA-YOLO achieved a better balance between detection accuracy and model complexity. These results demonstrate its effectiveness for accurate spacer defect detection in transmission lines. Full article
Show Figures

Figure 1

15 pages, 11675 KB  
Proceeding Paper
3D Models for Structural Analysis—Tests on Procedures and Point Cloud Processing
by Sara Gonizzi Barsanti
Eng. Proc. 2026, 149(1), 2; https://doi.org/10.3390/engproc2026149002 - 24 Jul 2026
Viewed by 80
Abstract
In recent years, Neural Radiance Fields (NeRFs) and 3D Gaussian Splatting (3DGS) have emerged as promising new approaches for 3D reconstruction. NeRFs rely on neural fields that generate a three-dimensional representation of a scene from photographs, estimating reflectance properties and reconstructing the underlying [...] Read more.
In recent years, Neural Radiance Fields (NeRFs) and 3D Gaussian Splatting (3DGS) have emerged as promising new approaches for 3D reconstruction. NeRFs rely on neural fields that generate a three-dimensional representation of a scene from photographs, estimating reflectance properties and reconstructing the underlying geometry. Since their introduction in 2020, NeRFs have attracted significant attention due to their wide range of potential applications. Conversely, 3D Gaussian Splatting (3DGS), introduced in 2023, employs Gaussian primitives to efficiently model objects and structures, offering a flexible and adaptive representation of 3D scenes. Starting from an established pipeline for the use of reality-based models for structural analysis, this paper investigates the performance of 3DGS in handling complex geometries and surfaces characterised by challenging acquisition conditions. Full article
Show Figures

Figure 1

20 pages, 16125 KB  
Article
KP-SLAM: Joint Flow-Pointmap Prior Synchronization for Robust Consistent Dense Mapping
by Song Gao, Xinyu Huang, Zheng Huang and Xinyu Wei
Symmetry 2026, 18(8), 1248; https://doi.org/10.3390/sym18081248 - 23 Jul 2026
Viewed by 161
Abstract
Monocular RGB dense SLAM remains challenging because depth and global metric scale are not directly observable from a single camera. Existing systems often combine optical-flow and monocular-geometry priors predicted by independently trained networks, which can provide inconsistent constraints to bundle adjustment (BA). Our [...] Read more.
Monocular RGB dense SLAM remains challenging because depth and global metric scale are not directly observable from a single camera. Existing systems often combine optical-flow and monocular-geometry priors predicted by independently trained networks, which can provide inconsistent constraints to bundle adjustment (BA). Our quantitative prior-consistency analysis indicates that this disagreement is an important contributor to unstable local optimization and reconstruction error rather than the sole cause of drift. We propose KP-SLAM, which predicts dense optical flow and paired pointmap priors from a shared representation and incorporates them into the same BA backend. We further introduce a Depth-Scale-Pose-to-Pointmap (DSPP) objective that relates optimized inverse depth, edge-wise relative scale, and camera pose to paired pointmap constraints. Experiments on ScanNet, TUM-RGBD, KITTI, Tanks-and-Temples, and dynamic sequences show improved tracking, depth, and rendering metrics over the compared RGB-only baselines under the reported settings. The results support the usefulness of synchronized priors while also revealing remaining limitations in highly dynamic, weakly textured, and large-scale scenes. Full article
(This article belongs to the Section A: Computer Science)
Show Figures

Figure 1

25 pages, 4394 KB  
Article
DCAF-Net: Density-Conditioned Attention Fusion Network for Single-Image Dehazing
by Nianfeng Li, Shaojie Liu, Hongjie Ding, Shenyan Gao, Zhiguo Xiao and Qian Liu
Sensors 2026, 26(14), 4656; https://doi.org/10.3390/s26144656 - 22 Jul 2026
Viewed by 143
Abstract
Single-image dehazing aims to recover clear scenes from degraded images affected by atmospheric scattering, serving as a critical preprocessing technique for improving the imaging quality of visual sensors. Existing deep learning-based dehazing methods exhibit limited generalization ability in real-world scenarios, primarily due to [...] Read more.
Single-image dehazing aims to recover clear scenes from degraded images affected by atmospheric scattering, serving as a critical preprocessing technique for improving the imaging quality of visual sensors. Existing deep learning-based dehazing methods exhibit limited generalization ability in real-world scenarios, primarily due to the spatial non-uniformity of haze and its coupling with illumination and texture degradation, as well as the scarcity of real paired data. To address these issues, this paper proposes a haze-density conditional attention fusion network (DCAF-Net). The network employs an adaptive haze density perception module to fuse priors such as the dark channel, local contrast, and saturation, generating a spatial haze density guidance map. This map is then embedded as conditional information into the multi-scale feature modulation and attention fusion process, enabling adaptive restoration of regions with different degradation levels. Furthermore, a residual dense cascaded feature enhancement module is designed to leverage feature reuse, gated fusion, and residual learning to enhance the representational capacity of deep features. Training adopts a joint optimization objective combining Charbonnier reconstruction loss, perceptual contrast loss, and structural similarity loss. Experimental results demonstrate that DCAF-Net achieves competitive performance against representative methods on multiple synthetic and real-world hazy datasets, and shows promising restoration performance on representative real-world hazy scenes, and can provide high-quality image preprocessing support for visual-sensor-based intelligent perception systems. Full article
(This article belongs to the Special Issue Intelligent Sensing and Digital Signal Processing in Smart Data)
Show Figures

Figure 1

40 pages, 90250 KB  
Perspective
Multi-Exposure HDR Imaging: A Review of Pixel-Level and Feature-Level Reconstruction Methods
by Qian Tao, Wei Wang, Chaobing Zheng and Zhengguo Li
Sensors 2026, 26(14), 4649; https://doi.org/10.3390/s26144649 - 22 Jul 2026
Viewed by 128
Abstract
Multi-exposure is an efficient way to capture real-world high-dynamic-range (HDR) scenes. However, HDR imaging suffers from severe ghosting artifacts in dynamic scenes due to the temporal gap between sequential exposures. In this article, we categorize the literature on two important topics on HDR [...] Read more.
Multi-exposure is an efficient way to capture real-world high-dynamic-range (HDR) scenes. However, HDR imaging suffers from severe ghosting artifacts in dynamic scenes due to the temporal gap between sequential exposures. In this article, we categorize the literature on two important topics on HDR imaging: multi-exposure fusion (MEF) and ghost removal. Conventional filter-based and data-driven methods are studied in pixel space and feature space. For popular deep learning-based approaches, we provide a granular taxonomy based on their alignment and fusion domains: pixel-space methods, which typically employ explicit motion compensation such as optical flow or spatial transformers, and feature-space methods, which leverage implicit alignment through deformable convolutions, attention mechanisms, or latent representation merging. Representative works are compared across different supervision settings, and key design principles are summarized. In addition, this survey summarizes commonly used datasets and evaluation metrics, discussing their applicability under diverse output forms. Finally, major bottlenecks and promising directions for future research are outlined. Full article
(This article belongs to the Special Issue Perspectives in Intelligent Sensors and Sensing Systems)
Show Figures

Figure 1

27 pages, 24217 KB  
Article
Examining Controlled Spaces in the Context of Environmental Psychology: A Case of One Flew over the Cuckoo’s Nest and the Shawshank Redemption
by Sonay Ayyıldız and Melike Yıldız
Buildings 2026, 16(14), 2912; https://doi.org/10.3390/buildings16142912 - 22 Jul 2026
Viewed by 237
Abstract
Controlled institutional environments such as prisons and psychiatric hospitals shape human behaviour through surveillance, discipline, and spatial regulation. Although these environments have been extensively examined from architectural and sociological perspectives, comparatively few studies have systematically investigated their psychological implications using an environmental psychology [...] Read more.
Controlled institutional environments such as prisons and psychiatric hospitals shape human behaviour through surveillance, discipline, and spatial regulation. Although these environments have been extensively examined from architectural and sociological perspectives, comparatively few studies have systematically investigated their psychological implications using an environmental psychology framework. This study examines how controlled institutional spaces influence environmental perception, privacy, belonging, proxemic relationships, crowding, environmental stress, and spatial discipline through a comparative analysis of One Flew Over the Cuckoo’s Nest (1975) and The Shawshank Redemption (1994). A qualitative comparative film analysis was conducted using a structured analytical framework derived from environmental psychology and supported by theories of surveillance, discipline, total institutions, and hegemony. Twenty-four candidate scenes were initially identified and independently coded by two researchers according to predefined observable spatial and behavioural indicators. Following the coding process, four scenes exhibiting overlapping analytical characteristics were excluded, resulting in a final dataset of twenty scenes. The findings demonstrate that architectural elements such as surveillance visibility, hierarchical spatial organization, controlled circulation, restricted privacy, and shared institutional environments shape psychological experiences including alienation, environmental stress, adaptation, resistance, and place attachment. At the same time, the comparative analysis shows that similar spatial conditions may generate different psychological responses depending on individuals’ interactions with institutional environments. The study contributes methodologically to qualitative film analysis through a systematic scene selection and coding procedure while providing a comprehensive framework for examining environmental psychology concepts through the cinematic representations of controlled institutional environments. The findings should be interpreted as analyses of cinematic representations rather than empirical evidence of real prison or psychiatric hospital environments. Full article
(This article belongs to the Section Architectural Design, Urban Science, and Real Estate)
Show Figures

Figure 1

18 pages, 3494 KB  
Article
Towards Rotated Object Detection with Pose-Aware and Dense Feature Modulation
by Dehua Bai, Donglin Jing and Meng Zhao
Algorithms 2026, 19(7), 605; https://doi.org/10.3390/a19070605 - 22 Jul 2026
Viewed by 168
Abstract
Remote sensing image object detection is a core task in computer vision, which plays a vital role in intelligent transportation, port monitoring, and infrastructure management. However, rotated dense objects in remote sensing scenes suffer from severe challenges, including arbitrary 0–360° pose variations, large-scale [...] Read more.
Remote sensing image object detection is a core task in computer vision, which plays a vital role in intelligent transportation, port monitoring, and infrastructure management. However, rotated dense objects in remote sensing scenes suffer from severe challenges, including arbitrary 0–360° pose variations, large-scale differences, dense spatial aggregation, and blurred boundaries. Traditional Convolutional Neural Networks rely on fixed sampling grids and receptive fields, failing to adaptively capture the dynamic morphological and pose features of tilted targets. Meanwhile, existing methods struggle to address feature coupling and boundary misjudgment among densely arranged objects, leading to degraded detection accuracy. To tackle these bottlenecks, we propose an adaptive detection framework named PDNet for rotated dense object detection. The framework integrates four key designs: First, a Pose-Aware Dynamic Sampling Mechanism (PDSM) is developed to estimate the target principal axis in real time and learn a deformable offset field, which dynamically adjusts the convolution sampling pattern and receptive field shape to adapt to target pose variations. Second, a Dense-Scene Feature Modulation Mechanism (DSFM) constructs a dynamic weight field based on local feature responses to enhance discriminative target features and suppress inter-target interference in dense regions. Third, the Strip-Based Context Attention (SCA) module fuses global and local contextual information to strengthen the representation of small and weak targets. Fourth, a boundary-aware rotation loss function is designed to optimize the regression accuracy of rotated bounding boxes via pixel-level supervision. Extensive experiments on DOTA-v1.0, AI-TOD achieve 79.86% mAP and 25.95% AP, outperforming state-of-the-art methods. Full article
(This article belongs to the Special Issue Advances in Deep Learning-Based Data Analysis)
Show Figures

Figure 1

40 pages, 52553 KB  
Article
An Adaptive Low-Light Image Enhancement Framework via Metaheuristic-Optimized Inverted Dehazing and Gamma Correction with Global Limits
by Cheng-Hsiung Hsieh, Xin-Rui Lin, Chia-Hsin Cheng, Yung-Hoh Sheu and Yung-Fa Huang
Electronics 2026, 15(14), 3210; https://doi.org/10.3390/electronics15143210 - 21 Jul 2026
Viewed by 132
Abstract
Low-light image enhancement (LLIE) is a fundamental task in computer vision, required for restoring luminance, contrast, and structural fidelity in images captured under suboptimal lighting environments. This paper introduces an optimization-driven, scene-adaptive LLIE framework, designated as [...] Read more.
Low-light image enhancement (LLIE) is a fundamental task in computer vision, required for restoring luminance, contrast, and structural fidelity in images captured under suboptimal lighting environments. This paper introduces an optimization-driven, scene-adaptive LLIE framework, designated as OMIDCPGCGL, which exploits the optical duality between low-light inversion and atmospheric scattering. The proposed methodology transforms low-light inputs into quasi-haze representations through an optical inversion process, followed by structural restoration using an Improved Dark Channel Prior (MIDCP) baseline. To refine the restored output, a Gamma Correction with Global Limits (GCGL) module is integrated as a boundary constraint to mitigate localized over-exposure and preserve chromatic consistency. A core novelty of this framework lies in the deployment of metaheuristic optimization algorithms (MOAs)—specifically the Gray Wolf Optimizer (GWO), Harris Hawks Optimization (HHO), and Marine Predators Algorithm (MPA)—to autonomously resolve optimal, image-specific parameter configurations. This search paradigm is guided by perception-driven fitness functions, namely the Patch-based Contrast Quality Index (PCQI) or the Blind/Referenceless Image Spatial Quality Evaluator (BRISQUE). Quantitative and qualitative evaluations across a comprehensive pool of 1092 benchmark images demonstrate that the proposed framework exhibits robust statistical resilience and cross-dataset generalization compared to four state-of-the-art deep learning methods. While data-driven deep learning architectures retain localized superiority under the extreme degradation boundaries of the DARK FACE dataset, the proposed physics-inspired optimization framework achieves the leading overall cross-dataset aggregate ranking (R¯=2.467) across diverse evaluation environments due to its per-image dynamic solution space mapping. Full article
Show Figures

Figure 1

23 pages, 2800 KB  
Article
A Topology-Aware Lane Instance Segmentation Network Based on Dynamic Snake Convolution and High-Resolution Feature Pyramid
by Sen Wang, Baohua Guo, Weifan Gu, Xiaoyu Zhang, Anthony Sigama and David Bassir
Electronics 2026, 15(14), 3213; https://doi.org/10.3390/electronics15143213 - 21 Jul 2026
Viewed by 148
Abstract
Accurate lane detection is essential for autonomous driving, but the elongated geometry of lane markings and the information loss caused by perspective projection still make continuous lane perception challenging. Existing Convolutional Neural Network (CNN) methods are limited by fixed receptive fields and repeated [...] Read more.
Accurate lane detection is essential for autonomous driving, but the elongated geometry of lane markings and the information loss caused by perspective projection still make continuous lane perception challenging. Existing Convolutional Neural Network (CNN) methods are limited by fixed receptive fields and repeated downsampling, which can weaken narrow curvilinear structures and distant lane details. To address these issues, this paper proposes an enhanced single-stage lane instance segmentation framework based on YOLOv11. Dynamic Snake Convolution (DSConv) is integrated into the feature extraction network as an adaptive sampling operator for finite-width lane-mask representations, enabling local sampling trajectories to better follow narrow curved lane structures. A high-resolution P2 branch (Stride = 4) is further constructed in the Feature Pyramid Network (FPN) to compensate for fine-grained spatial details lost in deep feature maps. A reproducible point-to-mask training and mask-to-point evaluation pipeline is also established for the sparse TuSimple annotations. Experiments on the TuSimple benchmark, which mainly represents structured daytime highway scenes, show that the proposed method improves mask Precision, Recall, and mAP50 by 8.5, 9.0, and 6.8 percentage points, respectively, compared with YOLOv11-seg, while reaching 103.09 FPS under the reported testing configuration. Qualitative results on a newly constructed CULane subset further indicate improved lane continuity under reflections, illumination changes, night scenes, and complex urban backgrounds, while failure cases reveal remaining sensitivity to lane-like linear structures and weak lane visibility. The results suggest a practical accuracy–efficiency trade-off within the YOLO-style segmentation framework, although robustness in broader real-world scenarios still requires further verification. Full article
Show Figures

Figure 1

25 pages, 2839 KB  
Article
UAV RF Signal Azimuth Estimation Using a UCA-8 and a Dual-Branch Circular-Regression Network
by Jingyang Wang, Jie Ma, Jiaxi Zhang, Zehan Li and Min Huang
Information 2026, 17(7), 705; https://doi.org/10.3390/info17070705 - 21 Jul 2026
Viewed by 169
Abstract
To address the problems that existing UAV RF signal azimuth estimation methods rely on idealized simulation data and lack accuracy and robustness in complex environments, a high-precision azimuth estimation method based on an improved ResNet, namely the Dual-Branch Circular-Regression Network, is proposed. Firstly, [...] Read more.
To address the problems that existing UAV RF signal azimuth estimation methods rely on idealized simulation data and lack accuracy and robustness in complex environments, a high-precision azimuth estimation method based on an improved ResNet, namely the Dual-Branch Circular-Regression Network, is proposed. Firstly, UCA-8 array data is generated from measured single-channel RF signals, and non-ideal factors such as channel mismatch and mutual coupling among array elements are incorporated to simulate the real RF receiving environment. Secondly, ResNet is improved from three aspects: input normalization, dynamic dual-branch (DDB) learning features and periodic angle regression. The input normalization strategy based on Per-Sample Complex Root Mean Square (PSCRMS) is adopted to improve the adaptability of the model to signal scale changes. The DDB structure is adopted to adaptively fuse I/Q spatiotemporal features with a spatial covariance statistical prior to enhance the spatial feature expression ability in complex scenes. A periodic angle regression method based on Unit Circular Vector Representation (UCVR) and the Huber Loss (GAH Loss) of geodesic angle distance is adopted to realize periodic angle continuous modeling and suppress abnormal angle errors. Finally, comparative and ablation experiments are conducted on the constructed UCA-8 dataset. The experimental results show that compared with the baseline ResNet, the Dual-Branch Circular-Regression Network achieves 85.8%, 95.3%, and 81.5% reductions in MAE, RMSE and P95, respectively, and maintains higher estimation accuracy and good robustness under low signal-to-noise ratio, hardware mismatch and co-frequency dual-source interference. Full article
Show Figures

Figure 1

Back to TopTop