Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

Search Results (315)

Search Parameters:
Keywords = filter spatial attention

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
23 pages, 2401 KB  
Article
A Multi-Site Probabilistic Water Quality Prediction Method Coupling Learnable Frequency-Domain Filtering and Multi-Residual Ensemble
by Wei Shao, Yuliang Wang and Lijuan Qiao
Water 2026, 18(17), 2060; https://doi.org/10.3390/w18172060 (registering DOI) - 22 Aug 2026
Abstract
Multi-site water quality sequences are jointly affected by seasonal periodicity, meteorological disturbances, and inter-site differences in the Jianghuai Watershed region. Conventional quality prediction models struggle to simultaneously achieve multi-scale feature extraction, spatial heterogeneity characterization, and prediction uncertainty expression. This study used daily-scale monitoring [...] Read more.
Multi-site water quality sequences are jointly affected by seasonal periodicity, meteorological disturbances, and inter-site differences in the Jianghuai Watershed region. Conventional quality prediction models struggle to simultaneously achieve multi-scale feature extraction, spatial heterogeneity characterization, and prediction uncertainty expression. This study used daily-scale monitoring data on dissolved oxygen (DO), pH, and ammonia nitrogen (NH3N) from 32 monitoring stations within the region in 2025 and proposed the FT-TransONet (Fourier-enhanced Temporal Transformer Operator Network) multi-site probabilistic water quality prediction model. Within a Transformer framework, the model employed a FourierTime learnable frequency-domain filtering module, a GeoBias (Geographic Bias) attention bias mechanism, and a multi-residual ensemble strategy composed of a multilayer perceptron (MLP), a gated recurrent unit (GRU), and a temporal convolutional network (TCN) combined with a mass conservation constraint, thereby achieving both point and interval prediction of key water quality indicators. The results showed that FT-TransONet achieved the lowest Macro_RMSE among all compared methods on the multi-site water quality prediction task. At a prediction horizon of three days, its Macro_RMSE reached 0.2255, which was 21.89% lower than that of the long short-term memory network and 5.57% lower than that of the strongest baseline MC-Dropout. For the three individual indicators, the model attained coefficients of determination of 0.9135, 0.9338, and 0.8660 for dissolved oxygen, pH, and ammonia nitrogen, with corresponding root-mean-square errors of 0.5145, 0.1098, and 0.0523, confirming its potential to characterize the temporal variation in the main water quality indicators. Under multi-step prediction, the error grew gently, with the Macro_RMSE rising only from 0.2255 to 0.2384 as the horizon extended from three to seven days, and the ablation experiments, together with the probabilistic prediction results, further supported the effectiveness of the proposed structural design. Validated on 32 water quality monitoring stations in the Jianghuai Watershed, the method improved multi-site prediction accuracy while accounting for stability and uncertainty quantification, providing a preliminary reference for regional water quality early warning and management. Full article
34 pages, 7079 KB  
Article
AI-Assisted Scan-to-BIM for Masonry Arch Bridges: From Point Cloud Segmentation to Parametric Heritage BIM Reconstruction
by Vincenzo Saverio Alfio, Massimiliano Pepe, Donato Palumbo, Ahmed Kamal Hamed Dewedar and Domenica Costantino
Appl. Sci. 2026, 16(16), 8248; https://doi.org/10.3390/app16168248 - 19 Aug 2026
Viewed by 106
Abstract
This paper shows an AI-assisted workflow supported by a Large Language Model for the geometric and informative reconstruction of masonry arch bridges from multi-source input data. The proposed methodology combines point cloud preprocessing, vegetation filtering, AI-based or semi-automatic segmentation, geometric feature extraction, 3D [...] Read more.
This paper shows an AI-assisted workflow supported by a Large Language Model for the geometric and informative reconstruction of masonry arch bridges from multi-source input data. The proposed methodology combines point cloud preprocessing, vegetation filtering, AI-based or semi-automatic segmentation, geometric feature extraction, 3D mesh generation, and parametric Heritage Building Information Modeling integration. The study specifically aims to examine how the quantity and type of input information influence the reconstruction process and the reliability of the resulting HBIM model. The input data considered include Point Clouds (PC), photographic images, and descriptive information related to the bridge geometry, construction features, and visible architectural components. Particular attention is given to the initial geometric characteristics of the point cloud, including its point density, spatial distribution, level of completeness, and overall number of points, to assess how point cloud numerosity affects the accuracy and level of detail of the reconstructed geometry. Starting from a dense 3D survey, the workflow identifies and reconstructs the main architectural and structural components of a masonry arch bridge, including the arch, intrados, parapets, masonry walls, roadway surface, abutments, and cutwaters. The extracted geometry is converted into a clean 3D mesh and subsequently structured into parametric HBIM objects suitable for documentation, conservation, structural assessment, and future monitoring activities. A Cloud-to-Mesh comparison is performed to evaluate the geometric accuracy of the reconstructed model with respect to the original point cloud. The results demonstrate the potential of hybrid AI and geometric approaches to improve the efficiency, repeatability, and reliability of Scan-to-BIM processes for historical masonry bridge heritage. Furthermore, they show that the geometric quality of the HBIM model depends primarily on the density, spatial distribution and completeness of the structural points, rather than on their total number. Well-distributed point clouds, in fact, allow for more reliable reconstructions than larger datasets characterised by uneven coverage. Full article
Show Figures

Figure 1

27 pages, 5829 KB  
Article
A U-Net-Based Hybrid Network with Large-Kernel Convolution and Spatially Reduced Self-Attention for Structural Plane Segmentation in Tunnel-Face Images
by Zhenglan Lu, Junfan Fu, Hongliang Liu and Jiakai Tian
Appl. Sci. 2026, 16(16), 8115; https://doi.org/10.3390/app16168115 - 14 Aug 2026
Viewed by 189
Abstract
The structural planes exposed on tunnel faces reflect the integrity and geological conditions of surrounding rock, and their accurate segmentation is important for geological logging and surrounding rock assessment. However, these targets are typically slender, discontinuous, and low-contrast, and are easily confused with [...] Read more.
The structural planes exposed on tunnel faces reflect the integrity and geological conditions of surrounding rock, and their accurate segmentation is important for geological logging and surrounding rock assessment. However, these targets are typically slender, discontinuous, and low-contrast, and are easily confused with dust, shadows, water seepage reflections, and blasting-induced textures. To address these challenges, this study proposes a hybrid network based on a U-shaped architecture (U-Net) for structural plane segmentation in tunnel-face images. The network combines a large-kernel detail enhancement module for preserving weak boundaries and fine linear features, a global dependency modeling module based on spatially reduced self-attention for capturing long-range relationships among discontinuous structural plane segments, and a semantic-guided filtering module for suppressing background interference in skip connections. Comparative and ablation experiments were conducted on a self-built tunnel-face structural plane image dataset containing 600 images. The results show that the proposed method achieves Dice, intersection over union (IoU), Precision, and Recall values of 69.28%, 53.00%, 68.92%, and 69.64%, respectively. Compared with the baseline U-Net, the Dice and IoU scores are improved by 6.70 and 7.49 percentage points, respectively. These results demonstrate that the proposed model achieves a favorable balance among local detail preservation, global connectivity modeling, and complex background suppression. It provides a reliable foundation for structural plane parameter extraction, surrounding rock assessment, and intelligent geological logging. Full article
(This article belongs to the Special Issue Advanced Tunnel and Underground Engineering Technology)
Show Figures

Figure 1

41 pages, 29978 KB  
Article
Attention-Guided Cross-Connected Filters Convolutional Neural Network with Surrogate-Based Interpretability for Image Splicing Forgery Detection
by Aruna Srinivasan, Surabhi Narayan and Aarnav Sandeep Deshmukh
Computers 2026, 15(8), 525; https://doi.org/10.3390/computers15080525 - 13 Aug 2026
Viewed by 181
Abstract
Background: Image splicing forgery detection is one of the most challenging problems in the field of image forensics as it involves identifying and localizing suspicious regions that are created by integrating contents from one or more different sources. The accurate detection and classification [...] Read more.
Background: Image splicing forgery detection is one of the most challenging problems in the field of image forensics as it involves identifying and localizing suspicious regions that are created by integrating contents from one or more different sources. The accurate detection and classification of splicing forgery still remains a difficult task because of the existence of overlapping image regions, which makes it complex to differentiate authentic and tampered images. This overlap causes a lack of feature representation, making it difficult to precisely detect the tampering in images. Also, the decision-making process of the model is often a black box, which makes it challenging to interpret and understand the rationale behind its decisions. Methods: To address these challenges, a Convolutional Block Attention Module (CBAM)–U-Net with Cross-Connected Filters–Convolutional Neural Network (CCF-CNN) is proposed to achieve precise detection and localization of spliced regions. The CBAM enhances spatial and channel-wise attention, enabling accurate localization of forged regions. The dual-phase CCF-CNN is incorporated with cross-connected filters to differentiate between the authentic and tampered regions by extracting global and local features. Additionally, a surrogate heatmap mechanism is introduced using intermediate decoder features to generate patch-level visual explanations, enabling precise localization of the spliced regions, thereby improving the model’s transparency in decision-making. Results: The proposed CCF-CNN obtains a high accuracy of 99.84% on the CASIA 2.0 dataset and an accuracy of 95.63% on the MISD. Conclusions: Compared to traditional CNNs such as VGG, ResNet and attention-based interpretability algorithms, the proposed model obtains higher performance in terms of detection and interpretability. Full article
(This article belongs to the Section ICT Infrastructures for Cybersecurity)
Show Figures

Figure 1

21 pages, 8105 KB  
Article
Bidirectional Cross-Level Feature Interaction and Context-Aware Multi-Scale Attention for Crowd Counting
by Zhifan Jin, Lin Zhou, He Wang, Sijia Chen, Liman Liu and Wenbing Tao
AI 2026, 7(8), 313; https://doi.org/10.3390/ai7080313 - 13 Aug 2026
Viewed by 250
Abstract
Crowd counting estimates the number and spatial distribution of people in images and videos, supporting smart city management and public safety. Existing methods often rely on intra-level feature refinement and simple cross-scale fusion, such as concatenation or addition, which limits interaction between fine-grained [...] Read more.
Crowd counting estimates the number and spatial distribution of people in images and videos, supporting smart city management and public safety. Existing methods often rely on intra-level feature refinement and simple cross-scale fusion, such as concatenation or addition, which limits interaction between fine-grained spatial details and high-level semantic representations. In addition, the limited receptive field of convolutional networks restricts global context modeling in scenes with heavy occlusion and extreme scale variation. To address these challenges, we propose a Hierarchical Context-Aware Multi-Scale Attention Network (HCMA). Its bidirectional cross-level interaction is realized through two complementary top-down decoding streams, where an attention-gating stream provides spatial guidance for the counting-oriented representations carried by a density-feature stream. HCMA includes three modules: the Selective Context-Aware Attention Module (SCAM), which performs context-dependent multi-scale filtering; Dynamic Positional Pooling (DPP), which introduces an image-level mean token and stochastic global-relation aggregation; and the Multi-Scale Enhancement Attention Module (MSEA), which refines high-level semantic features under scale variation. Experiments on ShanghaiTech, UCF-QNRF, and NWPU-Crowd show competitive counting accuracy across scenes with different density ranges, scale variation, and occlusion. In particular, HCMA achieves an MAE of 73.2 on NWPU-Crowd, 17.2% lower than that of DM-Count. Full article
(This article belongs to the Special Issue AI and Computer Vision in Real-World and Industrial Applications)
Show Figures

Figure 1

34 pages, 4665 KB  
Article
Dynamic Low-Rank Modulation and Frequency-Domain Collaboration for Scene-Adaptive Image Fusion Network
by Yao Zhang, Lin Tian and Sirui Huang
Sensors 2026, 26(16), 5106; https://doi.org/10.3390/s26165106 - 12 Aug 2026
Viewed by 285
Abstract
Infrared and visible image fusion aims to integrate the complementary information from heterogeneous sensors, thereby enhancing the robustness of visual perception in complex environments. Most existing methods, however, employ fixed network parameters, a characteristic that limits their adaptive modeling capabilities for cross-modal information [...] Read more.
Infrared and visible image fusion aims to integrate the complementary information from heterogeneous sensors, thereby enhancing the robustness of visual perception in complex environments. Most existing methods, however, employ fixed network parameters, a characteristic that limits their adaptive modeling capabilities for cross-modal information under scenarios such as drastic illumination changes, low light conditions, dense fog, and strong glare. To address this issue, we propose a scene-adaptive image fusion network, termed HL-Fuse, based on dynamic low-rank modulation and frequency-domain collaboration. For the spatial domain, the Hyper-LoRA is introduced via our designed SceneHyperNet, which mathematically constrains parameter variations within a low-rank subspace to adaptively calibrate attention mappings according to the global scene information. For the frequency domain, a tailored FAM is introduced to bridge spatial-domain feature aggregation and explicit spectrum reweighting by implementing targeted high- and low-frequency filtering, thereby enhancing edge and texture representation. Experiments conducted on the MSRS, TNO, M3FD, and FMB datasets demonstrate that HL-Fuse achieves competitive performance in terms of both multiple objective metrics and subjective visual quality, while the overall performance in the MSRS downstream object detection task is also enhanced. These results indicate the potential value of HL-Fuse for complex scene perception and remote sensing applications. Full article
Show Figures

Figure 1

23 pages, 7827 KB  
Article
A Visual Detection and Multi-Zone Personnel Safety Control Method for Firework Manufacturing Workshops
by Xiaoxi Yan, Hongwei Tao, Biao Xiong, Hui Wang, Wenhao Luo and Peiqiang Tian
Electronics 2026, 15(16), 3547; https://doi.org/10.3390/electronics15163547 - 10 Aug 2026
Viewed by 196
Abstract
Real-time visual detection is essential for personnel safety control in firework manufacturing workshops, where inadequate personnel-count control can increase safety risks in hazardous production zones. This study proposes a visual detection and multi-zone personnel safety control method that combines edge-oriented personnel detection, cross-camera [...] Read more.
Real-time visual detection is essential for personnel safety control in firework manufacturing workshops, where inadequate personnel-count control can increase safety risks in hazardous production zones. This study proposes a visual detection and multi-zone personnel safety control method that combines edge-oriented personnel detection, cross-camera identity association, and polygon-based boundary filtering. The detection module is built on YOLO26, which supports inference without non-maximum suppression (NMS). A Global Attention Mechanism (GAM) is adopted instead of the Convolutional Block Attention Module (CBAM) because its sequential channel-spatial attention preserves cross-dimensional interactions without the global pooling operations used in CBAM, thereby retaining weak spatial cues from small personnel targets. GAM is incorporated after the Spatial Pyramid Pooling-Fast (SPPF) module to improve detection under overhead views, dust, and occlusion. A cross-camera person re-identification (ReID) layer maintains identity consistency across workshops, while an Irregular Electronic Fence (IEF) excludes detections outside hazardous operating boundaries. On in-situ data collected from Deren Firework Co., Ltd., the proposed method achieves a 98.3% mean average precision at an intersection-over-union threshold of 0.5 (mAP@0.5) with a per-image central processing unit (CPU) processing time of 36.5 ms. Compared with YOLOv8n, this represents a 5.9 percentage point improvement in mAP@0.5 and a 54.6% latency reduction. A three-month field deployment detected 15 safety breaches and supported timely intervention in 4 critical overcrowding incidents, indicating the practical applicability of the method under the evaluated factory conditions. Full article
(This article belongs to the Section Computer Science & Engineering)
Show Figures

Figure 1

29 pages, 32553 KB  
Article
Dual-Module Bench-Line Extraction and Surface-Object Segmentation from UAV LiDAR Point Clouds in Open-Pit Mines Using Neighborhood Geometric Analysis and an Enhanced PointNet++ Network
by Shanfeng Ge, Nijia Qian, Jingxiang Gao, Xin Liu, Wenyuan Zhang, Yong Feng and Dehu Yang
Appl. Sci. 2026, 16(16), 7921; https://doi.org/10.3390/app16167921 - 8 Aug 2026
Viewed by 194
Abstract
Open-pit mines contain rapidly changing terrain, discontinuous bench structures, and mixed artificial–natural objects, which complicate automated three-dimensional mapping. This study presents a dual-module workflow for UAV LiDAR point clouds. Module A characterizes local geometry using normal and curvature descriptors, constructs local plane support [...] Read more.
Open-pit mines contain rapidly changing terrain, discontinuous bench structures, and mixed artificial–natural objects, which complicate automated three-dimensional mapping. This study presents a dual-module workflow for UAV LiDAR point clouds. Module A characterizes local geometry using normal and curvature descriptors, constructs local plane support through RANSAC fitting, and detects candidate bench-line points using an angular-gap criterion, followed by regional grouping and Kalman-filter refinement. Qualitative overlay with the orthophoto showed coherent correspondence with principal platform–slope transitions. Module B segments buildings, roads, and vegetation using a PointNet++ network enhanced by local Transformer self-attention and inverted residual feature transformation. Under a fixed spatial hold-out setting, the network achieved an overall accuracy of 97.6% and a mean intersection over union of 96.4%. It obtained the highest overall accuracy, mean intersection over union, and class-wise intersection over union among the selected baselines, whereas Point Transformer achieved a slightly higher mean class accuracy. The two independently operated modules provide complementary structural and semantic information for open-pit mine mapping. Broader applicability requires reference-based bench-line assessment and evaluation across additional mines and survey periods. Full article
Show Figures

Figure 1

30 pages, 5641 KB  
Article
A Method for Portal Crane Wire Rope Recognition Based on Improved PointNet++
by Xinyuan Li, Yujie Zhang and Yang Shen
J. Mar. Sci. Eng. 2026, 14(16), 1463; https://doi.org/10.3390/jmse14161463 - 8 Aug 2026
Viewed by 178
Abstract
In automated dry bulk terminal operations, accurate perception of the spatial pose of portal crane wire ropes is important for grab positioning and can provide geometric information for subsequent anti-sway control research. Vision-based measurements may be affected by metallic reflections, illumination variation, and [...] Read more.
In automated dry bulk terminal operations, accurate perception of the spatial pose of portal crane wire ropes is important for grab positioning and can provide geometric information for subsequent anti-sway control research. Vision-based measurements may be affected by metallic reflections, illumination variation, and dust occlusion, whereas inertial or mechanically coupled measurements may be affected by vibration and dynamic coupling. This study proposes a LiDAR-based method for wire rope point cloud segmentation and pose estimation using an improved PointNet++. Dual-LiDAR point clouds are aligned and filtered using a kinematic constraint-based Region of Interest (ROI) to reduce background redundancy. A Spatial Self-Attention (SSA) module is introduced to combine long-range semantic dependencies with local spatial weighting, improving the representation of sparse and fragmented wire rope points. The segmented wire rope points are separated by tk-means clustering and fitted with spatial lines for pose estimation. The complete acquisition comprises 11,348 annotated frames: a 9458-frame model development dataset from 1000 complete operating cycles, and a separately retained 1890-frame independent engineering test set from 200 condition-specific operating sequences. The development dataset was divided into mutually exclusive training and validation partitions at the level of complete operating cycles, and checkpoint selection was performed only on the validation set. Three independent training runs with fixed random seeds were conducted. On the independent test set, PointNet++ achieved an F1-score of 87.5 ± 0.2% and an mIoU of 79.0 ± 0.2%, whereas the complete proposed method achieved an F1-score of 92.8 ± 0.2% and an mIoU of 86.6 ± 0.2%. These results characterize performance on independent operating sequences collected from the crane and sensor configurations represented in the dataset. The standalone segmentation stage achieved 111.9 FPS, whereas the complete processing pipeline required slightly more than 2 s per frame because of frame-by-frame KD-ICP fine registration. Full article
Show Figures

Figure 1

18 pages, 19898 KB  
Article
Physics-Aware Deep Coupling Network for Extreme-Distance Infrared Ship Detection
by Ruiqi Wang, Ziquan Wang, Ling Guan and Zikai Zhang
Photonics 2026, 13(8), 748; https://doi.org/10.3390/photonics13080748 - 8 Aug 2026
Viewed by 229
Abstract
Detecting naval vessels at extreme distances using infrared search and track (IRST) systems presents severe physical challenges, notably the complete loss of geometric texture and the non-linear submersion of weak target signals within high-dynamic-range sea clutter. Traditional pure data-driven convolutional neural networks (CNNs) [...] Read more.
Detecting naval vessels at extreme distances using infrared search and track (IRST) systems presents severe physical challenges, notably the complete loss of geometric texture and the non-linear submersion of weak target signals within high-dynamic-range sea clutter. Traditional pure data-driven convolutional neural networks (CNNs) rely heavily on visual appearances and suffer from critical feature blind spots under such extreme physical degradation. To overcome this, we propose a Physics-Aware Deep Coupling Network that shifts the detection paradigm from appearance-based feature extraction to physics-guided attribute recognition. Our method deconstructs the degraded infrared signal into three complementary physical domains: an adaptive radiation energy mapping, corresponding to the energy domain, to rescue weak targets; a bio-inspired spatial saliency filtering mechanism, corresponding to the frequency domain, to maximize the signal-to-clutter ratio; and a PSF-coherent gradient topology framework, corresponding to the gradient domain, to discriminate genuine point targets from chaotic sun glints and island edges. These processed priors, alongside the raw image, are integrated into a 4-channel tensor and fused via a Cross-Domain Attention Module, ensuring deep network coupling. To evaluate this architecture, we conduct extensive experiments on the real-world Maritime-SIRST dataset. Since the original dataset provides only pixel-level segmentation masks, we generate axis-aligned bounding-box detection labels from these masks and retrain both the proposed method and a suite of state-of-the-art YOLO detectors under a unified detection paradigm. Extensive benchmarking demonstrates that our physics-aware methodology consistently outperforms these detectors, achieving a mAP50 of 0.923 and an F1 score of 89.92%, thus providing a highly interpretable and robust solution for maritime domain awareness under extreme physical constraints. Full article
Show Figures

Figure 1

25 pages, 5060 KB  
Article
Dynamics-Driven Dual-Stream Graph Neural Network with Adaptive Gated Fusion for Gearbox Fault Diagnosis
by Jiashuo Yu, Hanbin Xiao, Min Liu and Dinglong Zhu
Machines 2026, 14(8), 889; https://doi.org/10.3390/machines14080889 - 5 Aug 2026
Viewed by 302
Abstract
Conventional data-driven networks for gearbox fault diagnosis process multi-sensor streams as isolated sequences, failing to capture spatial-topological kinetic correlations and structural energy propagation pathways governed by multi-stage gearbox dynamics. To address these limitations, this study proposes a graph neural network-based fault diagnosis methodology [...] Read more.
Conventional data-driven networks for gearbox fault diagnosis process multi-sensor streams as isolated sequences, failing to capture spatial-topological kinetic correlations and structural energy propagation pathways governed by multi-stage gearbox dynamics. To address these limitations, this study proposes a graph neural network-based fault diagnosis methodology integrating multi-dimensional attention and dynamic topological priors (KT-GNN-CBAM). Mesh stiffness characteristics are analytically evaluated to initialize physical topology edge weights, while a convolutional block attention module filters spatio-temporal features to suppress background noise. Node features are subsequently aggregated through a parallel dual-stream architecture comprising a physics-prior kinetic stream and a data-driven attention stream, which are dynamically fused via an adaptive gated mechanism. Experimental validation on the HP-GBS-2023 testbed under mixed operations and 6 dB noise shows that the proposed framework achieves an optimal diagnostic accuracy of 98.87%. Ablation evaluations confirm that omitting the mechanics-driven prior branch induces a 282.30% relative surge in the model’s misclassification rate. Ultimately, embedding mechanical invariants as a physical inductive bias mitigates purely data-driven black-box constraints, offering an interpretable and robust solution for advanced intelligent fault diagnosis in complex gearbox systems. Full article
Show Figures

Figure 1

32 pages, 5193 KB  
Article
Frequency Decomposition and Spatial Dependency Mathematical Modeling for Small-Scale Open-World Object Detection
by Zhengbiao Jing, Qingjie Shi, Douping Bai, Baoyu Xiong and Donglin Jing
Algorithms 2026, 19(8), 644; https://doi.org/10.3390/a19080644 - 4 Aug 2026
Viewed by 306
Abstract
Intelligent transportation and aerial remote sensing scenes suffer from complex scene variations, abundant miniature targets and unpredictable out-of-distribution obstacles, which brings tough mathematical challenges to open-world detection tasks. Conventional detection algorithms lack rigorous frequency-domain separation and spatial constraint mathematical formulations, resulting in severe [...] Read more.
Intelligent transportation and aerial remote sensing scenes suffer from complex scene variations, abundant miniature targets and unpredictable out-of-distribution obstacles, which brings tough mathematical challenges to open-world detection tasks. Conventional detection algorithms lack rigorous frequency-domain separation and spatial constraint mathematical formulations, resulting in severe tiny-object feature attenuation, inefficient multimodal feature matching and catastrophic forgetting during incremental category iteration. To solve these mathematical bottlenecks, this paper constructs the TPCA-Net model built upon frequency decomposition and spatial dependency mathematical modelling. The entire framework consists of four fixed core modules: High-Frequency-Aware Multi-Scale Feature Enhancement (HSE), Reparameterized Adaptive Text–Visual Alignment (RTA), Double Wildcard Spatial Dependency Fusion (WSF), and Incremental Forgetting-Free Dual-Path Detection (DPD). From the mathematical perspective, the HSE module adopts discrete cosine transform-based filtering equations to split high-frequency object details from low-frequency background signals and establishes cross-attention spatial constraint formulas to make up for missing contextual information of small targets. The RTA module introduces low-rank decomposition mathematical optimization and reparameterized tensor fusion rules to realize domain-adaptive text embedding calibration and zero-cost cross-modal mapping at the inference stage. The WSF module constructs dual-wildcard self-supervised mathematical loss to finish unsupervised unknown-object identification and builds decoupled semantic–spatial fusion equations to improve the positioning precision of novel targets. The DPD module designs two sets of independent optimization objective functions and category-freezing incremental mathematical constraints to avoid conflicting parameter updates and eliminate forgetting defects in new-class expansion. Validated on COCO, DOTA and AI-TOD datasets, TPCA-Net achieves 56.0% AP on COCO, 79.30% mAP on DOTA, and 40.5% overall AP with 28.7% small-object AP on AI-TOD while delivering an inference throughput of 101.2 FPS on the Tesla T4 edge GPU. The proposed method outperforms existing mainstream open-world detection algorithms in tiny-object and rare-category recognition while maintaining efficient inference speed. Full article
(This article belongs to the Special Issue Advances in Deep Learning-Based Data Analysis)
Show Figures

Figure 1

25 pages, 2171 KB  
Article
TBSA: Tri-Domain Balanced Spectral–Spatial Attention with Deformable Frequency Filtering for Hyperspectral Image Classification
by Shuzhuan Tang and Xiaofei Yang
Mathematics 2026, 14(15), 2754; https://doi.org/10.3390/math14152754 - 3 Aug 2026
Viewed by 247
Abstract
Hyperspectral image classification (HSIC) requires a classifier to distinguish land-cover categories from densely sampled spectral signatures while preserving the spatial arrangement of local materials. Although convolutional networks, Transformer architectures, and recent state-space models have greatly improved spectral–spatial representation learning, three issues remain insufficiently [...] Read more.
Hyperspectral image classification (HSIC) requires a classifier to distinguish land-cover categories from densely sampled spectral signatures while preserving the spatial arrangement of local materials. Although convolutional networks, Transformer architectures, and recent state-space models have greatly improved spectral–spatial representation learning, three issues remain insufficiently resolved. First, spectral redundancy and local spatial textures are commonly modeled in the original feature domain, where low- and high-frequency responses are only implicitly separated. Second, fixed or weakly adaptive frequency operations cannot reflect the fact that different land-cover classes rely on different spectral smoothness, boundary, and texture cues. Third, spatial evidence, channel selectivity, and frequency responses are often fused by a uniform rule, which may be suboptimal under limited training samples and class imbalance. To address these issues, this paper proposes Tri-Domain Balanced Spectral–Spatial Attention(TBSA), a compact frequency-aware framework for HSIC. TBSA projects intermediate features into one-dimensional spectral, two-dimensional spatial, and three-dimensional spectral–spatial discrete cosine transform (DCT) domains, and it introduces a deformable frequency filter to adaptively separate low- and high-frequency components. Spatial–frequency and spatial–channel interaction form complementary evidence, while the final aggregation rule controls the balance between input-conditioned flexibility, numerical stability, and parameter cost. Experiments on Indian Pines, Houston 2013, and WHU-Hi-LongKou show competitive mean OA and strong class-wise or balanced-accuracy behavior. A five-run inferential analysis does not establish statistically significant OA superiority over the closest DCTN baseline, and the claims are therefore restricted to the observed mean and class-wise results. Full article
(This article belongs to the Special Issue Advances in Image Processing and Analysis)
Show Figures

Figure 1

15 pages, 606 KB  
Article
SPA-DETR: An Enhanced RT-DETR with Spatial-Preserving Attention and Adaptive Loss for UAV Spectrogram Signal Detection
by Conghao Fu, Lu Xu and Yijia Zhang
Sensors 2026, 26(15), 4846; https://doi.org/10.3390/s26154846 - 1 Aug 2026
Viewed by 328
Abstract
Rapid detection of unauthorized unmanned aerial vehicles (UAVs) via radio frequency (RF) spectrograms is critical for low-altitude security. However, standard object detectors struggle to locate transient, frequency-hopping UAV signals because their microscopic spatial footprints are easily discarded by conventional lossy downsampling and overwhelmed [...] Read more.
Rapid detection of unauthorized unmanned aerial vehicles (UAVs) via radio frequency (RF) spectrograms is critical for low-altitude security. However, standard object detectors struggle to locate transient, frequency-hopping UAV signals because their microscopic spatial footprints are easily discarded by conventional lossy downsampling and overwhelmed by complex background noise. To overcome this limitation, we propose SPA-DETR, a custom architecture based on the RT-DETR framework. The core of our design is the Spatial-Preserving Attention (SPA) block, which integrates Space-to-Depth Convolution (SPDConv) with a Parallel Patch-Aware Attention (PPA) module. By replacing traditional pooling mechanisms, the SPA block preserves the spatial details of weak signals without information loss, while the PPA module concurrently filters out ambient background interference. Furthermore, to address the severe foreground–background imbalance in RF spectrograms, we introduce an Adaptive Threshold Focal Loss (ATFL). Operating exclusively during training, ATFL prevents background noise gradients from dominating the learning process, forcing the network to focus on hard-to-detect signal patches without adding computational overhead during inference. Experiments on our public RFUAV dataset validate the approach. SPA-DETR achieves an mAP50:95 of 86.8% and an APS of 85.6%, improving upon the baseline RT-DETR-R18 by 4.9% and 5.2%, respectively. Operating at 235.2 FPS with only 23.74 M parameters, SPA-DETR outperforms contemporary detectors such as YOLOv10m, as well as heavier models like YOLOv8m and RT-DETR-R50, highlighting its efficiency and practical value for real-time low-altitude security applications. Full article
(This article belongs to the Special Issue Advanced Pattern Recognition: Intelligent Sensing and Imaging)
Show Figures

Figure 1

28 pages, 3215 KB  
Article
Self-Supervised Hyperspectral Image Clustering via Spatial–Frequency Interaction and Amplitude–Phase Decoupling
by Heng Yuan, Nan Huang, Qichao Liu, Pengfei Liu, Kang Ni and Zhizhong Zheng
Remote Sens. 2026, 18(15), 2494; https://doi.org/10.3390/rs18152494 - 31 Jul 2026
Viewed by 321
Abstract
Hyperspectral image (HSI) clustering assigns unlabeled pixels to land-cover groups by jointly exploiting spectral and spatial observations. Existing Vision Transformer-based deep clustering captures global dependencies through self-attention. However, the quadratic computational complexity of self-attention restricts practical applications in large HSI scenes. Furthermore, illumination [...] Read more.
Hyperspectral image (HSI) clustering assigns unlabeled pixels to land-cover groups by jointly exploiting spectral and spatial observations. Existing Vision Transformer-based deep clustering captures global dependencies through self-attention. However, the quadratic computational complexity of self-attention restricts practical applications in large HSI scenes. Furthermore, illumination variation and topographic shading shift spectral amplitude of co-class pixels toward divergent directions in feature space, enlarging intra-class distances and reducing inter-class separability in learned embeddings. To address the above limitations, we propose a self-supervised Spatial–Frequency Interaction and Amplitude–Phase Decoupling framework, termed SFI-APD, which integrates a High-Order Spatial–Frequency Interaction Module (HSFIM), a Frequency Feature Attention Block (FFAB), and a Frequency-Domain Vision Transformer (FreqViT) into a unified architecture. Specifically, HSFIM couples local convolutions with Fourier filtering to extract enriched spectral–spatial representations. FFAB then decouples amplitude and phase components to suppress brightness variations, yielding illumination-robust embeddings. Finally, FreqViT performs attention modulation across spectral channels, reducing token aggregation complexity from O(N2D) to O(NDlogN). On the Indian Pines, Salinas, Pavia University, and Yangzhou datasets, SFI-APD achieves OAs of 57.47%, 79.38%, 54.57%, and 64.11%, respectively, outperforming state-of-the-art self-supervised methods for large HSIs. Full article
(This article belongs to the Section Remote Sensing Image Processing)
Show Figures

Figure 1

Back to TopTop