Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (1,039)

Search Parameters:
Keywords = dual-scale attention

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
27 pages, 8204 KB  
Article
Dual-Level Spatial–Frequency Collaborative Detector for Oriented Object Detection in Remote Sensing Images
by Xuehuai Shi, Jingru Sun, Kun Yu, Zhihui Wei and Shangdong Zheng
Remote Sens. 2026, 18(16), 2845; https://doi.org/10.3390/rs18162845 - 21 Aug 2026
Viewed by 88
Abstract
Oriented object detection (OOD) in remote sensing images (RSIs) suffers from insufficient feature representation caused by arbitrary rotation angles and small spatial resolutions. Existing spatial–frequency fusion paradigms merely implement single-granularity feature interaction, either global image-level frequency compensation or local instance-level feature refinement, and [...] Read more.
Oriented object detection (OOD) in remote sensing images (RSIs) suffers from insufficient feature representation caused by arbitrary rotation angles and small spatial resolutions. Existing spatial–frequency fusion paradigms merely implement single-granularity feature interaction, either global image-level frequency compensation or local instance-level feature refinement, and fail to simultaneously capture global scene semantic consistency and local object fine-grained discriminability. In this paper, we propose a unified dual-level spatial–frequency collaborative detector (DSCDet) for remote sensing OOD tasks. Different from previous decoupled designs, the proposed DSCDet constructs a complete spatial–frequency collaborative fusion paradigm that shares a generic wavelet-based frequency extraction mechanism and cross-feature fusion module, which is adaptively deployed at both image-level and instance-level granularities. Specifically, our method introduces Haar wavelet transform to extract multi-scale frequency mutation features. On this basis, a generic cross-domain attention fusion (GCDAF) is constructed with granularity-dependent positional encoding constraints. The core difference between dual granularity fusion lies in geometric positional encoding, where image-level fusion adopts global scene positional embedding to maintain overall semantic stability, and instance-level fusion leverages local pairwise instance positional embedding to optimize fine-grained target feature interaction. The unified dual-level fusion architecture comprehensively integrates global semantic integrity and local target specificity, forming a robust and universal spatial–frequency feature representation system. Extensive experiments on three public remote sensing datasets, including DOTA-v1.0, DOTA-v1.5 and DIOR-R, demonstrate that the proposed DSCDet achieves competitive and superior performance against state-of-the-art OOD detectors. Full article
(This article belongs to the Section Remote Sensing Image Processing)
20 pages, 3895 KB  
Article
Attention-Enhanced Multi-Scale Feature-Wise Linear Modulation for Fine-Grained Poisonous Mushroom Image Recognition
by Yuan He, Haikun Lv, Chenyang Lu, Dengqi Yang, Xiaowei Li and Lina Zhang
J. Imaging 2026, 12(8), 398; https://doi.org/10.3390/jimaging12080398 - 21 Aug 2026
Viewed by 67
Abstract
Fine-grained poisonous mushroom recognition in natural scenes is challenging because of complex backgrounds, subtle morphological differences, and the limited interpretability of model decisions. To address these challenges, this paper proposes Att-FiLM, an attention-enhanced multi-scale Feature-Wise Linear Modulation network for poisonous mushroom image recognition. [...] Read more.
Fine-grained poisonous mushroom recognition in natural scenes is challenging because of complex backgrounds, subtle morphological differences, and the limited interpretability of model decisions. To address these challenges, this paper proposes Att-FiLM, an attention-enhanced multi-scale Feature-Wise Linear Modulation network for poisonous mushroom image recognition. The model adopts an asymmetric dual-backbone architecture in which a frozen ConvNeXt-Base branch provides global semantic priors, while a trainable EfficientNet-B0 branch learns local discriminative features. Rather than directly concatenating heterogeneous features, Att-FiLM generates scale and shift parameters from semantic features and performs channel-wise modulation on multi-scale EfficientNet features at Stage 2 and Stage 4. This mechanism enables global semantic information to guide local feature learning while reducing feature redundancy and semantic inconsistency. Experimental results show that Att-FiLM achieves an Accuracy of 95.58% and an F1-score of 0.9455 on the poisonous/edible binary classification task. On the 190-class species-level classification task, it achieves a Top-1 Accuracy of 93.63% and a Macro-F1 of 0.9347. Interpretability analysis further shows that decision-relevant responses are frequently associated with morphologically relevant regions, including gills, annuli, volvae, and cap textures. These results indicate that Att-FiLM provides effective recognition performance together with interpretable decision evidence for mushroom recognition in complex natural scenes. Full article
(This article belongs to the Section Image and Video Processing)
28 pages, 71265 KB  
Article
Sharing Cultural Values Through 3D Point-Cloud-Based Documentation of Transylvanian Heritage
by Alina Elena Voinea, Calin Neamtu and Virgil Pop
Remote Sens. 2026, 18(16), 2841; https://doi.org/10.3390/rs18162841 - 21 Aug 2026
Viewed by 141
Abstract
This paper presents a pilot educational workflow that couples 3D remote sensing with heritage-driven pedagogy by engaging architecture master’s students in the documentation and digital archiving of Transylvanian cultural sites. Using terrestrial and mobile 3D scanning, students documented multiple typologies—wooden churches (Târgușor, Tioltiur), [...] Read more.
This paper presents a pilot educational workflow that couples 3D remote sensing with heritage-driven pedagogy by engaging architecture master’s students in the documentation and digital archiving of Transylvanian cultural sites. Using terrestrial and mobile 3D scanning, students documented multiple typologies—wooden churches (Târgușor, Tioltiur), historical ensembles (Mociu, Coplean), industrial sites (1 Mai–Luduș, Vânătorilor–Luduș), and an urban street segment (Potaissa)—to generate dense point clouds that served as the basis for geometric reconstruction, semantic interpretation, and condition assessment. The study describes how the characteristics of different construction systems (timber, brick, stone, mixed structures) relate to point-cloud quality, survey coverage, and subsequent CAD/BIM drafting, with attention to the qualitative reading of minor deformations in wooden churches and of degradation patterns in masonry and industrial buildings. We also consider how artefacts in the data (noise, occlusions, registration errors) affect scene understanding and the interpretation of derived observations relevant to condition assessment and, prospectively, to monitoring. For the Tioltiur dual-sensor case, the TLS and SLAM datasets were compared through an internal CloudCompare registration check (final RMS 0.1121 on 50,000 points, fixed scale 1.0 and theoretical overlap 100%), surface-density displays (r = 0.005 for the Z+F dataset and for the GeoSLAM dataset), fitted-wall-plane readings (dip values around 89 deg. and 85 deg.) and a longitudinal section documenting roof/vault deformation. Beyond technical performance, the paper examines the self-reported formative impact on students’ digital skills and their understanding of cultural values, arguing that participation in 3D data acquisition, processing, and interpretation positions them as co-creators of a living digital archive. Pre- and post-workshop questionnaires (n = 13 each) are analysed descriptively—counts, percentages and medians with interquartile ranges—because the two instruments are unmatched and carry no shared identifier, so no paired test is applied; post-workshop self-ratings of technical competence, heritage understanding, archival awareness and collaboration were consistently high (medians 4–5), with uneven access to VR the main gap. By connecting point-cloud-based documentation workflows with heritage education, the project outlines a transferable, monitoring-ready baseline model in which 3D remote sensing supports both careful documentation and the transmission of regional identity and cultural meaning in architectural training. As an exploratory pilot with a small, self-reported sample, the study reports descriptive and qualitative findings rather than validated metric or statistical results. Full article
Show Figures

Figure 1

25 pages, 3738 KB  
Article
ESD-YOLO: A Method for Small-Target Termite Detection Under Complex Backgrounds
by Weiling Lu, Yuting Meng, Shan Wu and Hangjun Wang
Insects 2026, 17(8), 874; https://doi.org/10.3390/insects17080874 - 21 Aug 2026
Viewed by 73
Abstract
Timely and accurate termite detection is essential for effective termite control. To address the challenges posed by the small size of termite individuals and the susceptibility of target features to background texture interference under complex backgrounds, this study proposes ESD-YOLO, a fine-grained feature-enhanced [...] Read more.
Timely and accurate termite detection is essential for effective termite control. To address the challenges posed by the small size of termite individuals and the susceptibility of target features to background texture interference under complex backgrounds, this study proposes ESD-YOLO, a fine-grained feature-enhanced object detection model. Using YOLO11n as the baseline, ESD-YOLO redesigns the feature extraction, deep feature aggregation, and multi-scale feature fusion stages to improve the representation of small-scale termite targets under complex backgrounds. Specifically, the Efficient Multi-scale Attention (EMA) mechanism is incorporated into the C3k2 module to enhance feature discriminability between termite individuals and the background. A Spatial Pyramid Pooling-Fast with Dual Global Pooling (SPPF-DGP) module is employed to supplement deep features with global contextual information and salient response information. In addition, the DySample dynamic upsampling module is introduced to improve spatial alignment during multi-scale feature fusion and enhance boundary representation for small targets. Experimental results show that ESD-YOLO achieves Precision, Recall, mAP@0.5, and mAP@0.5:0.95 values of 95.39%, 96.33%, 97.85%, and 65.92%, respectively, with 2.67 M parameters and 6.68 G FLOPs. Compared with Faster R-CNN, RetinaNet, RT-DETR, and several YOLO-series models, ESD-YOLO demonstrates strong small-target detection and localization performance under the controlled complex-background conditions established in this study, providing a methodological reference for automated termite detection in practical settings. Full article
(This article belongs to the Special Issue AI and Cloud Computing for Insect Ecology and Management)
Show Figures

Figure 1

21 pages, 28113 KB  
Article
Cross-Scale Unified Semantic Space Learning for Small-Scale Pest and Disease Detection in Protected Agriculture
by Linmin Yu, Rongfang Qu, Qifeng Wu, Xiaofei An, Ruxiao Bai, Lingxian Zhang and Chunmei Zhu
AgriEngineering 2026, 8(8), 349; https://doi.org/10.3390/agriengineering8080349 - 21 Aug 2026
Viewed by 118
Abstract
In protected agriculture such as greenhouses, pest and disease monitoring via UAVs and fixed cameras suffers from extremely small object proportions owing to shooting altitude constraints, posing considerable detection challenges. Moreover, fine-grained annotation of numerous small-scale images incurs prohibitive costs. Targeting this bottleneck, [...] Read more.
In protected agriculture such as greenhouses, pest and disease monitoring via UAVs and fixed cameras suffers from extremely small object proportions owing to shooting altitude constraints, posing considerable detection challenges. Moreover, fine-grained annotation of numerous small-scale images incurs prohibitive costs. Targeting this bottleneck, this paper proposes a cross-scale unified semantic space learning framework and introduces an end-to-end DS-DETR detector based on DETR. Unlike existing methods relying on domain adaptation, multi-scale fusion, or super-resolution reconstruction, this work explicitly models instance-level cross-scale semantic correlation, transferring fine-grained semantics from large-scale close-up images to small-scale scene feature space. A Single-Point Dual-Shooting (SPDS) strategy is adopted to collect high-fidelity paired images via ordinary smartphones at low cost. A dual-stream encoder with cross-view attention and an instance-level contrastive loss align features of identical instances in a unified semantic space. A self-built CropScale-Det dataset covering three crop diseases is constructed in greenhouse scenarios. Experimental results show that DS-DETR achieves 42.5 ± 1.2% mAP@50 under limited annotations, outperforming YOLOv8-n by 11.2%, with small-target average precision reaching 26.8 ± 1.1%. Ablation experiments and feature visualization validate the effectiveness of the designed mechanism. This approach considerably reduces reliance on large-scale densely annotated data, establishing a data-efficient proof-of-concept for small-scale pest detection in protected agriculture. Full article
Show Figures

Figure 1

18 pages, 23282 KB  
Article
Research on an Improved YOLOv8-Based Object Detection Algorithm for Flame and Smoke Detection in Factory Environments
by Linlin Cao, Xinxin Chen, Sitong Guo, Jiaqi Wang, Duowen Chen, Fengyan Lun, Haoyu Zhang, Kaibao Wang and Jianyong Li
Appl. Sci. 2026, 16(16), 8325; https://doi.org/10.3390/app16168325 - 21 Aug 2026
Viewed by 71
Abstract
Overcoming complex background noise and poor small-target detection in industrial settings, this paper introduces YOLOv8-BBP2, an enhanced YOLOv8 model. To better extract dynamic features, the backbone integrates a BiFormer dual-level routing attention mechanism. Moreover, a learnable Bi-directional Feature Pyramid Network (BiFPN) replaces the [...] Read more.
Overcoming complex background noise and poor small-target detection in industrial settings, this paper introduces YOLOv8-BBP2, an enhanced YOLOv8 model. To better extract dynamic features, the backbone integrates a BiFormer dual-level routing attention mechanism. Moreover, a learnable Bi-directional Feature Pyramid Network (BiFPN) replaces the standard module, optimizing multi-scale feature integration. A P2 detection head is also added to accurately identify tiny objects, such as early flames and thin smoke. Tested on a custom factory fire dataset, YOLOv8-BBP2 yields 95.231% precision, 94.612% recall, and 89.677% mean average precision (mAP@0.5). These metrics represent respective gains of 3.31%, 4.934%, and 7.451% over the baseline YOLOv8s. Ultimately, with an inference speed of 20 ms per frame, the proposed network ensures highly robust, real-time performance. Full article
Show Figures

Figure 1

24 pages, 5349 KB  
Article
MarShip-DET: A Frequency-Aware Multi-Scale Fusion Algorithm for Ship Detection in Maritime Remote Sensing Imagery
by Keren Chen, Xufang Zhu, Zhikun Liu, Fuyan Zhao and Kang Wang
Sensors 2026, 26(16), 5295; https://doi.org/10.3390/s26165295 - 21 Aug 2026
Viewed by 178
Abstract
To address the challenges of multi-scale target variation, complex background interference, and insufficient feature fusion quality in maritime remote sensing ship detection, this paper proposes MarShip-DET, a frequency-aware multi-scale fusion detection algorithm based on YOLO11n. Three core modules are introduced: Channel-Decoupled Progressive Feature [...] Read more.
To address the challenges of multi-scale target variation, complex background interference, and insufficient feature fusion quality in maritime remote sensing ship detection, this paper proposes MarShip-DET, a frequency-aware multi-scale fusion detection algorithm based on YOLO11n. Three core modules are introduced: Channel-Decoupled Progressive Feature Extraction Module (CDPFEM), which employs asymmetric channel decoupling with dual-statistic channel attention and image-relative-position-encoded multi-head self-attention to enhance discriminative feature extraction; Edge-Aware Region Context Fusion Module (EARCFusion), which integrates learnable Sobel edge sensing and cross-attention correction to achieve precise foreground refinement; and Wavelet-guided Prototype Attention Module (WavePAM), which combines Haar wavelet frequency decomposition with prototype-guided spatial compression attention to strengthen deep semantic representation. Experiments on HRSC2016 demonstrate that MarShip-DET achieves an mAP50 of 94.9% and an mAP50-95 of 84.4%, improving by 3.7% and 4.3% over the baseline, respectively. Zero-shot experiments on HRSID and SSDD, including comparisons with YOLO11n, D-FINE-N, and YOLOv13n, provide additional evidence of cross-domain transferability under the evaluated protocol. Full article
(This article belongs to the Section Remote Sensors)
Show Figures

Figure 1

27 pages, 17769 KB  
Article
SFSMamba-DETR: Selective Feature Scanning with State Space Models and Dual-Scale Window Attention for Remote Sensing Object Detection
by Yuanli Cai, Junchao Zhao, Husheng Wu and Rui Ma
Remote Sens. 2026, 18(16), 2835; https://doi.org/10.3390/rs18162835 - 21 Aug 2026
Viewed by 197
Abstract
Object detection in remote sensing imagery remains challenging due to vast scale variations, complex backgrounds, and the prevalence of small, densely packed targets. Existing CNN-based detectors are limited by restricted receptive fields, while Transformer-based methods incur prohibitive computational overhead for high-resolution inputs. In [...] Read more.
Object detection in remote sensing imagery remains challenging due to vast scale variations, complex backgrounds, and the prevalence of small, densely packed targets. Existing CNN-based detectors are limited by restricted receptive fields, while Transformer-based methods incur prohibitive computational overhead for high-resolution inputs. In this paper, we propose SFSMamba-DETR, a detection framework that integrates state space models with Dual-Scale Window Attention for efficient and accurate remote sensing object detection. Specifically, we design a Selective Feature Scanning (SFS) module that uses the Mamba-based 2D Selective Scan mechanism to model long-range spatial dependencies with linear computational complexity. To capture both fine-grained local patterns and broader contextual cues simultaneously, we introduce a Dual-Scale Window Attention (DSWA) mechanism that operates at two complementary window scales with multi-kernel convolution bridging. These modules are orchestrated within a Cross-scale Feature Aggregation Module (CFAM) that performs hierarchical multi-scale fusion in a hybrid encoder. Extensive experiments on three primary benchmarks (MAR20, UCAS-AOD, and the Jilin-1 Satellite Aircraft Detection Dataset), together with supplementary results on DOTA and DIOR, demonstrate that SFSMamba-DETR achieves strong detection accuracy while maintaining competitive inference speed. Full article
Show Figures

Figure 1

26 pages, 9396 KB  
Article
Multi-Scale Spatiotemporal Graph ODE Networks for Marine Chlorophyll-a Prediction
by Xiaoyu He, Yijing Zhang, Xin Huang and Suixiang Shi
Remote Sens. 2026, 18(16), 2828; https://doi.org/10.3390/rs18162828 - 20 Aug 2026
Viewed by 110
Abstract
Chlorophyll-a concentration is a key indicator reflecting the growth status of phytoplankton, and its accurate prediction is of great significance for assessing the degree of water eutrophication. Although existing approaches have achieved good performance, they generally pay insufficient attention to multi-scale spatial information [...] Read more.
Chlorophyll-a concentration is a key indicator reflecting the growth status of phytoplankton, and its accurate prediction is of great significance for assessing the degree of water eutrophication. Although existing approaches have achieved good performance, they generally pay insufficient attention to multi-scale spatial information and show limitations in characterizing the continuous spatiotemporal dynamics. To address these issues, this paper proposes a multi-scale spatiotemporal graph ODE network (MGODE) for ocean chlorophyll-a prediction. The MGODE adopts a dual-layer structure, simultaneously processing chlorophyll-a concentration data at both the region level and node level to capture multi-scale spatial features, and it enables effective interaction of cross-scale features through dynamic transmission coefficients and a gated fusion mechanism. Meanwhile, the MGODE employs a dual-ODE architecture at both the node and region levels, utilizing spatiotemporal ODE blocks to continuously and deeply capture features, thereby simulating the continuous spatiotemporal dynamic evolution of chlorophyll-a. Experiments on real-world datasets from the Bohai Sea and South China Sea show that the proposed MGODE model achieves higher prediction accuracy than several current state-of-the-art models. Compared with the best baseline, the MGODE achieves reductions of 2.78% in MAE and 1.07% in RMSE on the Bohai Sea dataset and reductions of 1.19% in MAE and 1.38% in RMSE on the South China Sea dataset. These results demonstrate the potential of the MGODE to support marine chlorophyll-a forecasting and marine ecological monitoring. Full article
Show Figures

Figure 1

22 pages, 65601 KB  
Article
Dual-Domain Illumination Prior for Low-Light Remote Sensing Image Enhancement
by Chao Wang, Zhe Pan, Liangtian He, Jun Liu, Lin Mei, Rongsheng Lin, Hongming Chen and Chuansheng Yang
Remote Sens. 2026, 18(16), 2817; https://doi.org/10.3390/rs18162817 - 20 Aug 2026
Viewed by 158
Abstract
Low-light conditions degrade remote sensing imagery by reducing contrast, distorting color, and obscuring fine terrain structures and small objects critical for Earth observation. Accurate illumination adjustment under spatially varying scene content remains challenging for existing enhancement methods, and many prior-guided approaches operate exclusively [...] Read more.
Low-light conditions degrade remote sensing imagery by reducing contrast, distorting color, and obscuring fine terrain structures and small objects critical for Earth observation. Accurate illumination adjustment under spatially varying scene content remains challenging for existing enhancement methods, and many prior-guided approaches operate exclusively in either the spatial domain or the frequency domain. In this work, we propose a Dual-Domain Illumination Prior (DDIP), a trainable dual-domain illumination-prior module that is jointly optimized with each host backbone and exploits frequency-domain and spatial-domain illumination statistics. DDIP comprises three components: a Frequency-Domain Illumination Distribution Prior (FIDP) that performs per-color-channel amplitude calibration in Fourier space to improve global brightness; a Spatial-Domain Illumination Distribution Prior (SIDP), adapted from IDP-Net, that performs multi-scale sub-region statistical correction for local illumination adjustment; and a Selective Core Feature Fusion (SCFF) module that adaptively combines the frequency-domain output, the spatial-domain output, and the original input through an attention-based gating mechanism with dual pooling. DDIP is integrated with each host backbone while leaving its main restoration blocks unchanged. In the controlled reconstruction comparisons on iSAID-dark and the evaluated general low-light benchmarks, equipping the tested backbone networks with DDIP improves PSNR and SSIM over their corresponding baselines. Complementary LPIPS and CIELAB lightness measurements characterize perceptual similarity and lightness behavior, while a fixed-detector object-detection evaluation on the tested high-resolution iSAID-dark scenes examines the effect of the enhancement pipelines under the reported synthetic low-light conditions. The ablation studies further examine the contribution of the module components within the reported experimental settings. Full article
Show Figures

Figure 1

60 pages, 11445 KB  
Article
A Mamba-Driven Spatiotemporal Graph Neural Network for Fault Location in Low-Observability Active Distribution Networks
by Zhengying Hou, Jilong Ma and Xuguang Hu
Machines 2026, 14(8), 948; https://doi.org/10.3390/machines14080948 - 19 Aug 2026
Viewed by 163
Abstract
Accurate fault location in low-observability active distribution networks is hindered by uncertain inter-node relationships, underutilized early transients, and insufficient global context. To address these challenges, this paper proposes an adaptive Mamba-driven spatiotemporal graph neural network (AM-STGNN). It provides a unified task-driven spatiotemporal representation [...] Read more.
Accurate fault location in low-observability active distribution networks is hindered by uncertain inter-node relationships, underutilized early transients, and insufficient global context. To address these challenges, this paper proposes an adaptive Mamba-driven spatiotemporal graph neural network (AM-STGNN). It provides a unified task-driven spatiotemporal representation framework that progressively integrates fault-propagation modeling, global dependency modeling, transient-state learning, and topology-aware discriminative enhancement. Specifically, a prior-guided adaptive implicit topology is first learned to characterize task-dependent electrical coupling relationships among sparse observation nodes. Based on the resulting topology, topology-conditioned multi-order feature propagation and a dual-axis linear-attention module based on the spatiotemporal graph transformer (STGformer) are employed to capture local and global spatiotemporal dependencies. The resulting global spatiotemporal representation is subsequently processed by a Mamba selective state-space encoder to model input-dependent temporal evolution and emphasize informative fault transients. Finally, element-wise gated fusion, topology-aware differential output, and a margin constraint are employed to integrate the STGformer and Mamba representations and enhance the separability of adjacent faulted line sections with similar response characteristics. Extensive experiments demonstrate the effectiveness of AM-STGNN, while additional evaluations confirm its applicability to larger-scale networks, strongly phase-unbalanced conditions, and field-measured operating backgrounds. Robustness tests under individual and multi-level joint disturbances further demonstrate the practical relevance of the proposed architecture. Compared with the baseline models, AM-STGNN achieves consistent improvements in the macro-averaged F1 score (Macro-F1), exact accuracy, and one-hop accuracy under the clean IEEE 123-node condition. More importantly, it maintains clear performance advantages under identical mild, moderate, and severe joint disturbances, demonstrating improved robustness and practical relevance under simulated non-ideal operating conditions. Full article
Show Figures

Figure 1

23 pages, 11393 KB  
Article
Noise-Robust Multiclass Classification of Diesel-Engine Fault States Based on a Dual-Decoder Denoising Autoencoder
by Zeyu Yuan, Zhibin He and Qi Han
Appl. Sci. 2026, 16(16), 8238; https://doi.org/10.3390/app16168238 - 19 Aug 2026
Viewed by 112
Abstract
To address noise interference and feature overlap in multiclass diesel-engine fault-state classification, this study proposes a dual-decoder denoising autoencoder with dual-channel feature fusion. Multi-SNR denoising training pairs are encoded by a shared encoder and reconstructed by a general denoising decoder under reconstruction and [...] Read more.
To address noise interference and feature overlap in multiclass diesel-engine fault-state classification, this study proposes a dual-decoder denoising autoencoder with dual-channel feature fusion. Multi-SNR denoising training pairs are encoded by a shared encoder and reconstructed by a general denoising decoder under reconstruction and cross-SNR latent consistency constraints. With the encoder frozen, a normal-manifold decoder is trained on normal samples to generate a normal-state reference. The difference between the decoder outputs forms a normal-manifold residual that characterizes deviations from normal operation. Latent features are processed by a residual multilayer perceptron, whereas residuals are processed by multi-scale 1D convolution with efficient channel attention; gated fusion combines the two channels for classification. On the 3500-DEFault dataset, reconstruction MAE decreased by 93.75% at 0 dB and 90.86% at 15 dB. Anomaly detection based on normal-manifold residuals achieved an ROC-AUC of 0.9749. The proposed method attained an accuracy of 94.21% and a Macro-F1 of 94.28% on the independent test set, improving Macro-F1 by 3.07 percentage points over the DAE classifier. These results demonstrate the effectiveness of the proposed framework for noise-robust multiclass fault-state classification under the evaluated SNR conditions. Full article
(This article belongs to the Section Mechanical Engineering)
Show Figures

Figure 1

21 pages, 6311 KB  
Article
EP-Net: An Equipment-Guided Dual-Stream CNN–Transformer Network for Fine-Grained Sports Image Classification
by Xiaocui Sang, Changwei Gu and Lei Zhao
Appl. Sci. 2026, 16(16), 8229; https://doi.org/10.3390/app16168229 - 18 Aug 2026
Viewed by 209
Abstract
Fine-grained sports recognition from static images is challenging because visually similar sports often exhibit nearly identical human poses, whereas their decisive differences are encoded by small-scale equipment and subtle human–equipment interactions. Existing single-stream convolutional or Transformer-based models tend to emphasize either local appearance [...] Read more.
Fine-grained sports recognition from static images is challenging because visually similar sports often exhibit nearly identical human poses, whereas their decisive differences are encoded by small-scale equipment and subtle human–equipment interactions. Existing single-stream convolutional or Transformer-based models tend to emphasize either local appearance or global context, making them vulnerable to equipment-detail loss and background interference. To address this problem, we propose the Equipment-Primed Network (EP-Net), a heterogeneous dual-stream architecture that treats sports equipment as a primary semantic cue for action discrimination. EP-Net employs an EfficientNetV2-S-based Equipment Stream to capture localized equipment shapes and textures and a Swin-Tiny-based Behavior Stream to model the athlete’s spatial configuration and global scene context. We further introduce a Cross-Modal Channel Attention (CMCA) module that projects equipment features into the behavior-feature space and performs directional channel recalibration. Unlike simple feature concatenation, CMCA uses equipment information to enhance action-relevant channels while reducing the relative influence of background-dominated responses. Experiments on the Sports-100 dataset show that EP-Net achieves a Top-1 accuracy of 98.80%, outperforming EfficientNetV2-S and Swin-Tiny by 3.00 and 2.35 percentage points, respectively. It also improves on naive dual-stream concatenation by 0.88 percentage points. Grad-CAM visualizations further indicate that EP-Net attends more consistently to discriminative equipment and human–equipment interaction regions. These results suggest that equipment-guided local–global feature interaction provides an effective solution to pose ambiguity and background interference in static fine-grained sports recognition. Full article
Show Figures

Figure 1

27 pages, 19863 KB  
Article
CDF-DETR: Cross-Stage Attention and Dual-Scale Feature Calibration for Small-Object Detection in UAV Remote Sensing Imagery
by Rui Zou, Jinwei Guo, Jiaqi Liang, Kai Che, Yifan Deng and Binqi Chen
Remote Sens. 2026, 18(16), 2793; https://doi.org/10.3390/rs18162793 - 18 Aug 2026
Viewed by 251
Abstract
Small-object detection in unmanned aerial vehicle (UAV) remote sensing imagery is challenged by dense target distributions, substantial scale variation, complex ground backgrounds, and limited edge-computing resources. To address these challenges, we propose CDF-DETR, an end-to-end detector derived from the Real-Time Detection Transformer (RT-DETR). [...] Read more.
Small-object detection in unmanned aerial vehicle (UAV) remote sensing imagery is challenged by dense target distributions, substantial scale variation, complex ground backgrounds, and limited edge-computing resources. To address these challenges, we propose CDF-DETR, an end-to-end detector derived from the Real-Time Detection Transformer (RT-DETR). First, a Cross-Stage Partial Single-Head Attention Transformer (CSP-SHAT) backbone combines efficient local feature extraction with partial-channel global interaction to improve multi-scale representation while reducing the parameter count of the backbone. Second, a dual-scale feature calibration (DSFC) module sequentially performs contextual aggregation and deformable spatial alignment, thereby improving the consistency of shallow localization features and deep semantic features. Third, Focaler-MPDIoU integrates coordinate-sensitive regression with IoU-quality-based sample reweighting for dense small-object localization. Experiments on the VisDrone-2019 test set and the UAVDT and HIT-UAV validation sets demonstrate mAP50 improvements of 3.1, 1.4, and 3.0 percentage points, respectively, over the RT-DETR-R18 baseline. On the VisDrone-2019 validation set, CDF-DETR improves mAP5095 from 26.20% to 28.52%, corresponding to a gain of 2.32 percentage points, while reducing the parameter count by 25.7%. A compressed INT8 variant achieves 20.84 FPS for an offline image-level pipeline on an NVIDIA Jetson Orin Nano using ONNX and TensorRT. These results demonstrate improved detection accuracy with a reduced parameter footprint for UAV remote sensing image analysis. Full article
(This article belongs to the Section AI Remote Sensing)
Show Figures

Graphical abstract

44 pages, 60349 KB  
Article
YOLOv11–BiFPN–DAAF: An Object Detection Framework for Automated Surface Inspection of Balsa Wood Panels
by Cristian Zambrano-Vega, Washington Chiriboga-Casanova, Byron Oviedo, Efraín Díaz-Macías and Edgar Suárez Bardelline
Automation 2026, 7(4), 131; https://doi.org/10.3390/automation7040131 - 17 Aug 2026
Viewed by 192
Abstract
Automated surface inspection of balsa wood panels is challenging because defects may be small, elongated, weakly contrasted, or visually similar to natural grain patterns. This study proposes YOLOv11–BiFPN–DAAF, an enhanced object-detection architecture that combines bidirectional multi-scale feature fusion with adaptive dual-attention feature refinement. [...] Read more.
Automated surface inspection of balsa wood panels is challenging because defects may be small, elongated, weakly contrasted, or visually similar to natural grain patterns. This study proposes YOLOv11–BiFPN–DAAF, an enhanced object-detection architecture that combines bidirectional multi-scale feature fusion with adaptive dual-attention feature refinement. The task was formulated as single-class detection, with all anomalous surface regions labeled as Defect. The dataset comprised 508 manually annotated RGB images, including independent internal and external production test sets. A preliminary screening identified YOLOv11-m512 as the reference configuration, followed by a controlled 2×2 factorial ablation comprising the baseline, BiFPN, DAAF, and their combined integration. Each configuration was trained using five independent random seeds under identical experimental conditions. On the validation set, the combined architecture achieved a precision of 0.893±0.004, recall of 0.848±0.006, mAP@0.5 of 0.897±0.004, and mAP@0.5:0.95 of 0.389±0.004. Relative to the baseline, the largest improvement was obtained for mAP@0.5:0.95, with a relative gain of 9.89%, indicating improved localization under stricter IoU thresholds. The improvement was retained on the independent internal test set, where the proposed architecture reached mAP@0.5 and mAP@0.5:0.95 values of 0.892±0.005 and 0.384±0.006, respectively. On the external production test set, the corresponding values were 0.865±0.007 and 0.358±0.008, representing absolute improvements of 0.034 and 0.042 over the original YOLOv11 baseline. Under the matched experimental protocol, YOLOv11–BiFPN–DAAF also achieved the highest principal detection metrics among the evaluated representative detectors. Although BiFPN and DAAF introduced a moderate computational overhead, the architecture maintained an inference time of 9.6±0.3 ms per image. These findings support the potential of the proposed architecture for automated balsa wood panel inspection, while broader multi-site and hardware-level validation remains necessary before large-scale industrial deployment. Full article
Show Figures

Figure 1

Back to TopTop