Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

Search Results (695)

Search Parameters:
Keywords = dual multiscale attention

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
20 pages, 2047 KB  
Article
GSDAT-Net: Enhancing Image Super-Resolution with Grid-Spatial Dual Attention Hybrid Transformer
by Yunqiang Liu, Ting Wei and Jinhua Wang
Sensors 2026, 26(14), 4634; https://doi.org/10.3390/s26144634 - 22 Jul 2026
Abstract
Image super-resolution (SR) is a fundamental task in computer vision that aims to recover high-fidelity details from low-resolution images. To address the limitations of existing Transformer-based SR methods, this paper introduces GSDAT-Net, a grid-spatial dual-attention-driven residual hybrid Transformer. By integrating Grid Attention Block [...] Read more.
Image super-resolution (SR) is a fundamental task in computer vision that aims to recover high-fidelity details from low-resolution images. To address the limitations of existing Transformer-based SR methods, this paper introduces GSDAT-Net, a grid-spatial dual-attention-driven residual hybrid Transformer. By integrating Grid Attention Block (GAB), an Enhanced Spatial Attention (ESA) module, and SwinV2 Transformer layers (S2TL), GSDAT-Net extracts multi-scale features and improves local and global feature representation. Extensive experiments demonstrate that GSDAT-Net achieves competitive reconstruction accuracy and visual quality compared with the selected methods, with a maximum PSNR improvement of up to 0.15 dB on benchmark datasets. Full article
(This article belongs to the Special Issue Machine Learning in Image/Video Processing and Sensing)
Show Figures

Figure 1

22 pages, 15654 KB  
Article
A Method for Detecting Cattle Behaviors Based on RGB-Depth Dual-Modal Information Fusion
by Zihao Chen, Jiaxing Xie, Liang Mao, Qiuxia Chen and Linlin Wang
Animals 2026, 16(14), 2259; https://doi.org/10.3390/ani16142259 - 21 Jul 2026
Abstract
In large-scale cattle farming, accurate behavior recognition is central to achieving intensive health monitoring and animal welfare assessment. To address challenges such as background interference from fences, feed troughs, and stains in real-world barns, and overlapping of cattle coupled with the inability of [...] Read more.
In large-scale cattle farming, accurate behavior recognition is central to achieving intensive health monitoring and animal welfare assessment. To address challenges such as background interference from fences, feed troughs, and stains in real-world barns, and overlapping of cattle coupled with the inability of single-RGB modalities to capture physical spatial structure, which leads to issues like blurred detection boundaries and significant noise interference—we propose a cattle behavior detection method based on RGB-Depth dual-modal information fusion. This approach jointly models the texture information from RGB images and the spatial structural information from depth images. Within this framework, this paper constructs three collaborative optimization modules: first, the CDSAM module is developed, which evaluates neuron importance through a parameter-free attention mechanism and combines dynamic convolutions to adapt to the cattle’s variable postures, effectively suppressing complex background noise. Second, we propose the C2BRA module based on a two-layer routed attention mechanism. By adopting a two-stage modeling approach of “region-level routing—intra-region fine-grained attention,” it adapts to changes in target scale and enhances the model’s ability to represent spatial context for multi-scale semantic information. Finally, in the prediction stage, a lightweight shared convolutional detection head (LSCD) is introduced. By sharing convolutional parameters across scales and decoupling the classification and regression architectures, it reduces computational overhead while maintaining accuracy. Experimental results show that the improved model achieves a mAP@0.5 of 90.3% on our self-built cattle behavior dataset, representing a 4.3 percentage point increase compared to the baseline model, while reducing GFLOPs from 11.0 G to 9.6 G, a decrease of 12.7%; Visualization results indicate that the improved model can focus more accurately on cattle body contours and key behavioral regions, thereby reducing false negatives and enhancing detection accuracy. Concurrently, the model achieves an optimal balance between detection performance and computational complexity, providing robust technical support for automated cattle behavior monitoring on smart farms. Full article
Show Figures

Figure 1

45 pages, 4136 KB  
Article
TRT-GLA: Tri-Representation Transformers with Global–Local Attention for High-Fidelity Multi-Modal MRI Super-Resolution
by Suhaila Abuowaida, Hamza Abu Owida, Tareq Hamadneh, Nawaf Alshdaifat, Hamza A. Mashagba, Mwaffaq Abu Alhaija and Azlan B. Abd Aziz
Algorithms 2026, 19(7), 603; https://doi.org/10.3390/a19070603 - 21 Jul 2026
Abstract
The super-resolution (SR) of Magnetic Resonance Imaging (MRI) is essential for utilizing clinical scans with limited resolution, noise, and anisotropic sampling, such as multi-modal brain tumor imaging. In this work, we propose a Tri-Representation hybrid framework for MRI SR, TRT-GLA, that redefines the [...] Read more.
The super-resolution (SR) of Magnetic Resonance Imaging (MRI) is essential for utilizing clinical scans with limited resolution, noise, and anisotropic sampling, such as multi-modal brain tumor imaging. In this work, we propose a Tri-Representation hybrid framework for MRI SR, TRT-GLA, that redefines the MRI SR task as a joint spatial–spectral–structural high-resolution image generation problem. TRT-GLA utilizes (i) spatial global–local attentions for modeling the spatial anatomy, (ii) a Fourier spectral transfer mechanism for upholding spectral consistency, and (iii) multi-scale hierarchical spectral decomposition for improved edge details. To adapt the learning framework to medical imaging characteristics, we introduce a tri-representation consistent loss function that explicitly combines pixel-wise, spectral, and edge structure priors from the high-resolution ground-truth, as well as a progressive resolution learning strategy. Our large-scale brain tumor experiments, on the IXI, BraTS 2019, 2020, and 2023 datasets, show that TRT-GLA achieves state-of-the-art results at upsampling factors of ×2, ×4, and ×8, respectively, achieving substantial improvements across CNN, GAN, and transformer-based methods in Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Multi-scale Structural Similarity Index (MS-SSIM). We further demonstrate how SR benefits brain tumor segmentation through the downstream task evaluation of a dual-branch segmentation framework. TRT-GLA produces highly accurate tumor segmentation results from low-resolution inputs, improving over native high-resolution inputs at ×8 in critical tumor boundary regions and in small tumor regions. There remains a small gap between native, high-resolution imaging and SR-enhanced performance, which TRT-GLA nearly closes under realistic scenarios. Our results highlight the importance of synthesizing unified priors over spatial, spectral, and structural domains within a transformer for anatomically faithful reconstructions. Importantly, we also establish the utility of TRT-GLA in supporting quantitative analysis through a downstream tumor segmentation experiment that is clinically relevant. Full article
(This article belongs to the Special Issue Artificial Intelligence in Sustainable Development)
26 pages, 17140 KB  
Article
Oil Palm Fruit Maturity Classification Algorithm Based on LAB Dual-Branch Feature Fusion
by Junli Wu, Xuewei Liu, Junjie Song, Kai Zhang and Yan Zhao
Appl. Sci. 2026, 16(14), 7292; https://doi.org/10.3390/app16147292 - 21 Jul 2026
Abstract
Accurate classification of oil palm fruit maturity is critical for optimizing harvest schedules and maximizing yield. To address the limited discrimination accuracy caused by subtle color variations, this paper proposes an LAB dual-branch feature fusion method for five-class oil palm fruit maturity classification. [...] Read more.
Accurate classification of oil palm fruit maturity is critical for optimizing harvest schedules and maximizing yield. To address the limited discrimination accuracy caused by subtle color variations, this paper proposes an LAB dual-branch feature fusion method for five-class oil palm fruit maturity classification. A classification model named CMF-Net is constructed, in which the luminance (L) and chrominance (ab) components are decoupled into two parallel branches to separately extract structural and color information. An input adaptation module is introduced to perform channel mapping and feature enhancement. During feature fusion, a gated fusion mechanism is employed for adaptive integration of the dual-branch features, followed by a BiFPN module for multi-scale feature interaction. A multi-level attention enhancement module composed of coordinate attention and self-attention is designed to strengthen key-region perception and global contextual modeling. An auxiliary color regression constraint and ordinal supervision are incorporated to improve the model’s sensitivity to subtle chromatic changes and the progressive relationship among maturity stages. The proposed algorithm provides a color-decoupled and order-aware representation framework for gradual fruit maturity classification, offering a useful reference for fine-grained agricultural visual classification tasks involving continuous color transitions. Experimental results on a public dataset show that CMF-Net achieves a validation accuracy of 94.66%, outperforming mainstream models by 1.3–6.1 percentage points, while also yielding superior recall and F1-score. Full article
(This article belongs to the Section Agricultural Science and Technology)
Show Figures

Figure 1

30 pages, 1129 KB  
Review
Radiation-Induced Defect Engineering in REBCO High-Temperature Superconductors: Defect Morphology, Vortex Pinning, and Technological Reliability
by Sanat Tolendiuly, Karakat Bolatzhan, Nursultan Rakhym, Sergey Fomenko, Kaster Kamunur, Beibit Karibayev, Aigerim Sovet and Sharafkhan Assylkhan
Sci 2026, 8(7), 178; https://doi.org/10.3390/sci8070178 - 20 Jul 2026
Viewed by 181
Abstract
Radiation-induced defect engineering is an effective approach for modifying the vortex-pinning landscape in high-temperature superconductors, particularly REBCO-coated conductors and Bi-based cuprates. This review critically summarizes the relationship between irradiation parameters, defect morphology, and superconducting performance. The discussion covers point defects, defect clusters, columnar [...] Read more.
Radiation-induced defect engineering is an effective approach for modifying the vortex-pinning landscape in high-temperature superconductors, particularly REBCO-coated conductors and Bi-based cuprates. This review critically summarizes the relationship between irradiation parameters, defect morphology, and superconducting performance. The discussion covers point defects, defect clusters, columnar tracks, planar defects, and displacement cascades generated by electrons, gamma rays, light ions, heavy ions, and neutrons. Special attention is given to the dual role of irradiation: moderate defect concentrations can enhance the critical current density by introducing artificial pinning centers, whereas excessive disorder suppresses the superconducting transition temperature and degrades current transport. The review also discusses the relevance of irradiation effects for fusion magnets, space technologies, accelerator systems, and high-field applications. Finally, the review identifies key challenges for future HTS radiation engineering, including cryogenic in situ irradiation, coupled radiation–strain–field experiments, damage metrics beyond dpa, and multi-scale models capable of linking atomic defect production, oxygen disorder, vortex pinning, and macroscopic Jc/Tc degradation. Full article
(This article belongs to the Section Materials Science)
Show Figures

Figure 1

46 pages, 17142 KB  
Article
Topological Continuity-Enforced Retinal Vessel Segmentation via Frequency-Aware Decomposition and Prototype Refinement
by Feng Li and Yaoyao Feng
Symmetry 2026, 18(7), 1228; https://doi.org/10.3390/sym18071228 - 20 Jul 2026
Viewed by 71
Abstract
Automated and accurate segmentation of retinal vessels in fundus images provides pivotal evidence for ophthalmologists to effectively and non-invasively diagnose prevalent ocular and systemic diseases. However, existing methods often struggle to maintain the topological continuity of fine-diameter capillaries, leading to severe vascular discontinuity [...] Read more.
Automated and accurate segmentation of retinal vessels in fundus images provides pivotal evidence for ophthalmologists to effectively and non-invasively diagnose prevalent ocular and systemic diseases. However, existing methods often struggle to maintain the topological continuity of fine-diameter capillaries, leading to severe vascular discontinuity and fragmented segmentation results in challenging scenarios such as complex, irregular microvascular branches, pathological lesions, and high-noise conditions. To address these limitations, we developed a novel symmetric dual-branch network with frequency-aware decomposition and prototype refinement (FDPR-DBNet). Specifically, the network initially utilizes the discrete wavelet transform (DWT) to decompose input retinal images into high-frequency and low-frequency components, which are then processed by a structurally symmetric dual-branch encoder. In the high-frequency branch, the parallel atrous convolution activation (PACA) module is designed to explore fine-grained contour and edge patterns related to vessel terminals and microvessels. Concurrently, within the low-frequency branch, the spatial-frequency characteristic activation (SFCA) unit is constructed by introducing the selective state-space model (S6) and Fourier transform to extract salient structural backbones. Moreover, the spatial attention residual fusion (SARF) module and cross-frequency fusion (CFF) block are designed to establish a symmetric guidance mechanism, effectively reinforcing bidirectional feature interaction and alignment across different frequency spectra to eliminate vascular fragmentation. Furthermore, by embedding global and local window self-attention into the Transformer, we formulated the cross-scale enhancement (CSE) module, comprising global semantic enhancement (GSE) and local detail enhancement (LDE), to model multi-scale contextual semantic correlations and enhance the adaptive recognition of vessel structures. Ultimately, we embedded the multi-wise prototype characteristic refinement (MPCR) component into the decoder to correct cross-scale semantic features through a dynamic calibration mechanism, while introducing a new connectivity loss to strictly enforce topological continuity. Experimental results on four publicly available retinal image datasets (DRIVE, CHASE_DB1, STARE, and IOSTAR) demonstrate that the proposed model achieves competitive performance and effectively preserves vascular integrity even in the presence of fundus lesions and noise. Full article
(This article belongs to the Section Computer)
Show Figures

Figure 1

18 pages, 7020 KB  
Article
DDFNet: A Dual-Stream Decoupled Feature Alignment Network for Multimodal Medical Image Fusion
by Pengquan Han, Manyuan Cheng, Cuiyin Liu and Bo Liang
Appl. Sci. 2026, 16(14), 7245; https://doi.org/10.3390/app16147245 - 20 Jul 2026
Viewed by 154
Abstract
Multimodal Medical Image Fusion (MMIF) aims to exploit the correlation and complementarity of different imaging modalities to integrate cross-modal information and provide more potential support for clinical diagnosis. However, effectively preserving both shallow and deep features of each modality while integrating multi-source representations [...] Read more.
Multimodal Medical Image Fusion (MMIF) aims to exploit the correlation and complementarity of different imaging modalities to integrate cross-modal information and provide more potential support for clinical diagnosis. However, effectively preserving both shallow and deep features of each modality while integrating multi-source representations remains challenging, often leading to blurred edges and loss of fine details in fused images. To address this issue, this paper proposes a novel Dual-stream Decoupled Feature parallel alignment fusion network (DDFNet). Built upon an autoencoder (AE) framework, the encoder adopts a dual-branch multi-scale design, consisting of a Detail Feature Extraction Module (DFEM) and a Global Context Feature Extraction Module (GCFEM). These two parallel branches leverage the complementary advantages of CNNs and Transformers to capture local details and global contextual dependencies, respectively, while operating independently without feature interference. In addition, an Improved Multi-scale Cross-alignment and Spatial Attention Fusion module (IMSC-SAF) is introduced to enhance feature integration. It performs complementary feature alignment and cross-attention to strengthen multiscale interactions, while a spatial attention mechanism highlights spatially salient regions. The decoder then reconstructs the fused representation and generates the final fused image. Extensive experiments on publicly available datasets from Harvard Medical School demonstrate that the proposed method achieves competitive fusion performance. In particular, DDFNet consistently obtains the highest MI, VIF, and SSIM values across the evaluated datasets while maintaining competitive results in the remaining objective metrics. Qualitative comparisons further demonstrate that the proposed method effectively preserves complementary anatomical and functional information, producing visually balanced fusion results. Full article
Show Figures

Figure 1

25 pages, 5239 KB  
Article
An Ultra-Short-Term Wind Farm Power Forecasting Method Incorporating Spatial Features for Sustainable Energy Integration
by Yanxia Wang, Weilong Yu, Minghan Ma, Yongqiang Kang, Yunyun Yun, Xiping Ma and Shuaibing Li
Sustainability 2026, 18(14), 7387; https://doi.org/10.3390/su18147387 - 19 Jul 2026
Viewed by 229
Abstract
Accurate wind power forecasting is imperative for ensuring grid stability and facilitating the large-scale integration of renewable energy—both central pillars of the global energy transition and the Dual Carbon strategic goals. However, existing methods often fail to fully capture the spatial heterogeneity and [...] Read more.
Accurate wind power forecasting is imperative for ensuring grid stability and facilitating the large-scale integration of renewable energy—both central pillars of the global energy transition and the Dual Carbon strategic goals. However, existing methods often fail to fully capture the spatial heterogeneity and interdependencies among individual turbines, limiting their effectiveness for sustainable grid operation. To address this gap, this paper proposes an ultra-short-term wind power forecasting framework that incorporates explicit multi-dimensional spatial features. At the feature level, a 12-dimensional spatial feature system is constructed to quantify the microscale topology of wind farms. These static spatial attributes are seamlessly fused with dynamic temporal data using a dimensionality-balance factor strategy. Finally, a hybrid deep learning network comprising a multi-scale CNN, a multi-layer BiLSTM, and a multi-head self-attention mechanism is developed to capture complex spatiotemporal patterns. Experimental results on three real-world datasets show that the proposed method significantly outperforms baseline models, reducing the Mean Absolute Percentage Error by up to 11.09% and improving the coefficient of determination R2 up to 0.9120. By improving forecast accuracy and robustness, the method directly supports more reliable grid dispatching, reduces curtailment of wind energy, and thus contributes to the sustainable utilization of renewable resources. These findings demonstrate that incorporating explicit spatial correlation effectively enhances the accuracy and robustness of ultra-short-term wind power forecasting, providing robust decision support for power grid dispatching and advancing the sustainability of modern power systems. Full article
(This article belongs to the Section Energy Sustainability)
Show Figures

Figure 1

27 pages, 42347 KB  
Article
Multi-Scale Feature Refinement Road Crack Detection Algorithm Based on Improved RT-DETR
by Wenxuan Xu, Yong Wang and Hairui Zhang
Appl. Sci. 2026, 16(14), 7177; https://doi.org/10.3390/app16147177 - 17 Jul 2026
Viewed by 200
Abstract
Road crack detection is an important part of the operation and maintenance of intelligent transportation infrastructure. However, the traditional convolution method has the problem of geometric mismatch when dealing with the slender linear structure of cracks, and the multi-scale adaptability and anti-interference ability [...] Read more.
Road crack detection is an important part of the operation and maintenance of intelligent transportation infrastructure. However, the traditional convolution method has the problem of geometric mismatch when dealing with the slender linear structure of cracks, and the multi-scale adaptability and anti-interference ability of the detection model under a complex road background are still insufficient. Aiming at key problems such as the difficulty of linear feature extraction, poor multi-scale adaptability, and complex background interference, this paper proposes a road crack detection algorithm based on an improved RT-DETR multi-dimensional feature fusion method. This method improves the detection performance by introducing three core innovative modules. Firstly, a lightweight directional decoupled dynamic convolution (D3Conv) is designed, which makes the convolution kernel fit the crack direction through direction prediction and adaptive sampling, so as to improve the recall rate by 0.012 (from 0.647 to 0.659) and reduce the computational burden. Secondly, a multi-scale cross-attention enhancement module (MCAA) was proposed to fuse multi-scale convolution and direction-aware strip convolution, and the dual attention mechanism was combined to enhance the perception of crack morphology and scale. Furthermore, a context-guided feature reconstruction (CGFR) module was constructed, which effectively aggregated global semantics and local details through dynamic feature selection and a multi-branch refining mechanism to improve the accuracy of boundary location. Experiments on the public dataset SVRDD2024 show that the proposed algorithm achieves an mAP@50 of 0.724, which is 2.2% higher than the baseline RT-DETR-R18 in terms of mAP@50. Especially, the detection performance is significantly improved for small cracks and in complex environments. It provides a reliable and efficient solution for the automatic inspection of intelligent transportation infrastructure. Full article
Show Figures

Figure 1

27 pages, 69728 KB  
Article
SAG-DeepLabV3+: An Enhanced Deep Learning Model for High-Precision Detection of Mining-Induced Ground Fissures from UAV Imagery
by Bo Xu, Di Cai, Jintao Shi, Kelin Sui, Wentai Tang and Chuangchuang Liu
Remote Sens. 2026, 18(14), 2388; https://doi.org/10.3390/rs18142388 - 17 Jul 2026
Viewed by 224
Abstract
To address the challenges of low detection accuracy and weak generalization in identifying mining-induced ground fissures from UAV imagery, caused by their slender and discontinuous morphology, complex background clutter, and multi-scale surface features, this paper proposes an enhanced deep semantic segmentation model, SAG-DeepLabV3+ [...] Read more.
To address the challenges of low detection accuracy and weak generalization in identifying mining-induced ground fissures from UAV imagery, caused by their slender and discontinuous morphology, complex background clutter, and multi-scale surface features, this paper proposes an enhanced deep semantic segmentation model, SAG-DeepLabV3+ (with Spatial Vision Transformer, Attention mechanisms, and Adaptive Gated Fusion). Specifically, to enhance global context modeling and fine boundary delineation, we introduce a Spatial Vision Transformer (SVT) branch within the Atrous Spatial Pyramid Pooling (ASPP) module. We further employ a dual attention mechanism, sequentially combining Squeeze-and-Excitation (SE) and a Convolutional Block Attention Module (CBAM), for progressive channel and spatial feature refinement. Moreover, an Adaptive Gated Fusion (AGF) module is designed to dynamically optimize the fusion of multi-level decoder features. Experiments on a dedicated UAV-based mining fissure dataset comprising 1280 annotated images show that SAG-DeepLabV3+ achieves a state-of-the-art mean Intersection over Union (mIoU) of 79.52% (with Xception backbone) and 79.19% (with lightweight MobileNetV2 backbone), surpassing DeepLabV3+, U-Net, and PSPNet by a significant margin. Furthermore, by leveraging transfer learning (pre-training on the public CrackVision12K dataset and fine-tuning on our mining fissure dataset), the model’s mIoU is further elevated to 82.04%, demonstrating superior generalization capability. The proposed SAG-DeepLabV3+ effectively balances high accuracy with operational efficiency, fulfilling the potential demand for lightweight automated fissure monitoring under resource-limited field deployments, and lays a foundation for subsequent real-time on-site deployment verification. Full article
Show Figures

Figure 1

24 pages, 11916 KB  
Article
Symmetry-Aware Stock Prediction Based on Optimized Multi-Module Collaborative Features with LSTM-CBAM-Time2Vec-KAN
by Huiyong Wu and Xiufeng Hong
Symmetry 2026, 18(7), 1198; https://doi.org/10.3390/sym18071198 - 16 Jul 2026
Viewed by 209
Abstract
This study proposes a hybrid deep learning model named LSTM-CBAM-Time2Vec-KAN based on symmetry awareness and optimized multi-module collaborative features, aiming to improve the accuracy and stability of stock price prediction. To address common shortcomings in traditional forecasting models such as insufficient feature extraction, [...] Read more.
This study proposes a hybrid deep learning model named LSTM-CBAM-Time2Vec-KAN based on symmetry awareness and optimized multi-module collaborative features, aiming to improve the accuracy and stability of stock price prediction. To address common shortcomings in traditional forecasting models such as insufficient feature extraction, difficulties in parameter optimization, and inadequate utilization of temporal characteristics, the research innovatively exploits the symmetry inherent in financial time series, particularly their temporal periodicity and cross-dimensional feature consistency, to construct an intelligent prediction framework that integrates multiple modules. First, wavelet transform is applied to perform multi-scale decomposition and signal reconstruction on the raw stock price sequence, effectively extracting high signal-to-noise ratio features. Second, the Northern Goshawk Optimization (NGO) algorithm is employed to jointly optimize key hyperparameters of the model, including the LSTM hidden layer dimension and CBAM compression ratio, thereby resolving the challenge of parameter coupling across modules. Third, the CBAM attention mechanism enhances the importance of temporal features extracted by LSTM through a dual mechanism of channel and spatial attention, enabling the model to focus on critical price movement points. Meanwhile, Time2Vec encoding transforms temporal information into embedding representations with periodic properties, effectively capturing cyclical patterns at daily, weekly, and monthly trading intervals. Finally, the Kolmogorov–Arnold network (KAN) fuses multimodal features and produces precise predictive outputs. Experimental results show that the proposed model significantly outperforms all baseline models in four evaluation metrics, namely mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), and coefficient of determination (R2), which verifies its superior prediction accuracy and robustness. Furthermore, analyses of stock price forecasting under different time spans and simulated trading performance under various trading strategies further demonstrate that this study provides a feasible and effective technical solution for financial time-series forecasting, with important theoretical research value and practical application value. Full article
Show Figures

Figure 1

34 pages, 18987 KB  
Article
SFE-FM: A Dual-Branch Network with Spectral Feature Enhancement and Feature Mixing for Hyperspectral Image Classification
by Xiangsuo Fan, Guilan Huang, Yong Mei and Peng Li
Remote Sens. 2026, 18(14), 2362; https://doi.org/10.3390/rs18142362 - 15 Jul 2026
Viewed by 242
Abstract
Hyperspectral image (HSI) classification remains a challenging task due to its high-dimensional spectral characteristics and the complex spatial heterogeneity of remote sensing scenes. Although significant progress has been made with convolutional neural networks (CNNs) and Transformer methods, existing approaches often fail to adequately [...] Read more.
Hyperspectral image (HSI) classification remains a challenging task due to its high-dimensional spectral characteristics and the complex spatial heterogeneity of remote sensing scenes. Although significant progress has been made with convolutional neural networks (CNNs) and Transformer methods, existing approaches often fail to adequately model fine-grained variations in spectral curves, whilst struggling to strike a good balance between preserving local texture and modelling global context. To this end, this paper proposes a spatial–spectral dual-branch network (SFE-FM) for HSI classification, which combines spectral feature enhancement with a lightweight feature fusion mechanism to improve feature representational capacity. Specifically, a Spectral Feature Enhancement (SFE) module is designed to explicitly model spectral trends through first- and second-order differential operations, thereby enhancing potential discriminative information; at the same time, a lightweight Feature Mixing (FM) module is introduced into the network to model global dependencies across channels. In the dual-branch architecture, the spatial branch and the spectral branch extract spatial texture features and spectral semantic features respectively, and these features are adaptively reweighted using a cross-branch-guided fusion strategy (CSSF) to facilitate the effective integration of the two types of information. In addition, a multi-scale attention optimisation module (MSAO) has been introduced to enhance the response in key areas and improve the robustness of feature representations. Experiments conducted on the four benchmark datasets—WHU-Hi-HanChuan, Qingyun, Salinas and Pavia University, the proposed method achieved overall classification accuracies of 99.62%, 98.70%, 99.97% and 99.81%, respectively. Whilst maintaining a relatively low number of parameters (0.631M), it delivered competitive performance, demonstrating a good balance between classification accuracy and computational efficiency. Full article
Show Figures

Figure 1

27 pages, 2043 KB  
Article
Bio-Inspired Enhanced Adaptive Centered Collision Optimizer for Hyperparameter Optimization of Multi-Scale Spatio-Temporal ConvNeXt in Boxing Action Recognition
by Tianyue Liu
Biomimetics 2026, 11(7), 497; https://doi.org/10.3390/biomimetics11070497 - 15 Jul 2026
Viewed by 303
Abstract
Accurate boxing action recognition is critical for intelligent combat training, action quality assessment, and sports injury prevention. However, existing deep learning approaches face three key challenges: limited feature extraction for high-speed non-rigid boxing motions, weak robustness against background interference and occlusion, and performance [...] Read more.
Accurate boxing action recognition is critical for intelligent combat training, action quality assessment, and sports injury prevention. However, existing deep learning approaches face three key challenges: limited feature extraction for high-speed non-rigid boxing motions, weak robustness against background interference and occlusion, and performance instability from labor-intensive manual hyperparameter tuning. Furthermore, the original Centered Collision Optimizer (CCO), a biomimetic algorithm inspired by celestial collision dynamics, suffers from insufficient population diversity, poor adaptive regulation, and premature convergence in high-dimensional hyperparameter optimization tasks. To address these issues, this paper proposes a novel biomimetic optimization-driven boxing action recognition framework, where an Enhanced Adaptive Centered Collision Optimizer (EACCO) automatically optimizes the hyperparameters of a Multi-Scale Spatio-Temporal Adaptive ConvNeXt (MSTA-ConvNeXt) network. First, the MSTA-ConvNeXt backbone integrates multi-scale dynamic deformable convolution, a Bi-GRU spatio-temporal fusion module, and a dual-channel attention mechanism to enhance fine-grained feature extraction and temporal modeling. Second, three biomimetic improvements are introduced to CCO: Tent chaotic elite opposition-based initialization, adaptive nonlinear convergence factor with dynamic weight guidance, and adaptive Gaussian-Cauchy hybrid mutation, which balance exploration and exploitation and avoid local optima. Experiments on two public benchmark datasets show that the proposed framework achieves 96.1% accuracy, 95.9% precision, 95.7% recall, and 95.8% F1-score on the Boxing Jab Skeleton Dataset, and 95.4% accuracy, 95.2% precision, 94.9% recall, and 95.0% F1-score on the Olympic Boxing dataset, outperforming all state-of-the-art methods. Ablation studies validate the effectiveness of each EACCO component and confirm that this biomimetic hyperparameter optimization approach outperforms manual tuning and other popular optimizers. This work provides an effective biomimetic optimization solution for intelligent sports action recognition. Full article
(This article belongs to the Special Issue Bio-Inspired Computation and Its Applications)
Show Figures

Figure 1

21 pages, 2637 KB  
Article
Hybrid Transformer–CNN with Boundary-Aware Attention for Accurate Multi-Modal Brain Tumor Segmentation
by Jamshid Khamzaev, Jakhongir Karimberdiyev, Mekhriddin Rakhimov, Islambek Saymanov, Shavkat Otamurodov, Odiljon Rikhsimboev, Ilin Dmitriy, Alpamis Kutlimuratov and Fazliddin Makhmudov
BioMedInformatics 2026, 6(4), 46; https://doi.org/10.3390/biomedinformatics6040046 - 14 Jul 2026
Viewed by 277
Abstract
Background: Accurate segmentation of brain tumors from multi-modal magnetic resonance imaging (MRI) is essential for diagnosis, treatment planning, and therapy monitoring. However, this task remains challenging due to tumor heterogeneity, irregular boundaries, and the complex anatomical structure of surrounding tissues. In particular, precise [...] Read more.
Background: Accurate segmentation of brain tumors from multi-modal magnetic resonance imaging (MRI) is essential for diagnosis, treatment planning, and therapy monitoring. However, this task remains challenging due to tumor heterogeneity, irregular boundaries, and the complex anatomical structure of surrounding tissues. In particular, precise delineation of tumor sub-regions—including whole tumor, tumor core, and enhancing tumor—continues to be a major limitation of existing automated methods. Methods: In this study, we propose a novel hybrid CNN–Transformer framework that integrates local feature extraction with global contextual modeling for improved brain tumor segmentation. The architecture consists of three main components: a dual-pathway encoder for capturing fine-grained and contextual features, a multi-scale feature fusion module based on spatial pyramid pooling with dense connections, and a boundary-aware attention decoder designed to enhance segmentation accuracy around tumor edges. The model utilizes four MRI modalities (T1, T1ce, T2, and FLAIR) to capture complementary tumor characteristics. In addition, a hybrid loss function combining Dice, focal Tversky, and boundary losses is employed to address class imbalance and improve boundary precision. Results: Experimental results on the BraTS 2023 dataset demonstrate superior performance, achieving Dice scores of 92.3%, 88.7%, and 84.5% for whole tumor, tumor core, and enhancing tumor, respectively, while maintaining high computational efficiency. Conclusion: The proposed framework achieves accurate and robust brain tumor segmentation by effectively integrating local and global features, demonstrating its potential for automated multi-modal MRI analysis in clinical practice. Full article
Show Figures

Figure 1

32 pages, 28977 KB  
Article
Acoustic Emission-Based Offshore Pipeline Valve Leakage Detection Toward Enhanced Process Safety
by Hongdong Qin, Xingshuang Hao, Zhenhao Zhu, Weizhe Ren, Xiaolong Qiu, Yuchen Lu, Hongbing Liu and Yuxuan Zhang
Sensors 2026, 26(14), 4451; https://doi.org/10.3390/s26144451 - 13 Jul 2026
Viewed by 277
Abstract
Valve leakage in marine oil and gas pipelines is a critical failure mode that threatens operational safety, ecological integrity and production economic benefits, creating an urgent demand for accurate, real-time and robust fault diagnosis systems. Acoustic Emission (AE) technology captures transient acoustic signatures [...] Read more.
Valve leakage in marine oil and gas pipelines is a critical failure mode that threatens operational safety, ecological integrity and production economic benefits, creating an urgent demand for accurate, real-time and robust fault diagnosis systems. Acoustic Emission (AE) technology captures transient acoustic signatures generated by leakage to enable non-intrusive online monitoring, while deep learning supports intelligent analysis through automatic signal feature extraction. Nevertheless, traditional AE-based leakage diagnosis methods rely heavily on manual feature engineering and fixed signal processing rules. Existing AE-driven deep learning methods fail to simultaneously deliver high detection accuracy, low inference latency and strong noise immunity, hindering their practical deployment on offshore platforms. To address these limitations, this paper proposes a Parameter-free Star-shaped Attention Fusion Network (SAFNet) for lightweight valve leakage localization using AE signals. Centered on the Temporal Pyramid Encoder (TPE) and Progressive Lightweight Star-shaped Attention (PLSA) module, SAFNet integrates Dual Bilinear Star Mapping (DBSM), Energy-Driven Feature Refiner (EDFR) and Multi-Scale Gated Attention Fusion (MS-GAF) modules. This architecture achieves efficient multi-scale temporal feature extraction, parameter-free nonlinear enhancement, noise-resistant refined feature processing and adaptive hierarchical feature fusion. The proposed method is applicable to valve leakage diagnosis of marine oil and gas pipelines under variable pressure and complex marine noise conditions. Comprehensive experiments are conducted on a dataset constructed by combining laboratory controlled leakage signals with real marine background noise recorded from the Liwan 3-1 offshore platform. The experimental results reveal that SAFNet balances high detection accuracy, compact model size and low inference latency simultaneously. Specifically, the network maintains a stable detection accuracy above 95% under pipeline pressures ranging from 2 MPa to 5 MPa, and exhibits excellent stability under extreme heavy noise environments. Ablation experiments further validate the synergistic performance gain brought by all core modules. The presented network delivers an efficient lightweight solution for valve leakage localization under simulated marine acoustic conditions, promotes the development of intelligent monitoring technologies for marine pipeline systems, and comprehensively improves offshore operational safety and marine ecological protection capacity. Full article
(This article belongs to the Section Physical Sensors)
Show Figures

Figure 1

Back to TopTop