Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (95)

Search Parameters:
Keywords = dual-stream CNN

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
30 pages, 754 KB  
Article
Joint Multi-Channel Dual-Polarization Autoencoder for Scalable End-to-End Long-Haul Optical Transmission
by Abid Iqbal, Waqas A. Imtiaz, Muhammad Ismail Mohmand and Muhammad Kamran Abbasi
Photonics 2026, 13(9), 859; https://doi.org/10.3390/photonics13090859 - 11 Sep 2026
Viewed by 67
Abstract
This paper introduces a Joint 4-channel wavelength-division multiplexed (WDM) Dual-Polarization Autoencoder (J-4WDM-DPAE) framework to address the scaling limitations of existing end-to-end learning architectures for long-haul coherent optical transmission. The proposed approach employs a unified one-dimensional residual convolutional neural network (CNN) decoder that processes [...] Read more.
This paper introduces a Joint 4-channel wavelength-division multiplexed (WDM) Dual-Polarization Autoencoder (J-4WDM-DPAE) framework to address the scaling limitations of existing end-to-end learning architectures for long-haul coherent optical transmission. The proposed approach employs a unified one-dimensional residual convolutional neural network (CNN) decoder that processes all eight complex symbol streams (4 WDM channels × 2 polarizations) at the symbol rate, enabling simultaneous exploitation of inter-channel and inter-polarization correlations with low inference latency. The transceiver is trained through a fully differentiable dual-polarization Manakov split-step Fourier method (SSFM) model including span-wise amplified spontaneous emission (ASE) noise and an effective combined transmitter–local-oscillator phase-noise process, enabling co-optimization of a shared geometric constellation shaping (GCS) encoder under realistic nonlinear and linewidth constraints. Robustness is further enhanced by randomized launch powers and signal-to-noise ratio (SNR) conditions during training. Additional robustness is assessed by cross-SPS evaluation (SSFM-resolution mismatch) and SSFM convergence checks, indicating that the reported achievable-rate trends are not an artifact of the baseline SSFM discretization. Evaluations over standard single-mode fiber (SSMF) for 16-, 32-, and 64-quadrature amplitude modulation (QAM) show that at 1000 km, the learned constellations achieve generalized mutual information close to the dual-polarization limits, with pre-FEC bit-error rates remaining below an adopted threshold of 2×102 (used as a representative soft-decision FEC operating target) up to 4000 km when the model is re-trained for each distance. Complexity analysis further indicates that the unified WDM-aware decoder provides a quantitative performance–computational trade-off compared to existing counterparts under matched link conditions. Full article
(This article belongs to the Special Issue Machine Learning and Artificial Intelligence for Optical Networks)
22 pages, 1810 KB  
Article
Dual-Stream Spatial–Spectral Network with Nested Attention for Hyperspectral Image Classification
by Jianing Wang, Fanghao Li, Liang Chen, Shijie Liu, Wanjiao Zhang, Lijun Jiang and Chuanjie Zhang
Remote Sens. 2026, 18(18), 3126; https://doi.org/10.3390/rs18183126 - 11 Sep 2026
Viewed by 94
Abstract
Hyperspectral image classification (HSI) requires a model to distinguish subtle spectral differences while preserving the spatial structure of land-cover regions. CNN-based methods are effective for local spectral–spatial extraction, but their limited receptive fields can weaken broader context modelling. Transformer-based methods improve long-range dependency [...] Read more.
Hyperspectral image classification (HSI) requires a model to distinguish subtle spectral differences while preserving the spatial structure of land-cover regions. CNN-based methods are effective for local spectral–spatial extraction, but their limited receptive fields can weaken broader context modelling. Transformer-based methods improve long-range dependency modelling, yet fixed patch partitioning may reduce their sensitivity to fine local structures. To address these limitations, this study proposes the Dual-Stream Spatial–Spectral Network with Nested Attention (DSSN), which separates local spectral–spatial feature extraction from multi-scale spatial-context modelling before adaptive fusion. The DSSN combines a cascaded 3D-CNN spectral stream, a nested Transformer spatial stream with pixel-level and patch-level interactions, and a channel-attention-based adaptive fusion module. Experiments on Indian Pines, Pavia University and Salinas show DSSN achieves overall accuracies of 98.11%%, 99.88% and 99.82%, respectively, outperforming other baselines. The ablation experiments confirm that each major component contributes to the final performance. Although the model requires more parameters and longer inference time than several compared baselines, its inference time remains at the millisecond level. These results suggest that decoupled spatial–spectral representation and adaptive multi-scale fusion can improve hyperspectral image classification under the evaluated benchmark settings. Full article
(This article belongs to the Section Remote Sensing Image Processing)
24 pages, 21810 KB  
Article
A Dual-Stream Network with Dynamic Graph Convolution and Attention-Based BiGRU for IGBT Open-Circuit Fault Diagnosis in T-NPC Three-Level Inverters
by Lin Bai, Bo Guo, Weiye Jing, Feng Li and Wei Luo
Energies 2026, 19(17), 4227; https://doi.org/10.3390/en19174227 - 7 Sep 2026
Viewed by 193
Abstract
Existing CNN, TCN, residual, and lightweight network methods have achieved good performance in IGBT open-circuit fault diagnosis, but they often overlook the non-Euclidean relationships among signals. To address this limitation, this paper proposes a parallel graph–temporal network for T-NPC three-level inverters. A shared [...] Read more.
Existing CNN, TCN, residual, and lightweight network methods have achieved good performance in IGBT open-circuit fault diagnosis, but they often overlook the non-Euclidean relationships among signals. To address this limitation, this paper proposes a parallel graph–temporal network for T-NPC three-level inverters. A shared CNN extracts compact features from the three-phase currents and voltages, while the Sinkhorn–Wasserstein distance constructs a sample-level weighted dynamic graph for GCN-based relationship extraction. In parallel, BiGRU with global attention captures temporal information. Unlike fixed or equally weighted graphs, the proposed method adapts signal connections to different fault conditions. Furthermore, simulation models are constructed in MATLAB/Simulink, and the T-NPC converter operation is emulated on a real-time simulator Starsim MT6060. The proposed dual-stream model classifies 21 fault states, achieving an average validation accuracy of 99.88%, while maintaining high accuracy under severe noise. Full article
Show Figures

Figure 1

17 pages, 2940 KB  
Article
A Dual-Stream Deep Learning Framework Fusing Fractional-Order and Multifractal Features for Bridge Damage Detection
by Shuai Teng, Qinghua Xu, Zhifang Sun, Xiaosan Yin and Yingjie Wang
Algorithms 2026, 19(9), 756; https://doi.org/10.3390/a19090756 - 4 Sep 2026
Viewed by 179
Abstract
Deep learning has shown promise for vibration-based damage detection, yet its practical deployment is hindered by data scarcity, environmental sensitivity, and limited interpretability. To address these issues, we propose a dual-stream framework that explicitly embeds fractional-order and multifractal physical priors. The Grünwald–Letnikov fractional [...] Read more.
Deep learning has shown promise for vibration-based damage detection, yet its practical deployment is hindered by data scarcity, environmental sensitivity, and limited interpretability. To address these issues, we propose a dual-stream framework that explicitly embeds fractional-order and multifractal physical priors. The Grünwald–Letnikov fractional derivative (ν = 1.3) sharpens damage-related singularities, while multifractal detrended fluctuation analysis (MF-DFA) produces spatial maps of multiscale complexity across the sensor array. A temporal stream (1D-CNN + BiLSTM) processes the fractional-order enhanced signals, and a spatial stream (2D-CNN) processes MF-DFA maps derived from the same enhanced signals. Under the evaluation protocol adopted in this study, the framework achieves 97.3% accuracy on Z24 and 94.8% in the independent within-dataset KW51 experiment. In a separate cross-bridge experiment, the Z24-pretrained model obtains 86.7% accuracy on KW51 without target-domain updating and 94.6% after fine-tuning on 10% of the target data. Full article
Show Figures

Figure 1

21 pages, 6311 KB  
Article
EP-Net: An Equipment-Guided Dual-Stream CNN–Transformer Network for Fine-Grained Sports Image Classification
by Xiaocui Sang, Changwei Gu and Lei Zhao
Appl. Sci. 2026, 16(16), 8229; https://doi.org/10.3390/app16168229 - 18 Aug 2026
Viewed by 305
Abstract
Fine-grained sports recognition from static images is challenging because visually similar sports often exhibit nearly identical human poses, whereas their decisive differences are encoded by small-scale equipment and subtle human–equipment interactions. Existing single-stream convolutional or Transformer-based models tend to emphasize either local appearance [...] Read more.
Fine-grained sports recognition from static images is challenging because visually similar sports often exhibit nearly identical human poses, whereas their decisive differences are encoded by small-scale equipment and subtle human–equipment interactions. Existing single-stream convolutional or Transformer-based models tend to emphasize either local appearance or global context, making them vulnerable to equipment-detail loss and background interference. To address this problem, we propose the Equipment-Primed Network (EP-Net), a heterogeneous dual-stream architecture that treats sports equipment as a primary semantic cue for action discrimination. EP-Net employs an EfficientNetV2-S-based Equipment Stream to capture localized equipment shapes and textures and a Swin-Tiny-based Behavior Stream to model the athlete’s spatial configuration and global scene context. We further introduce a Cross-Modal Channel Attention (CMCA) module that projects equipment features into the behavior-feature space and performs directional channel recalibration. Unlike simple feature concatenation, CMCA uses equipment information to enhance action-relevant channels while reducing the relative influence of background-dominated responses. Experiments on the Sports-100 dataset show that EP-Net achieves a Top-1 accuracy of 98.80%, outperforming EfficientNetV2-S and Swin-Tiny by 3.00 and 2.35 percentage points, respectively. It also improves on naive dual-stream concatenation by 0.88 percentage points. Grad-CAM visualizations further indicate that EP-Net attends more consistently to discriminative equipment and human–equipment interaction regions. These results suggest that equipment-guided local–global feature interaction provides an effective solution to pose ambiguity and background interference in static fine-grained sports recognition. Full article
Show Figures

Figure 1

18 pages, 11820 KB  
Article
RTGNet: A Dual-Branch Network Integrating Recurrent Texture and Temporal Dynamics from sEMG for Lower-Limb Joint Angle Prediction
by Zhiwei Hu, Quansheng Xu, Shaowei Su, Yinggan Tang and Yonghong Xu
Sensors 2026, 26(16), 5144; https://doi.org/10.3390/s26165144 - 14 Aug 2026
Viewed by 361
Abstract
Accurate continuous prediction of lower-limb joint angles from surface electromyography (sEMG) remains challenging because of the nonlinear, non-stationary, and subject-specific nature of sEMG signals, which can reduce robustness and lead to degraded prediction accuracy during highly dynamic gait phases. In this study, we [...] Read more.
Accurate continuous prediction of lower-limb joint angles from surface electromyography (sEMG) remains challenging because of the nonlinear, non-stationary, and subject-specific nature of sEMG signals, which can reduce robustness and lead to degraded prediction accuracy during highly dynamic gait phases. In this study, we propose RTGNet, a dual-branch deep learning framework for lower-limb joint-angle prediction from multichannel sEMG. The method constructs two feature views from the same sEMG stream: recurrence-plot (RP)-based representations for nonlinear texture characterization and time-series sequences for long-term temporal dependency modeling. These views are processed by a convolutional neural network (CNN) with a convolutional block attention module (CBAM) and a bidirectional long short-term memory network (BiLSTM), respectively, and integrated through an adaptive gated fusion mechanism. An enhanced Huber-TopK loss is further employed to emphasize samples with large prediction errors. Experiments on the SIAT-LLMD dataset under an offline cross-subject evaluation setting show that RTGNet achieves a mean absolute error (MAE) of 3.87°, a root mean square error (RMSE) of 5.25°, and an R2 of 0.81 during walking, as well as an MAE of 4.53°, an RMSE of 6.47°, and an R2 of 0.84 during stair ascent. The proposed framework outperforms temporal-only and RP-based baselines, and ablation results further support the effectiveness of the gated fusion strategy and CBAM attention. Overall, these results suggest that integrating recurrence texture and temporal dynamics is a promising strategy for sEMG-driven joint-angle prediction and provides a useful basis for future exoskeleton control-oriented studies. Full article
Show Figures

Figure 1

24 pages, 72650 KB  
Article
Real-Time Road Crack Detection on Smartphones Through ConvLSTM-Based Temporal Knowledge Distillation from a CNN-KAN and VMamba Dual-Path Network
by Mengzhao Nie, Hua Huang, Mengxue Guo, Mingxia Dang and Ming Tang
Sensors 2026, 26(16), 5071; https://doi.org/10.3390/s26165071 - 10 Aug 2026
Viewed by 370
Abstract
Road crack images captured by smartphones suffer from low resolution, uneven illumination, and complex background interference. Mobile devices also have limited resources for real-time high-accuracy segmentation. A two-stage framework combines a high-accuracy dual-path teacher model with a knowledge-distilled lightweight student model. The teacher [...] Read more.
Road crack images captured by smartphones suffer from low resolution, uneven illumination, and complex background interference. Mobile devices also have limited resources for real-time high-accuracy segmentation. A two-stage framework combines a high-accuracy dual-path teacher model with a knowledge-distilled lightweight student model. The teacher model integrates a CNN-KAN path for local texture extraction and a VMamba path for global context modeling at linear complexity. A dedicated KAN-based fusion module learns adaptive nonlinear mappings between the two feature streams. On public crack datasets, the teacher model achieves an mIoU of 0.8087 and an mDice of 0.9028. It is then transferred to a self-constructed smartphone dataset built from continuous 30 fps video, where it reaches an mIoU of 0.7084 with strong robustness to illumination and blur. A GAN-based super-resolution strategy further improves the mIoU by 4.01%. A ConvLSTM-based knowledge distillation framework compresses the teacher into a lightweight MobileViT student model. This reduces the parameter count from 57.80 M to 1.57 M and cuts the GPU inference time from 99.56 ms to 1.46 ms, while retaining an mIoU of 0.7078. The deployed student model runs at 15 to 20 frames per second on an Android smartphone. An ablation study confirms that the ConvLSTM-based temporal distillation contributes beyond standard distillation. This framework provides a practical solution for real-time road crack monitoring on smartphones. Full article
(This article belongs to the Section Sensing and Imaging)
Show Figures

Figure 1

21 pages, 947 KB  
Article
A Stock Market Price Prediction Model Integrating a CNN–Transformer Dual-Channel Dynamic Attention Architecture
by Chengcheng Han, Jingwei Guo and Xingyu Feng
Mathematics 2026, 14(16), 2888; https://doi.org/10.3390/math14162888 - 10 Aug 2026
Viewed by 498
Abstract
Stock market price prediction remains a persistent challenge owing to the non-stationarity, high noise content, and intricate spatiotemporal dependencies that characterize financial time series. Existing approaches typically excel at either local pattern extraction or long-range dependency modeling, yet seldom reconcile both within a [...] Read more.
Stock market price prediction remains a persistent challenge owing to the non-stationarity, high noise content, and intricate spatiotemporal dependencies that characterize financial time series. Existing approaches typically excel at either local pattern extraction or long-range dependency modeling, yet seldom reconcile both within a unified framework. This paper introduces a CNN–Transformer dual-channel architecture equipped with a dynamic attention fusion module for stock price forecasting. The convolutional channel applies hierarchical dilated convolutions to distill fine-grained local patterns from multi-indicator sequences while suppressing high-frequency noise. Simultaneously, the Transformer channel employs multi-head self-attention to capture long-distance temporal correlations and regime-shift dynamics. A learnable gating mechanism then fuses the two feature streams by adaptively weighting local detail against global trend information according to market conditions. Experiments conducted on four real-world stock datasets spanning the S&P 500, CSI 300, NASDAQ Composite, and Hang Seng Index show that the proposed model reduces mean absolute error by 9.7–15.3% and root mean square error by 9.5–13.8% relative to competitive baselines including LSTM, CNN–LSTM, Informer, and PatchTST. Ablation studies further indicate that both channels and the fusion module contribute to prediction accuracy, and the architecture remains effective across markets with differing volatility profiles. Full article
Show Figures

Figure 1

18 pages, 7020 KB  
Article
DDFNet: A Dual-Stream Decoupled Feature Alignment Network for Multimodal Medical Image Fusion
by Pengquan Han, Manyuan Cheng, Cuiyin Liu and Bo Liang
Appl. Sci. 2026, 16(14), 7245; https://doi.org/10.3390/app16147245 - 20 Jul 2026
Viewed by 441
Abstract
Multimodal Medical Image Fusion (MMIF) aims to exploit the correlation and complementarity of different imaging modalities to integrate cross-modal information and provide more potential support for clinical diagnosis. However, effectively preserving both shallow and deep features of each modality while integrating multi-source representations [...] Read more.
Multimodal Medical Image Fusion (MMIF) aims to exploit the correlation and complementarity of different imaging modalities to integrate cross-modal information and provide more potential support for clinical diagnosis. However, effectively preserving both shallow and deep features of each modality while integrating multi-source representations remains challenging, often leading to blurred edges and loss of fine details in fused images. To address this issue, this paper proposes a novel Dual-stream Decoupled Feature parallel alignment fusion network (DDFNet). Built upon an autoencoder (AE) framework, the encoder adopts a dual-branch multi-scale design, consisting of a Detail Feature Extraction Module (DFEM) and a Global Context Feature Extraction Module (GCFEM). These two parallel branches leverage the complementary advantages of CNNs and Transformers to capture local details and global contextual dependencies, respectively, while operating independently without feature interference. In addition, an Improved Multi-scale Cross-alignment and Spatial Attention Fusion module (IMSC-SAF) is introduced to enhance feature integration. It performs complementary feature alignment and cross-attention to strengthen multiscale interactions, while a spatial attention mechanism highlights spatially salient regions. The decoder then reconstructs the fused representation and generates the final fused image. Extensive experiments on publicly available datasets from Harvard Medical School demonstrate that the proposed method achieves competitive fusion performance. In particular, DDFNet consistently obtains the highest MI, VIF, and SSIM values across the evaluated datasets while maintaining competitive results in the remaining objective metrics. Qualitative comparisons further demonstrate that the proposed method effectively preserves complementary anatomical and functional information, producing visually balanced fusion results. Full article
Show Figures

Figure 1

21 pages, 5616 KB  
Article
Dual-Stream SPP-CNN for High-Precision sEMG Gesture Recognition in Human–Machine Interfaces
by Zebin Li, Gang Zhang, Lifu Gao, Wenming Wang, Wei Lu, Guocai Liu and Jinzhong Zhang
Biomimetics 2026, 11(7), 508; https://doi.org/10.3390/biomimetics11070508 - 19 Jul 2026
Viewed by 415
Abstract
Surface electromyography (sEMG) signals directly reflect movement intention and are therefore promising for natural human–machine interaction. However, their inherent non-stationarity and high inter-subject variability remain major obstacles to robust feature extraction and model generalization. To address these challenges, this study proposes a dual-stream [...] Read more.
Surface electromyography (sEMG) signals directly reflect movement intention and are therefore promising for natural human–machine interaction. However, their inherent non-stationarity and high inter-subject variability remain major obstacles to robust feature extraction and model generalization. To address these challenges, this study proposes a dual-stream spatial pyramid pooling convolutional neural network (DSSCNN). In this framework, one-dimensional sEMG segments are transformed into two complementary image representations, continuous wavelet transform (CWT) spectrograms and Gramian angular difference field (GADF) images, forming a dual-channel input that jointly preserves time–frequency dynamics and temporal correlation structures. A dual-stream convolutional architecture then extracts discriminative features from each modality, after which a spatial pyramid pooling (SPP) layer aggregates multi-scale representations, enhancing the network’s capacity to capture robust spatiotemporal patterns. Extensive experiments demonstrate that DSSCNN achieves an average gesture recognition accuracy of 97.88% with low inter-subject variance under intra-subject random split, and 96.59% under leave-one-subject-out (LOSO) protocol. The practical viability of the proposed approach is further validated through real-time control of an unmanned ground vehicle (UGV). These results not only indicate that the dual-stream framework combined with SPP layer provides an effective strategy for high-precision sEMG-based gesture recognition but also provides a promising technical pathway toward next-generation natural human–machine interaction. Full article
(This article belongs to the Section Bioinspired Sensorics, Information Processing and Control)
Show Figures

Figure 1

44 pages, 1844 KB  
Article
LiveCH-VVC: Latency-Aware Dynamic Bitrate Ladder Prediction for VVC/LL-DASH Live Streaming
by Reka Sandaruwan Gallena Watthage and Anil Fernando
Signals 2026, 7(4), 64; https://doi.org/10.3390/signals7040064 - 7 Jul 2026
Viewed by 599
Abstract
Adaptive bitrate streaming over HTTP relies on carefully constructed bitrate ladders and ordered sets of bitrate–resolution pairs to deliver optimal perceptual quality under fluctuating network conditions. While content-aware methods based on convex hull optimisation have substantially improved ladder efficiency for Video-on-Demand, they require [...] Read more.
Adaptive bitrate streaming over HTTP relies on carefully constructed bitrate ladders and ordered sets of bitrate–resolution pairs to deliver optimal perceptual quality under fluctuating network conditions. While content-aware methods based on convex hull optimisation have substantially improved ladder efficiency for Video-on-Demand, they require exhaustive multi-resolution pre-encoding that is computationally prohibitive under the real-time constraints of live streaming. This challenge is compounded by the H.266/Versatile Video Coding (VVC) standard, which offers approximately 50% compression gains over HEVC at 8–10× the encoding complexity. This paper presents LiveCH-VVC, a latency-aware dynamic bitrate ladder prediction framework for VVC-encoded live streaming over Low-Latency DASH (LL-DASH) with CMAF packaging. The framework introduces four integrated modules: (i) a Lightweight Dual-Path CNN (LDP-CNN), obtained via teacher–student knowledge distillation (∼5 M parameters, 148 ms GPU inference), that jointly extracts spatial–temporal features from raw frames and compression-domain statistics from a fast VVC probe encode; (ii) an adaptive scene change detector with exponential moving average thresholding (F1 = 0.925) that triggers ladder updates only upon significant complexity shifts; (iii) a temporally augmented XGBoost multi-label classifier that predicts latency-constrained Pareto-optimal bitrate–resolution pairs; and (iv) an online adaptation engine that integrates Common Media Client Data (CMCD) feedback from CDN edge servers for continuous closed-loop refinement. Comprehensive evaluation on 81 UHD sequences (∼4050 CMAF segments) from three benchmark datasets demonstrates an average BD-Rate of +0.68% relative to the per-segment oracle convex hull 5.4× better than the state-of-the-art ARTEMIS framework (+3.67%) while achieving 73.3% encoding time savings, 2.37 s end-to-end latency, and a QoE score of 81.6 in live simulation with 100 concurrent clients. Ablation analysis confirms that the dual-path compression-domain branch (+0.44 pp) and temporal context augmentation (+0.35 pp) are the primary performance drivers, while the online adaptation mechanism provides 42% relative improvement over extended streaming sessions. Full article
Show Figures

Figure 1

19 pages, 17897 KB  
Article
S2M-Net: Dynamic Hyperspectral Unmixing Network Integrating Spectral Sequence Mamba and Local Spatial–Spectral Awareness
by Yongqing Yang, Mengmeng Xu, Weidong Zhang, Ji Zhang and Yuquan Gan
Remote Sens. 2026, 18(13), 2228; https://doi.org/10.3390/rs18132228 - 6 Jul 2026
Viewed by 518
Abstract
Hyperspectral unmixing aims to extract pure endmembers and their corresponding abundance from mixed pixels. Existing deep learning-based unmixing methods predominantly rely on convolutional neural networks (CNNs) or Transformer architectures. However, CNNs suffer from limited receptive fields and struggle to capture long-range spectral dependencies [...] Read more.
Hyperspectral unmixing aims to extract pure endmembers and their corresponding abundance from mixed pixels. Existing deep learning-based unmixing methods predominantly rely on convolutional neural networks (CNNs) or Transformer architectures. However, CNNs suffer from limited receptive fields and struggle to capture long-range spectral dependencies across the entire spectral sequence. While Transformers possess global modeling capabilities, they are constrained by quadratic computational complexity and lack the ability to adaptively filter redundant noise in consecutive spectral bands. To address these limitations, this paper proposes a dynamic hyperspectral unmixing network integrating a spectral sequence Mamba with local spatial–spectral awareness. Specifically, the network features a novel asymmetric dual-stream collaborative architecture. The first branch, the spectral sequence Mamba, models hyperspectral data as a one-dimensional continuous sequence and employs the selective state space model to perform global scanning with linear complexity. This adaptively filters redundant spectral bands to accurately extract high-purity global spectral semantics. The second branch, dedicated to local spatial–spectral awareness, uses an attention-augmented CNN to capture local continuous spectral variations and spatial textures, providing fine-grained geometric boundary constraints for abundance estimation. Furthermore, a spatially adaptive gated fusion module is designed to dynamically balance global spectral semantics and local spatial–spectral details according to the pixel mixing complexity of varying spatial regions. Extensive experiments on multiple public hyperspectral datasets demonstrate that the proposed method achieves significant improvements in unmixing accuracy over comparative methods. Full article
Show Figures

Figure 1

27 pages, 11691 KB  
Article
GoldFormer: A Texture-Aware Vision Transformer-Based Algorithm for Detecting Near-Identical Images
by Zobeir Raisi
Algorithms 2026, 19(7), 530; https://doi.org/10.3390/a19070530 - 1 Jul 2026
Viewed by 526
Abstract
Distinguishing authentic gold products from high-quality counterfeits is a challenging fine-grained computer vision problem; counterfeit items are engineered to replicate surface texture, hallmark engravings, color, and geometry with remarkable fidelity, making visual discrimination unreliable even for trained professionals. In this paper, we address [...] Read more.
Distinguishing authentic gold products from high-quality counterfeits is a challenging fine-grained computer vision problem; counterfeit items are engineered to replicate surface texture, hallmark engravings, color, and geometry with remarkable fidelity, making visual discrimination unreliable even for trained professionals. In this paper, we address the problem of visual gold authentication from unconstrained smartphone imagery in three main contributions. First, we introduce GoldNet, a public benchmark dataset designed for this task, comprising 2127 real-world images of authentic and counterfeit gold items collected under diverse real-world conditions. Second, we evaluate fourteen classification architectures spanning classical handcrafted texture descriptors, convolutional neural networks (CNNs), and vision transformers under a rigorous transfer learning protocol, establishing the first comprehensive baseline for this problem. Third, we propose GoldFormer, a hybrid dual-stream algorithm that combines the local texture representations of ResNet-50 with the global contextual modeling capability of the Swin Transformer (Swin-T) through a newly designed Texture-Aware Attention Gate (TAAG) module. The TAAG dynamically modulates Swin feature dimensions using CNN-derived texture energy, providing improved discriminability and per-prediction interpretability without requiring post hoc attribution. Experimental results show that, under matched-resolution 5-fold cross-validation, the proposed GoldFormer attains the highest overall accuracy (95.02%, F1-score 0.9502) at roughly half the FLOPs of its higher-resolution setting, statistically tied with the strongest individual backbone (ViT-B/16, 94.31%; McNemar p=0.23) and on par with a training-free soft-voting ensemble (94.92%), while significantly improving on its own Swin-T backbone (93.65%) and adding built-in, attribution-free texture-gate interpretability. GoldFormer surpasses trained human-expert performance (89.80%) by approximately 5 percentage points. Full article
(This article belongs to the Section Algorithms for Multidisciplinary Applications)
Show Figures

Figure 1

21 pages, 8652 KB  
Article
Benign Share Benefit to Malignant: Balanced Mixing on Feature Space for Imbalanced Breast Cancer Classification
by Farchan Hakim Raswa, Muhammad Fadlurrohman, Bach-Tung Pham, Ika Candradewi, Afiahayati, Ming-Hsiang Su, Chung-I Huang, Kuo-Chen Li, Shih-Lun Chen, Yung-Hui Li and Jia-Ching Wang
Bioengineering 2026, 13(7), 765; https://doi.org/10.3390/bioengineering13070765 - 30 Jun 2026
Viewed by 814
Abstract
A deep learning model with an imbalanced mammography dataset can bias models toward common benign BI-RADS categories and reduce recognition of less frequent malignant or high-risk categories. To address this issue, we propose B2M (Benign Share Benefit to Malignant), a model-agnostic framework for [...] Read more.
A deep learning model with an imbalanced mammography dataset can bias models toward common benign BI-RADS categories and reduce recognition of less frequent malignant or high-risk categories. To address this issue, we propose B2M (Benign Share Benefit to Malignant), a model-agnostic framework for imbalance-aware multi-class BI-RADS classification in C-View mammography. B2M uses a two-phase training strategy that combines dual sampling with feature-space mixing. In Phase I, the model is trained with dual sampling, integrating instance-based and class-balanced sampling to increase minority-class representation while preserving majority-class diversity. In Phase II, the model is fine-tuned with feature-space mixing using samples from the two sampling streams. A soft-target regularization objective supervises the mixed features using labels from both streams, encouraging smoother decision boundaries across BI-RADS categories. We evaluated B2M on an imbalanced mammography cohort from the C-View EMBED dataset using stratified 5-fold cross-validation across multiple CNN backbones. C-View is a synthesized 2D mammographic image generated from 3D digital breast tomosynthesis data, capturing DBT-derived structural information while requiring less memory and computation than processing the full 3D image volume. Among these experiments, ResNeXt-50 with B2M achieved the highest balanced accuracy and Macro-F1 scores compared with the evaluated oversampling and mixing-based methods. This improvement requires an offline training-time overhead of approximately 2.81×, but it does not increase inference cost. Overall, the results suggest that B2M may be useful for imbalanced multi-class BI-RADS classification in C-View mammography. However, the findings are based on the EMBED cohort, and further validation, including external and prospective evaluation, is needed before clinical use. Full article
Show Figures

Figure 1

27 pages, 18725 KB  
Article
Physics-Guided Dual-Stream Fusion for Extreme Few-Shot Fault Diagnosis Under Massive Domain Shifts
by Shiqian Wu, Weiming Zhang, Huiyu Liu, Yuchen Lu and Yuxuan Zhang
Processes 2026, 14(12), 2012; https://doi.org/10.3390/pr14122012 - 20 Jun 2026
Viewed by 392
Abstract
Reliable fault diagnosis of rotating machinery is critical for averting serious failures in modern industrial systems. While data-driven deep learning has advanced condition monitoring, its success is fundamentally predicated on the availability of independent and identically distributed (I.I.D.) datasets. In realistic operational environments, [...] Read more.
Reliable fault diagnosis of rotating machinery is critical for averting serious failures in modern industrial systems. While data-driven deep learning has advanced condition monitoring, its success is fundamentally predicated on the availability of independent and identically distributed (I.I.D.) datasets. In realistic operational environments, machinery frequently experiences massive domain shifts induced by varying rotational speeds. Concurrently, acquiring high-fidelity fault instances is limited compared to abundant healthy baseline data, often resulting in a long-tailed distribution. Under such data-starved conditions, conventional few-shot domain adaptation (FSDA) methodologies often may be affected by distributional erasure; global alignment objectives are mainly driven by the healthy majority, causing sparse fault signatures to be erroneously absorbed as noise and leading to severe diagnostic performance degradation. To address this setting, this study develops a physics-guided dual-stream fusion framework for extreme few-shot cross-domain fault diagnosis. The method does not treat the Laplace wavelet, STFT, CNNs, or AdaBN as newly introduced techniques. Instead, it integrates these components into a unified diagnostic pipeline designed for long-tailed target support sets under large speed shifts. A learnable Laplace wavelet convolution is used in the temporal branch to emphasize transient impact responses, while STFT spectrograms provide a complementary time-frequency representation for the two-dimensional branch. The two feature streams are then fused for target fault classification. For domain adaptation, a Strict AdaBN strategy is applied using only the target support set, rather than the target test data or a large unlabeled target pool. Under the evaluated 50 healthy + 12 fault support condition, the healthy samples provide target-domain operating-background statistics for BN recalibration, while the limited fault samples are used for supervised classifier adjustment. Experiments on the HUSTbearing and Torino DIRG datasets show that the proposed integrated framework achieves stable performance under the evaluated few-shot cross-speed settings. These results suggest that combining physics-guided Laplace convolution, time-frequency representations, and support-set-restricted BN recalibration can be useful for bearing fault diagnosis when target fault samples are limited. Full article
(This article belongs to the Section Manufacturing Processes and Systems)
Show Figures

Figure 1

Back to TopTop