Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (158)

Search Parameters:
Keywords = dual-head Transformer

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
16 pages, 11729 KB  
Article
A Large-Field Photoacoustic-OCT Dual-Modal Imaging System Based on Temporal Medium Separation and Hardware-Based Coordinate Locking
by Hai Lin, Yuqian Liu, Yutong Wu, Yidan Zhang, Tianyang Deng and Yubin Liu
Photonics 2026, 13(8), 788; https://doi.org/10.3390/photonics13080788 - 19 Aug 2026
Viewed by 143
Abstract
Optical coherence tomography (OCT) and photoacoustic imaging (PAI) provide complementary structural and absorption contrasts but require different coupling conditions: 1310 nm swept-source OCT is attenuated by water, whereas PAI requires acoustic coupling. We developed a large-field dual-modal imaging system combining temporal medium separation [...] Read more.
Optical coherence tomography (OCT) and photoacoustic imaging (PAI) provide complementary structural and absorption contrasts but require different coupling conditions: 1310 nm swept-source OCT is attenuated by water, whereas PAI requires acoustic coupling. We developed a large-field dual-modal imaging system combining temporal medium separation with hardware-based coordinate locking. The OCT head, linear-array ultrasound transducer, and photoacoustic excitation fiber bundle were mounted on a rigid common platform, and a one-time calibration established a two-dimensional affine transformation between the modality coordinate systems. OCT was acquired in air and PAI in deionized water within a common large-field coordinate range. In five paired air–water measurements with an approximately 23 mm water path, the displayed OCT peak level decreased from 98.4 ± 1.5 dB in air to 79.4 ± 1.8 dB in water, corresponding to a mean reduction of 19.0 ± 1.4 dB. Quantitative registration was evaluated using a 5 × 5 dual-modal landmark phantom, with nine landmarks used for affine calibration and 16 excluded landmarks reserved for independent validation. The mean two-dimensional validation error was 0.235 ± 0.128 mm, with an RMSE of 0.266 mm and a maximum error of 0.446 mm. Five additional medium-switching cycles performed without recalibration yielded an overall registration error of 0.369 ± 0.163 mm across 80 validation measurements. PA spatial resolution was further characterized using six thin hair targets, yielding lateral and axial FWHM values of 0.342 ± 0.069 mm and 0.394 ± 0.073 mm, respectively. These results demonstrate reproducible two-dimensional en face OCT–PA coordinate mapping under modality-specific coupling conditions and support the proposed workflow as a phantom-based technical validation for large-field multimodal imaging. Full article
(This article belongs to the Special Issue Photoacoustic Imaging: Methods, Systems, and Applications)
Show Figures

Figure 1

27 pages, 19863 KB  
Article
CDF-DETR: Cross-Stage Attention and Dual-Scale Feature Calibration for Small-Object Detection in UAV Remote Sensing Imagery
by Rui Zou, Jinwei Guo, Jiaqi Liang, Kai Che, Yifan Deng and Binqi Chen
Remote Sens. 2026, 18(16), 2793; https://doi.org/10.3390/rs18162793 - 18 Aug 2026
Viewed by 251
Abstract
Small-object detection in unmanned aerial vehicle (UAV) remote sensing imagery is challenged by dense target distributions, substantial scale variation, complex ground backgrounds, and limited edge-computing resources. To address these challenges, we propose CDF-DETR, an end-to-end detector derived from the Real-Time Detection Transformer (RT-DETR). [...] Read more.
Small-object detection in unmanned aerial vehicle (UAV) remote sensing imagery is challenged by dense target distributions, substantial scale variation, complex ground backgrounds, and limited edge-computing resources. To address these challenges, we propose CDF-DETR, an end-to-end detector derived from the Real-Time Detection Transformer (RT-DETR). First, a Cross-Stage Partial Single-Head Attention Transformer (CSP-SHAT) backbone combines efficient local feature extraction with partial-channel global interaction to improve multi-scale representation while reducing the parameter count of the backbone. Second, a dual-scale feature calibration (DSFC) module sequentially performs contextual aggregation and deformable spatial alignment, thereby improving the consistency of shallow localization features and deep semantic features. Third, Focaler-MPDIoU integrates coordinate-sensitive regression with IoU-quality-based sample reweighting for dense small-object localization. Experiments on the VisDrone-2019 test set and the UAVDT and HIT-UAV validation sets demonstrate mAP50 improvements of 3.1, 1.4, and 3.0 percentage points, respectively, over the RT-DETR-R18 baseline. On the VisDrone-2019 validation set, CDF-DETR improves mAP5095 from 26.20% to 28.52%, corresponding to a gain of 2.32 percentage points, while reducing the parameter count by 25.7%. A compressed INT8 variant achieves 20.84 FPS for an offline image-level pipeline on an NVIDIA Jetson Orin Nano using ONNX and TensorRT. These results demonstrate improved detection accuracy with a reduced parameter footprint for UAV remote sensing image analysis. Full article
(This article belongs to the Section AI Remote Sensing)
Show Figures

Graphical abstract

25 pages, 10657 KB  
Article
MLP-LSTM-Attention Algorithm for DAS Cable Intrusion Detection Based on Multi-Domain Feature Fusion
by Li Yuan, Jun Xing, Bowen Shen, Yuancheng Du, Wenchi Wei and Xicheng Rao
Photonics 2026, 13(8), 768; https://doi.org/10.3390/photonics13080768 - 14 Aug 2026
Viewed by 155
Abstract
Underground cables are critical infrastructure for electrical power and communication transmission, and their reliable operation is of paramount importance to urban public safety. Although Distributed Acoustic Sensing (DAS) enables wide-range, continuous, and real-time monitoring, traditional DAS signal processing methods suffer from poor intrusion [...] Read more.
Underground cables are critical infrastructure for electrical power and communication transmission, and their reliable operation is of paramount importance to urban public safety. Although Distributed Acoustic Sensing (DAS) enables wide-range, continuous, and real-time monitoring, traditional DAS signal processing methods suffer from poor intrusion discrimination and weak anti-interference capability. To address these limitations, we propose a dual-branch network based on multi-domain feature fusion, integrating a Multilayer Perceptron, a Long Short-Term Memory network (LSTM), and an attention mechanism. Vibration signals corresponding to four representative high-risk intrusion events were acquired through controlled field experiments, and a standardized, category-balanced dataset was constructed accordingly. Time-domain, frequency-domain and joint time-frequency features were extracted and mapped through a time-frequency weighting transformation to form one branch of the network, while the parallel branch employed an LSTM to capture long-range temporal dependencies. A multi-head attention mechanism enables deep adaptive fusion of two types of modal information and overcomes the limitations of conventional simple feature concatenation. Comparative experiments against KNN, 1D-CNN and LSTM baselines demonstrate that the proposed model achieves a test accuracy of 98.89%, outperforming all reference methods. Ablation studies further validate the necessity and effectiveness of each constituent module within the proposed architecture. The results indicate that this approach provides reliable support for DAS-based online monitoring of power cables against external damage. Full article
(This article belongs to the Special Issue Recent Advances in Infrared Lasers and Applications)
Show Figures

Figure 1

35 pages, 42203 KB  
Article
Wind Direction Retrieval from X-Band Marine Radar Images Using 2D-DTCWT–CSC and Maximum-Energy Radial Rings
by Jie Xiao, Hui Wang, Zhizhong Lu, Baotian Wen and Yanbo Wei
Remote Sens. 2026, 18(16), 2728; https://doi.org/10.3390/rs18162728 - 13 Aug 2026
Viewed by 264
Abstract
Under moderate-to-high wind conditions, low-frequency wind direction modulation signals in X-band marine radar images are strongly coupled with wave textures, sea clutter, and blind-zone interference, which degrades wind direction retrieval accuracy. To address this problem, this study proposes a wind direction retrieval method [...] Read more.
Under moderate-to-high wind conditions, low-frequency wind direction modulation signals in X-band marine radar images are strongly coupled with wave textures, sea clutter, and blind-zone interference, which degrades wind direction retrieval accuracy. To address this problem, this study proposes a wind direction retrieval method based on two-dimensional dual-tree complex wavelet transform (2D-DTCWT), convolutional sparse coding (CSC), and maximum-energy radial rings. First, 2D-DTCWT is used to suppress wave textures and local noise in the wavelet domain while enhancing low-frequency wind direction modulation signals. Then, K–singular value decomposition (K-SVD) learns the energy distribution characteristics of wind signals, and CSC obtains the spatial response distribution of wind energy in radar images. Finally, the maximum-energy radial ring is adaptively identified, and azimuthal energy statistics within this ring are fitted using a cosine-squared function. The proposed method was evaluated using X-band marine radar data collected during sea trials in the coastal waters of Zhejiang, China. On the 900-sample main validation dataset, the proposed method achieved the highest correlation coefficient (CC) of 0.85 and an overall root mean square error (RMSE) of 4.24°, reducing the RMSE by 43.0% and 66.7% compared with conventional single-curve fitting and extended-bow-heading DWT, respectively. The results demonstrate improved robustness under both upwind and downwind blind-zone conditions. Full article
(This article belongs to the Special Issue Feature Paper Special Issue on Ocean Remote Sensing (Third Edition))
Show Figures

Graphical abstract

21 pages, 947 KB  
Article
A Stock Market Price Prediction Model Integrating a CNN–Transformer Dual-Channel Dynamic Attention Architecture
by Chengcheng Han, Jingwei Guo and Xingyu Feng
Mathematics 2026, 14(16), 2888; https://doi.org/10.3390/math14162888 - 10 Aug 2026
Viewed by 310
Abstract
Stock market price prediction remains a persistent challenge owing to the non-stationarity, high noise content, and intricate spatiotemporal dependencies that characterize financial time series. Existing approaches typically excel at either local pattern extraction or long-range dependency modeling, yet seldom reconcile both within a [...] Read more.
Stock market price prediction remains a persistent challenge owing to the non-stationarity, high noise content, and intricate spatiotemporal dependencies that characterize financial time series. Existing approaches typically excel at either local pattern extraction or long-range dependency modeling, yet seldom reconcile both within a unified framework. This paper introduces a CNN–Transformer dual-channel architecture equipped with a dynamic attention fusion module for stock price forecasting. The convolutional channel applies hierarchical dilated convolutions to distill fine-grained local patterns from multi-indicator sequences while suppressing high-frequency noise. Simultaneously, the Transformer channel employs multi-head self-attention to capture long-distance temporal correlations and regime-shift dynamics. A learnable gating mechanism then fuses the two feature streams by adaptively weighting local detail against global trend information according to market conditions. Experiments conducted on four real-world stock datasets spanning the S&P 500, CSI 300, NASDAQ Composite, and Hang Seng Index show that the proposed model reduces mean absolute error by 9.7–15.3% and root mean square error by 9.5–13.8% relative to competitive baselines including LSTM, CNN–LSTM, Informer, and PatchTST. Ablation studies further indicate that both channels and the fusion module contribute to prediction accuracy, and the architecture remains effective across markets with differing volatility profiles. Full article
Show Figures

Figure 1

33 pages, 10301 KB  
Article
An Explainable Multi-Task Deep Learning Framework for Service-Gap Identification and Emerging Urban Prediction
by Abdulelah Algosaibi
Electronics 2026, 15(15), 3470; https://doi.org/10.3390/electronics15153470 - 6 Aug 2026
Viewed by 249
Abstract
Rapid urbanization has intensified pressure on public services, infrastructure systems, and spatial equity, highlighting the need for integrated approaches that jointly assess service deficits and emerging urban growth. This study proposes an explainable urban analytics framework for identifying service-gap risk and emerging urban [...] Read more.
Rapid urbanization has intensified pressure on public services, infrastructure systems, and spatial equity, highlighting the need for integrated approaches that jointly assess service deficits and emerging urban growth. This study proposes an explainable urban analytics framework for identifying service-gap risk and emerging urban patterns using harmonized spatial, service, population, and digital-readiness indicators. The framework integrates an MCP-enabled data harmonization pipeline, composite service-availability and population-adjusted service-stress features, and a Dual-Head MLP architecture that supports shared representation learning across two related prediction tasks. SHAP-based explainability is employed to interpret the relative contribution of service, stress, digital-readiness, and spatial-context features. The framework was evaluated using Saudi district-level data and proxy-based external city datasets to assess cross-city transferability. Compared with classical machine-learning and single-task neural baselines, the Dual-Head MLP demonstrated stronger task-wise predictive performance, while the external evaluation indicated stable but context-dependent transfer across heterogeneous urban settings. The findings suggest that the proposed framework can support urban planners and municipal decision-makers in prioritizing infrastructure investment, identifying underresourced areas, and interpreting early urban transformation patterns. However, the outputs should be regarded as proxy-based decision-support indicators rather than official administrative measures of service adequacy. Full article
Show Figures

Figure 1

16 pages, 1328 KB  
Article
DR-Transformer: A Dual-Regularized Transformer Combining Sparse Attention and Supervised Contrastive Learning for Interpretable Stress Detection in Social Media Text
by Mehdi Chrifi Alaoui, Nour-Eddine Joudar and Mohamed Ettaouil
AI 2026, 7(8), 300; https://doi.org/10.3390/ai7080300 - 4 Aug 2026
Viewed by 439
Abstract
Automatic detection of stress in social media text holds promise for supporting digital mental health, but most existing Transformer-based approaches are opaque and computationally demanding. This work presents DR-Transformer, a Dual-Regularized Transformer that combines two complementary mechanisms: (i) a group sparsity penalty ( [...] Read more.
Automatic detection of stress in social media text holds promise for supporting digital mental health, but most existing Transformer-based approaches are opaque and computationally demanding. This work presents DR-Transformer, a Dual-Regularized Transformer that combines two complementary mechanisms: (i) a group sparsity penalty (L2,1/L2 elastic net) applied to the query and key projection matrices of every attention head, which encourages whole-row sparsity, producing more concentrated and inspectable attention patterns; (ii) a supervised contrastive loss on the [CLS] projection, which organizes the latent space according to the stress label. The architecture is intentionally lightweight (six layers, eight heads, 256-dim embeddings; ∼9.5 M parameters) and runs entirely on consumer-grade hardware (NVIDIA GTX 1660, 6 GB). Experiments on the publicly available Dreaddit dataset (binary stress classification, 2838 train/715 test segments) compare DR-Transformer against Logistic Regression, BiLSTM, a Standard Transformer of identical architecture, and MentalBERT. Across five seeded runs, DR-Transformer (Full) reaches F1=0.876 (bootstrap 95% CI 0.8520.898), outperforming the Standard Transformer (F1=0.842; McNemar p<0.001 with Bonferroni correction) and performing comparably to the much larger MentalBERT (F1=0.879; p=0.421). Sparse regularization increases the fraction of near-zero attention weights (below 0.01) from 0.215 to 0.682, while the supervised contrastive loss improves the silhouette score of [CLS] embeddings from 0.312 to 0.483. Dual regularization thus combines accuracy, efficiency, and structurally induced attention concentration in a single model which can be trained without specialized infrastructure. We use the term “interpretable” throughout in this restricted, structural sense—to refer to concentrated and inspectable attention—rather than in the sense of established causal or mechanistic faithfulness; this is only partially and indirectly supported by our token deletion analysis. Full article
Show Figures

Figure 1

31 pages, 22136 KB  
Article
Swarm Intelligence-Guided Hybrid Transfer Learning for Gastrointestinal Polyp Classification
by Una Tuba, Mladen Veinovic, Eva Tuba, Adis Alihodzic and Milan Tuba
Biomimetics 2026, 11(8), 541; https://doi.org/10.3390/biomimetics11080541 - 3 Aug 2026
Viewed by 276
Abstract
Colorectal cancer remains a leading cause of cancer-related mortality worldwide, with automated polyp classification from endoscopic images offering a promising avenue for improving early detection. Existing approaches rely on single convolutional neural network (CNN) backbones with manually designed classification heads, limiting both representational [...] Read more.
Colorectal cancer remains a leading cause of cancer-related mortality worldwide, with automated polyp classification from endoscopic images offering a promising avenue for improving early detection. Existing approaches rely on single convolutional neural network (CNN) backbones with manually designed classification heads, limiting both representational capacity and deployment flexibility. This paper presents a swarm intelligence-augmented multi-backbone deep learning framework for eight-class gastrointestinal lesion classification on the Kvasir benchmark. Four CNN backbones (ResNet50, DenseNet121, MobileNetV2, EfficientNetB3) are independently fine-tuned using a two-phase transfer learning protocol and their penultimate-layer features concatenated into a 5888-dimensional representation, reduced to 256 dimensions via PCA. Five swarm intelligence algorithms—Particle Swarm Optimization, Artificial Bee Colony, JADE, L-SHADE, and CMA-ES—are benchmarked on the classification head architecture search task; all independently converge to tanh activation, a consistent pattern across independently initialized algorithms that is suggestive of, though not conclusive evidence for, particular geometric properties of PCA-transformed deep feature spaces. The PSO-optimized single-layer head (284 units, tanh) outperforms a manually designed three-layer baseline by 0.75% while using 67% fewer parameters. SI-guided class weight optimization yields targeted F1 improvements on the two most clinically significant classes (polyps: +0.015, ulcerative-colitis: +0.013). The fixed-head classifier trained on fused four-backbone features achieves 91.08% accuracy on Kvasir v2 (multi-seed mean 91.47% ± 0.49 across nine converging seeds; one seed failed to converge and is disclosed rather than excluded), below end-to-end DenseNet121 (92.25%; Wilcoxon p = 0.31, not statistically significant), while enabling classifier updates in under 30 s; a three-backbone subset dropping the weakest backbone (EfficientNetB3) reaches 92.33%, exceeding the full four-backbone fusion. Cross-dataset evaluation on Kvasir v1-to-v2 confirms near-zero generalization gaps across dataset scales; a restricted two-class evaluation on HyperKvasir (the only two of eight classes with usable labeled data) reaches 96.28% accuracy, and dual Grad-CAM with SI minimal sufficient region analysis, validated quantitatively against Kvasir-SEG ground-truth masks, provides spatially grounded, clinically interpretable explanations. Full article
Show Figures

Figure 1

32 pages, 1113 KB  
Article
Hyperspectral Image Classification Based on a Spatial–Spectral Dual-Branch Mamba Architecture
by Jialing Li, Shangbo Zhou, Yawen Liu, Guiwen Hu and Xiaojuan Liu
Remote Sens. 2026, 18(15), 2526; https://doi.org/10.3390/rs18152526 - 2 Aug 2026
Viewed by 258
Abstract
Hyperspectral image classification is a core task in remote sensing image analysis and understanding. Existing Transformer-based methods have achieved excellent performance but are limited by the quadratic computational complexity of the self-attention mechanism, while the high-dimensional redundancy of hyperspectral data and the difficulty [...] Read more.
Hyperspectral image classification is a core task in remote sensing image analysis and understanding. Existing Transformer-based methods have achieved excellent performance but are limited by the quadratic computational complexity of the self-attention mechanism, while the high-dimensional redundancy of hyperspectral data and the difficulty in deeply integrating spatial–spectral features also restrict further performance improvement. To address these issues, we introduce the Mamba architecture based on state-space models into hyperspectral image classification and propose the DFMamba model. The main innovations include (1) constructing a Hyperspectral Spatial Attention Embed (HSAE) to achieve efficient channel compression and feature extraction via adaptive grouped convolution, depth-wise separable convolution, and spatial attention; (2) proposing a spatial–spectral dual-branch collaborative modeling mechanism, EnhancedBothMamba, which separately models global dependencies in the spatial and spectral branches and integrates their outputs through softmax-normalized learnable global weights together with a learnable residual scaling factor; and (3) building an improved classification head, ClsHead, with a multi-scale branch fusion strategy to fully exploit local and global feature information. The experimental results on four standard hyperspectral datasets demonstrate that DFMamba achieves overall accuracy (OA) of 97.41% on the Pavia University dataset, 92.25% on the HanChuan dataset, 95.12% on the HongHu dataset, and 94.98% on the Houston dataset. Under the adopted evaluation protocol, DFMamba obtains higher mean OA than MambaHSI and the other compared methods while retaining favorable computational efficiency. Full article
(This article belongs to the Section Remote Sensing Image Processing)
Show Figures

Figure 1

32 pages, 4667 KB  
Article
Reinforcement Learning-Based Soft–Hard Damage Cooperative Task Allocation and Optimization Method for Multi-Laser Systems Against UAV Swarms
by Jingyi Zhang, Lin Zhang, Bo Zhang, Wenfeng Wang, Wei Liu, Lan Yao, Lin Cui and Mingang Zhang
Aerospace 2026, 13(8), 692; https://doi.org/10.3390/aerospace13080692 - 30 Jul 2026
Viewed by 242
Abstract
To address the dynamic target allocation and resource scheduling problems of multiple High-Energy Laser Systems (HELSs) in defending critical infrastructure against large-scale heterogeneous UAV swarm penetration, this paper proposes a soft–hard damage cooperative task allocation and optimization method, termed Attention-based Centralized Training and [...] Read more.
To address the dynamic target allocation and resource scheduling problems of multiple High-Energy Laser Systems (HELSs) in defending critical infrastructure against large-scale heterogeneous UAV swarm penetration, this paper proposes a soft–hard damage cooperative task allocation and optimization method, termed Attention-based Centralized Training and Decentralized Execution proximal policy optimization (Attn-CTDE-PPO), which integrates an attention mechanism with multi-agent reinforcement learning. First, a cooperative multi-agent sequential decision-making model for multi-HELS engagement against UAV swarms is constructed by considering continuous laser irradiation, system energy consumption, thermal accumulation limits, and soft–hard damage mechanisms. Second, a multi-head attention set encoder based on a Transformer is introduced to extract global situational features at the state level. This design enables the policy network to handle a time-varying number of targets and mitigates the explosion of hybrid action spaces. At the action level, a discrete–continuous dual-head network with a masking mechanism is proposed to achieve target allocation and laser-parameter scheduling. Furthermore, a multi-dimensional threat assessment module is developed, which can prune infeasible actions in the action space via physical rule-based masking, thereby accelerating the decision-making speed of the method. Monte Carlo simulation experiments verified the defense reliability, energy management efficiency and real-time responsiveness of the proposed method. The experimental results demonstrate that, in heterogeneous swarm scenarios with different numbers of UAVs, the proposed method effectively suppresses UAV penetration, improves the defense success rate, and reduces total system energy consumption compared with baseline algorithms. Full article
(This article belongs to the Section Aeronautics)
Show Figures

Figure 1

43 pages, 8097 KB  
Article
Toward Reliable Diabetic Retinopathy Screening
by Hendrio Bragança, Ítalo P. Caliari, Wington L. Vital, Antonio Fontenele, Sergio Cavalcante and Glaucio Messias
Sensors 2026, 26(14), 4515; https://doi.org/10.3390/s26144515 - 16 Jul 2026
Viewed by 549
Abstract
Diabetic retinopathy (DR) grading requires reliable five-grade severity assessment under substantial acquisition variability and cross-dataset distribution shift. We propose PRISM-DR, a multi-objective five-grade DR grading framework trained under a gradient-partitioned strategy. The architecture is organized as a feedforward pipeline: a data-driven preprocessing stage [...] Read more.
Diabetic retinopathy (DR) grading requires reliable five-grade severity assessment under substantial acquisition variability and cross-dataset distribution shift. We propose PRISM-DR, a multi-objective five-grade DR grading framework trained under a gradient-partitioned strategy. The architecture is organized as a feedforward pipeline: a data-driven preprocessing stage followed by a ConvNeXtV2-Base backbone, a Recurrent BiFPN neck for multi-scale feature fusion, a Frequency-Aware Fusion module, a lightweight multi-scale reasoning transformer, dual classification heads with gradient-isolated pathways (categorical and ordinal), and a prototype memory module for embedding regularization. The CORAL ordinal head operates through a dedicated projection layer and is gradient-isolated from the backbone; the backbone is shaped by the cross-entropy, prototype contrastive, and view-consistency objectives, which carry indirect ordinal signal through severity-weighted class penalties and grade-indexed cluster regularization. The model is trained in a multi-crop setting with a phased loss curriculum designed for severely imbalanced DR datasets. Evaluated across six datasets under Fixed-Source, Multi-Target (FSMT) protocols, PRISM-DR trained on EyePACS + DDR achieves QWK of 0.835 on IDRiD, 0.865 on APTOS2019, and 0.720 on Messidor-2, with in-domain QWK = 0.920 and AUC-PR = 0.941 on EyePACS, outperforming RETFound, RETFound-Green, and MedGemma-4B in AUC-PR across all evaluated datasets. Quantitative interpretability evaluation against 755 expert-annotated lesion images yields 8.0× Energy Ratio Enrichment and a FAF gate retention ratio of 4.4× inside lesion regions, confirming that anatomically plausible spatial priors emerge from grade-level supervision alone, without pixel-level annotation. PRISM-DR establishes a superior accuracy–robustness–capacity trade-off for scalable, automated DR screening. Full article
(This article belongs to the Section Biomedical Sensors)
Show Figures

Figure 1

23 pages, 2748 KB  
Article
LDS-Net: A Lightweight Dual-Branch Network for Slender Tree Branch Segmentation in Complex Natural Scenes
by Xinyan Zhang, Tianlong Deng, Yin Wu, Wenjie Wu and Yanyi Liu
Forests 2026, 17(7), 811; https://doi.org/10.3390/f17070811 - 10 Jul 2026
Viewed by 306
Abstract
Accurate semantic segmentation of slender curvilinear structures, such as tree branches, in complex natural scenes remains challenging. The main difficulties arise from frequent occlusions, ambiguous boundaries, and limited edge computing resources. To address these issues, we propose LDS-Net, a lightweight dual-branch network designed [...] Read more.
Accurate semantic segmentation of slender curvilinear structures, such as tree branches, in complex natural scenes remains challenging. The main difficulties arise from frequent occlusions, ambiguous boundaries, and limited edge computing resources. To address these issues, we propose LDS-Net, a lightweight dual-branch network designed for thin and continuous branch structures. The Detail-Aware Branch uses the Dynamic Snake Convolution (DSConv) to model irregular local geometry, while the Context-Aware Branch uses a Spatial Efficient Separable Pyramid module (SESP) to capture multi-scale context. To improve segmentation under occlusion and boundary ambiguity, we further integrate a Global Topology Transformer module (GTT) and a boundary guidance mechanism (BG). These features are fused via a Pixel-Wise Attention Fusion module (PAF) and optimized using a multi-head compound loss. Experiments on USTD and N-ABSD show that LDS-Net outperforms six representative networks. It also requires only 14.19 G FLOPs, a 24% reduction compared with PIDNet-s. These results suggest that LDS-Net has potential for future deployment on resource-constrained agricultural and ecological monitoring platforms. Full article
Show Figures

Figure 1

21 pages, 6420 KB  
Article
Attention-Driven CNNs as a Strong Default for HER2 Prediction from DCE-MRI: A Comparison with Transformer Architectures
by Naomi Fridman and Anat Goldstein
Bioengineering 2026, 13(7), 788; https://doi.org/10.3390/bioengineering13070788 - 8 Jul 2026
Viewed by 553
Abstract
Background: HER2 status guides targeted therapy in breast cancer but is currently determined by invasive biopsy. Imaging-based HER2 prediction from dynamic contrast-enhanced MRI (DCE-MRI) could provide a non-invasive adjunct decision-support signal, but published models are typically single-center with heterogeneous preprocessing that limits reproducibility. [...] Read more.
Background: HER2 status guides targeted therapy in breast cancer but is currently determined by invasive biopsy. Imaging-based HER2 prediction from dynamic contrast-enhanced MRI (DCE-MRI) could provide a non-invasive adjunct decision-support signal, but published models are typically single-center with heterogeneous preprocessing that limits reproducibility. Methods: We trained a Triple-Head Dual-Attention ResNet (THDA-ResNet) that processes three DCE phases (pre-contrast, early post-contrast, and late post-contrast) on the multicenter BreastDCEDL dataset (n = 1149, I-SPY trials), and we compared it with Vision Transformer (ViT) and Convolutional Vision Transformer (CvT) baselines, all ImageNet-pretrained. We benchmarked 14 preprocessing strategies, with and without N4 bias-field correction. External validation used the independent BreastDCEDL_AMBL cohort (43 lesions). AUC confidence intervals used stratified bootstrap; model comparisons used DeLong’s test. Results: THDA-ResNet achieved the highest AUC, 0.74 (95% CI 0.65–0.83), versus 0.66 for ViT and 0.63 for CvT, with the advantage reaching borderline significance over CvT (p=0.054) and not significant over ViT (p=0.14). At a threshold of 0.7, it retained discrimination (sensitivity 0.41, specificity 0.86), while transformers collapsed to near-trivial classifiers. External AUC was 0.66 (0.49–0.81). N4 correction did not improve performance. Conclusions: Attention-driven CNNs are a strong default for HER2 prediction from DCE-MRI on medium-sized cohorts, and N4 correction can be omitted, simplifying the pipeline. Full article
(This article belongs to the Special Issue AI-Driven Imaging and Analysis for Biomedical Applications)
Show Figures

Figure 1

33 pages, 14316 KB  
Article
IMU-Sequence-Based GNSS Short Outage Compensation and Hybrid Positioning Strategy
by Ziyong Lei, Luyao Du and Zelong Lian
Mathematics 2026, 14(13), 2423; https://doi.org/10.3390/math14132423 - 6 Jul 2026
Viewed by 932
Abstract
Pure inertial dead reckoning during short GNSS outages causes rapid drift on low-cost MEMS GNSS/IMU platforms. Most learning-based compensators upgrade a single predictor and rarely address late-outage drift or cross-domain bias mismatch. This paper proposes two enhancements over a Transformer baseline (TF-Base) plus [...] Read more.
Pure inertial dead reckoning during short GNSS outages causes rapid drift on low-cost MEMS GNSS/IMU platforms. Most learning-based compensators upgrade a single predictor and rarely address late-outage drift or cross-domain bias mismatch. This paper proposes two enhancements over a Transformer baseline (TF-Base) plus a lightweight inference-time fusion strategy. MGTR (Motion-Guided Transformer with Tail-aware Readout) adds residual motion gating and a tail-aware readout for hard-segment and late-outage response. TAMS (Temporal Attention Multi-Scale) replaces global average pooling with learnable temporal attention and a short-window dual head. Delayed-Switch selects among TF-Base, MGTR, and TAMS without retraining backbones; its classifier needs a one-pass target-domain calibration, so it is not zero-shot. On real 5 Hz GNSS/IMU recordings under a three-tier protocol, where dead reckoning yields a 40.02 m mean RMSE on cross-domain segments, MGTR cuts the 90th-percentile 2D-RMSE by 20.3% over TF-Base, and Delayed-Switch reaches 30.32 m mean RMSE (24.2% below dead reckoning, 9.5% below TF-Base), within 0.51 m of the better-of-two upper bound. Against two recent baselines under the same protocol, only the AT-LSTM gain is significant after multiple-comparison correction; the margins over the strongest predictors are numerically favorable but not significant at this sample size, with gains concentrated on a few hard segments. Full article
Show Figures

Figure 1

35 pages, 12331 KB  
Article
A Physics-Aware Dual-Branch CNN-MLP Fusion Framework for Stage-Aware Bearing Degradation Monitoring and RUL Prognosis from Vibration Signals
by Bowen Dong, Xinyu Zhang, Yifan Feng, Weiyan Zhu, Chaoya Yan and Lingmin Hou
Electronics 2026, 15(13), 2910; https://doi.org/10.3390/electronics15132910 - 2 Jul 2026
Viewed by 337
Abstract
Rolling element bearing degradation monitoring is critical for predictive maintenance in rotating machinery systems. Existing methods predominantly address fault classification and remaining useful life (RUL) estimation as separate tasks, thereby failing to capture the progressive and multistage nature of bearing deterioration. This paper [...] Read more.
Rolling element bearing degradation monitoring is critical for predictive maintenance in rotating machinery systems. Existing methods predominantly address fault classification and remaining useful life (RUL) estimation as separate tasks, thereby failing to capture the progressive and multistage nature of bearing deterioration. This paper proposes a physics-aware multi-modal fusion framework for continuous RUL prediction from vibration signals, organized around a stage-aware representation of the bearing life cycle. The proposed pipeline integrates two complementary preprocessing branches: Hilbert envelope demodulation followed by short-time Fourier transform (STFT) to generate degradation-sensitive time–frequency spectrograms, and handcrafted statistical feature extraction to yield compact global severity descriptors. A dual-head convolutional neural network-multilayer perceptron (CNN-MLP) architecture is designed to learn discriminative representations from both modalities and fuse them for end-to-end normalized RUL regression. The bearing life cycle is further partitioned into four ordered degradation stages based on normalized life–progress ratios, providing an interpretable health representation that complements the continuous prognosis target. Experiments conducted on the PRONOSTIA/FEMTO-ST benchmark dataset demonstrate that the proposed framework achieves an RMSE of 0.1597, an MAE of 0.1328, and an R2 of 0.7487 on normalized RUL prediction, with stable error behavior across most of the life cycle. Feature importance analysis confirms that the CNN branch captures localized low-to-mid-frequency spectral evolution while the MLP branch encodes amplitude variability and impulsive indicators, validating the complementarity of the dual-branch design. The proposed method offers a unified, interpretable, and engineering-relevant solution for intelligent bearing condition monitoring and prognostic health management. Full article
Show Figures

Figure 1

Back to TopTop