Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (3,137)

Search Parameters:
Keywords = global–local attention

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
29 pages, 6755 KB  
Article
Research on Intelligent Diagnosis of DC Magnetic Bias of Power Transformers Based on Vibration Signals and Improved 2DWT-CNN-Transformer Framework
by Huida Duan, Zhipeng Gao, Song Bai, Yihan Wang, Shihao Zhao and Ying Zhao
Electronics 2026, 15(17), 3789; https://doi.org/10.3390/electronics15173789 - 24 Aug 2026
Abstract
DC bias will cause the magnetization working point of the transformer core to shift and cause local saturation, and generate abnormal vibration through the magnetostrictive effect, which threatens the safe operation of the transformer. Aiming at the problem that the time–frequency characteristics of [...] Read more.
DC bias will cause the magnetization working point of the transformer core to shift and cause local saturation, and generate abnormal vibration through the magnetostrictive effect, which threatens the safe operation of the transformer. Aiming at the problem that the time–frequency characteristics of transformer vibration signals under DC bias are complex and the adjacent bias levels are difficult to distinguish, this paper proposes a 2DWT-CNN-Transformer diagnostic method that combines two-dimensional discrete wavelet transform, a convolutional neural network, and Transformer Encoder. Firstly, the multi-physical-field finite element model of three-phase three-column transformer is established, and the L0–L5 six-class DC bias dataset is constructed. Secondly, the one-dimensional vibration signal is reconstructed into a two-dimensional matrix, and the multi-subband time–frequency features of LL, LH, HL, and HH are extracted by two-dimensional discrete wavelet transform. The local texture features are extracted by the CNN, and the multi-head self-attention mechanism of Transformer Encoder is introduced to establish the global dependence and enhance the discrimination ability of adjacent bias levels. Compared with the traditional time–frequency-feature deep learning model, the proposed method achieves higher accuracy, especially in the high-noise environment of 15 dB, where it can still maintain accuracy of 96.23%. The visualization results further show that the model can form a more compact intra-class aggregation and a clearer inter-class boundary. This also provides an effective solution for the identification and evaluation of transformer DC bias states based on vibration signals in complex environments in the future. Full article
Show Figures

Figure 1

21 pages, 2641 KB  
Article
CA-MC-Transformer: An Operating Condition-Adaptive and Multi-Scale Convolution-Enhanced Transformer Architecture for Furnace Temperature Prediction
by Jiayang Dai, Zhen Chen, Shenwang Li and Thomas Wu
Electronics 2026, 15(17), 3784; https://doi.org/10.3390/electronics15173784 - 24 Aug 2026
Abstract
Regenerative aluminum melting serves as a core process in recycled aluminum production. In the regenerative aluminum melting process, the furnace temperature is a key variable which affects product performance and energy costs. The extreme in-furnace temperature necessitates sensors equipped with protective jackets, which [...] Read more.
Regenerative aluminum melting serves as a core process in recycled aluminum production. In the regenerative aluminum melting process, the furnace temperature is a key variable which affects product performance and energy costs. The extreme in-furnace temperature necessitates sensors equipped with protective jackets, which increases measurement costs and severely compromises real-time monitoring capability. Accordingly, accurate furnace temperature prediction is highly valuable for regenerative aluminum melting. In regenerative aluminum melting furnaces, periodic burner nozzle commutation and frequent material charging and discharging lead to complex and time-varying operating conditions, posing considerable challenges to high-precision furnace temperature prediction. To address these issues, a condition-adaptive multi-scale convolution-enhanced Transformer (CA-MC-Transformer) model is proposed for furnace temperature prediction. Firstly, an agglomerative hierarchical clustering algorithm based on the weighted dynamic time warping (WDTW) distance is designed to perform unsupervised clustering on historical process data, thereby extracting physically interpretable prior labels for macroscopic operating conditions. Secondly, multi-scale dilated causal convolutions are utilized to capture local dynamic features at diverse temporal resolutions. A soft attention mechanism is further introduced to dynamically assign fusion weights to condition embeddings and local features, enabling condition-adaptive feature reconstruction. Finally, the fused adaptive features are fed into an encoder-only Transformer network to capture the global long-range temporal dependencies and achieve accurate furnace temperature prediction. Comparative experiments conducted on real operational datasets from an aluminum plant verify that the proposed method effectively eliminates the inherent tracking lag of conventional deep learning models, and substantially improves prediction accuracy and anti-noise robustness under complex and variable operating conditions. Full article
(This article belongs to the Special Issue AI Driven Digital Twinning: A Trend Challenging the Future)
Show Figures

Figure 1

24 pages, 19539 KB  
Article
Early Prediction of Lithium-Ion Battery Remaining Useful Life Using a GWO-Optimized CNN–Transformer–BiGRU Network
by Chongyang Wei, Xinfu Pang, Jingran Sheng, Hongxia Yu, Zedong Zheng and Pengwei Yu
Batteries 2026, 12(9), 320; https://doi.org/10.3390/batteries12090320 - 24 Aug 2026
Abstract
Lithium-ion batteries are widely used in various energy sectors, and accurately predicting their early remaining useful life (RUL) is crucial for shortening battery evaluation time and accelerating battery commercialization. However, information on degradation during the early cycling stages of batteries is limited, and [...] Read more.
Lithium-ion batteries are widely used in various energy sectors, and accurately predicting their early remaining useful life (RUL) is crucial for shortening battery evaluation time and accelerating battery commercialization. However, information on degradation during the early cycling stages of batteries is limited, and it is difficult to fully characterize their lifespan. This study proposes a CNN–Transformer–BiGRU-based method for predicting the early RUL of lithium-ion batteries using Grey Wolf Optimization (GWO). First, using only the first 100 cycles of each battery in the MIT dataset, early degradation features are extracted from the dimensions of capacity and internal resistance, and then standardized. Second, a CNN is employed to extract local degradation features, while the Transformer’s self-attention mechanism is used to capture global correlations, and BiGRU is utilized to further extract bidirectional temporal dependency information. Building on this foundation, GWO is introduced to perform joint optimization of the model’s key hyperparameters to obtain optimal network parameters. Finally, the effectiveness of the proposed method is validated through ablation and comparison experiments. The experimental results show that the proposed model achieved an R2 of 0.9633, with RMSE, MAE, and MAPE values of 80.5608 cycles, 63.2524 cycles, and 7.29%, respectively, demonstrating overall prediction performance superior to that of the comparison models. This method can effectively mine degradation information related to battery life from limited early-cycle data, providing an effective approach for the accurate prediction of the early RUL of lithium-ion batteries. Full article
(This article belongs to the Section Lithium-Ion and Solid-State Batteries)
Show Figures

Figure 1

19 pages, 6270 KB  
Article
Semi-Supervised Acoustic Impedance Inversion Based on a Hybrid Deep Learning Network
by Yan Huang, Xiangfei Nie, Wei Huang, Gang Fang, Weiwei Li and Wenliang Nie
Appl. Sci. 2026, 16(17), 8401; https://doi.org/10.3390/app16178401 - 24 Aug 2026
Abstract
Accurate estimation of subsurface acoustic impedance is fundamental to quantitative reservoir characterization in seismic exploration. Nevertheless, a single network architecture cannot adequately represent both the local details and the global trends of seismic records within a unified framework, while the severe scarcity of [...] Read more.
Accurate estimation of subsurface acoustic impedance is fundamental to quantitative reservoir characterization in seismic exploration. Nevertheless, a single network architecture cannot adequately represent both the local details and the global trends of seismic records within a unified framework, while the severe scarcity of annotated well-log data substantially constrains the generalization capability and predictive accuracy of deep-learning-based inversion approaches. To overcome these limitations, a semi-supervised acoustic impedance inversion framework based on a hybrid deep learning architecture is proposed. The framework employs a cascaded architecture consisting of a multi-scale depthwise separable convolution with channel attention (MSDSE) module and a convolution-augmented Transformer encoder. Seismic data are first processed by the MSDSE module to extract local multi-scale temporal features, and are subsequently passed to the convolution-augmented Transformer encoder, which captures global long-range sequence dependencies while retaining complementary local temporal information. The two modules progress hierarchically and jointly achieve a feature representation that spans from local details to global trends, and the initial low-frequency model is fused with the network output via channel-wise concatenation. Meanwhile, an initial-model constraint together with a physical-consistency constraint are simultaneously imposed within the loss function, thereby improving training stability while fully leveraging the physical information embedded in unlabeled traces. Experiments on both synthetic and field data confirm the effectiveness of the proposed method. The results show that, even with a small number of labels, the method produces stable impedance estimates and outperforms conventional deep learning methods in both generalization and prediction accuracy. Full article
(This article belongs to the Section Earth Sciences)
Show Figures

Figure 1

24 pages, 386 KB  
Review
Carbapenem-Resistant Klebsiella pneumoniae in Healthcare-Associated Infections: Global and Regional Epidemiology, Resistance Mechanisms, and Therapeutic Strategies, with Particular Attention to Romania and Eastern Europe (2020–2025)
by Oana-Elena Ioniţă, Roxana-Carmen Cernat, Nicola-Maria Militaru, Maria-Elena Vodarici, Maria Fulina, Daniela Pițigoi, Elena Mocanu, Beatrice Severin, Claudia-Simona Cambrea and Irina-Magdalena Dumitru
Microorganisms 2026, 14(9), 1870; https://doi.org/10.3390/microorganisms14091870 - 23 Aug 2026
Abstract
Healthcare-associated infections (HAIs) caused by multidrug-resistant Klebsiella spp. represent a critical and escalating global public health threat. Carbapenem-resistant Klebsiella pneumoniae (CRKP) has been designated a critical-priority pathogen by the World Health Organization, and in the 2024 WHO Bacterial Priority Pathogens List, it was [...] Read more.
Healthcare-associated infections (HAIs) caused by multidrug-resistant Klebsiella spp. represent a critical and escalating global public health threat. Carbapenem-resistant Klebsiella pneumoniae (CRKP) has been designated a critical-priority pathogen by the World Health Organization, and in the 2024 WHO Bacterial Priority Pathogens List, it was the top-ranked pathogen overall. The convergence of carbapenem resistance with hypervirulence in emerging strains has further complicated therapeutic decision-making. This review provides a narrative synthesis of the evidence published between 2020 and 2025 on the prevalence, resistance mechanisms, molecular epidemiology, clinical outcomes, and therapeutic strategies for Klebsiella pneumoniae infections acquired in healthcare settings, with particular attention to the Eastern European and Romanian context. PubMed/MEDLINE, Embase, Web of Science, and the Cochrane Library were searched for relevant publications from January 2020 to June 2025, supplemented by WHO and ECDC surveillance reports. Studies were selected narratively for their relevance to the themes addressed. No new quantitative pooling was undertaken; all summary estimates reported below are cited from the published meta-analyses and surveillance reports that generated them. In the most recent global meta-analysis of hospital-acquired CRKP infection, which pooled 61 studies and 513,307 patients from 14 countries, the global prevalence of CRKP among nosocomial K. pneumoniae infections was 28.69% (95% CI: 26.53–30.86%), with pronounced regional variation from 14.29% in high-income North America to 66.04% in South Asia, and 42.05% in Western Europe. Pooled mortality among patients infected with CRKP has been estimated in a separate meta-analysis at 42.14%, compared with 21.16% among patients infected with carbapenem-susceptible strains, rising to 54.30% in bloodstream infections. Surveillance data place Romania third in Europe for carbapenem resistance among invasive K. pneumoniae isolates, at 50.30%, with a distinctive predominance of NDM plus OXA-48-like co-producers. Ceftazidime-avibactam is recommended for KPC- and OXA-48-producing strains, whereas metallo-beta-lactamase producers require aztreonam-containing combinations. CRKP in HAIs constitutes a global epidemiological emergency characterised by marked regional heterogeneity in carbapenemase distribution, high attributable mortality and rapidly evolving molecular profiles. Locally adapted surveillance, rapid molecular diagnostics, and stewardship programmes are required since empirical therapy cannot be standardised across regions. Full article
(This article belongs to the Section Public Health Microbiology)
22 pages, 3121 KB  
Article
TriAIF-RWKV: A Physiology-Guided Spatiotemporal Framework for Robust Arterial Input Function Selection in CT Perfusion Imaging
by Lei Lei, Yu Shen, Dawei Wang, Feng Xi, Yixin He, Chaochao Wang and Jiandong Liu
Entropy 2026, 28(9), 944; https://doi.org/10.3390/e28090944 - 22 Aug 2026
Abstract
Accurate delineation of infarct core and ischemic penumbra in acute ischemic stroke primarily relies on computed tomography perfusion (CTP), where the arterial input function (AIF) is essential for reliable perfusion quantification. However, reliable and fast AIF selection remains challenging in clinical practice due [...] Read more.
Accurate delineation of infarct core and ischemic penumbra in acute ischemic stroke primarily relies on computed tomography perfusion (CTP), where the arterial input function (AIF) is essential for reliable perfusion quantification. However, reliable and fast AIF selection remains challenging in clinical practice due to noise, vascular heterogeneity, and inter-patient variability in bolus dynamics. In this study, we propose TriAIF-RWKV, a three-stage framework for robust and automated AIF extraction. Specifically, ACSANet is first employed for spatial vascular localization using axial and channel-aware attention mechanisms, thereby narrowing the candidate arterial region and reducing the AIF search space. Then, a Dilated-RWKV network is introduced to model temporal intensity dynamics from a global sequence perspective, allowing robust identification of AIF-consistent patterns. Finally, a physiology-informed scoring strategy is used to select the optimal AIF by evaluating baseline stability, peak enhancement, and washout characteristics. Extensive experiments on CTP datasets were conducted from multiple perspectives, including AIF waveform fidelity, perfusion parameter estimation, and lesion-level analysis. The results demonstrate that the proposed method achieved high agreement with expert-selected AIFs, with a global waveform PCC of 0.973, peak correlation of 0.942, and TTP correlation of 0.973 with a mean error of 0.923 s. Furthermore, the proposed method provides more consistent downstream perfusion quantification, achieving higher consistency of CTP-derived parameters and improved lesion-to-normal tissue discrimination compared with existing approaches. These results highlight its potential for reliable clinical perfusion assessment. Full article
(This article belongs to the Special Issue Entropy in Image, Video and Signal Processing)
Show Figures

Figure 1

19 pages, 2404 KB  
Article
Image Design Feature Enhancement Method Based on Transformer and Multimodal Neural Networks
by Qianning Xu and Jun Wang
Electronics 2026, 15(16), 3758; https://doi.org/10.3390/electronics15163758 - 21 Aug 2026
Viewed by 152
Abstract
With the rapid development of artificial intelligence technology, image design feature enhancement plays a crucial role in computer vision, image processing, and multimedia applications. Traditional feature enhancement methods often have limitations when dealing with complex and changeable image data. Therefore, the study proposes [...] Read more.
With the rapid development of artificial intelligence technology, image design feature enhancement plays a crucial role in computer vision, image processing, and multimedia applications. Traditional feature enhancement methods often have limitations when dealing with complex and changeable image data. Therefore, the study proposes an innovative MN-T (Multimodal Neural Network enhanced by Transformer) strategy to overcome these challenges. The MN-T strategy combines the attention mechanism of Transformer with the cross-modal learning capabilities of multimodal neural networks to more accurately capture and enhance key features in images. The Transformer attention mechanism enables MN-T to efficiently process global and local information in images, while the cross-modal learning capability of multimodal neural networks further enhances its ability to understand and express image features. The results show that MN-T strategy has excellent performance in processing different image data. Compared with the existing image design feature enhancement methods, MN-T strategy has achieved significant improvement in the root mean square error (RMSE), mean absolute error (MAE), peak signal-to-noise ratio (PSNR), and other key performance indicators. The research results of this paper not only provide a new idea and method for image design feature enhancement methods, but also provide a new theoretical support and experimental basis for the research of related fields. Full article
Show Figures

Figure 1

27 pages, 8211 KB  
Article
Dual-Level Spatial–Frequency Collaborative Detector for Oriented Object Detection in Remote Sensing Images
by Xuehuai Shi, Jingru Sun, Kun Yu, Zhihui Wei and Shangdong Zheng
Remote Sens. 2026, 18(16), 2845; https://doi.org/10.3390/rs18162845 - 21 Aug 2026
Viewed by 88
Abstract
Oriented object detection (OOD) in remote sensing images (RSIs) suffers from insufficient feature representation caused by arbitrary rotation angles and small spatial resolutions. Existing spatial–frequency fusion paradigms merely implement single-granularity feature interaction, either global image-level frequency compensation or local instance-level feature refinement, and [...] Read more.
Oriented object detection (OOD) in remote sensing images (RSIs) suffers from insufficient feature representation caused by arbitrary rotation angles and small spatial resolutions. Existing spatial–frequency fusion paradigms merely implement single-granularity feature interaction, either global image-level frequency compensation or local instance-level feature refinement, and fail to simultaneously capture global scene semantic consistency and local object fine-grained discriminability. In this paper, we propose a unified dual-level spatial–frequency collaborative detector (DSCDet) for remote sensing OOD tasks. Different from previous decoupled designs, the proposed DSCDet constructs a complete spatial–frequency collaborative fusion paradigm that shares a generic wavelet-based frequency extraction mechanism and cross-feature fusion module, which is adaptively deployed at both image-level and instance-level granularities. Specifically, our method introduces Haar wavelet transform to extract multi-scale frequency mutation features. On this basis, a generic cross-domain attention fusion (GCDAF) is constructed with granularity-dependent positional encoding constraints. The core difference between dual granularity fusion lies in geometric positional encoding, where image-level fusion adopts global scene positional embedding to maintain overall semantic stability, and instance-level fusion leverages local pairwise instance positional embedding to optimize fine-grained target feature interaction. The unified dual-level fusion architecture comprehensively integrates global semantic integrity and local target specificity, forming a robust and universal spatial–frequency feature representation system. Extensive experiments on three public remote sensing datasets, including DOTA-v1.0, DOTA-v1.5 and DIOR-R, demonstrate that the proposed DSCDet achieves competitive and superior performance against state-of-the-art OOD detectors. Full article
(This article belongs to the Section Remote Sensing Image Processing)
Show Figures

Figure 1

20 pages, 3894 KB  
Article
Attention-Enhanced Multi-Scale Feature-Wise Linear Modulation for Fine-Grained Poisonous Mushroom Image Recognition
by Yuan He, Haikun Lv, Chenyang Lu, Dengqi Yang, Xiaowei Li and Lina Zhang
J. Imaging 2026, 12(8), 398; https://doi.org/10.3390/jimaging12080398 - 21 Aug 2026
Viewed by 67
Abstract
Fine-grained poisonous mushroom recognition in natural scenes is challenging because of complex backgrounds, subtle morphological differences, and the limited interpretability of model decisions. To address these challenges, this paper proposes Att-FiLM, an attention-enhanced multi-scale Feature-Wise Linear Modulation network for poisonous mushroom image recognition. [...] Read more.
Fine-grained poisonous mushroom recognition in natural scenes is challenging because of complex backgrounds, subtle morphological differences, and the limited interpretability of model decisions. To address these challenges, this paper proposes Att-FiLM, an attention-enhanced multi-scale Feature-Wise Linear Modulation network for poisonous mushroom image recognition. The model adopts an asymmetric dual-backbone architecture in which a frozen ConvNeXt-Base branch provides global semantic priors, while a trainable EfficientNet-B0 branch learns local discriminative features. Rather than directly concatenating heterogeneous features, Att-FiLM generates scale and shift parameters from semantic features and performs channel-wise modulation on multi-scale EfficientNet features at Stage 2 and Stage 4. This mechanism enables global semantic information to guide local feature learning while reducing feature redundancy and semantic inconsistency. Experimental results show that Att-FiLM achieves an Accuracy of 95.58% and an F1-score of 0.9455 on the poisonous/edible binary classification task. On the 190-class species-level classification task, it achieves a Top-1 Accuracy of 93.63% and a Macro-F1 of 0.9347. Interpretability analysis further shows that decision-relevant responses are frequently associated with morphologically relevant regions, including gills, annuli, volvae, and cap textures. These results indicate that Att-FiLM provides effective recognition performance together with interpretable decision evidence for mushroom recognition in complex natural scenes. Full article
(This article belongs to the Section Image and Video Processing)
Show Figures

Figure 1

25 pages, 3738 KB  
Article
ESD-YOLO: A Method for Small-Target Termite Detection Under Complex Backgrounds
by Weiling Lu, Yuting Meng, Shan Wu and Hangjun Wang
Insects 2026, 17(8), 874; https://doi.org/10.3390/insects17080874 - 21 Aug 2026
Viewed by 73
Abstract
Timely and accurate termite detection is essential for effective termite control. To address the challenges posed by the small size of termite individuals and the susceptibility of target features to background texture interference under complex backgrounds, this study proposes ESD-YOLO, a fine-grained feature-enhanced [...] Read more.
Timely and accurate termite detection is essential for effective termite control. To address the challenges posed by the small size of termite individuals and the susceptibility of target features to background texture interference under complex backgrounds, this study proposes ESD-YOLO, a fine-grained feature-enhanced object detection model. Using YOLO11n as the baseline, ESD-YOLO redesigns the feature extraction, deep feature aggregation, and multi-scale feature fusion stages to improve the representation of small-scale termite targets under complex backgrounds. Specifically, the Efficient Multi-scale Attention (EMA) mechanism is incorporated into the C3k2 module to enhance feature discriminability between termite individuals and the background. A Spatial Pyramid Pooling-Fast with Dual Global Pooling (SPPF-DGP) module is employed to supplement deep features with global contextual information and salient response information. In addition, the DySample dynamic upsampling module is introduced to improve spatial alignment during multi-scale feature fusion and enhance boundary representation for small targets. Experimental results show that ESD-YOLO achieves Precision, Recall, mAP@0.5, and mAP@0.5:0.95 values of 95.39%, 96.33%, 97.85%, and 65.92%, respectively, with 2.67 M parameters and 6.68 G FLOPs. Compared with Faster R-CNN, RetinaNet, RT-DETR, and several YOLO-series models, ESD-YOLO demonstrates strong small-target detection and localization performance under the controlled complex-background conditions established in this study, providing a methodological reference for automated termite detection in practical settings. Full article
(This article belongs to the Special Issue AI and Cloud Computing for Insect Ecology and Management)
Show Figures

Figure 1

43 pages, 11529 KB  
Article
Enhancing End-to-End Graphite Ore Grade Detection via Boundary-Aware Refinement, Bidirectional Fusion, and Difficulty-Aware Distillation
by Yanwu Yi, Binghui Wei, Zeyang Qiu, Chen Yang and Xueyu Huang
Appl. Sci. 2026, 16(16), 8332; https://doi.org/10.3390/app16168332 - 21 Aug 2026
Viewed by 199
Abstract
Graphite ore grade sorting is a key step toward intelligent mineral processing; however, it faces three representational contradictions: ambiguous classification posteriors at grade boundaries, asymmetric multi-scale feature interaction, and the mismatch between class-agnostic self-distillation assignment and sample-level difficulty. Targeting these, this paper adopts [...] Read more.
Graphite ore grade sorting is a key step toward intelligent mineral processing; however, it faces three representational contradictions: ambiguous classification posteriors at grade boundaries, asymmetric multi-scale feature interaction, and the mismatch between class-agnostic self-distillation assignment and sample-level difficulty. Targeting these, this paper adopts D-FINE as the baseline and introduces three decoupled improvements at its decoder, encoder, and criterion layers. (1) Boundary-Grade-aware Distribution Refinement (BG-FDR) online identifies boundary samples via the Top-2 classification score gap and modulates regression-distribution refinement, yielding +2.69 percentage points in mAP@0.5 with zero additional trainable parameters. (2) Bidirectional Feature Pyramid with Global–Local Spatial Attention (BiFPN-GLSA) builds a learnable weighted bidirectional multi-scale fusion path. (3) Difficulty-Aware Decoupled Distillation with Wise-Inner-Shape-IoU (DADD+Wise-IoU) imposes class- and sample-level difficulty-aware constraints. In the integrated full model, this increases Precision from 66.21% to 71.43% (+5.22 pp), F1 from 73.57% to 77.57%, and mean IoU from 97.81% to 98.35%, while false positives drop by 19.6%; the only parameter overhead (+3.84M) comes from BiFPN-GLSA, with BG-FDR and DADD adding effectively no network weights. Ablation on a self-built 3800-image dataset reveals a non-monotonic AP–Precision relationship: the mAP-optimal configuration (BG-FDR+BiFPN-GLSA, 94.17%) and the Precision-optimal one (DADD+Wise-IoU, 77.54%) do not coincide, providing a quantitative basis for objective-driven module selection in industrial sorting. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

25 pages, 3707 KB  
Article
ESNformer: A Hybrid Reservoir–Transformer Architecture for Interpretable, Position-Aware Classification of Structured Assessment Data, with a Braille-Literacy Case Study
by Cesar H. Valencia-Niño, Rafael A. Nuñez-Rodriguez, Marley M. B. R. Vellasco and Jeison Marin
Technologies 2026, 14(8), 517; https://doi.org/10.3390/technologies14080517 - 21 Aug 2026
Viewed by 153
Abstract
We present ESNformer, a hybrid architecture that couples an Echo State Network (ESN) reservoir with a Transformer encoder for classification of structured, multi-indicator assessment data: a fixed-order vector of complementary indicators per assessment instance rather than a repeated-measures time series. The reservoir acts [...] Read more.
We present ESNformer, a hybrid architecture that couples an Echo State Network (ESN) reservoir with a Transformer encoder for classification of structured, multi-indicator assessment data: a fixed-order vector of complementary indicators per assessment instance rather than a repeated-measures time series. The reservoir acts as a fixed nonlinear feature map over the indicator vector, while self-attention, made position-aware over the fixed column order, learns how each indicator’s evidence contributes to the final decision, so the two components, together, capture local, indicator-level detail and global, cross-indicator interactions within a single, end-to-end trainable model. Interpretability is treated as a first-class design requirement rather than an afterthought: the architecture is paired with an explainability layer combining SHAP feature attribution (reported both globally and per class), the model’s own attention weights, a deletion/insertion faithfulness test that quantitatively verifies which inputs the model actually relies on, and counterfactual maps that translate a prediction into an actionable, inspectable recommendation. We evaluate the architecture on a concrete case study, classifying Braille-literacy instructional recommendations from 15 pedagogical indicators grouped into three categories (Mangold’s, ABKL, and Progresar), using a benchmark of 900 real assessment instances (630 used, together with a class-conditional augmentation procedure, to build a 2100-instance training set) with validation and test partitions (135 instances each) kept exclusively real. On this benchmark, the tuned model reached 85.33% accuracy, 85.90% macro-precision, 85.33% macro-recall, an F1 score of 85.25%, and an AUC of 0.95 on the real test set. SHAP attribution, attention weights, and the faithfulness test converge on the same two dominant indicators (response time and error count): removing them alone collapses accuracy to chance, while retaining only them recovers most of the model’s accuracy. We report this transparently alongside a comparison against ESN-only, Transformer-only, and tabular baselines (logistic regression, decision tree, random forest, XGBoost, and an MLP) on the same data and discuss what the hybrid architecture and its explainability pipeline add beyond what the two dominant indicators already explain and how the approach generalizes to other tabular and mixed-granularity assessment settings that require both predictive accuracy and a verifiable account of what drove each decision. Full article
Show Figures

Figure 1

27 pages, 2817 KB  
Article
A Controlled Evaluation of Dual-Channel Feature Enhancement and Multi-Level Knowledge Distillation for Lightweight Plant Disease Recognition
by Xin Lei, Yonghuai Liu, Ardhendu Behera, Reena Reena, Yang Sun, Fuzhong Li, Wuping Zhang and Chao Lei
Agriculture 2026, 16(16), 1790; https://doi.org/10.3390/agriculture16161790 - 21 Aug 2026
Viewed by 187
Abstract
Plant disease symptoms combine local texture changes with patterns distributed across a leaf, while practical recognition models must remain compact. We introduce DC-FEN, a MobileNetV3-based design that models spatial-token relations and channel interactions in parallel and injects them through gated residual fusion. We [...] Read more.
Plant disease symptoms combine local texture changes with patterns distributed across a leaf, while practical recognition models must remain compact. We introduce DC-FEN, a MobileNetV3-based design that models spatial-token relations and channel interactions in parallel and injects them through gated residual fusion. We also examine output-distribution, direct-feature, and token-relation transfer under same-backbone and heterogeneous teachers. PlantVillage and Plant Pathology 2021 (FGVC8) are evaluated with duplicate-audited, group-aware 70/15/15 splits, an explicit unresolved-leaf sensitivity check, validation-only selection, five training seeds, class-sensitive metrics, and paired seed-wise descriptive summaries. On PlantVillage, the no-additional-attention student, DC-FEN teacher, and DC-FEN joint student obtain macro F1 scores of 96.46±0.91%, 96.90±0.40%, and 96.55±0.25%. On FGVC8, the corresponding scores are 87.29±0.63%, 87.14±0.52%, and 87.20±0.26%. At the prespecified FGVC8 threshold of 0.5, DCAB changed sample-wise F1 by 0.02±0.55 percentage points relative to the unmodified backbone; validation-selected global and label-specific thresholds changed this contrast to +0.28±0.55 and +0.55±0.29 points, while threshold-free macro mAP remained essentially unchanged. A duplicate-audited PlantDoc pressure test reduced frozen-checkpoint accuracy to 30.34±1.10% and 29.57±1.10%, showing that external generalization remains unestablished. A ResNet50 teacher gives logit-only students 97.42±0.51% macro F1 on PlantVillage and 89.82±0.43% sample-wise F1 on FGVC8. After separately weighting the direct and relation terms, the corresponding joint students obtain 97.37±0.56% and 89.94±0.27%, recovering the degradation seen with unit internal weights while remaining close to logit-only transfer. Thus, the study evaluates the benefits and limits of explicit spatial–channel interaction and shows that adding intermediate transfer constraints does not guarantee a stronger student. Full article
(This article belongs to the Section Artificial Intelligence and Digital Agriculture)
Show Figures

Figure 1

22 pages, 1236 KB  
Article
ACSE-RNformer: Amplitude-Calibrated Sequence Embedding and Response-Normalized Transformer for Vibration-Based Rotating Machinery Fault Diagnosis
by Yan Yan, Ting Shang, Kun Zeng, Songnan Yang, Haiyan Cheng and Wei Quan
Sensors 2026, 26(16), 5275; https://doi.org/10.3390/s26165275 - 20 Aug 2026
Viewed by 195
Abstract
To address the insufficient representation of fault characteristics in rotating machinery vibration signals, the sensitivity of conventional Transformers to variations in input response amplitudes, and the limited ability of fixed sequence embedding to preserve continuous temporal information, a rotating machinery fault diagnosis method [...] Read more.
To address the insufficient representation of fault characteristics in rotating machinery vibration signals, the sensitivity of conventional Transformers to variations in input response amplitudes, and the limited ability of fixed sequence embedding to preserve continuous temporal information, a rotating machinery fault diagnosis method based on Amplitude-Calibrated Sequence Embedding (ACSE) and a Response-Normalized Transformer (RNformer) is proposed. First, ACSE is designed to construct local temporal feature representations through continuous convolutional mapping, while an amplitude response estimation and adaptive amplitude calibration mechanism is employed to dynamically recalibrate the response intensity at different temporal positions. Rather than simply rescaling the signal amplitude range, amplitude calibration adaptively strengthens the feature contribution of regions associated with fault-induced impacts according to the vibration response intensity, thereby highlighting fault-sensitive information while suppressing the influence of noncritical amplitude fluctuations. In this way, continuous temporal characteristics are preserved while fault-relevant information is enhanced. Second, RNformer is constructed by incorporating a response normalization mechanism into the Transformer encoder to mitigate the interference of abnormal amplitude responses with global feature modeling, thereby improving the stability and robustness of feature representations under complex operating conditions. Finally, a lightweight channel attention mechanism is introduced to further enhance critical fault features and perform fault classification. Experiments were conducted on the Paderborn University bearing dataset and the University of Connecticut gear dataset. The proposed method achieved average diagnostic accuracies of 99.36% and 99.28%, respectively, outperforming the best-performing baseline methods by 1.82 and 1.56 percentage points. These results demonstrated the effectiveness of the proposed method for fault diagnosis. Full article
(This article belongs to the Special Issue Intelligent Sensors and Signal Processing in Industry—2nd Edition)
Show Figures

Figure 1

22 pages, 65601 KB  
Article
Dual-Domain Illumination Prior for Low-Light Remote Sensing Image Enhancement
by Chao Wang, Zhe Pan, Liangtian He, Jun Liu, Lin Mei, Rongsheng Lin, Hongming Chen and Chuansheng Yang
Remote Sens. 2026, 18(16), 2817; https://doi.org/10.3390/rs18162817 - 20 Aug 2026
Viewed by 158
Abstract
Low-light conditions degrade remote sensing imagery by reducing contrast, distorting color, and obscuring fine terrain structures and small objects critical for Earth observation. Accurate illumination adjustment under spatially varying scene content remains challenging for existing enhancement methods, and many prior-guided approaches operate exclusively [...] Read more.
Low-light conditions degrade remote sensing imagery by reducing contrast, distorting color, and obscuring fine terrain structures and small objects critical for Earth observation. Accurate illumination adjustment under spatially varying scene content remains challenging for existing enhancement methods, and many prior-guided approaches operate exclusively in either the spatial domain or the frequency domain. In this work, we propose a Dual-Domain Illumination Prior (DDIP), a trainable dual-domain illumination-prior module that is jointly optimized with each host backbone and exploits frequency-domain and spatial-domain illumination statistics. DDIP comprises three components: a Frequency-Domain Illumination Distribution Prior (FIDP) that performs per-color-channel amplitude calibration in Fourier space to improve global brightness; a Spatial-Domain Illumination Distribution Prior (SIDP), adapted from IDP-Net, that performs multi-scale sub-region statistical correction for local illumination adjustment; and a Selective Core Feature Fusion (SCFF) module that adaptively combines the frequency-domain output, the spatial-domain output, and the original input through an attention-based gating mechanism with dual pooling. DDIP is integrated with each host backbone while leaving its main restoration blocks unchanged. In the controlled reconstruction comparisons on iSAID-dark and the evaluated general low-light benchmarks, equipping the tested backbone networks with DDIP improves PSNR and SSIM over their corresponding baselines. Complementary LPIPS and CIELAB lightness measurements characterize perceptual similarity and lightness behavior, while a fixed-detector object-detection evaluation on the tested high-resolution iSAID-dark scenes examines the effect of the enhancement pipelines under the reported synthetic low-light conditions. The ablation studies further examine the contribution of the module components within the reported experimental settings. Full article
Show Figures

Figure 1

Back to TopTop