Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

Search Results (457)

Search Parameters:
Keywords = cross-modal attention mechanism

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
41 pages, 13249 KB  
Article
A Gated Multi-Source Signal Fusion Method for Bearing Fault Diagnosis with a Fusion Negative-Transfer Suppression Mechanism
by Tianhao Gao, Ke Zhang, Nan Wang, Yang Hong and Shijie Wang
Machines 2026, 14(8), 940; https://doi.org/10.3390/machines14080940 - 14 Aug 2026
Abstract
Multi-source information fusion is regarded as a key approach for improving bearing fault diagnosis. However, due to the heterogeneity of multi-source information, asymmetric information contributions, and imbalanced discriminative features, negative transfer may occur during fusion. To address this issue, this paper proposes a [...] Read more.
Multi-source information fusion is regarded as a key approach for improving bearing fault diagnosis. However, due to the heterogeneity of multi-source information, asymmetric information contributions, and imbalanced discriminative features, negative transfer may occur during fusion. To address this issue, this paper proposes a negative-transfer-suppression diagnosis framework based on physical-information guidance and adversarially disentangled representation. First, an adaptive preprocessing mechanism guided by acoustic–vibration cross-correlation and mutual information entropy is constructed to extract intrinsic cross-modal correlations, enabling source-end feature reconstruction and commonality enhancement. Second, an attention-based spatial feature extraction operator and an adversarial common-domain representation model are developed to suppress modality-specific interference and disentangle cross-modal shared features. On this basis, sparse coding is employed to fuse common-domain and modality-specific features. Furthermore, a classification effectiveness evaluation index based on fuzzy clustering is introduced into the loss function to dynamically constrain sparse coding weights, thereby reducing interference features and suppressing negative transfer under strong-noise conditions. Experimental results demonstrate that the proposed method effectively achieves the design objective that “fusion outperforms non-fusion,” exhibiting strong noise robustness and high diagnostic accuracy under complex operating conditions. Full article
Show Figures

Figure 1

21 pages, 3149 KB  
Article
Heterogeneous SNN-ANN Multimodal Fusion Framework for Comprehensive Fruit Quality Assessment
by Weibin Tang, Qi Sun, Yunfan Guo and Zhen Cao
Electronics 2026, 15(16), 3613; https://doi.org/10.3390/electronics15163613 - 13 Aug 2026
Abstract
Reliable fruit quality assessment is crucial for ensuring food safety and value in modern agriculture. However, many current approaches still rely heavily on visual cues, making it difficult to assess internal quality indicators such as sweetness or internal decay. To address this limitation, [...] Read more.
Reliable fruit quality assessment is crucial for ensuring food safety and value in modern agriculture. However, many current approaches still rely heavily on visual cues, making it difficult to assess internal quality indicators such as sweetness or internal decay. To address this limitation, we propose HSAF-Net, a heterogeneous multimodal fusion framework integrating spiking neural networks (SNNs) and artificial neural networks (ANNs) for comprehensive, non-destructive fruit quality assessment. Specifically, the SNN encodes near-infrared (NIR) spectral signals to extract internal sugar-related features, whereas the ANN-based TH-YOLOv8 model detects external surface defects from high-resolution RGB images. A microsecond-level synchronous acquisition scheme is implemented to ensure precise alignment between the NIR and RGB modalities. To effectively combine heterogeneous features, we design a Heterogeneous Modality Attention (HMA) mechanism that dynamically fuses multi-source information based on task-specific relevance. Compared with image-only detection, the proposed framework explicitly separates internal biochemical sensing from external defect localization and then integrates their complementary decisions in a unified grading pipeline. Experimental results on 616 pear samples demonstrate that the HSAF-Net achieves 95.2% classification accuracy, 95.1% mAP95, and an internal defect miss rate as low as 7.5%, outperforming conventional single-modality and early-fusion baselines by a notable margin. The system maintains a real-time inference speed of 55 ms per sample on the Ascend Atlas 200DK A2 edge platform, validating its deployment potential. The current evaluation is based on crisp pear samples collected under controlled acquisition conditions; therefore, broader cross-variety and cross-season validation remains necessary before large-scale commercial deployment. This study presents an end-to-end multimodal SNN-ANN fusion architecture tailored for fruit grading and provides a scalable, high-precision solution for post-harvest quality assessment with broad applicability to other agricultural products. Full article
14 pages, 2487 KB  
Article
CM-FuseNet: An Attention-Augmented Hybrid EEG–EMG Cognitive–Motor Fusion Network with Soft Actor-Critic Reinforcement Learning for Adaptive Lower-Limb Exoskeleton Control
by Yong-Deok Park, Dae-seob Shin and Hun-kee Kim
Appl. Sci. 2026, 16(16), 8042; https://doi.org/10.3390/app16168042 - 12 Aug 2026
Abstract
Population aging and the rising prevalence of motor disorders are driving demand for assistive lower-limb robotic systems capable of decoding user intention rather than merely providing mechanical support. We present CM-FuseNet, an attention-augmented hybrid Brain–Computer–Muscle Interface (BCMI) that simultaneously fuses cortical concentration indices [...] Read more.
Population aging and the rising prevalence of motor disorders are driving demand for assistive lower-limb robotic systems capable of decoding user intention rather than merely providing mechanical support. We present CM-FuseNet, an attention-augmented hybrid Brain–Computer–Muscle Interface (BCMI) that simultaneously fuses cortical concentration indices extracted from electroencephalography (EEG) and lower-limb intention patterns derived from electromyography (EMG) to adaptively control a 4-DOF assistive lower-limb exoskeleton. To eliminate the burden of human-subject ethics review and to ensure reproducibility of the proposed methodology, all validation is performed exclusively on (i) permissively licensed open-access biomedical datasets, (ii) high-fidelity OpenSim 4.5 and MuJoCo 3.1 musculoskeletal–exoskeleton co-simulation, and (iii) limited self-experimentation by the corresponding author with non-invasive consumer-grade devices. Three components are introduced: (i) a log-tanh normalized concentration index CI in (0, 1) derived from the (PSMR+PMidBeta)/PTheta ratio; (ii) a bidirectional Cross-Modal Transformer (CMT) with eight-head self- and cross-attention; and (iii) a Soft Actor-Critic (SAC) reinforcement-learning controller that adaptively tunes four servo PID gains using a concentration-weighted state. Experiments on the PhysioNet EEGMMIDB, Ninapro DB2/DB7, HuMoD and WAY-EEG-GAL datasets (combining N = 162 trial sessions, 47,520 windows, and five-fold cross-validation) yield a gait-phase classification accuracy of 96.84 ± 1.18%, torque-tracking RMSE of 0.072 ± 0.008 N·m, information transfer rate of 38.6 bits/min, end-to-end latency of 9.4 ms, and a 27.4% reduction in simulated metabolic cost over an EMG-only PID baseline (one-way ANOVA: F(4, 75) = 47.83, p < 0.001; Tukey HSD: p < 0.01 against all baselines). Under high cognitive load, CM-FuseNet preserves accuracy with only a 4.63 percentage-point degradation versus 13.22 percentage points for the EMG-only baseline. Full article
(This article belongs to the Section Robotics and Automation)
Show Figures

Figure 1

24 pages, 13353 KB  
Article
Morphological Buffering in Neuro-Architecture: Interactive Effects of Color, Shape, and Area Proportion on Affective Processing
by Xiaoxiao Dou, Yan Zhang, Yannan Zhang, Mengyao Kang, Qiangqiang Fan, Xiaona Xie, Yufan Sun and Mengyao Li
Behav. Sci. 2026, 16(8), 1364; https://doi.org/10.3390/bs16081364 - 9 Aug 2026
Viewed by 149
Abstract
While architectural visual features significantly influence human emotion, the spatiotemporal neuro-mechanisms underlying their interactive effects remain poorly under-researched. This study systematically decoupled the dynamics of indoor spatial perception by investigating the interactive impacts of color combination, color shape, and area proportion to find [...] Read more.
While architectural visual features significantly influence human emotion, the spatiotemporal neuro-mechanisms underlying their interactive effects remain poorly under-researched. This study systematically decoupled the dynamics of indoor spatial perception by investigating the interactive impacts of color combination, color shape, and area proportion to find out the effects on participants’ emotion. By integrating high-resolution electroencephalography (EEG), specifically analyzing the N200, late positive potential (LPP), Frontal Alpha Asymmetry (FAA), and Beta bands with Self-Assessment Manikin (SAM) ratings and Liking scores in immersive virtual indoor environments, data from 61 valid participants were analyzed. The results revealed a hierarchical neuro-affective processing structure: morphology dominated early cognitive processing, where curved geometries acted as a pre-attentive cognitive buffer, acting as a cognitive buffer by significantly attenuating the early-stage visual stress induced by angular, high-contrast configurations. Conversely, color combination primarily drove late-stage sustained environmental arousal, tracked by the LPP. Notably, area proportion demonstrated a significant amplification effect, significantly exacerbating pre-existing neuro-affective biases. Crucially, cross-modal triangulation revealed a selective inter-modal coupling: while sustained cortical activation (ΔLPP) significantly predicted subjective emotional arousal, early physiological buffering (ΔN200/ΔFAA) was functionally dissociated from subjective aesthetic preference (ΔLiking). This decoupling provides preliminary evidence consistent with a dual-system processing framework, suggesting that while curves may facilitate a preconscious biological safe baseline, ultimate aesthetic appreciation is likely subject to top-down cognitive over-riding. Ultimately, these findings transition indoor spatial color design from intuitive practice to evidence-based neuro-aesthetics, highlighting the critical role of morphological buffering and EEG metrics in optimizing human emotional well-being. Full article
(This article belongs to the Section Cognition)
Show Figures

Figure 1

32 pages, 13437 KB  
Article
MSFusion: Multi-Scale Cross-Modal Fusion with Adaptive Attention for Multimodal Medical Image Fusion
by Liu Wang, Yang Zhou, Wenjia Li and Lijuan Shi
Biosensors 2026, 16(8), 423; https://doi.org/10.3390/bios16080423 - 6 Aug 2026
Viewed by 229
Abstract
Multimodal medical image fusion integrates complementary information from heterogeneous imaging modalities to provide comprehensive visual support for clinical analysis. Most existing methods adopt an “encode–fuse–decode” paradigm that applies a single fusion rule only at the deepest network layer, often discarding shallow detail features [...] Read more.
Multimodal medical image fusion integrates complementary information from heterogeneous imaging modalities to provide comprehensive visual support for clinical analysis. Most existing methods adopt an “encode–fuse–decode” paradigm that applies a single fusion rule only at the deepest network layer, often discarding shallow detail features and yielding blurred outputs with poor textural fidelity. To address this limitation, we propose MSFusion, a novel hierarchical framework that distributes adaptive fusion throughout the entire decoder stage. By leveraging skip connections to align decoder layers with corresponding encoder features, MSFusion enables full-scale integration of multi-resolution representations. The architecture employs a dual-branch convolutional encoder and introduces two core modules in the decoder: (1) the Multi-Scale Adaptive Fusion (MSAF) module, which addresses insufficient exploitation of cross-scale complementarity by dynamically weighting features via learnable attention, thereby balancing fine details and global semantics, and (2) the Multi-Scale Cross-Modal Cooperative Fusion (MSCMCF) module, which mitigates semantic misalignment through a cross-modal interactive attention mechanism that establishes robust inter-modality correspondences and promotes deep feature alignment. Additionally, a Vision RWKV (VRWKV) block is integrated to efficiently model both local and global spatial dependencies with linear computational complexity. Extensive experiments on public CT–MRI, PET–MRI, and SPECT–MRI datasets are evaluated using six standard metrics. On the CT–MRI benchmark, our method achieves MSE (↓) = 0.0423, CC = 0.8198, and SCD = 0.7023, outperforming state-of-the-art approaches. These results—combined with superior visual quality—demonstrate that MSFusion sets a new standard for accurate, detailed, and clinically meaningful multimodal image fusion. Full article
(This article belongs to the Special Issue The Smart Biosensors Era: AI in Cancer Detection and Imaging)
Show Figures

Figure 1

31 pages, 24356 KB  
Article
PCFD-Net: A Parallel Collaborative Fusion-Detection Network for SAR and Optical Imagery
by Yixuan An, Ning Wang, Haixiao Wu, Yuchen Wu and Tao Liu
Remote Sens. 2026, 18(15), 2595; https://doi.org/10.3390/rs18152595 - 5 Aug 2026
Viewed by 197
Abstract
Synthetic aperture radar (SAR)–optical image fusion and object detection are two closely related tasks in remote sensing. Fusion can provide richer texture and structural cues for downstream detection, while detection can, in turn, provide object-level location and semantic information to improve fusion. However, [...] Read more.
Synthetic aperture radar (SAR)–optical image fusion and object detection are two closely related tasks in remote sensing. Fusion can provide richer texture and structural cues for downstream detection, while detection can, in turn, provide object-level location and semantic information to improve fusion. However, effectively integrating these two tasks within a unified training framework remains challenging. Their optimization objectives are inherently different: fusion emphasizes cross-modal information preservation and structural fidelity, whereas detection focuses more on discriminative target representation. As a result, direct joint training often leads to mutual interference rather than mutual reinforcement. In addition, most existing joint frameworks remain serial or unidirectional, limiting effective bidirectional knowledge transfer between fusion and detection. To address these issues, we propose PCFD-Net (Parallel Collaborative Fusion-Detection Network), which consists of a fusion branch, a detection branch, and a bidirectional interaction branch, and unifies fused image generation and oriented object detection within a single training framework through explicit bidirectional interaction. The fusion branch employs dual ResNet-50 encoders, a multi-scale attention fusion module, and a progressive decoder, while the detection branch is built on YOLOv8. The key component of the proposed framework is the bidirectional interaction branch. On the one hand, the multi-scale fused features generated by the fusion branch are injected into the detection backbone to enhance the exploitation of cross-modal intermediate representations. On the other hand, we develop CSMDE (Category Semantic–Modality Disentangled Embedding), which disentangles category-discriminative and modality-preference semantics to map detector category outputs into instance-level semantic embeddings. These embeddings, together with object locations, are further fed into a dual-discriminator mechanism to reversely constrain the fusion branch, thereby strengthening SAR-discriminative target preservation and optical background structure consistency. Experiments on the M4-SAR and OGSOD1.0 datasets demonstrate that PCFD-Net consistently outperforms representative fusion and detection methods, achieving superior fusion quality and stronger downstream detection performance. Full article
Show Figures

Figure 1

28 pages, 5504 KB  
Article
Multimodal Heterogeneous CNN with Adaptive Modality Fusion for Intelligent Fault Diagnosis of Bearings
by Chang Sun, Chenkun Wang, Shiwei Huang and Tianci Zhang
Machines 2026, 14(8), 875; https://doi.org/10.3390/machines14080875 - 1 Aug 2026
Viewed by 289
Abstract
In industrial equipment fault diagnosis, vibration and acoustic signals are highly complementary yet exhibit significant differences in frequency distribution and noise sensitivity. Traditional multimodal methods generally rely on homogeneous feature extractors and direct feature concatenation, which may fail to capture modality-specific characteristics and [...] Read more.
In industrial equipment fault diagnosis, vibration and acoustic signals are highly complementary yet exhibit significant differences in frequency distribution and noise sensitivity. Traditional multimodal methods generally rely on homogeneous feature extractors and direct feature concatenation, which may fail to capture modality-specific characteristics and introduce irrelevant information during fusion. To address this, we propose a novel multimodal heterogeneous convolutional neural network framework. Specifically, separate 1D CNN branches are designed for vibration and acoustic signals. Their architectural differences are determined by the characteristics of each sensing modality. The vibration branch focuses on extracting high-level discriminative fault features, including impulse responses and modulated components from vibration signals, while the acoustic branch is designed to preserve fragile high-frequency details of acoustic signals. Furthermore, an adaptive cross-attention fusion module is introduced to dynamically model cross-modal dependencies, assigning Softmax-based weights to enhance dominant features and suppress noise. Experiments based on bearing fault experimental data demonstrate that the proposed heterogeneous architecture significantly outperforms traditional homogeneous models. The dynamic weighting mechanism effectively prevents inferior noisy modalities from degrading overall performance, achieving high diagnostic accuracy. Although validated on rolling bearing fault diagnosis, the proposed heterogeneous multimodal framework is not restricted to bearings and can be readily extended to other intelligent condition monitoring tasks involving heterogeneous sensor fusion, such as gearboxes, motors, and other rotating machinery. Full article
Show Figures

Figure 1

29 pages, 4881 KB  
Article
An Explainable Multimodal Framework for Breast Ultrasound Report Generation Using Vision-Language Transformers
by Prashanth Gowda Attahalli Shivakumar, Azhar Mahmood and Shaheen Khatoon
J. Imaging 2026, 12(8), 338; https://doi.org/10.3390/jimaging12080338 - 27 Jul 2026
Viewed by 426
Abstract
Breast cancer remains one of the leading causes of cancer-related mortality among women worldwide, where early and accurate diagnosis is critical for effective treatment. Although recent advances in deep learning have enabled automated radiology report generation from breast ultrasound images, most existing approaches [...] Read more.
Breast cancer remains one of the leading causes of cancer-related mortality among women worldwide, where early and accurate diagnosis is critical for effective treatment. Although recent advances in deep learning have enabled automated radiology report generation from breast ultrasound images, most existing approaches function as black-box systems, limiting clinical trust and interpretability. This study proposes a trustworthy and explainable framework for automated breast ultrasound report generation that combines Vision-Language Modelling (VLM) with multi-level Explainable Artificial Intelligence (XAI). The proposed architecture integrates a Swin Transformer for visual feature extraction, BioBERT/ClinicalBERT for clinical text representation, and a GPT-2-based decoder for report generation through a dual cross-attention fusion mechanism. The framework is evaluated on benchmark breast ultrasound datasets paired with expert-annotated radiology reports using standard natural language generation metrics, including BLEU, ROUGE-L, METEOR, and CIDEr. Experimental results demonstrate that the multimodal architecture significantly improves report quality, clinical consistency, and semantic accuracy compared with conventional image-only and single-modal baselines. To address transparency and trustworthiness, the framework provides dual-level explanations through Grad-CAM visual heatmaps and LIME/SHAP-based token attribution analysis, enabling clinicians to understand both image regions and textual features influencing generated reports. Qualitative assessment further indicates strong alignment between model explanations and radiologist-identified diagnostic findings. Full article
(This article belongs to the Section AI in Imaging)
Show Figures

Figure 1

21 pages, 2146 KB  
Article
A Multi-Stage Dual Encoder–Decoder Network Based on Event Image Cross-Modal Fusion for Image Deblurring
by Yan Liu, Yanfei Jia, Sheng Qiang, Yongpei Lin and Liquan Zhao
Sensors 2026, 26(15), 4762; https://doi.org/10.3390/s26154762 - 27 Jul 2026
Viewed by 303
Abstract
Most existing event-driven image deblurring methods ignore inherent differences between the two modalities and lack explicit alignment strategies, leading to cross-modal mismatches and degraded feature reconstruction. To address this issue, a multi-stage dual encoder–decoder image deblurring method based on event image cross-modal fusion [...] Read more.
Most existing event-driven image deblurring methods ignore inherent differences between the two modalities and lack explicit alignment strategies, leading to cross-modal mismatches and degraded feature reconstruction. To address this issue, a multi-stage dual encoder–decoder image deblurring method based on event image cross-modal fusion is proposed. The proposed network consists of an encoder and a decoder. The encoder employs dilated convolutional residual modules for feature extraction. It also integrates a cross-modal feature fusion module and a local scoring mechanism. These components combine event features with frame image features while suppressing noise. The decoder reconstructs image features via two directional decoding sub-networks. It also incorporates a feedback attention module. This module selects informative features along the feedback path. As a result, the image reconstruction quality is enhanced. In addition to the standard loss, mean absolute error, structural similarity, and frequency reconstruction losses are used to optimize deblurring performance. Extensive experiments are conducted on the GoPro, REBlur, and RwEvent datasets. For PSNR, our method exceeds REFID by 0.23 dB, 0.20 dB, and 0.49 dB on the three datasets. For SSIM, our model achieves gains of 0.002, 0.002, and 0.017 against REFID. In terms of computational cost and inference speed, our network adds only 3.4 M parameters and 105.2 GFLOPs, with an FPS reduction of only 2.12. This trivial efficiency loss delivers significant improvements in both pixel and structural restoration performance. Ablation experiments verify the independent positive contribution of each designed module. Both qualitative visual comparisons and quantitative metrics demonstrate that the proposed network has stronger deblurring and generalization capabilities. Full article
(This article belongs to the Special Issue AI-Based Sensing and Imaging Applications)
Show Figures

Figure 1

24 pages, 2424 KB  
Article
FraudDebate-Agent: A Multi-Agent LLM Framework with an Evidence-Based Debate Mechanism for Financial Statement Fraud Detection
by Xinran Yue, Jingyun Yang and Wenhe Liu
Mathematics 2026, 14(15), 2695; https://doi.org/10.3390/math14152695 - 27 Jul 2026
Viewed by 404
Abstract
Financial statement fraud inflicts large and recurring losses on capital markets, yet the dominant detection paradigm still relies on single, black-box classifiers (e.g., RUSBoost) trained on structured accounting ratios alone. Two limitations follow: (i) the rich, unstructured Management Discussion and Analysis (MD&A) narrative [...] Read more.
Financial statement fraud inflicts large and recurring losses on capital markets, yet the dominant detection paradigm still relies on single, black-box classifiers (e.g., RUSBoost) trained on structured accounting ratios alone. Two limitations follow: (i) the rich, unstructured Management Discussion and Analysis (MD&A) narrative of the 10-K filing is discarded, and (ii) the resulting scores are difficult for auditors to trust because they carry no transparent, standards-aligned rationale. Recent large language model (LLM) systems have shown that multi-agent collaboration is more robust than a single LLM for anomaly detection, but no study has systematically transferred this paradigm to listed-company statement fraud. We propose FraudDebate-Agent, a four-role multi-agent system in which a Quantitative Analyst agent scores 28 raw accounting items and 14 ratios with gradient-boosted and tabular attention models, a Narrative Auditor agent quantifies tone, linguistic uncertainty, and year-over-year textual novelty of the MD&A with FinBERT, and an Industry Peer agent uses retrieval-augmented generation to measure industry-relative anomaly. A Critic–Debate agent then orchestrates a pair-wise Evidence-based Multi-Agent Debate (EMAD) that reconciles disagreement across modalities and arbitrates a reconciled fraud-risk assessment, which is aggregated over a tri-modal evidence graph. Our contributions are as follows: (1) the first use of an evidence-grounded debate mechanism for accounting fraud, which materially reduces LLM hallucination; (2) a numerical–textual–peer evidence graph that fuses heterogeneous signals; and (3) an explainable report aligned with the PCAOB AS 2401 fraud-risk taxonomy. On AAER-labelled firm-years linked across a SEC financial dataset and EDGAR-CORPUS, FraudDebate-Agent improves the area under the ROC curve and the rare-event ranking metric NDCG@k over the strongest single-modality and single-LLM baselines while producing substantially more faithful explanations. We frame the system as a fraud-risk screening and risk-ranking tool for AAER-labelled misstatement risk rather than a determination of fraudulent intent. We report results over multiple seeds to reflect real-world stochasticity and discuss limitations and cross-domain applications. Full article
Show Figures

Figure 1

19 pages, 13011 KB  
Article
Defect Target Detection Network Integrating Depth Spatial Information and Adaptive Progressive Feature Fusion for Power Transmission and Distribution Lines
by Junsheng Lin, Jinchao Guo, Gao Liu, Feng Zhang, Changyu Li, Kaipeng Gao, Benxi Tian, Zhenbing Zhao and Haopeng Li
Energies 2026, 19(15), 3519; https://doi.org/10.3390/en19153519 - 26 Jul 2026
Viewed by 242
Abstract
Transmission and distribution networks are evolving into distributed smart grids, making accurate defect localization of line components increasingly important for power inspection. However, traditional manual inspection is inefficient and unreliable in complex scenarios, while most existing deep learning methods rely only on two-dimensional [...] Read more.
Transmission and distribution networks are evolving into distributed smart grids, making accurate defect localization of line components increasingly important for power inspection. However, traditional manual inspection is inefficient and unreliable in complex scenarios, while most existing deep learning methods rely only on two-dimensional visible-light images and struggle to capture spatial structure and occlusion relationships. To address these limitations, this paper proposes a defect detection network that integrates monocular relative-depth information with adaptive progressive feature fusion. The framework adopts a dual-branch architecture, where visible-light images provide appearance information and relative-depth maps generated from the corresponding RGB frames using Depth Anything V2 encode auxiliary spatial relationships. Because both modalities originate from the same image frame, no additional depth sensor or cross-sensor temporal synchronization is required. An adaptive progressive fusion module performs coarse-to-fine cross-modal interaction, and a coordinate attention mechanism is introduced to enhance positional encoding and suppress background interference. Experiments on a self-constructed dataset of 9838 RGB images paired with estimated relative-depth maps covering five typical defect categories show that the proposed method achieves 92.4% mAP@50 and 67.1% mAP@50–95. The results demonstrate its effectiveness in improving defect localization, small-object detection, and robustness across the evaluated transmission and distribution line scenes. Full article
Show Figures

Figure 1

34 pages, 5212 KB  
Review
Text-to-Image Generation via Deep Learning: A Comprehensive Review of Models, Architectures, and Future Directions
by Abdussalam Elhanashi, Siham Essahraui, Qinghe Zheng and Sergio Saponara
Appl. Sci. 2026, 16(15), 7430; https://doi.org/10.3390/app16157430 - 24 Jul 2026
Viewed by 311
Abstract
Text-to-image generation is an increasingly fast-paced field of generative artificial intelligence, consisting of synthesizing images of high quality and semantic consistency based on natural language descriptions. In this paper, we give an extensive overview of the approach to text-to-image generation using deep learning, [...] Read more.
Text-to-image generation is an increasingly fast-paced field of generative artificial intelligence, consisting of synthesizing images of high quality and semantic consistency based on natural language descriptions. In this paper, we give an extensive overview of the approach to text-to-image generation using deep learning, including the most common core model families, architecture designs, training approaches, and evaluation systems. We discuss the paradigms of the generative adversarial networks (GANs), variational autoencoders (VAEs), transformer-based designs, and diffusion models, with the last one representing the state of the art in image generation models. The review also discusses key aspects of pipelines such as text encoding, cross-modal alignment, mechanisms of attention, and decoding images. Popular datasets, methods, and metrics of evaluation, including Fréchet Inception Distance (FID) and CLIP-based similarity, are discussed. The application domains that involve creative content creation, medical imaging, education and industrial design are critically discussed. Despite significant advances, various issues still exist, such as low stability in training, excessive computational complexity, amplification of bias, generated images, and text–image alignment errors. Moral and social issues, such as misinformation, intellectual property, and equity, are critically examined. Lastly, we present future research directions to more controllable, more efficient and more interpretable text-to-image systems, focusing on multimodal foundation models and human–AI collaborative design. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

24 pages, 2235 KB  
Article
An Improved Recognition Technique for Ship Targets Based on Dual-Path Cooperative Fusion Mechanism
by Sifan Su, Wei Yang, Shiwen Lei, Xiaozhang Zhu, Jing Tian and Haoquan Hu
Remote Sens. 2026, 18(15), 2447; https://doi.org/10.3390/rs18152447 - 24 Jul 2026
Viewed by 342
Abstract
With the increasing complexity of the electromagnetic environment, traditional radar target recognition methods face severe challenges. High-resolution range profile (HRRP) and Radar Cross Section (RCS), as two important radar features, each has its own advantages in target recognition but also exhibits limitations. To [...] Read more.
With the increasing complexity of the electromagnetic environment, traditional radar target recognition methods face severe challenges. High-resolution range profile (HRRP) and Radar Cross Section (RCS), as two important radar features, each has its own advantages in target recognition but also exhibits limitations. To enhance radar target recognition performance in complex scenarios such as low signal-to-noise ratio (SNR), this paper proposes a recognition method based on heterogeneous multi-modal feature fusion. The proposed method constructs a three-channel parallel encoding network, which utilizes Convolutional Long Short-Term Memory (ConvLSTM), One-Dimensional Convolutional Gated Recurrent Unit (Conv1D-GRU), and Gated Recurrent Unit (GRU) to extract deep discriminative features from raw HRRP sequences, RCS sequences, and HRRP statistical features, respectively. Furthermore, it innovatively designs a dual-path cooperative fusion mechanism, achieving explicit inter-modal correlation modeling through a cross-attention module and dynamically learning the importance of each modality through an adaptive weight fusion layer, thereby realizing deep complementarity and enhancement of multi-modal information. Experimental results demonstrate that under various signal-to-noise ratios and polarization conditions, the proposed method achieves a maximum average recognition accuracy of over 99% for 6 ship targets. Compared with the traditional three-channel fixed-weight fusion method, the recognition accuracy of the proposed method increases from 90.01% to 99.42% under co-polarization, and from 81.14% to 98.15% under cross-polarization, fully validating the effectiveness and superiority of the dual-path fusion mechanism. Full article
(This article belongs to the Section Engineering Remote Sensing)
Show Figures

Figure 1

38 pages, 3295 KB  
Article
MAF-SleepNet: A Multimodal Attention-Enhanced Fusion Network for Automatic Multi-Class Sleep Disorder Classification from Polysomnography
by Suleyman Yaman, Hasan Guler and Abdul Hafeez-Baig
Diagnostics 2026, 16(15), 2317; https://doi.org/10.3390/diagnostics16152317 - 23 Jul 2026
Viewed by 341
Abstract
Background/Objectives: Sleep disorders are heterogeneous conditions with diverse neural, muscular, and ocular manifestations, making polysomnography (PSG) the gold standard for accurate diagnosis. Artificial intelligence-based approaches, particularly deep learning (DL) models capable of integrating heterogeneous information, offer a promising solution for reliable decision-making in [...] Read more.
Background/Objectives: Sleep disorders are heterogeneous conditions with diverse neural, muscular, and ocular manifestations, making polysomnography (PSG) the gold standard for accurate diagnosis. Artificial intelligence-based approaches, particularly deep learning (DL) models capable of integrating heterogeneous information, offer a promising solution for reliable decision-making in such clinical scenarios. However, most existing DL studies have focused on a single disorder, relied on limited datasets, or employed epoch-level labeling strategies that overlook the episodic nature of sleep pathophysiology, thereby limiting clinical applicability. To address these gaps, we propose a novel multimodal attention-enhanced fusion network (MAF-SleepNet) for automatic multi-class sleep disorder classification based on the International Classification of Sleep Disorders. Methods: MAF-SleepNet jointly processes electroencephalography (EEG), electrooculography (EOG), and leg electromyography (EMG) signals through modality-specific feature extraction and adaptive attention mechanisms, capturing both intra- and inter-modality dependencies. The model was evaluated on a combined dataset of 141 recordings from three public databases, including five PSG-requiring disorders and a healthy class. Results: Experimental results demonstrated that MAF-SleepNet achieved 86.07 ± 3.66% accuracy and 82.67 ± 4.46% macro-F1 under a strict subject-independent cross-validation, and 98.96 ± 0.66% accuracy and 98.93 ± 0.60% macro-F1 under subject-dependent cross-validation. Conclusions: These results demonstrate that the proposed approach provides a more reliable and clinically meaningful assessment compared to many existing studies that rely on subject-dependent evaluation or epoch-level labeling. The findings highlight the effectiveness of adaptive multimodal fusion for robust and clinically relevant sleep disorder classification. Future work should investigate the integration of respiratory and autonomic modalities and validation on larger multi-center cohorts. Full article
(This article belongs to the Section Machine Learning and Artificial Intelligence in Diagnostics)
Show Figures

Figure 1

24 pages, 2860 KB  
Article
CGF-Net: A Multi-View Contrastive Learning Model for Encrypted Traffic Classification
by Yanlin He and Ning Hu
Electronics 2026, 15(15), 3249; https://doi.org/10.3390/electronics15153249 - 23 Jul 2026
Viewed by 286
Abstract
For the task of encrypted traffic classification, existing approaches commonly rely on a single feature view, such as side-channel characteristics or raw packet bytes, often incorporating techniques inspired by natural language processing and computer vision for modeling and classification. In recent years, multi-view [...] Read more.
For the task of encrypted traffic classification, existing approaches commonly rely on a single feature view, such as side-channel characteristics or raw packet bytes, often incorporating techniques inspired by natural language processing and computer vision for modeling and classification. In recent years, multi-view learning has gained increasing attention due to its ability to enhance discriminative power and generalization performance by capturing complementary information from different perspectives. However, heterogeneous feature views often exhibit distributional discrepancies, which makes direct multi-view integration difficult. To address this issue, we propose CGF-Net, a multi-view contrastive learning framework for encrypted traffic classification. The proposed model is inspired by cross-modal contrastive learning and employs lightweight adaptation to learn representations from both behavioral and content views. During pre-training, CGF-Net performs instance-level cross-view contrastive learning by treating the behavioral and content views of the same network flow as a positive pair, thereby aligning heterogeneous representations at the flow-instance level. In addition, a lightweight fine-tuning module together with a gating-based fusion mechanism is introduced to improve the collaborative modeling capability of multi-view representations. Extensive experiments on four public datasets show that CGF-Net achieves ACC scores of 95.81%, 93.38%, 96.38%, and 95.72% on CSTNET-TLS1.3, CipherSpectrum, ISCXVPN2016, and ISCXTor2016, respectively. Compared with the best-performing baseline on each dataset, CGF-Net improves the average ACC and F1-score by 0.78 and 0.69 percentage points, respectively, demonstrating the effectiveness of the proposed model. Full article
(This article belongs to the Section Networks)
Show Figures

Figure 1

Back to TopTop