Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (1,610)

Search Parameters:
Keywords = multi-modality data fusion

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
37 pages, 59861 KB  
Article
MC-OIWQR: Multimodal Contrastive Learning for Optically Inactive Water Quality Retrieval
by Weixuan Li, Fangling Pu, Jiehao Xue, Yue Dai, Lin Cong and Xin Xu
Remote Sens. 2026, 18(16), 2736; https://doi.org/10.3390/rs18162736 - 14 Aug 2026
Abstract
Retrieving optically inactive nutrients from satellite observations remains challenging because TN, TP, and NH3-N are only indirectly linked to water reflectance and may respond to different environmental contexts. Here, we propose MC-OIWQR, a multimodal framework that combines spatiotemporal contrastive learning from [...] Read more.
Retrieving optically inactive nutrients from satellite observations remains challenging because TN, TP, and NH3-N are only indirectly linked to water reflectance and may respond to different environmental contexts. Here, we propose MC-OIWQR, a multimodal framework that combines spatiotemporal contrastive learning from unlabeled HLS Sentinel-2 imagery with meteorological, land-use, and nighttime-light information through cross-attention fusion. Evaluated on long-term in situ observations from Lake Ontario, the framework was analyzed through modality ablation and SHAP-based attribution. Across 10 repeated stratified data partitions, MC-OIWQR achieved mean test R2 values of 0.9208, 0.8663, and 0.9409 for TN, TP, and NH3-N, respectively, obtaining the highest mean R2 and lowest mean RMSE among the evaluated baselines. SHAP-based analysis suggested that TN predictions were associated with land-use and nighttime-light proxies of watershed anthropogenic activity, TP predictions with hydrometeorological forcing related to precipitation and wind, and NH3-N predictions with multivariate environmental context, indicating parameter-specific attribution patterns in the trained model. Long-term retrieval maps from 2016 to 2025 revealed persistent nearshore–offshore nutrient gradients and event-driven variability in Quinte Bay. Cross-lake experiments on Lake Huron and Lake Erie provided preliminary evidence that MC-OIWQR may be adapted through lake-specific re-pretraining for TN and TP retrieval. These results suggest that optically inactive nutrient retrieval benefits from parameter-specific multimodal information rather than a uniform optical regression strategy. Full article
(This article belongs to the Section Remote Sensing in Geology, Geomorphology and Hydrology)
Show Figures

Figure 1

27 pages, 13812 KB  
Article
Multimodal Data Fusion for a Self-Adaptive, Smart, Serious-Game Ecosystem Under Development
by Xiya Tao, Peng Chen and Martina Eckert
Sensors 2026, 26(16), 5132; https://doi.org/10.3390/s26165132 - 13 Aug 2026
Abstract
This article presents the implementation and technical feasibility evaluation of the multimodal sensing and feature-level fusion layer of BLEXER v3, a broader serious-game ecosystem under development for upper-limb rehabilitation. The implemented framework integrates Kinect-based motion tracking, Polar H10 and Bangle.js physiological sensing, wearable [...] Read more.
This article presents the implementation and technical feasibility evaluation of the multimodal sensing and feature-level fusion layer of BLEXER v3, a broader serious-game ecosystem under development for upper-limb rehabilitation. The implemented framework integrates Kinect-based motion tracking, Polar H10 and Bangle.js physiological sensing, wearable accelerometer data, and facial affective cues within a middleware-based architecture. Heterogeneous sensor streams are locally preprocessed, temporally aligned, and transformed into a common quality-aware multimodal feature representation containing motion, heart-rate and heart-rate-variability-related descriptors, affective information, availability indicators, signal-quality metadata, and freshness descriptors. The fused representation is additionally mapped, using predefined rules, to heuristic operational descriptors, including low demand, moderate stable, active engagement, physical load, affective activation, high demand, and uncertain. These descriptors are not intended as clinical diagnoses, independently validated user states, or final adaptation decisions. An exploratory K-means analysis of 12,675 complete multimodal windows reveals partial correspondence between the data-driven cluster structure and the predefined operational descriptors. Some descriptors show comparatively concentrated cluster patterns, whereas others exhibit overlap and internal heterogeneity. The results demonstrate the technical feasibility of generating structured multimodal representations that can provide input for subsequent context-aware reasoning. Independent validation of the operational descriptors, completion and evaluation of the whole system, and clinical validation with rehabilitation patients remain future work. Full article
(This article belongs to the Special Issue Smart Sensing System for Intelligent Human–Computer Interaction)
26 pages, 3594 KB  
Article
Master Mix Localization Algorithm for Autonomous Systems in Indoor Environments
by Zakaryae Ezzouine, Adil Salbi, Mohamed Abouzahir, Ilham Elmourabit, Adil Brouri and Sébastien Roy
Entropy 2026, 28(8), 903; https://doi.org/10.3390/e28080903 - 12 Aug 2026
Abstract
Reliable navigation in GPS-denied environments remains a critical challenge for autonomous vehicles (AVs), particularly in complex indoor and urban settings. GPS-based localization systems often fail under these conditions, highlighting the need for resilient multimodal solutions. In this article, we present a radar-assisted tracking [...] Read more.
Reliable navigation in GPS-denied environments remains a critical challenge for autonomous vehicles (AVs), particularly in complex indoor and urban settings. GPS-based localization systems often fail under these conditions, highlighting the need for resilient multimodal solutions. In this article, we present a radar-assisted tracking system that integrates LiDAR and inertial measurements within a sensor-fusion architecture to achieve robust navigation. The principal methodological contribution is a unified tracking and prediction framework that combines Bayesian state estimation with learning-based temporal prediction, enabling accurate tracking while continuously forecasting the slave robot’s short-term future state from mapping observations generated by the master robot, with a typical end-to-end perception-to-action latency of 20–60 ms. The communication and prediction forecasting module operates with an update interval below 35 ms, enabling real-time cooperative robotic operation. Sensor data are fused through a pipeline incorporating Gaussian Mixture Models (GMMs) for post-processing, which helps mitigate the limitations associated with individual sensors during edge processing. Moreover, Kalman filtering is employed to mitigate sensor noise and drift, thereby improving state estimation accuracy through trajectory smoothing. The fused spatiotemporal information is subsequently exploited by a Convolutional Recurrent Neural Network (CRNN) coupled with a Nonlinear Autoregressive model with eXogenous Inputs (NARX) to model the robot’s motion dynamics and provide short-horizon state prediction. Through simulations and real-world indoor experiments conducted in GPS-denied environments, we validate the system’s ability to provide accurate and continuous pose estimation with low localization errors. Experimental results show that the proposed framework achieves root-mean-square errors of 0.12 m, 0.15 m, and 0.28 m along the X, Y, and Z axes, respectively, while maintaining sub-meter maximum position deviations throughout the evaluated trajectories. These results confirm that the proposed framework provides reliable localization and predictive state estimation for cooperative robotic navigation in indoor GPS-denied environments. Future work will investigate outdoor validation and extend the framework to additional data-driven decision-making models for future robotic services. Full article
(This article belongs to the Special Issue Topics from the 2025 Biennial Symposium on Communications)
Show Figures

Figure 1

19 pages, 753 KB  
Article
Breaking the Sign Symmetry of Attention: A Conflict-Aware Vision–Language Fusion Framework for Privacy-Sensitive Information Detection
by Ming Lian, Yuanyuan Li and Teng Li
Symmetry 2026, 18(8), 1352; https://doi.org/10.3390/sym18081352 - 11 Aug 2026
Viewed by 90
Abstract
As image data are shared ever more openly, they increasingly leak privacy-sensitive information, yet existing detectors seldom model how the text embedded in an image relates to its visual content and tend to fail precisely when the two modalities disagree. We note that [...] Read more.
As image data are shared ever more openly, they increasingly leak privacy-sensitive information, yet existing detectors seldom model how the text embedded in an image relates to its visual content and tend to fail precisely when the two modalities disagree. We note that the conventional softmax attention used for multimodal fusion carries an implicit sign symmetry: Every source token contributes only additively, so conflicting evidence is averaged away rather than resolved. We propose a symmetric dual-source fusion framework whose decoder deliberately breaks this sign symmetry. Image and text are first encoded by a Swin Transformer and a policy knowledge-enhanced BERT (KL-BERT) and projected into a common space to form a permutation-symmetric dual-source memory. A Signed Cross-attention Auto-compressing Decoder (SCAD) then fuses the two sources through an attention map that factorizes into a sign-symmetric (even) magnitude term and a sign-antisymmetric (odd) polarity term, allowing the model to either reinforce or actively subtract cross-modal evidence. Experiments on a self-constructed privacy-sensitive image dataset show that the proposed method attains an accuracy of 97.69% and an F1-score of 96.70%, with its largest gains on the face category, the most conflict-prone class in our data, suggesting that controlled symmetry breaking is an effective principle for cross-modal fusion. Full article
(This article belongs to the Special Issue Symmetry in Fault Diagnosis: Methods, Models, and Applications)
Show Figures

Figure 1

27 pages, 25544 KB  
Article
AOPQ-Net Acoustic–Optical Proposal Query Network for Underwater Multimodal Object Detection
by Yanze Lu, Zhengyan Zhang, Shuoshuo Ding, Haochen Hu, Chih-Yung Wen and Tiedong Zhang
Remote Sens. 2026, 18(16), 2703; https://doi.org/10.3390/rs18162703 - 11 Aug 2026
Viewed by 82
Abstract
Optical cameras and imaging sonars are widely used sensors in autonomous underwater vehicles. However, their different imaging mechanisms introduce substantial cross-modal discrepancies in the acquired data. In addition, underwater optical images are often degraded by low illumination, scattering, and turbidity, whereas sonar images [...] Read more.
Optical cameras and imaging sonars are widely used sensors in autonomous underwater vehicles. However, their different imaging mechanisms introduce substantial cross-modal discrepancies in the acquired data. In addition, underwater optical images are often degraded by low illumination, scattering, and turbidity, whereas sonar images commonly suffer from speckle noise and low spatial resolution. As a result, object detection based on a single optical or acoustic modality is often insufficient in challenging underwater environments. To address this problem, this paper proposes an acoustic–optical fusion network for underwater object detection, termed an Acoustic–Optical Proposal Query Network (AOPQ-Net). First, a Sonar Position Encoding (SPE) module is designed to explicitly encode the geometric priors in sonar images. Second, a Bi-directional Discrepancy-aware Spatial Alignment (BDSA) module is introduced to alleviate spatial misalignment between the two modalities at the feature level. Third, a Proposal Query Transformer (PQT) module performs target-oriented cross-modal interaction at the proposal level. Furthermore, this study constructs a dedicated dataset for underwater acoustic–optical fusion object detection, named Haiqin Underwater Fusion (HUF), and conducts systematic experiments on this dataset. The experimental results show that AOPQ-Net outperforms single-modality baselines and representative multimodal fusion methods in both optical and acoustic image spaces, which demonstrate the effectiveness of the proposed method. Full article
Show Figures

Figure 1

19 pages, 161996 KB  
Article
DBCS-T: A Dual-Branch Cross-Attention Synergistic Transformer for Multimodal Image Fusion and Semantic Segmentation
by Yiming Liu, Bo Gao, Xiao Yang, Hang Li, Weixing Yu and Huangrong Xu
Remote Sens. 2026, 18(16), 2700; https://doi.org/10.3390/rs18162700 - 11 Aug 2026
Viewed by 71
Abstract
Spectro-polarimetric imaging systems can simultaneously acquire spatial, spectral, and polarimetric information during remote sensing, yet the multimodal fusion data are often constrained in practical applications by insufficient exploitation of complementary information across different modalities. To address this issue, we propose a multimodal image [...] Read more.
Spectro-polarimetric imaging systems can simultaneously acquire spatial, spectral, and polarimetric information during remote sensing, yet the multimodal fusion data are often constrained in practical applications by insufficient exploitation of complementary information across different modalities. To address this issue, we propose a multimodal image fusion method based on a Dual-Branch Cross-Attention Synergistic Transformer (DBCS-T). In our method, three complementary feature components are independently extracted, i.e., a Characteristic Polarization Image (CPI), a Characteristic Spectral Image (CSI) and a Characteristic Intensity Image (CII). For CPI, it is derived from Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) inputs via DBCS-T, which integrates a Cross-Channel Transposed Attention (CCTA) module for cross-modal interaction and a Multi-Scale Polarization Feature Adaptive Modulation (MPAM) module for local feature enhancement. CSI is obtained by leveraging maximum-divergence spectral band differences guided by prior spectral radiance curves. CII is computed from the Stokes parameter S0. These three components are then fused via Principal Component Analysis (PCA) into a Multimodal Fusion Image (MFI). To demonstrate the effectiveness of our method, an experiment was conducted on a scene that contains real vegetation, artificial foliage, and same-color metallic objects. The experimental results show that the proposed method achieves effective semantic segmentation of all three target categories. Furthermore, quantitative evaluation demonstrates that the MFI attains the lowest Kullback–Leibler (KL) and Jensen–Shannon (JS) Divergence values among all evaluated modalities, with image entropy exceeding that of individual source inputs. These results validate the complementarity of the extracted multimodal features, significantly enhance the interpretation performance for complex scenes, and demonstrate the broad application potential of the proposed fusion framework in remote sensing and multidimensional imaging. Full article
(This article belongs to the Section Remote Sensing Image Processing)
Show Figures

Figure 1

32 pages, 3945 KB  
Article
A Real-Time Edge-Enabled IoT Framework with Federated Differential Privacy for Multi-Modal Crowd Monitoring in Mega Events Using SmartCrowd IoT
by Saleh Alharbi
Electronics 2026, 15(16), 3570; https://doi.org/10.3390/electronics15163570 - 11 Aug 2026
Viewed by 99
Abstract
Mega-events present acute challenges in crowd safety, requiring sub-second monitoring, heterogeneous sensing, and strict privacy compliance at scale. We present SmartCrowd-IoT, a multi-modal crowd analytics framework built on a three-tier (sensor, edge, coordination) architecture incorporating (i) temporally aligned, reliability-aware weighted fusion across RGB, [...] Read more.
Mega-events present acute challenges in crowd safety, requiring sub-second monitoring, heterogeneous sensing, and strict privacy compliance at scale. We present SmartCrowd-IoT, a multi-modal crowd analytics framework built on a three-tier (sensor, edge, coordination) architecture incorporating (i) temporally aligned, reliability-aware weighted fusion across RGB, thermal, WiFi/BLE, acoustic, and RFID streams; and (ii) lightweight edge inference with federated differential privacy, enabling continuous model improvement without raw data leaving the venue. Evaluated on PETS2009, UCY, Mall, and a custom 61.3-h multi-modal corpus across three controlled mega-event simulations, SmartCrowd-IoT achieves 92.6% crowd-density accuracy, 77 ms end-to-end latency, 92.9% anomaly detection precision, and 83.4% backbone bandwidth reduction. Ablation studies confirm that both temporal alignment and reliability-aware fusion contribute significantly to these gains. The framework provides a deployable, privacy-by-design solution for mega-event crowd safety that scales to 200 edge nodes and 3000 sensors while maintaining sub-100 ms emergency response. Full article
Show Figures

Figure 1

28 pages, 2988 KB  
Article
Structure-Aware Heterogeneous Dual-Stream Network with Wavelet-Guided Fusion for UAV Infrared–Visible Object Detection
by Weijian Jia, Fenghua Wang, Haiwen Zheng, Penglei Hu, Xiaobing Wang, Yao Zhao, Pengdong Zhang and Yufei Gao
Drones 2026, 10(8), 614; https://doi.org/10.3390/drones10080614 - 11 Aug 2026
Viewed by 129
Abstract
To address the susceptibility of single-modality approaches to illumination variations and imaging conditions in low-altitude unmanned aerial vehicle (UAV) detection, this study proposes a structure-aware heterogeneous dual-stream detection network based on infrared and visible-light fusion. First, a multimodal UAV detection dataset oriented toward [...] Read more.
To address the susceptibility of single-modality approaches to illumination variations and imaging conditions in low-altitude unmanned aerial vehicle (UAV) detection, this study proposes a structure-aware heterogeneous dual-stream detection network based on infrared and visible-light fusion. First, a multimodal UAV detection dataset oriented toward complex low-altitude scenarios is constructed, providing a data foundation for cross-modal detection research. Then, a structure-aware heterogeneous dual-stream feature extraction framework is designed to enable collaborative modeling of visible-light and infrared features through modality-specific encoding. In the visible-light branch, a Structure-Aware Gated Enhancement Block (SAGE Block) is introduced to enhance the representation of fine-grained structural and edge information. In the cross-modal fusion stage, a Bidirectional Wavelet-Guided Fusion Module (BWFM) is proposed to decouple structural semantics and detailed information in the frequency domain. Adaptive fusion is further achieved through low-frequency cross-modal interaction and high-frequency detail-preservation strategies. Finally, the proposed method is experimentally validated on the proposed Multispectral UAV Detection Dataset (MUDD) and the Multi-scenario Multi-Modality Fusion Dataset (M3FD).. The experimental results show that the proposed method achieves an mAP@0.5 of 0.9680 and an mAP@0.5:0.95 of 0.6794 on the proposed MUDD, as well as an mAP@0.5:0.95 of 0.6072 on the M3FD dataset, demonstrating competitive detection accuracy and generalization capability. Ablation experiments further indicate that the SAGE Block, BWFM, and the low-frequency cross-modal fusion and high-frequency detail-preservation strategies within BWFM all contribute positively to performance improvement. Full article
Show Figures

Figure 1

30 pages, 2943 KB  
Article
A Quality-Aware Multimodal Reliability Framework for Health Assessment and Remaining Useful Life Prediction of Cold-Region Tunnels
by Boyang Liu, Jing Guan, Yi Yang and Wuer Ha
Infrastructures 2026, 11(8), 283; https://doi.org/10.3390/infrastructures11080283 - 10 Aug 2026
Viewed by 147
Abstract
This study proposes a quality-aware multimodal framework for health-state assessment and remaining useful life (RUL) prediction of cold-region tunnels. The framework integrates structural-response, environmental, apparent-defect, and engineering-inspectiondata, with the apparent-defect pathway jointly encoding raw images through a convolutional neural network and structured defect [...] Read more.
This study proposes a quality-aware multimodal framework for health-state assessment and remaining useful life (RUL) prediction of cold-region tunnels. The framework integrates structural-response, environmental, apparent-defect, and engineering-inspectiondata, with the apparent-defect pathway jointly encoding raw images through a convolutional neural network and structured defect variables. Five data-quality dimensions-completeness, accuracy, consistency, timeliness, and traceability are incorporated intoreliability-guided multimodal fusion. Their base weights were re-audited through two rounds of expert consultation, each comprising 323 valid questionnaires. The Cr-weighted group analytic hierarchy process yielded weights of 0.0548, 0.1326, 0.1372, 0.2279, and 0.4474, respectively, with a group consistency ratio of 0.0455; the ranking remained stable under one-at-a-time +10% perturbations. In the primary tunnel case study, the framework achieved 89.7% health-state accuracy, a 6.3% RUL mean absolute percentage error, and 84.1% accuracy under Gaussian perturbation of standardized numerical inputs at a noise scale of 0.15. To further examine the reliability contribution of data-quality information, an independent field panel comprising 600 segment-month observations from 25 segments across three operational tunnels was evaluated using target-excluded specifications, two-way fixed effects, leave-one-tunnel-out validation, multiple baseline models, and five fixed random seeds. A one-standard-deviation increase in lagged quality instability was associated with a 0.0151 increase in the subsequent state-error index (95% CI: 0.0118-0.0184; p < 0.001). In cross-tunnel random-forest tests, incorporating quality information increased mean R2 from 0.8277 to 0.8323 for state-error prediction and from 0.8517 to 0.8673 for RUL-contraction prediction, with both improvements significant in paired tests (p < 0.001). Split-conformal intervals achieved mean cross-tunnel coverage of 95.8% and 95.9%, respectively. These findings demonstrate that data-quality information provides a modest but statistically supported improvement in cross-tunnel reliability, whilethe principal contribution lies in integrating auditable data governance, reliability-aware fusion, and engineering decision support within a unified tunnel health-management framework. Full article
Show Figures

Graphical abstract

27 pages, 8086 KB  
Article
Small-Sample Motor Fault Identification via Fusion of Fixed-Resolution and Multiscale Time–Frequency Features
by Jingyu Yang, Jikai Xu, Li Peng, Longfu Luo, Wanting Li and Hengrui Ma
Machines 2026, 14(8), 916; https://doi.org/10.3390/machines14080916 - 10 Aug 2026
Viewed by 147
Abstract
Motor fault identification is often constrained by scarce labeled samples and the limited representation capability of a single time–frequency transform. Conventional CNN–Softmax models may also produce unstable decision boundaries under small-sample conditions. To address these issues, this paper proposes a motor fault identification [...] Read more.
Motor fault identification is often constrained by scarce labeled samples and the limited representation capability of a single time–frequency transform. Conventional CNN–Softmax models may also produce unstable decision boundaries under small-sample conditions. To address these issues, this paper proposes a motor fault identification method based on the fusion of fixed-resolution and multiscale time–frequency features. Each vibration segment is transformed into short-time Fourier transform (STFT) and synchrosqueezed wavelet transform (SWT) maps. Two parallel convolutional branches extract complementary features, which are fused by element-wise addition and classified using a radial basis function support vector machine. Experiments on the HUST motor multimodal fault dataset show that the proposed method achieves 100% accuracy under the conventional 70%/30% train–test split. When the training proportion is reduced to 20%, 15%, 10%, and 5%, the corresponding accuracies remain at 99.46%, 99.10%, 98.78%, and 96.77%, respectively. Across operating speeds of 5, 10, 20, and 30 Hz, the average accuracies reach 98.75% and 94.61% under the 20% and 5% training conditions. The model also maintains 100% accuracy at signal-to-noise ratios of 15 dB and above. These results demonstrate that complementary time–frequency feature fusion combined with maximum-margin classification improves identification accuracy and decision-boundary stability under limited training data. Full article
(This article belongs to the Section Electrical Machines and Drives)
Show Figures

Figure 1

18 pages, 1133 KB  
Article
Bimodal Speech Emotion Recognition Using a Hybrid CNN-LSTM Architecture with Sentiment Fusion
by Tze-Syn Yap and Lee-Yeng Ong
Future Internet 2026, 18(8), 421; https://doi.org/10.3390/fi18080421 - 10 Aug 2026
Viewed by 130
Abstract
Speech emotion recognition (SER) is a fundamental task in affective computing; however, traditional unimodal approaches often struggle to capture the complex emotional cues present in spontaneous conversational speech. Bimodal frameworks that integrate acoustic and textual information have therefore emerged to provide complementary semantic [...] Read more.
Speech emotion recognition (SER) is a fundamental task in affective computing; however, traditional unimodal approaches often struggle to capture the complex emotional cues present in spontaneous conversational speech. Bimodal frameworks that integrate acoustic and textual information have therefore emerged to provide complementary semantic and acoustic representations. This study proposes a bimodal SER framework based on a hybrid convolutional neural network–long short-term memory (CNN–LSTM) architecture. Using the Multimodal EmotionLines Dataset (MELD), the framework combines temporal acoustic features, statistical acoustic features, and predicted textual sentiment. Experimental results indicate that the proposed model achieves reliable recognition of majority emotion classes but exhibits limited performance on underrepresented minority classes due to severe class imbalance. To better understand the contribution of each modality, feature sufficiency and feature necessity analyses were conducted. Furthermore, an evaluation of alternative fusion strategies showed that the expressive attention networks did not provide meaningful performance improvements over simple feature concatenation. These findings suggest that class imbalance, rather than fusion complexity, remains the primary limitation in conversational SER, highlighting the importance of addressing data imbalance before pursuing more sophisticated multimodal architectures. Full article
(This article belongs to the Special Issue Artificial Intelligence (AI) and Natural Language Processing (NLP))
Show Figures

Graphical abstract

18 pages, 2606 KB  
Article
Multimodal Fusion Interpolation Method for Missing Monitoring Data in Deep-Buried Tunnel Rockburst Prediction
by Xianfeng Duan and Jianxi Wang
Buildings 2026, 16(16), 3156; https://doi.org/10.3390/buildings16163156 - 8 Aug 2026
Viewed by 191
Abstract
Missing monitoring data reduce the reliability of rockburst early warning in deep-buried tunnel engineering. This study proposes a multimodal fusion interpolation framework that combines LSTM-based temporal estimation, Pearson’s/Spearman’s/MIC correlation analysis, and Whale Optimization Algorithm (WOA)-based weight allocation for missing monitoring indicators. A case [...] Read more.
Missing monitoring data reduce the reliability of rockburst early warning in deep-buried tunnel engineering. This study proposes a multimodal fusion interpolation framework that combines LSTM-based temporal estimation, Pearson’s/Spearman’s/MIC correlation analysis, and Whale Optimization Algorithm (WOA)-based weight allocation for missing monitoring indicators. A case dataset containing 251 consecutive samples from one tunnel in southwestern China was used to examine the engineering feasibility of the method. The interpolated data were further used in GRU, CNN-LSTM, N-BEATS, and deep fully connected prediction models. The results indicate that the interpolated dataset improves downstream rockburst prediction in this case study. Because the dataset is limited to one tunnel project, the conclusions are now restricted to similar deep-buried tunnel conditions and should not be interpreted as universal proof of superiority. Full article
Show Figures

Figure 1

34 pages, 7253 KB  
Review
From Multisensor Fusion to Intelligent Geospatial Monitoring: Emerging Architectures for Geotechnical Hazard Assessment
by Meghdad Bagheri, Thalosang Tshireletso and Seyed Ali Ghorashi
Remote Sens. 2026, 18(16), 2669; https://doi.org/10.3390/rs18162669 - 8 Aug 2026
Viewed by 258
Abstract
Geotechnical hazards such as landslides, subsidence, slope instability, and infrastructure deformation threaten rapidly urbanising and environmentally stressed regions worldwide, intensifying the need for scalable and intelligent monitoring systems capable of continuously observing complex Earth surface dynamics. Although multisensor remote sensing fusion has substantially [...] Read more.
Geotechnical hazards such as landslides, subsidence, slope instability, and infrastructure deformation threaten rapidly urbanising and environmentally stressed regions worldwide, intensifying the need for scalable and intelligent monitoring systems capable of continuously observing complex Earth surface dynamics. Although multisensor remote sensing fusion has substantially expanded the observational capabilities of modern geotechnical monitoring through the integration of Synthetic Aperture Radar (SAR), optical imagery, Light Detection and Ranging (LiDAR), and environmental data, existing fusion pipelines remain subject to several well-documented constraints, including weak semantic alignment, limited temporal reasoning, and poor transferability across heterogeneous environmental conditions. This review synthesises the emerging transition from conventional sensor-centric fusion toward intelligent geospatial monitoring architectures centred on deep multimodal representation learning, transformer-based temporal reasoning, self-supervised learning, and geospatial foundation models. Particular emphasis is placed on how recent architectures are designed to better preserve coherent spatial, temporal, and contextual environmental relationships within unified latent representation spaces rather than through downstream handcrafted integration. The review further examines the growing role of multimodal transformers, masked autoencoders, contrastive learning, and large-scale geospatial foundation models in enabling scalable environmental reasoning, adaptive multimodal learning, and transferable geospatial intelligence across sensing modalities and geographic domains. Finally, remaining challenges involving uncertainty, explainability, computational scalability, and environmental generalisation are discussed alongside future research directions involving continual learning, physics-aware artificial intelligence, and autonomous geotechnical monitoring systems. Together, the reviewed literature suggests that multimodal Earth observation is evolving from passive environmental sensing toward adaptive geospatial intelligence systems capable of scalable hazard reasoning and autonomous environmental understanding. Full article
(This article belongs to the Section Engineering Remote Sensing)
Show Figures

Figure 1

32 pages, 847 KB  
Review
A Review of Adversarial Example Detection in IoT Sensor Networks: Methods, Evaluation, and Edge Deployment Constraints
by Wenqiang Xu and Jian Li
Sensors 2026, 26(16), 5044; https://doi.org/10.3390/s26165044 - 8 Aug 2026
Viewed by 132
Abstract
Deep learning has been widely deployed in critical scenarios such as the Internet of Things (IoT), industrial sensing, network intrusion detection, and cyber-physical system monitoring, where model inference directly affects system security, operational reliability, and service continuity. However, existing adversarial example detection studies [...] Read more.
Deep learning has been widely deployed in critical scenarios such as the Internet of Things (IoT), industrial sensing, network intrusion detection, and cyber-physical system monitoring, where model inference directly affects system security, operational reliability, and service continuity. However, existing adversarial example detection studies remain insufficient for practical IoT deployment, as their validation often overlooks endpoint resource constraints, heterogeneous data modalities, physical environmental interference, communication protocol specifications, adaptive attacks, and adversary capability models. Moreover, detection outcomes are rarely connected with deployment locations, computational overhead, formal security assurance, and subsequent response strategies, which limits their engineering applicability. To address these limitations, this review systematically synthesizes recent representative studies in adversarial example detection and constructs a unified analytical framework integrating detection evidence, IoT deployment feasibility, and adaptive-attack evaluation. Based on the source of detection evidence, existing methods are categorized into input-consistency-based, feature-statistics-based, predictive-uncertainty-based, model-reconstruction-based, runtime-context-aware, and multi-strategy fusion detection, while formal certification is discussed as an independent security-assurance dimension. The review further analyzes the principles, applicable conditions, limitations, compatibility conflicts with IoT deployment constraints, and typical failure modes of these methods. The analysis identifies four key challenges: the lack of IoT-native adaptive evaluation, limited anomaly-boundary identification and cross-modal generalization, insufficient deployment-time security assurance, and weak coordination between detection decisions and security responses. Future research should therefore emphasize feasible attack paradigms, hierarchical lightweight detection, reliable multimodal fusion, certifiable operational boundaries, and auditable end-to-end response mechanisms, thereby supporting the evaluation and deployment of adversarial example detection in IoT scenarios. Full article
Show Figures

Figure 1

23 pages, 6534 KB  
Article
State of Health Estimation of Large-Capacity Energy Storage Batteries Based on Mechanical–Electrical–Thermal Multi-Modal Features
by Rong He, Jiang He, Lu Wang, Meng Wei and Sijia Yang
Batteries 2026, 12(8), 295; https://doi.org/10.3390/batteries12080295 - 8 Aug 2026
Viewed by 187
Abstract
This paper proposes an SOH estimation method that fuses mechanical–electrical–thermal multi-modal features by introducing expansion force monitoring. Aging tests on 16 prismatic 530 Ah LiFePO4 batteries from two brands are conducted at 25 and 45 °C. Each full cycle is divided into [...] Read more.
This paper proposes an SOH estimation method that fuses mechanical–electrical–thermal multi-modal features by introducing expansion force monitoring. Aging tests on 16 prismatic 530 Ah LiFePO4 batteries from two brands are conducted at 25 and 45 °C. Each full cycle is divided into charge, post-charge rest, discharge, and post-discharge rest, with SOH defined by the capacity ratio. From cycle-level data, 37 candidate features are extracted and cleaned using local median and median absolute deviation. Using only training cells, Spearman correlation eliminates highly redundant features, and 12 key features are retained via internal validation. Under 4-fold cross-validation with complete battery grouping, Random Forest, XGBoost, LightGBM, and LSTM are compared. LightGBM achieves the best performance with an average MAE of 0.0032, RMSE of 0.0037, and R2 of 91.36%. Ablation shows multi-modal fusion outperforms single-type features; five-category fused features reduce RMSE by ~75.57% versus electrical-only features. Removing expansion force features increases RMSE to 0.0064 and drops R2 to 75.66%. These findings confirm that expansion force supplies critical mechanical degradation information, significantly improving SOH estimation for large-capacity energy storage batteries. Full article
(This article belongs to the Section Energy Storage System Aging, Diagnosis and Safety)
Show Figures

Graphical abstract

Back to TopTop