Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (23)

Search Parameters:
Keywords = efficient multimodal late fusion

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
24 pages, 5385 KB  
Article
A Late-Fusion Multimodal Approach for Safety-Aware Workspace Modeling in Collaborative Robotic Systems
by Kevin David Ortega-Quiñones, Elias Escobar-Pereira, Michael Felipe Cifuentes-Molano, Germán Andrés Holguín-Londoño and Mauricio Holguín-Londoño
Robotics 2026, 15(7), 127; https://doi.org/10.3390/robotics15070127 - 30 Jun 2026
Viewed by 279
Abstract
Ensuring safe coexistence between human operators and industrial robot manipulators is a critical challenge in collaborative manufacturing environments. Existing approaches rely either on dedicated safety-rated hardware, which is expensive and difficult to retrofit, or on purely vision-based classifiers that discard the precise kinematic [...] Read more.
Ensuring safe coexistence between human operators and industrial robot manipulators is a critical challenge in collaborative manufacturing environments. Existing approaches rely either on dedicated safety-rated hardware, which is expensive and difficult to retrofit, or on purely vision-based classifiers that discard the precise kinematic state available from the robot controller, leading to unresolved visual ambiguities when different joint configurations produce similar appearances from fixed camera viewpoints. Kinematics-only approaches, while precise, lack the spatial context needed to disambiguate configurations near workspace boundaries. We propose RGBJointsNet, a late-fusion multimodal deep learning classifier that combines RGB visual features extracted by a frozen EfficientNet-B2 convolutional backbone with a compact kinematic stream processing the 12-dimensional joint angle vector of a dual-UR5 robotic cell. The model maps each observation to one of five mutually exclusive workspace zones: rest (C0), nominal (C1), extended (C2), shared/collision-risk (C3), and joint-limit/singularity (C4). A dedicated simulation environment built on ROS 2 Humble Hawksbill and Gazebo Classic 11 was used to generate a labelled dataset of 54,309 frames and 162,927 RGB images from three calibrated overhead cameras, with analytic ground-truth labels derived from closed-form forward kinematics. Training on a CPU with a feature-caching strategy brings the per-epoch wall-clock time to seconds, making the approach tractable without GPU hardware. On the held-out test set, the model achieves 87.1% overall accuracy and a macro-averaged F1 score of 90.0%, with near-perfect recall of 99.3% for the safety-critical shared zone C3. The trained classifier is integrated as an ROS 2 inference node capable of running at 10 Hz on a standard workstation. Our results demonstrate that joint angle information is a decisive complement to RGB imagery for fine-grained, safety-oriented workspace classification in simulation-derived settings. Full article
Show Figures

Graphical abstract

22 pages, 1755 KB  
Article
TriDA: Privacy-Aware and Efficient Multimodal AI for Disaster Assessment
by Md Abdullahil Oaphy, Adeel Khalid, Da Hu and Honghui Xu
Mathematics 2026, 14(12), 2064; https://doi.org/10.3390/math14122064 - 10 Jun 2026
Viewed by 354
Abstract
As disaster imagery and social media reports become vital for crisis response, automated assessment systems must address challenges of multimodal integration, privacy-aware learning, and computational efficiency. To address these challenges, we propose TriDA, a privacy-aware and efficiency-conscious multimodal disaster classification framework that fuses [...] Read more.
As disaster imagery and social media reports become vital for crisis response, automated assessment systems must address challenges of multimodal integration, privacy-aware learning, and computational efficiency. To address these challenges, we propose TriDA, a privacy-aware and efficiency-conscious multimodal disaster classification framework that fuses image features with text representations through a late-fusion design. A classifier-head DP-SGD stage is used to report training-record-level differential privacy accounting for paired image–text samples under the stated private optimization protocol. To study efficiency-oriented simplification, structured neuron pruning reduces redundant capacity in the classification head while preserving predictive utility. Experiments on the multimodal damage identification dataset show that TriDA maintains strong classification performance, exhibits a controlled privacy–utility trade-off under increasing DP noise, and achieves quantifiable classifier-head parameter and MAC reductions through pruning. These findings position TriDA as a controlled empirical framework for privacy-aware and resource-conscious multimodal disaster assessment. Full article
Show Figures

Figure 1

49 pages, 2508 KB  
Review
Sensing the Action: Rethinking Sensor Modalities and Multi-Modal Fusion in Vision–Language–Action Models for Robotic Manipulation
by Byoung Chul Ko
Sensors 2026, 26(11), 3541; https://doi.org/10.3390/s26113541 - 3 Jun 2026
Viewed by 1059
Abstract
Recent Vision–Language–Action (VLA) models have rapidly emerged as general-purpose robotic policies that integrate language understanding, visual perception, and robot control. However, prior studies and surveys have primarily emphasized backbone architectures, action decoders, training recipes, and benchmark performance, whereas relatively limited systematic attention has [...] Read more.
Recent Vision–Language–Action (VLA) models have rapidly emerged as general-purpose robotic policies that integrate language understanding, visual perception, and robot control. However, prior studies and surveys have primarily emphasized backbone architectures, action decoders, training recipes, and benchmark performance, whereas relatively limited systematic attention has been given to sensor modality selection, heterogeneous signal alignment and fusion, and their connection to action generation, all of which are critical to the performance and safety of real-world robotic manipulation. This survey addresses this gap by reinterpreting VLA within the framework of a sensor–fusion–action pipeline. This study first presents a systematic taxonomy of major sensor modalities, including RGB, depth, tactile sensing, force/torque, proprioception and inertial measurement unit, multi-spectral/thermal, and event-based vision, and compares them in terms of the physical information they provide, their characteristic failure modes, and their deployment constraints. This survey further reviews teleoperation-, human video-, and simulation-based data collection pipelines, together with representative dataset configurations, and analyzes the multi-modal design space from a sensor-centric perspective, including early and late fusion, cross-attention, token-level fusion, adapters, mixture of experts, and multi-rate action representations. In addition, this study identifies a strong bias in existing benchmarks toward RGB-centric inputs and single success-rate metrics and emphasizes the need for a multidimensional evaluation framework incorporating robustness, worst-case performance, safety, latency, and efficiency. By shifting the focus away from a model-centric narrative and explicitly accounting for real-world sensor complexity, this survey seeks to establish a sensor-centered foundation for the next generation of Physical AI. Full article
(This article belongs to the Special Issue Feature Review Papers in Sensors and Robotics)
Show Figures

Figure 1

12 pages, 800 KB  
Article
Construction of an Accurate Evaluation Model for Apple Flowering Period Based on Multimodal Data
by Ruoxin Qi, Zeyu Ye, Xuanzhang Tang, Desheng Jin, Dong Liang and Hui Xia
Agronomy 2026, 16(11), 1103; https://doi.org/10.3390/agronomy16111103 - 3 Jun 2026
Viewed by 321
Abstract
Flowering period management is a critical component of orchard production, significantly influencing the accuracy and timeliness of agricultural decisions such as flower and fruit thinning, yield stabilization, improvement in fruit commodity value, and control of mold core disease. Aiming at the problems of [...] Read more.
Flowering period management is a critical component of orchard production, significantly influencing the accuracy and timeliness of agricultural decisions such as flower and fruit thinning, yield stabilization, improvement in fruit commodity value, and control of mold core disease. Aiming at the problems of traditional flowering period judgment relying on manual experience, strong subjectivity, low efficiency, and difficulty in large-scale implementation, this study proposes an accurate evaluation model for apple flowering period based on near–far view multimodal visual data. A dedicated near–far view combined vision acquisition system was built to synchronously obtain panoramic images of fruit tree canopies and high-definition close-up images of single flowers/clusters, constructing a multimodal dataset covering the canopy spatial structure and fine floral organ morphology. YOLOv5s and ResNet-50 were employed to extract macro flowering proportion features from far views and micro morphological features from near views, respectively. A feature fusion strategy was introduced to realize the deep fusion of macro–micro features, and finally, a multimodal flowering period classification model was constructed to accurately divide the apple flowering period into four stages: bud stage, initial bloom stage, full bloom stage and late bloom stage. The overall recognition accuracy of the model reached 95.7%. The accurate apple flowering period evaluation system built based on this model has realized the paradigm shift in flowering period judgment from “qualitative manual experience” to “accurate quantification by machine vision”, providing a scientific time window basis for core orchard operations such as pre-flower re-pruning, flowering pollination, fruit setting evaluation and fruit thinning and bagging, and effectively promoting the intelligent and operational development of orchard management. Full article
(This article belongs to the Section Precision and Digital Agriculture)
Show Figures

Figure 1

20 pages, 1844 KB  
Article
AI-Enhanced Prognostic Model for Predicting Polyp Recurrence and Guiding Post-Polypectomy Surveillance Intervals Using the ERCPMP-V5 Dataset
by Sri Harsha Boppana, Sachin Sravan Kumar Komati, Ritwik Raj, Gautam Maddineni, Raja Chandra Chakinala, Pradeep Yarra, Venkata C. K. Sunkesula and Cyrus David Mintz
J. Clin. Med. 2026, 15(9), 3303; https://doi.org/10.3390/jcm15093303 - 26 Apr 2026
Viewed by 638
Abstract
Introduction: Colorectal cancer remains a leading cause of cancer-related morbidity and mortality, with adenomatous polyps representing a common precursor. Post-polypectomy polyp recurrence represents a significant risk of colorectal cancer, driving periodic colonoscopy surveillance and polypectomy as needed. In this study, we explore a [...] Read more.
Introduction: Colorectal cancer remains a leading cause of cancer-related morbidity and mortality, with adenomatous polyps representing a common precursor. Post-polypectomy polyp recurrence represents a significant risk of colorectal cancer, driving periodic colonoscopy surveillance and polypectomy as needed. In this study, we explore a multimodal machine learning approach that integrates endoscopic imaging with clinical and pathology data to improve recurrence risk prediction and support individualized surveillance planning. Methods: We developed and evaluated a multimodal artificial intelligence (AI) model to predict post-polypectomy colorectal polyp recurrence using the ERCPMP-v5 dataset. The cohort included 217 patients with 796 high-resolution endoscopic RGB images and 21 endoscopic videos; video data were converted to still frames at 2 frames per second. Images and frames were resized to 224 × 224 pixels and normalized. Patient-level demographic, morphological (Paris, Kudo Pit, JNET), anatomical, and pathological variables were encoded using standard scaling for continuous features and one-hot encoding for categorical features. Visual representations were extracted using a pretrained Vision Transformer backbone (ViT-Base-Patch16-224) with frozen weights. Structured metadata (79 variables) was encoded using a multilayer perceptron. A late fusion framework used image and metadata representations to generate a recurrence probability via a sigmoid classifier; probabilities were thresholded at 0.5 for binary prediction. Model performance was evaluated on a held-out test set using accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve (AUC). We additionally compared fusion performance with image-only and metadata-only baselines. Predicted probabilities were translated to surveillance recommendations using risk tiers: low risk (0.00 ≤ p < 0.20), moderate risk (0.20 ≤ p < 0.50), and high risk (p ≥ 0.50). Results: On the test set, the multimodal fusion model achieved 90.4% accuracy, 86.7% precision, 83.1% recall, 84.9% F1-score, and an AUC of 0.920. The image-only model achieved 84.6% accuracy (AUC 0.880), and the metadata-only model achieved 81.9% accuracy (AUC 0.850), indicating improved performance with multimodal fusion. Risk stratification enabled surveillance recommendations of 1–3 years for low risk, 6–12 months for moderate risk, and 3–6 months for high risk. Conclusions: A late-fusion multimodal model integrating endoscopic imaging with structured clinical and pathology variables demonstrated excellent performance for predicting post-polypectomy recurrence and generated actionable risk-based surveillance intervals. This approach may support individualized follow-up planning and more efficient allocation of surveillance resources, while prioritizing timely evaluation for patients at higher predicted risk. Full article
Show Figures

Graphical abstract

13 pages, 2368 KB  
Article
DGE-YOLO: Dual-Branch Gathering and Attention for Efficient Accurate UAV Object Detection
by Kunwei Lv, Zhiren Xiao, Hang Ren, Xiali Li and Ping Lan
Appl. Sci. 2026, 16(8), 4004; https://doi.org/10.3390/app16084004 - 20 Apr 2026
Cited by 1 | Viewed by 832
Abstract
The rapid proliferation of unmanned aerial vehicles (UAVs) has amplified the need for robust and efficient object detection in diverse aerial environments. However, detecting small objects under complex conditions (e.g., low illumination, cluttered backgrounds, and thermal–visual discrepancies) remains challenging. While many existing detectors [...] Read more.
The rapid proliferation of unmanned aerial vehicles (UAVs) has amplified the need for robust and efficient object detection in diverse aerial environments. However, detecting small objects under complex conditions (e.g., low illumination, cluttered backgrounds, and thermal–visual discrepancies) remains challenging. While many existing detectors emphasize real-time inference, they often rely on weak or late fusion strategies, resulting in suboptimal utilization of complementary multi-modal cues. To address this limitation, we propose DGE-YOLO, an enhanced YOLO-based framework for effective infrared–visible (IR–RGB) multi-modal fusion in UAV object detection. DGE-YOLO adopts a dual-branch architecture for modality-specific feature extraction, preserving modality-aware representations before fusion. To strengthen cross-scale semantics, we introduce an Efficient Multi-scale Attention (EMA) module that improves feature discrimination across spatial resolutions. Furthermore, we replace the conventional neck with a Gather-and-Distribute module to reduce information loss during feature aggregation and improve multi-scale feature propagation. Extensive experiments on the DroneVehicle dataset demonstrate that DGE-YOLO consistently outperforms state-of-the-art baselines, confirming its effectiveness and practicality as an applied multi-modal detection solution for UAV scenarios. Full article
(This article belongs to the Special Issue Applied Multimodal AI: Methods and Applications Across Domains)
Show Figures

Figure 1

25 pages, 4141 KB  
Article
CARYPAR: A Multimodal Decision-Support Framework Integrating Satellite Bio-Environmental Reanalysis and Proximal Edge-Intelligence for Hylocereus spp. Health Monitoring
by Carlos Diego Rodríguez-Yparraguirre, Abel José Rodríguez-Yparraguirre, Cesar Moreno-Rojo, Wendy Akemmy Castañeda-Rodríguez, Iván Martin Olivares-Espino, Andrés David Epifania-Huerta, María Adriana Vilchez-Reyes, Dany Paul Gonzales-Romero, Enrique Jannier Boy-Vásquez and Wilson Arcenio Maco-Vasquez
Sustainability 2026, 18(8), 3928; https://doi.org/10.3390/su18083928 - 15 Apr 2026
Viewed by 565
Abstract
Pitahaya (Hylocereus spp.) production is increasingly affected by climatic factors, as well as by phytopathogens and abiotic stress, leading to delays in agronomic interventions and reduced productivity. The objective was to design, implement, and validate a multimodal system (CARYPAR) that enables early [...] Read more.
Pitahaya (Hylocereus spp.) production is increasingly affected by climatic factors, as well as by phytopathogens and abiotic stress, leading to delays in agronomic interventions and reduced productivity. The objective was to design, implement, and validate a multimodal system (CARYPAR) that enables early disease detection and agile decision-making, characterized by low latency and reduced dependence on cloud connectivity. The methodology integrates climate reanalysis from NASA POWER, biophysical remote sensing variables derived from Sentinel-1/2, and proximal computer vision captured via mobile devices using a late fusion architecture and an optimized convolutional neural network, EfficientNet-V2B0, which discriminates between optimal and pathological conditions in vegetative tissues and fruit. The results of the experimental validation carried out in 160 georeferenced units achieved an overall accuracy of 80.0% and an F1 score of 0.8645 for Bad Fruit. The McNemar test and the operational agreement with agro-industrial experts yielded a Cohen’s Kappa index of κ = 0.6831, with an inference latency reduced to 22.00 ms. It is concluded that the multimodal integration of satellite bio-environmental data with edge computer vision achieves substantial agreement with agronomic expert judgment under heterogeneous field conditions (Cohen’s κ = 0.6831), supporting its role as a decision-support tool rather than a replacement for expert assessment. Therefore, its adoption can enhance real-time irrigation management and crop protection, while contributing to traceability and sustainable resource management in agricultural regions with limited connectivity. Full article
(This article belongs to the Section Sustainable Agriculture)
Show Figures

Figure 1

21 pages, 56996 KB  
Article
Comprehensive Analysis of Multimodal Fusion Techniques for Ocular Disease Detection
by Veena K. M., Pragya Gupta, Ruthvik Avadhanam, Rashmi Naveen Raj, Sulatha V. Bhandary, Varadraj Gurupur and Veena Mayya
AI 2026, 7(4), 126; https://doi.org/10.3390/ai7040126 - 1 Apr 2026
Viewed by 1563
Abstract
Accurate and early identification of ocular diseases is essential to prevent vision impairment and enable timely medical intervention. In routine clinical practice, ophthalmologists rely on a structured diagnostic workflow that incorporates multiple imaging modalities to manually assess and diagnose ocular diseases. However, interpreting [...] Read more.
Accurate and early identification of ocular diseases is essential to prevent vision impairment and enable timely medical intervention. In routine clinical practice, ophthalmologists rely on a structured diagnostic workflow that incorporates multiple imaging modalities to manually assess and diagnose ocular diseases. However, interpreting each modality requires significant clinical experience and can be time-consuming. These limitations can be effectively addressed through the application of AI (Artificial intelligence)-driven multimodal fusion techniques. In this study, we conducted an empirical investigation to assess the impact of different fusion strategies—including early, intermediate, and late fusion—on diagnostic performance, training requirements, and interpretability. The proposed methodology was evaluated using three publicly available datasets: FFA-Fundus (Fundus fluorescein angiography), GAMMA (Glaucoma Analysis and Multi-Modal Assessment), and OLIVES (Ophthalmic Labels to Investigate Visual Eye Semantics). Experimental results demonstrate that multimodal feature fusion improves disease detection performance. Although fused models typically required an increase in training parameters compared to single-modality models, they provided interpretability on par with that of individual single-modal networks. However, inference time increased by approximately 50% for multimodal architectures. These findings underscore the value of integrating diverse ophthalmic imaging modalities to enhance diagnostic accuracy in automated disease detection systems. At the same time, the results highlight that unimodal models containing highly discriminative features can also perform competitively, particularly when a single modality is sufficient for disease identification. Multimodal fusion provides the greatest benefit in scenarios where complementary information across modalities contributes distinct and non-redundant features. Furthermore, fusing all available modalities may not be optimal due to increased computational cost and reduced inference efficiency; thus, selective modality integration and lightweight fusion strategies are essential to balance accuracy, interpretability, and efficiency in clinical deployment. Full article
Show Figures

Figure 1

28 pages, 7980 KB  
Article
Smart Predictive Maintenance: A TCN-Based System for Early Fault Detection in Industrial Machinery
by Abuzar Khan, Ahmad Junaid, Muhammad Farooq Siddique, Abid Iqbal, Husam S. Samkari, Mohammed F. Allehyani and Ghassan Husnain
Machines 2026, 14(2), 164; https://doi.org/10.3390/machines14020164 - 1 Feb 2026
Cited by 11 | Viewed by 2471
Abstract
Modern factories still struggle with unexpected machine failures because traditional maintenance systems depend on fixed rules and threshold-based alerts. These older approaches often overlook subtle or complex patterns in multimodal sensor data, causing them to miss early signs of wear and leading to [...] Read more.
Modern factories still struggle with unexpected machine failures because traditional maintenance systems depend on fixed rules and threshold-based alerts. These older approaches often overlook subtle or complex patterns in multimodal sensor data, causing them to miss early signs of wear and leading to late or incorrect maintenance decisions. As a result, production can slow down, costs increase and equipment reliability suffers. To address this challenge, this study introduces a smart and interpretable fault diagnosis and predictive maintenance framework designed to detect wear, degradation and potential failures before they disrupt operations. The proposed framework integrates multiscale feature extraction, multimodal sensor fusion and cross-sensor correlation analysis with advanced temporal modeling using a Temporal Convolutional Network (TCN). By jointly performing tool-health classification and Remaining Useful Life (RUL) estimation, the framework provides a comprehensive assessment of machine condition. When evaluated on the NASA Ames milling dataset, the model achieved an overall accuracy of 86%, correctly classifying healthy and failed tools in more than 88% of cases and worn tools in over 75%, demonstrating consistent performance across different stages of tool wear. Explainable artificial intelligence (XAI) techniques, including attention-based visualizations and SHAP-based feature attribution, reveal that electrical and vibration signals are the most influential early indicators of tool degradation. The proposed framework exhibits low computational latency and minimal memory requirements, making it suitable for real-time fault diagnosis and deployment on industrial edge devices. Overall, the framework balances predictive accuracy, interpretability and practical applicability, enabling proactive and reliable maintenance decisions that enhance machine uptime and support efficient smart manufacturing operations. Full article
Show Figures

Figure 1

23 pages, 3037 KB  
Article
Depth Matters: Geometry-Aware RGB-D-Based Transformer-Enabled Deep Reinforcement Learning for Mapless Navigation
by Alpaslan Burak İnner and Mohammed E. Chachoua
Appl. Sci. 2026, 16(3), 1242; https://doi.org/10.3390/app16031242 - 26 Jan 2026
Cited by 2 | Viewed by 1087
Abstract
Autonomous navigation in unknown environments demands policies that can jointly perceive semantic context and geometric safety. Existing Transformer-enabled deep reinforcement learning (DRL) frameworks, such as the Goal-guided Transformer Soft Actor–Critic (GoT-SAC), rely on temporal stacking of multiple RGB frames, which encodes short-term motion [...] Read more.
Autonomous navigation in unknown environments demands policies that can jointly perceive semantic context and geometric safety. Existing Transformer-enabled deep reinforcement learning (DRL) frameworks, such as the Goal-guided Transformer Soft Actor–Critic (GoT-SAC), rely on temporal stacking of multiple RGB frames, which encodes short-term motion cues but lacks explicit spatial understanding. This study introduces a geometry-aware RGB-D early fusion modality that replaces temporal redundancy with cross-modal alignment between appearance and depth. Within the GoT-SAC framework, we integrate a pixel-aligned RGB-D input into the Transformer encoder, enabling the attention mechanism to simultaneously capture semantic textures and obstacle geometry. A comprehensive systematic ablation study was conducted across five modality variants (4RGB, RGB-D, G-D, 4G-D, and 4RGB-D) and three fusion strategies (early, parallel, and late) under identical hyperparameter settings in a controlled simulation environment. The proposed RGB-D early fusion achieved a 40.0% success rate and +94.1 average reward, surpassing the canonical 4RGB baseline (28.0% success, +35.2 reward), while a tuned configuration further improved performance to 54.0% success and +146.8 reward. These results establish early pixel-level multimodal fusion (RGB-D) as a principled and efficient successor to temporal stacking, yielding higher stability, sample efficiency, and geometry-aware decision-making. This work provides the first controlled evidence that spatially aligned multimodal fusion within Transformer-based DRL significantly enhances mapless navigation performance and offers a reproducible foundation for sim-to-real transfer in autonomous mobile robots. Full article
Show Figures

Figure 1

43 pages, 6570 KB  
Article
A Multimodal Phishing Website Detection System Using Explainable Artificial Intelligence Technologies
by Alexey Vulfin, Alexey Sulavko, Vladimir Vasiliev, Alexander Minko, Anastasia Kirillova and Alexander Samotuga
Mach. Learn. Knowl. Extr. 2026, 8(1), 11; https://doi.org/10.3390/make8010011 - 4 Jan 2026
Cited by 2 | Viewed by 4062
Abstract
The purpose of the present study is to improve the efficiency of phishing web resource detection through multimodal analysis and using methods of explainable artificial intelligence. We propose a late fusion architecture in which independent specialized models process four modalities and are combined [...] Read more.
The purpose of the present study is to improve the efficiency of phishing web resource detection through multimodal analysis and using methods of explainable artificial intelligence. We propose a late fusion architecture in which independent specialized models process four modalities and are combined using weighted voting. The first branch uses CatBoost for URL features and metadata; the second uses CNN1D for symbolic-level URL representation; the third uses a Transformer based on a pretrained CodeBERT for the homepage HTML code; and the fourth uses EfficientNet-B7 for page screenshot analysis. SHAP, Grad-CAM, and attention matrices are used to interpret decisions; a local LLM generates a consolidated textual explanation. A prototype system based on a microservice architecture, integrated with the SOC, has been developed. This integration enables streaming processing and reproducible validation. Computational experiments using our own updated dataset and the public MTLP dataset show high performance: F1-scores of up to 0.989 on our own dataset and 0.953 on MTLP; multimodal fusion consistently outperforms single-modal baseline models. The practical significance of this approach for zero-day detection and false positive reduction, through feature alignment across modalities and explainability, is demonstrated. All limitations and operational aspects (data drift, adversarial robustness, LLM latency) of the proposed prototype are presented. We also outline areas for further research. Full article
(This article belongs to the Section Safety, Security, Privacy, and Cyber Resilience)
Show Figures

Figure 1

24 pages, 1689 KB  
Article
Safeguarding Brand and Platform Credibility Through AI-Based Multi-Model Fake Profile Detection
by Vishwas Chakranarayan, Fadheela Hussain, Fayzeh Abdulkareem Jaber, Redha J. Shaker and Ali Rizwan
Future Internet 2025, 17(9), 391; https://doi.org/10.3390/fi17090391 - 29 Aug 2025
Cited by 4 | Viewed by 2270
Abstract
The proliferation of fake profiles on social media presents critical cybersecurity and misinformation challenges, necessitating robust and scalable detection mechanisms. Such profiles weaken consumer trust, reduce user engagement, and ultimately harm brand reputation and platform credibility. As adversarial tactics and synthetic identity generation [...] Read more.
The proliferation of fake profiles on social media presents critical cybersecurity and misinformation challenges, necessitating robust and scalable detection mechanisms. Such profiles weaken consumer trust, reduce user engagement, and ultimately harm brand reputation and platform credibility. As adversarial tactics and synthetic identity generation evolve, traditional rule-based and machine learning approaches struggle to detect evolving and deceptive behavioral patterns embedded in dynamic user-generated content. This study aims to develop an AI-driven, multi-modal deep learning-based detection system for identifying fake profiles that fuses textual, visual, and social network features to enhance detection accuracy. It also seeks to ensure scalability, adversarial robustness, and real-time threat detection capabilities suitable for practical deployment in industrial cybersecurity environments. To achieve these objectives, the current study proposes an integrated AI system that combines the Robustly Optimized BERT Pretraining Approach (RoBERTa) for deep semantic textual analysis, ConvNeXt for high-resolution profile image verification, and Heterogeneous Graph Attention Networks (Hetero-GAT) for modeling complex social interactions. The extracted features from all three modalities are fused through an attention-based late fusion strategy, enhancing interpretability, robustness, and cross-modal learning. Experimental evaluations on large-scale social media datasets demonstrate that the proposed RoBERTa-ConvNeXt-HeteroGAT model significantly outperforms baseline models, including Support Vector Machine (SVM), Random Forest, and Long Short-Term Memory (LSTM). Performance achieves 98.9% accuracy, 98.4% precision, and a 98.6% F1-score, with a per-profile speed of 15.7 milliseconds, enabling real-time applicability. Moreover, the model proves to be resilient against various types of attacks on text, images, and network activity. This study advances the application of AI in cybersecurity by introducing a highly interpretable, multi-modal detection system that strengthens digital trust, supports identity verification, and enhances the security of social media platforms. This alignment of technical robustness with brand trust highlights the system’s value not only in cybersecurity but also in sustaining platform credibility and consumer confidence. This system provides practical value to a wide range of stakeholders, including platform providers, AI researchers, cybersecurity professionals, and public sector regulators, by enabling real-time detection, improving operational efficiency, and safeguarding online ecosystems. Full article
Show Figures

Figure 1

29 pages, 7018 KB  
Article
Real-Time Efficiency Prediction in Nonlinear Fractional-Order Systems via Multimodal Fusion
by Biao Ma and Shimin Dong
Fractal Fract. 2025, 9(8), 545; https://doi.org/10.3390/fractalfract9080545 - 19 Aug 2025
Viewed by 1041
Abstract
Rod pump systems are complex nonlinear processes, and conventional efficiency prediction methods for such systems typically rely on high-order fractional partial differential equations, which severely constrain real-time inference. Motivated by the increasing availability of measured electrical power data, this paper introduces a series [...] Read more.
Rod pump systems are complex nonlinear processes, and conventional efficiency prediction methods for such systems typically rely on high-order fractional partial differential equations, which severely constrain real-time inference. Motivated by the increasing availability of measured electrical power data, this paper introduces a series of prediction models for nonlinear fractional-order PDE systems efficiency based on multimodal feature fusion. First, three single-model predictions—Asymptotic Cross-Fusion, Adaptive-Weight Late-Fusion, and Two-Stage Progressive Feature Fusion—are presented; next, two ensemble approaches—one based on a Parallel-Cascaded Ensemble strategy and the other on Data Envelopment Analysis—are developed; finally, by balancing base-learner diversity with predictive accuracy, a multi-strategy ensemble prediction model is devised for online rod pump system efficiency estimation. Comprehensive experiments and ablation studies on data from 3938 oil wells demonstrate that the proposed methods deliver high predictive accuracy while meeting real-time performance requirements. Full article
(This article belongs to the Special Issue Artificial Intelligence and Fractional Modelling for Energy Systems)
Show Figures

Figure 1

20 pages, 5700 KB  
Article
Multimodal Personality Recognition Using Self-Attention-Based Fusion of Audio, Visual, and Text Features
by Hyeonuk Bhin and Jongsuk Choi
Electronics 2025, 14(14), 2837; https://doi.org/10.3390/electronics14142837 - 15 Jul 2025
Cited by 7 | Viewed by 4254
Abstract
Personality is a fundamental psychological trait that exerts a long-term influence on human behavior patterns and social interactions. Automatic personality recognition (APR) has exhibited increasing importance across various domains, including Human–Robot Interaction (HRI), personalized services, and psychological assessments. In this study, we propose [...] Read more.
Personality is a fundamental psychological trait that exerts a long-term influence on human behavior patterns and social interactions. Automatic personality recognition (APR) has exhibited increasing importance across various domains, including Human–Robot Interaction (HRI), personalized services, and psychological assessments. In this study, we propose a multimodal personality recognition model that classifies the Big Five personality traits by extracting features from three heterogeneous sources: audio processed using Wav2Vec2, video represented as Skeleton Landmark time series, and text encoded through Bidirectional Encoder Representations from Transformers (BERT) and Doc2Vec embeddings. Each modality is handled through an independent Self-Attention block that highlights salient temporal information, and these representations are then summarized and integrated using a late fusion approach to effectively reflect both the inter-modal complementarity and cross-modal interactions. Compared to traditional recurrent neural network (RNN)-based multimodal models and unimodal classifiers, the proposed model achieves an improvement of up to 12 percent in the F1-score. It also maintains a high prediction accuracy and robustness under limited input conditions. Furthermore, a visualization based on t-distributed Stochastic Neighbor Embedding (t-SNE) demonstrates clear distributional separation across the personality classes, enhancing the interpretability of the model and providing insights into the structural characteristics of its latent representations. To support real-time deployment, a lightweight thread-based processing architecture is implemented, ensuring computational efficiency. By leveraging deep learning-based feature extraction and the Self-Attention mechanism, we present a novel personality recognition framework that balances performance with interpretability. The proposed approach establishes a strong foundation for practical applications in HRI, counseling, education, and other interactive systems that require personalized adaptation. Full article
(This article belongs to the Special Issue Explainable Machine Learning and Data Mining)
Show Figures

Figure 1

33 pages, 17535 KB  
Article
MultiScaleFusion-Net and ResRNN-Net: Proposed Deep Learning Architectures for Accurate and Interpretable Pregnancy Risk Prediction
by Amna Asad, Madiha Sarwar, Muhammad Aslam, Edore Akpokodje and Syeda Fizzah Jilani
Appl. Sci. 2025, 15(11), 6152; https://doi.org/10.3390/app15116152 - 30 May 2025
Cited by 2 | Viewed by 2258
Abstract
Women exhibit marked physiological transformations in pregnancy, mandating regular and holistic assessment. Maternal and fetal vitality is governed by a spectrum of clinical, demographic, and lifestyle factors throughout this critical period. The existing maternal health monitoring techniques lack precision in assessing pregnancy-related risks, [...] Read more.
Women exhibit marked physiological transformations in pregnancy, mandating regular and holistic assessment. Maternal and fetal vitality is governed by a spectrum of clinical, demographic, and lifestyle factors throughout this critical period. The existing maternal health monitoring techniques lack precision in assessing pregnancy-related risks, often leading to late interventions and adverse outcomes. Accurate and timely risk prediction is crucial to avoid miscarriages. This research proposes a deep learning framework for personalized pregnancy risk prediction using the NFHS-5 dataset, and class imbalance is addressed through a hybrid NearMiss-SMOTE approach. Fifty-one primary features are selected via the LASSO to refine the dataset and enhance model interpretability and efficiency. The framework integrates a multimodal model (NFHS-5, fetal plane images, and EHG time series) along with two core architectures. ResRNN-Net further combines Bi-LSTM, CNNs, and attention mechanisms to capture sequential dependencies. MultiScaleFusion-Net leverages GRU and multiscale convolutions for effective feature extraction. Additionally, TabNet and MLP models are explored to compare interpretability and computational efficiency. SHAP and Grad-CAM are used to ensure transparency and explainability, offering both feature importance and visual explanations of predictions. The proposed models are trained using 5-fold stratified cross-validation and evaluated with metrics including accuracy, precision, recall, F1-score, and ROC–AUC. The results demonstrate that MultiScaleFusion-Net balances accuracy and computational efficiency, making it suitable for real-time clinical deployment, while ResRNN-Net achieves higher precision at a slight computational cost. Performance comparisons with baseline machine learning models confirm the superiority of deep learning approaches, achieving over 80% accuracy in pregnancy complication prediction. Full article
(This article belongs to the Special Issue Application of Artificial Intelligence in Biomedical Informatics)
Show Figures

Figure 1

Back to TopTop