Topic Editors

Institute of Artificial Intelligence and Blockchain, Guangzhou University, Guangzhou 510006, China
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China
Prof. Dr. Jianjun Li
School of Digital and Intelligent Industry, Inner Mongolia University of Science & Technology, Baotou 014010, China
Dr. Jin Liu
School of Computer Science and Engineering, Central South University, Changsha 410083, China
Dr. Jia Wang
College of Information Science and Engineering, Xinjiang University, Urumqi 830046, China

AI and Data-Driven Advancements in Industry 4.0, 2nd Edition

Abstract submission deadline
8 September 2026
Manuscript submission deadline
8 December 2026
Viewed by
65977

Topic Information

Dear Colleagues,

Our society is awash with diverse forms of data—pictures, point clouds, text, audio, video, and beyond. Data-driven artificial intelligence offers the potential to extract meaningful insights from this deluge, unlocking extraordinary opportunities in both theory and application. Recent years have witnessed a proliferation of AI theories and algorithms, now trusted and employed across sectors like finance, security, education, art, neuroscience, and even music. Inspired by these advancements, substantial AI-based techniques are being implemented in fields as varied as autonomous driving, virtual reality, human–computer interaction, remote sensing, and artistic creation.

This focus area seeks to collect and highlight the latest breakthroughs in next-generation artificial intelligence and its applications within the context of Industry 4.0. Our interest spans the full spectrum of AI research, from compact on-device models to expansive large-scale models, across diverse domains. We are particularly keen on sophisticated framework designs, training strategies, optimization techniques, and ensuring the trustworthiness and robustness of AI systems, along with their real-world applications.

Topics include (but are not limited to) the following:

  • Computer vision, natural language processing, reinforcement learning;
  • Large-scale language model, large-scale vision model, prompt learning, retrieval-augmented generation;
  • Multi-modal learning, object recognition, detection, segmentation, tracking;
  • Graph neural network, knowledge graph, recommendation system;
  • Pattern recognition and intelligent system;
  • Blockchain theory or application, smart contract;
  • Artificial intelligence security, data security and privacy;
  • Remote sensing image interpretation;
  • Art generation and creation, art analysis and understanding;
  • Virtual reality, robotics, edge computing, on-device models;
  • Autonomous driving.

Dr. Teng Huang
Dr. Yan Pang
Prof. Dr. Qiong Wang
Prof. Dr. Jianjun Li
Dr. Jin Liu
Dr. Jia Wang
Topic Editors

Keywords

  • computer vision
  • natural language processing
  • large-scale language model
  • large-scale vision model
  • scheduling optimization
  • pattern recognition and intelligent system
  • blockchain
  • security and privacy
  • network security
  • remote sensing image interpretation

Participating Journals

Journal Name Impact Factor CiteScore Launched Year First Decision (median) APC
AI
ai
6.5 7.3 2020 20.4 Days CHF 1800 Submit
Drones
drones
5.2 10.0 2017 21.1 Days CHF 2600 Submit
Electronics
electronics
2.9 7.0 2012 14.8 Days CHF 2400 Submit
Mathematics
mathematics
2.3 5.4 2013 17.4 Days CHF 2600 Submit
Sensors
sensors
4.0 9.4 2001 17.8 Days CHF 2600 Submit

Preprints.org is a multidisciplinary platform offering a preprint service designed to facilitate the early sharing of your research. It supports and empowers your research journey from the very beginning.

MDPI Topics is collaborating with Preprints.org and has established a direct connection between MDPI journals and the platform. Authors are encouraged to take advantage of this opportunity by posting their preprints at Preprints.org prior to publication:

  1. Share your research immediately: disseminate your ideas prior to publication and establish priority for your work.
  2. Safeguard your intellectual contribution: Protect your ideas with a time-stamped preprint that serves as proof of your research timeline.
  3. Boost visibility and impact: Increase the reach and influence of your research by making it accessible to a global audience.
  4. Gain early feedback: Receive valuable input and insights from peers before submitting to a journal.
  5. Ensure broad indexing: Web of Science (Preprint Citation Index), Google Scholar, Crossref, SHARE, PrePubMed, Scilit and Europe PMC.

Published Papers (43 papers)

Order results
Result details
Journals
Select all
Export citation of selected articles as:
19 pages, 3576 KB  
Article
Deep Fusion Modeling for Lithium Batteries SOC-SOH Joint Estimation
by Jian Wang, Shu Cheng and Lulin Zhang
Mathematics 2026, 14(14), 2505; https://doi.org/10.3390/math14142505 - 11 Jul 2026
Viewed by 346
Abstract
This paper proposes a SOC-SOH joint estimation method based on adaptive weighted multi-channel LSTM Transformer fusion network (MLTA-Net). The proposed method constructs a battery health factor set which covers multi-level features, and the aging trend of batteries can be characterized from multiple dimensions. [...] Read more.
This paper proposes a SOC-SOH joint estimation method based on adaptive weighted multi-channel LSTM Transformer fusion network (MLTA-Net). The proposed method constructs a battery health factor set which covers multi-level features, and the aging trend of batteries can be characterized from multiple dimensions. The MLTA-Net model adopts a multi-channel parallel architecture, which can analyze the different types of battery data characteristics. Short-term temporal dependencies are captured by LSTM encoder, and global operating characteristics are analyzed using Transformer multi head self-attention mechanism. Based on adaptive weighted fusion layer for feature fusion, high-precision estimation of battery state can be achieved. Experimental results on CATL-1 and CATL-2 datasets show that the proposed MLTA-Net achieves superior SOH estimation accuracy, with RMSE values of 0.286 and 0.287, MAE values of 0.151 and 0.162, MAPE values of 0.053 and 0.056, and R2 values of 0.997 and 0.997, respectively. Compared with CNN-GRU, MLP-Attention, Transformer, MLP, RNN, and SVR models, the proposed method exhibits lower prediction errors and better robustness. Full article
Show Figures

Figure 1

21 pages, 1086 KB  
Article
Linking Tea Aroma Chemistry to Quality Grades via a Single MOS Gas Sensor: Classical Machine Learning vs. Deep Learning
by Ahmet Turan Tasdemir, Erkan Caner Ozkat, Gozde Yalcin Ozkat and Fatih Gul
Sensors 2026, 26(12), 3877; https://doi.org/10.3390/s26123877 - 18 Jun 2026
Cited by 2 | Viewed by 553
Abstract
Black tea quality is governed by aroma chemistry: terpene alcohols (linalool, geraniol, nerolidol), methyl salicylate, and short-chain aldehydes whose abundance and release kinetics from the polyphenol-rich leaf matrix shape perceived grade. Grade information lies not only in the average headspace concentration but in [...] Read more.
Black tea quality is governed by aroma chemistry: terpene alcohols (linalool, geraniol, nerolidol), methyl salicylate, and short-chain aldehydes whose abundance and release kinetics from the polyphenol-rich leaf matrix shape perceived grade. Grade information lies not only in the average headspace concentration but in the temporal shape of volatile organic compound (VOC) release under controlled heating. Conventional electronic noses obscure this signal: they rely on multi-sensor arrays, compress each response into summary statistics, and report accuracy only at the level of individual measurements. Whether a single low-cost metal–oxide–semiconductor (MOS) gas sensor can recover grade-defining aroma chemistry, and whether waveform-level modeling can exploit it, was therefore investigated. A portable electronic nose built around a Bosch BME688 sensor recorded 90 time series, each comprising four directly measured channels (temperature, humidity, pressure, gas sensor resistance) and a derived indoor-air-quality (IAQ) proxy computed from them by the on-chip BSEC library, from 16 commercial Turkish black teas across three quality grades. Two representations were compared on the same data: a feature-based pipeline reducing 25 statistical descriptors to seven principal components for six classifiers (best F1-macro = 0.624, MLP), and a raw-waveform Multi-Scale 1D-CNN with Squeeze–Excitation and temporal self-attention (MS-CNN-Attention). Under product-grouped cross-validation, the deep model reached F1-macro = 0.811 (+30%) and graded 14 of 16 products correctly by majority vote, against 11 of 16 for the MLP, with the largest gain in the medium grade (F1: 0.52 → 0.79), where summary-statistic compression destroys the release-kinetic signal. The contributions are threefold: one programmable MOS sensor operated as a thermal-desorption profiler rather than a sensor array; a direct comparison of feature-based classical learning against raw-waveform deep learning on the same small, non-normally distributed dataset; and a product-level decision-consistency metric suited to batch screening. Pairing a low-cost MOS sensor with waveform-level modeling offers a rapid, non-destructive route to aroma-chemistry-based tea quality screening. Full article
Show Figures

Figure 1

33 pages, 4102 KB  
Article
Real-Time Explanation Intrusion Detection: An XAI-Enriched Hybrid CNN-LSTM Architecture for Operational Cybersecurity
by Ayman Alnsour, Jamal Zarqou and Ahmad Shalaldeh
Mathematics 2026, 14(11), 1977; https://doi.org/10.3390/math14111977 - 3 Jun 2026
Viewed by 577
Abstract
Deep learning-based intrusion detection systems offer world-class accuracy in threat classification. They are also generally not easily explainable to security analysts, which represents a major hurdle in their use in real-world Security Operations Centers (SOCs) where explainability and trust are critical. This operational [...] Read more.
Deep learning-based intrusion detection systems offer world-class accuracy in threat classification. They are also generally not easily explainable to security analysts, which represents a major hurdle in their use in real-world Security Operations Centers (SOCs) where explainability and trust are critical. This operational challenge is tackled with a systems-engineered approach combining the CNN-LSTM architecture with the computationally optimized SHAP and LIME approaches for enabling real-time, interpretable threat detection. Unlike novel mathematical formulations, we concentrate on practical innovations in systems engineering that we believe are required to generate explanations in real-time: quantization of the numbers to INT8, execution of explanation algorithms in parallel, asynchronously, and caching of similar traffic patterns. CNN-LSTM combines the convolutional function to capture spatial dependencies and the recurrent function to capture temporal dynamics of network traffic, and SHAP and LIME capture global and local feature attributions, respectively. One of the major innovations is the parallel execution which brings the latency of explanation down from 117 ms (sequential SHAP + LIME) to 46 ms (parallel, cache-miss) and 39 ms (average with caching) and 46 ms (without caching), which is sufficient for operational “real-time” requirements. The framework is evaluated on CICIDS2017 and NSL-KDD benchmark datasets, and results show that it can achieve 98.7% accuracy with 98.6% F1-score and sub-50 ms explanation latency. The results here show that explainability and operational efficiency can be attained with the same level of accuracy in the detection of abnormal events, through careful systems engineering. This paper presents a systems-engineered framework demonstrating the feasibility of real-time, interpretable IDS for deployment in Security Operations Centers (SOCs) and addresses the challenges of combining high-performance deep learning with operational transparency in cybersecurity. Full article
Show Figures

Figure 1

30 pages, 3473 KB  
Article
AEConvNeXt: An Attention-Enhanced ConvNeXt Framework for Imbalanced Photovoltaic Fault Classification with Explainable Feature Analysis
by Ehtisham Lodhi and Lin Qiu
AI 2026, 7(6), 182; https://doi.org/10.3390/ai7060182 - 22 May 2026
Viewed by 655
Abstract
Background: Solar energy provides a sustainable and environmentally friendly alternative to fossil fuels, and photovoltaic (PV) systems are increasingly deployed worldwide. However, their operational reliability is often compromised by various fault conditions, which reduce power output and shorten system lifespan. Although automated image-based [...] Read more.
Background: Solar energy provides a sustainable and environmentally friendly alternative to fossil fuels, and photovoltaic (PV) systems are increasingly deployed worldwide. However, their operational reliability is often compromised by various fault conditions, which reduce power output and shorten system lifespan. Although automated image-based deep learning methods have shown promise for PV fault classification, their performance is often limited by severe class imbalance and subtle, low-contrast defect patterns. This study aims to address these challenges by proposing an improved deep learning framework for robust PV fault classification. Method: An attention-enhanced convolutional neural network framework, termed AEConvNeXt, is proposed for PV fault classification. The model is built on a ConvNeXt-Tiny backbone and incorporates a dropout-regularized Convolutional Block Attention Module (CBAM) to enhance localized feature refinement. To further improve learning under imbalanced data conditions, a hybrid loss function combining Cross-Entropy Loss and Focal Loss is employed. Results: Experimental evaluations demonstrate that AEConvNeXt achieves an overall accuracy of 94.37% and a macro F1-score of 94.43%, outperforming the strongest baseline model, ResNet-50, by more than 3%. Grad-CAM visualizations further confirm that the model effectively focuses on fault-relevant regions, improving interpretability. The proposed framework also shows consistent and robust performance across all six PV fault categories under varying conditions. Conclusions: The proposed AEConvNeXt framework provides an accurate and explainable solution for real-time PV fault detection, effectively addressing class imbalance and improving minority fault recognition. Full article
Show Figures

Figure 1

16 pages, 1770 KB  
Article
A Hybrid AI Approach for Intelligent Group Buying and Digital Marketing Strategy Optimization Based on Machine Learning and Evolutionary Algorithms
by Zhansaya Abildaeva, Raissa Uskenbayeva, Zhuldyz Kalpeyeva, Aizhan Kassymova, Aigul Dauitbayeva and Adranova Asselkhan
Mathematics 2026, 14(10), 1755; https://doi.org/10.3390/math14101755 - 20 May 2026
Viewed by 440
Abstract
This study considers the digital transformation of Kazakhstan’s agro-industrial complex, which has created an urgent need for scientifically grounded methods that can optimize marketing strategies under conditions of resource limitations, production seasonality, and heterogeneous consumer behavior. This study proposes a hybrid decision-support framework [...] Read more.
This study considers the digital transformation of Kazakhstan’s agro-industrial complex, which has created an urgent need for scientifically grounded methods that can optimize marketing strategies under conditions of resource limitations, production seasonality, and heterogeneous consumer behavior. This study proposes a hybrid decision-support framework integrating a modified NSGA-III algorithm with machine learning techniques for optimizing digital marketing strategies in the agro-industrial complex of Kazakhstan. The model considers three objectives: maximizing channel efficiency and audience reach while minimizing marketing costs. Experimental results based on a dataset of N = 1200 observations demonstrate that the proposed approach improves the composite performance indicator by 12.4% compared to baseline single-objective optimization methods. Pareto front analysis reveals three distinct clusters of strategies, corresponding to (1) high-impact integrated digital TV strategies, (2) cost-efficient traditional channel strategies, and (3) high-risk high-return allocations. The clustering validity is confirmed by a silhouette score of 0.624, indicating strong separation between strategy groups. The results highlight the practical significance of adaptive budget allocation and demonstrate the effectiveness of combining evolutionary optimization with machine learning for decision support in complex marketing environments. Full article
Show Figures

Graphical abstract

20 pages, 18950 KB  
Article
Multi-View Industrial Image Super-Resolution via Hierarchical Multi-Scale Data Fusion
by Wenqin Zhao, Carman Ka Man Lee, Da Li and Benny Chi Fai Cheung
AI 2026, 7(5), 172; https://doi.org/10.3390/ai7050172 - 16 May 2026
Viewed by 599
Abstract
Machine vision plays a pivotal role in precision engineering for high-precision measurement that relies on high-resolution images. The highly reflective nature of metal surfaces and the need for high-quality images pose significant challenges in image processing. Although existing research has made significant progress [...] Read more.
Machine vision plays a pivotal role in precision engineering for high-precision measurement that relies on high-resolution images. The highly reflective nature of metal surfaces and the need for high-quality images pose significant challenges in image processing. Although existing research has made significant progress in enhancing the resolution of natural images, super-resolution methods specifically tailored for multi-view metal images remain unexplored areas. To fill this gap, this paper focuses on developing a deep learning-based super-resolution algorithm, focusing on detail recovery on under multi-view metal images. The proposed super-resolution model utilizes a hybrid-resolution input that combines light field super-resolution at the image level and reference-based super-resolution at the feature level, demonstrating the effectiveness for achieving a large-scale multi-view metal image super-resolution. An experiment using a public metal object image dataset is conducted, and a comparison has been carried out with Bicubic, LFhybridSR and ERVSR. The proposed method demonstrates superior SSIM and achieves average PSNR improvements of 4.45 dB and 1.18 dB on synthetic data and real-world data. The results demonstrate that the method can improve the resolution and detail representation of metal images in terms of PSNR/SSIM and address the problem of super-resolution in multi-view metal images. Furthermore, applying the proposed SR method as preprocessing reduces the absolute relative error in depth estimation from approximately 0.5 to 0.1. Full article
Show Figures

Figure 1

35 pages, 14998 KB  
Article
A Unified Deep Learning-Based Corridor Following with Image-Based Obstacle Avoidance for Autonomous Wheelchair Navigation
by A. H. Abdul Hafez
Mathematics 2026, 14(10), 1698; https://doi.org/10.3390/math14101698 - 15 May 2026
Viewed by 684
Abstract
Autonomous wheelchair navigation requires both reliable global guidance and safe local interaction with the environment, typically addressed using separate perception and control strategies. This paper presents a unified vision-based control framework that combines learning-based corridor following with image-based obstacle avoidance under a common [...] Read more.
Autonomous wheelchair navigation requires both reliable global guidance and safe local interaction with the environment, typically addressed using separate perception and control strategies. This paper presents a unified vision-based control framework that combines learning-based corridor following with image-based obstacle avoidance under a common visual servoing perspective. This work provides a unified interpretation of learning-based and analytical control as complementary realizations of visual servoing. A convolutional neural network (CNN) is employed to directly predict steering commands from monocular images, enabling robust corridor following without explicit feature extraction. In parallel, obstacle avoidance is formulated as an image-based visual servoing (IBVS) task, where detected obstacles are represented as image features and regulated toward safe regions. A supervisory control strategy coordinates these components by prioritizing safety-critical avoidance when necessary, while maintaining nominal navigation otherwise. The system is implemented using a single monocular camera and deployed on a low-cost embedded platform. Experimental results demonstrate that the CNN-based module maintains stable performance under challenging visual conditions, while the IBVS controller provides predictable and reliable avoidance behavior. The proposed framework highlights the complementary roles of learning-based and analytical visual servoing, offering a practical and scalable solution for assistive autonomous mobility. Full article
Show Figures

Figure 1

14 pages, 1597 KB  
Article
Physics-Informed POD-PINN for Fast Wake Prediction of Twin Vertical-Axis Hydroturbine Arrays
by Ai Shan, Hu Chao and Ma Yong
Mathematics 2026, 14(10), 1579; https://doi.org/10.3390/math14101579 - 7 May 2026
Viewed by 504
Abstract
Accurate prediction of wake interactions in twin vertical-axis hydroturbine (VAHT) arrays is important for dense tidal-farm layout assessment but remains computationally expensive when based directly on Computational Fluid Dynamics (CFD) reference simulations. While simplified analytical models offer speed, they fail to capture the [...] Read more.
Accurate prediction of wake interactions in twin vertical-axis hydroturbine (VAHT) arrays is important for dense tidal-farm layout assessment but remains computationally expensive when based directly on Computational Fluid Dynamics (CFD) reference simulations. While simplified analytical models offer speed, they fail to capture the non-axisymmetric wake characteristics of VAHT arrays, and standard Physics-Informed Neural Networks (PINNs) often struggle with convergence in small-sample, high-dimensional flow settings. To address this challenge, this study proposes a Physics-Informed POD-PINN framework for predicting configuration-wise time-averaged wake fields. The hybrid architecture combines Proper Orthogonal Decomposition (POD) for dimensionality reduction with a dual-branch neural network: a global POD branch captures dominant flow structures, while a lightweight spatial correction branch acts as a continuity-informed regularization on the predicted field. Trained on CFD-generated reference data covering diverse longitudinal and lateral spacing configurations, the model learns to map geometric parameters to a three-component wake field represented on a regularized 3D grid. Results show that the proposed framework achieves the lowest mean streamwise error among the tested surrogate models while maintaining millisecond-level inference speed. This study provides an efficient and physics-aware surrogate tool for repeated wake-field evaluation in twin-hydroturbine configuration exploration. Full article
Show Figures

Figure 1

22 pages, 80280 KB  
Article
Research on Precise Detection Methods for the Maturity of Pleurotus ostreatus in Complex Mushroom Cultivation Environments
by Jun Yu, Changshou Luo, Qingfeng Wei, Yang Lu and Yaming Zheng
Sensors 2026, 26(9), 2583; https://doi.org/10.3390/s26092583 - 22 Apr 2026
Cited by 1 | Viewed by 704
Abstract
Addressing the challenges of complex background interference, low lighting conditions, small target recognition, and difficulty in maturity grading in the automated detection of Pleurotus ostreatus, this study proposes a lightweight improved scheme based on color feature enhancement. By collecting 4779 images from [...] Read more.
Addressing the challenges of complex background interference, low lighting conditions, small target recognition, and difficulty in maturity grading in the automated detection of Pleurotus ostreatus, this study proposes a lightweight improved scheme based on color feature enhancement. By collecting 4779 images from five developmental stages in three typical planting environments, including greenhouses and mushroom houses, an HSV hue analysis database was established to determine key hue intervals [4°, 38°] or [110°, 155°] for different environments. Secondly, based on the hue interval distribution of Pleu-rotus ostreatus, YOLOv13 was used as the base model, with the addition of an HSV hue mask as the fourth channel to improve the input layer. The custom ColorWeight module was used to enhance color feature expression; the hypergraph computation module was improved to enhance feature correlation; and the neck network incorporated the StockenAttention module to improve the ability to capture maturity features. The accuracy of the improved model was increased to 89.5% in mAP@0.5 (+3.3%), surpassing the mainstream YOLOv8n-12n series. Efficiency optimization achieved real-time detection at 12.58 FPS on the RTX3090Ti platform. In practical applications, the accuracy of maturity recognition was significantly improved, with a 73.6% decrease in the misclassification rate of maturity and a reduction in missed detections, achieving an F1 score of 0.91. In conclusion, through the deep integration of Hue features and deep learning models, while ensuring lightweight deployment (with only a 10.5% increase in parameter count), the accuracy and practicality of Pleurotus ostreatus detection were significantly improved, providing an effective solution for intelligent mushroom house management. Full article
Show Figures

Graphical abstract

18 pages, 2600 KB  
Article
Fourier Neural Operator for Turbine Wake Flow Prediction with Out-of-Distribution Generalization
by Shan Ai, Chao Hu and Yong Ma
Mathematics 2026, 14(8), 1275; https://doi.org/10.3390/math14081275 - 11 Apr 2026
Viewed by 717
Abstract
Amid the global transition to carbon neutrality, tidal current energy has become a strategic sustainable energy resource due to its high predictability, power density, and environmental compatibility. Horizontal-axis turbines show great potential for marine energy harvesting, yet the large-scale commercialization of tidal turbines [...] Read more.
Amid the global transition to carbon neutrality, tidal current energy has become a strategic sustainable energy resource due to its high predictability, power density, and environmental compatibility. Horizontal-axis turbines show great potential for marine energy harvesting, yet the large-scale commercialization of tidal turbines is severely hindered by complex wake dynamics and the lack of reliable, efficient prediction tools for out-of-distribution (OOD) operating conditions. Traditional high-fidelity CFD methods are computationally prohibitive for engineering optimization, while conventional data-driven surrogate models suffer from poor extrapolation performance, extrapolation collapse near training parameter boundaries, and the absence of uncertainty quantification. To address these bottlenecks, this study focuses on the OOD extrapolation of wake flow prediction across tip speed ratio (TSR) distributions for a single horizontal-axis tidal turbine. A CFD-generated spatiotemporal benchmark dataset is constructed for comparative OOD evaluation across various TSR conditions with 9504 total samples. A novel physics-constrained Fourier neural operator framework named TSR-FNO is proposed to improve OOD generalization. The model integrates TSR–Lipschitz regularization to suppress extrapolation collapse and Monte Carlo Dropout to provide reliable uncertainty estimation. Extensive experiments demonstrate that the proposed method effectively reduces prediction error in unseen TSR regimes, mitigates performance degradation in far-field extrapolation, and produces well-calibrated uncertainty estimates consistent with actual prediction confidence. This work provides a data-driven surrogate modeling strategy for fast and reliable wake prediction on a common CFD-generated benchmark, supporting the efficient design, array layout optimization, and engineering deployment of tidal current energy systems. Full article
Show Figures

Figure 1

16 pages, 841 KB  
Article
DSAK: Distillation of Self-Adaptive Knowledge for Membership Privacy Protection
by Qian Sheng, Jiaming Liang, Xinyu Li and Yan Huang
Mathematics 2026, 14(8), 1249; https://doi.org/10.3390/math14081249 - 9 Apr 2026
Viewed by 423
Abstract
The utilization of machine learning models is extensive in a wide array of significant applications. However, their vulnerability to security and privacy attacks is a serious concern, for example, for the protection of financially sensitive data such as account flow. Particularly troubling is [...] Read more.
The utilization of machine learning models is extensive in a wide array of significant applications. However, their vulnerability to security and privacy attacks is a serious concern, for example, for the protection of financially sensitive data such as account flow. Particularly troubling is the threat of membership inference, which enables attackers to determine whether a given data sample is included in the training set of a targeted machine-learning model. Existing knowledge distillation techniques have shown promise in balancing model performance with data privacy. However, achieving superior privacy during the training process of the target model is challenging due to the teacher model’s performance limitations and the scarcity of unlabeled benchmark data. To address this issue, we propose a novel framework called Distillation of Self-Adaptive Knowledge (DSAK). DSAK utilizes self-duplicated teacher and noise-generative models to introduce specialized self-adaptive noise for privacy training in the target model. By incorporating new data features derived from this noise, DSAK improves model performance and reduces the risk of memorizing member data. Experimental results demonstrate DSAK’s effectiveness in defending against existing attack schemes across multiple datasets while surpassing other membership inference defense schemes in terms of efficiency. Full article
Show Figures

Figure 1

33 pages, 736 KB  
Article
Analysis of Chip Electronic Components’ Typical Yield in Taping Process Based on Virtual Metrology
by Shiqi Zhang, Lizhen Chen, Jiangcheng Fu, Chenghu Yang and Guangli Chen
Sensors 2026, 26(8), 2292; https://doi.org/10.3390/s26082292 - 8 Apr 2026
Viewed by 662
Abstract
This study addresses virtual metrology (VM) for the taping process of chip electronic components, in which partial observability, unmeasured disturbances, and severe label imbalance make direct batch-wise yield prediction unstable. Rather than proposing a new standalone learning algorithm, we develop a data-centric VM [...] Read more.
This study addresses virtual metrology (VM) for the taping process of chip electronic components, in which partial observability, unmeasured disturbances, and severe label imbalance make direct batch-wise yield prediction unstable. Rather than proposing a new standalone learning algorithm, we develop a data-centric VM framework that reformulates the task as the prediction of operating-condition-level typical yield. First, physically relevant features are retained based on process knowledge and analyzed using Pearson correlation, Spearman correlation, and mutual information. We then perform multidimensional equal-frequency binning to partition the observable feature space into locally homogeneous operating condition groups, and define the within-bin median yield as the typical yield, thereby constructing an operating condition dictionary. Based on this dictionary-based representation, low-yield-oriented sample weighting is combined with nested cross-validation and Bayesian optimization for model comparison and hyperparameter tuning. Using desensitized production data from an electronic component taping process, the results under this representation show more stable prediction than direct modeling on unbinned batch samples while also improving tail-oriented fitting relative to unweighted baselines. These findings suggest that, for partially observable manufacturing data, operating condition stratification provides a practical basis for stabilizing VM prediction, while low-yield-oriented sample weighting further improves sensitivity to the low-yield tail, supporting picture yield early warning and process-level decision making. Full article
Show Figures

Figure 1

22 pages, 4848 KB  
Article
A Lightweight Improved RT-DETR for Stereo-Vision-Based Excavator Posture Recognition
by Yunlong Hou, Ke Wu, Yuhan Zhang, Mengying Zhou, Jiasheng Lu and Zhao Zhang
Mathematics 2026, 14(7), 1226; https://doi.org/10.3390/math14071226 - 7 Apr 2026
Cited by 1 | Viewed by 669
Abstract
In intelligent excavator applications, traditional excavator posture recognition methods face two major challenges: limited recognition accuracy and insufficient computing resources on edge devices. To address these issues, this study proposes an excavator posture recognition method based on an improved Real-Time Detection Transformer (RT-DETR). [...] Read more.
In intelligent excavator applications, traditional excavator posture recognition methods face two major challenges: limited recognition accuracy and insufficient computing resources on edge devices. To address these issues, this study proposes an excavator posture recognition method based on an improved Real-Time Detection Transformer (RT-DETR). First, a new backbone network is designed based on the Reparameterized Vision Transformer to improve feature utilization efficiency while reducing computational demands. Next, the overall architecture is optimized by introducing lightweight Dynamic Upsamplers, which reduce information loss during upsampling and enhance multi-scale feature fusion. In addition, a Cross-Attention Fusion Module is adopted to strengthen local feature extraction while retaining the global modeling capability of the Transformer, thereby improving the discrimination between foreground and background. Finally, a Multi-Scale Fusion Network is introduced to further enhance the multi-scale feature representation ability of RT-DETR. Experimental results show that the proposed method achieves a mean average precision (mAP) of 94.29% for small object detection, which is 7.96% higher than that of the baseline RT-DETR, while reducing the number of model parameters by 34.95%. Compared with YOLO-series models, the proposed method improves mAP by 8.62% to 12.75%. These results indicate that the proposed method outperforms existing methods in both detection accuracy and computational efficiency and provides an efficient and feasible solution for real-time excavator posture recognition. Full article
Show Figures

Figure 1

28 pages, 1152 KB  
Article
Enhanced Solution for the Advection–Diffusion–Reaction Equation Using the Physics-Informed Neural Network Technique
by Thabo Lekaba, Ndivhuwo Ndou, Kizito Muzhinji and Simiso Moyo
Mathematics 2026, 14(7), 1194; https://doi.org/10.3390/math14071194 - 2 Apr 2026
Cited by 1 | Viewed by 1297
Abstract
This study focuses on the use of Physics-Informed Neural Networks (PINNs) to solve the 1D Advection–Diffusion–Reaction (ADR) equation. The performance of the PINN model is evaluated in comparison with the classical Crank–Nicolson Finite Difference Method (CNFDM) and validated against analytical solutions to assess [...] Read more.
This study focuses on the use of Physics-Informed Neural Networks (PINNs) to solve the 1D Advection–Diffusion–Reaction (ADR) equation. The performance of the PINN model is evaluated in comparison with the classical Crank–Nicolson Finite Difference Method (CNFDM) and validated against analytical solutions to assess improvements in accuracy, robustness, and flexibility. Quantitative analysis reveals that the PINN achieved a high level of accuracy with absolute errors ranging from approximately 2.13×104 to 1.17×103 across the spatial domain. The study utilizes a neural network architecture with two hidden layers of 80 neurons each, optimized through a two-stage training process involving Adam and L-BFGS optimizers. This work contributes to the growing field of physics-informed machine learning by demonstrating the strengths and quantitative reliability of the PINN technique for solving complex partial differential equations in transport phenomena. Full article
Show Figures

Figure 1

28 pages, 14283 KB  
Article
FSD-YOLO: A Fusion Framework for Region Segmentation and Deformable Object Detection in Container Yards
by Linghao Dai, Zhihong Liang, Qi Feng, Shihuan Xie and Hongxu Li
Sensors 2026, 26(7), 2029; https://doi.org/10.3390/s26072029 - 24 Mar 2026
Viewed by 738
Abstract
Safety monitoring in container hoisting operations within rail-road intermodal logistics parks is a critical task in industrial safety management. Such scenarios are characterized by complex environments, large variations in target scales, deformable object shapes, and frequent occlusions, which pose significant challenges to visual [...] Read more.
Safety monitoring in container hoisting operations within rail-road intermodal logistics parks is a critical task in industrial safety management. Such scenarios are characterized by complex environments, large variations in target scales, deformable object shapes, and frequent occlusions, which pose significant challenges to visual perception systems. Conventional single-task models suffer from inherent limitations in handling low recall rates for distant small targets and insufficient adaptability to geometric deformations, making them inadequate for high-precision, real-time safety warning applications. To address these challenges, this study proposes a unified visual analysis framework that integrates semantic segmentation and object detection to enhance the recognition performance of small and deformable targets in complex operational environments, enabling real-time perception and safety warning of key objects and hazardous regions within container yards. Specifically, we introduce FSD-YOLO, a fusion-based architecture composed of the following key components. First, a SegFormer-based semantic segmentation module is employed to achieve pixel-level delineation of different operational regions. Second, an improved object detection network is developed based on the YOLOv8n architecture, incorporating: (1) the integration of C2f modules in the shallow layers of the backbone to enhance high-resolution feature extraction; (2) the embedding of C2fDCN modules within the detection head to improve modeling capability for deformable objects via deformable convolution; (3) the adoption of CARAFE upsampling operators to optimize multi-scale feature fusion; and (4) a dynamic loss-weighting strategy for small objects, where loss weights are adaptively adjusted according to target area to increase training emphasis on small-scale targets. Finally, a decision-level fusion strategy is applied to combine segmentation and detection outputs, enabling real-time safety judgment based on semantic rules. Experimental results on a self-constructed container yard dataset demonstrate that the proposed detection model achieves an mAP50-95 of 0.6433 and an mAP50 of 0.9565, significantly outperforming the baseline YOLOv8n model (mAP50-95: 0.5394, mAP50: 0.8435), thereby validating the effectiveness of the proposed framework. Full article
Show Figures

Figure 1

33 pages, 8140 KB  
Article
Diagnosing Shortcut Learning in CNN-Based Photovoltaic Fault Recognition from RGB Images: A Multi-Method Explainability Audit
by Bogdan Marian Diaconu
AI 2026, 7(3), 94; https://doi.org/10.3390/ai7030094 - 4 Mar 2026
Cited by 1 | Viewed by 1243
Abstract
Convolutional neural networks (CNNs) can achieve high accuracy in photovoltaic (PV) fault recognition from RGB imagery, yet their decisions may rely on shortcut cues induced by heterogeneous backgrounds, viewpoints, and class imbalance. This work presents a multi-method explainability audit on the Kaggle PV [...] Read more.
Convolutional neural networks (CNNs) can achieve high accuracy in photovoltaic (PV) fault recognition from RGB imagery, yet their decisions may rely on shortcut cues induced by heterogeneous backgrounds, viewpoints, and class imbalance. This work presents a multi-method explainability audit on the Kaggle PV Panel Defect Dataset (six classes), comparing five architectures (Baseline CNN, VGG16, ResNet50, InceptionV3, EfficientNetB0). Explanations are obtained with LIME superpixel surrogates (reported together with kernel-weighted surrogate fidelity), occlusion sensitivity (quantified via IoU@Top10% against consistent proxy masks, Shannon entropy, and Hoyer sparsity), and Integrated Gradients evaluated by deletion–insertion faithfulness and a Faithfulness Gap. While ResNet50 yields the best predictive performance, EfficientNetB0 shows the most consistent faithfulness evidence and stable panel-centered attributions. The analysis highlights class-dependent vulnerability to context cues, especially for the Clean and damaged classes, and supports using quantitative explainability diagnostics during model selection and dataset curation to mitigate shortcuts in vision-based PV monitoring. Full article
Show Figures

Figure 1

20 pages, 4824 KB  
Article
CIR-SQL: A Dual-Model Intent Recognition Framework for Chinese Text-to-SQL
by Yao Wang, Huiyong Lv and Yurong Qian
AI 2026, 7(3), 91; https://doi.org/10.3390/ai7030091 - 4 Mar 2026
Cited by 1 | Viewed by 1685
Abstract
In Industry 4.0 environments, operators and production managers frequently query industrial databases for production monitoring, quality control, and equipment maintenance using natural language. Existing Chinese NL2SQL systems often process semantic, program, and schema information in a single encoder, which leads to semantic-program interference [...] Read more.
In Industry 4.0 environments, operators and production managers frequently query industrial databases for production monitoring, quality control, and equipment maintenance using natural language. Existing Chinese NL2SQL systems often process semantic, program, and schema information in a single encoder, which leads to semantic-program interference and frequent structural or schema errors in the generated SQL. We present CIR-SQL, a dual-model framework that separates intent recognition from SQL generation via structured intermediate representations, decoupling semantic understanding from program synthesis. CIR-SQL employs a seven-category intent classification system (simple_select, count_query, filter_query, max_min_query, sort_query, join_query, group_by_query) and leverages large language models for intent recognition and structured information extraction. A three-level hierarchical backtracking strategy (SQL, context, intent) further improves robustness by correcting different error types. The architecture is particularly suited to Industry 4.0 scenarios where Chinese-speaking operators interact with complex industrial databases containing production data, quality metrics, and equipment status information. Full article
Show Figures

Figure 1

16 pages, 511 KB  
Article
A Comparative Study of Machine Learning and Deep Learning Models for Real-Time UAV Positioning Error Estimation
by Mei Yang, Hua Zhuo, Jun-Gang Ma, Guo-Hui Niu, Zulmira Mamtimin, Mei Tao, Ya-Qiong Zhu, Jun Li, Murat Abdughani and Aihemaitijiang Sidike
Drones 2026, 10(3), 172; https://doi.org/10.3390/drones10030172 - 2 Mar 2026
Viewed by 1110
Abstract
Accurate real-time positioning of Unmanned Aerial Vehicles (UAVs) is critical for navigation and mapping but remains challenging in complex environments due to signal blockages and multipath effects. This study presents a comparative framework for real-time error prediction of the Global Navigation Satellite System [...] Read more.
Accurate real-time positioning of Unmanned Aerial Vehicles (UAVs) is critical for navigation and mapping but remains challenging in complex environments due to signal blockages and multipath effects. This study presents a comparative framework for real-time error prediction of the Global Navigation Satellite System (GNSS), evaluating two machine learning models (Random Forest and XGBoost) and a deep learning model (Long Short-Term Memory network) against an Extended Kalman Filter baseline. A high-precision total station provides ground-truth coordinates, enabling the derivation of positioning error labels from synchronized GNSS raw data. Among the evaluated models, the tree-based XGBoost model achieves a significantly lower Mean Squared Error (MSE) and a considerably higher Coefficient of Determination (R2) score than other models in predicting positioning deviations. The high-accuracy error predictions from the optimal model establish the core of a software-only solution for positioning integrity. The framework demonstrates that reliable, real-time error estimates can be derived directly from observation data, providing the essential input required for future compensation systems without necessitating additional hardware. Full article
Show Figures

Figure 1

24 pages, 1964 KB  
Article
Research on Remaining Useful Life Prediction of Equipment Based on Digital Twins
by Jiaju Wu, Yuanlin Zhou, Xiaodong Wang, Chuan Chen, Yongqi Ma and Chunrui Zhang
Sensors 2026, 26(4), 1240; https://doi.org/10.3390/s26041240 - 13 Feb 2026
Cited by 2 | Viewed by 1600
Abstract
Remaining Useful Life (RUL) prediction is a key factor in fault diagnosis, prediction, and health management (PHM) during equipment operation and service. Its purpose is to predict the time interval from the current moment to the complete failure of the equipment, serving as [...] Read more.
Remaining Useful Life (RUL) prediction is a key factor in fault diagnosis, prediction, and health management (PHM) during equipment operation and service. Its purpose is to predict the time interval from the current moment to the complete failure of the equipment, serving as the basis for condition-based maintenance strategies. Effective RUL prediction enables the scheduling of maintenance plans in advance, thereby reducing equipment downtime and safety incidents. The RUL prediction of equipment and its critical components is an important means of fault diagnosis and prediction. Real-time and accurate RUL prediction results are prerequisites for implementing preventive maintenance, condition-based maintenance, and failure-based maintenance strategies, allowing the identification of optimal maintenance timing. This constitutes a crucial aspect of precise equipment support. The real-time, high-efficiency communication of digital twin technology can support real-time online RUL prediction for equipment. This paper introduces digital twin technology and constructs a digital twin-based RUL prediction model for equipment. The study proposes an integrated learning-based RUL prediction method for equipment, validated through experiments to demonstrate its accuracy and robustness. Finally, this paper presents an engineering implementation plan for online RUL prediction of equipment based on digital twins. Full article
Show Figures

Figure 1

21 pages, 27867 KB  
Article
An Adaptive Attention DropBlock Framework for Real-Time Cross-Domain Defect Classification
by Shailaja Pasupuleti, Ramalakshmi Krishnamoorthy and Hemalatha Gunasekaran
AI 2026, 7(2), 56; https://doi.org/10.3390/ai7020056 - 3 Feb 2026
Viewed by 1172
Abstract
The categorization of real-time defects in heterogeneous domains is a long-standing challenge in the field of industrial visual inspection systems, primarily due to significant visual variations and the lack of labelled information in real-world inspection settings. This work presents the Adaptive Attention DropBlock [...] Read more.
The categorization of real-time defects in heterogeneous domains is a long-standing challenge in the field of industrial visual inspection systems, primarily due to significant visual variations and the lack of labelled information in real-world inspection settings. This work presents the Adaptive Attention DropBlock (AADB) framework, a lightweight deep learning framework that was developed to promote cross-domain defect detection using attention-guided regularization. The proposed architecture integrates the Convolutional Block Attention Module (CBAM) and an organized DropBlock-based regularization scheme, creating a unified and robust framework. Although CBAM-based approaches improve localization of defect-related areas and traditional DropBlock provides a generic spatial regularization, neither of them alone is specifically designed to reduce domain overfitting. To address this limitation, AADB combines attention-directed feature refinement with a progressive, transfer-aware dropout policy that promotes the learning of domain-invariant representations. The proposed model is built on a MobileNetV2 base and trained through a two-phase transfer learning regime, where the first phase consists of pretraining on a source domain and the second phase consists of adaptation to a visually dissimilar target domain with constrained supervision. The overall analysis of a metal surface defect dataset (source domain) and an aircraft surface defect dataset (target domain) shows that AADB outperforms CBAM-only, DropBlock-only, and conventional MobileNetV2 models, with an overall accuracy of 91.06%, a macro-F1 of 0.912, and a Cohen’s k of 0.866. Improved feature separability and localization of error are further described by qualitative analyses using Principal Component Analysis (PCA) and Grad-CAM. Overall, the framework provides a practical, interpretable, and edge-deployable solution to the classification of cross-domain defects in the industrial inspection setting. Full article
Show Figures

Figure 1

12 pages, 4449 KB  
Article
Modeling Extreme Rainfall Using the Generalized Extreme Value Distribution and Exceedance Analysis in Colima, Mexico
by Raúl Renteria, Raúl Aquino and Mayrén Polanco
Sensors 2026, 26(2), 532; https://doi.org/10.3390/s26020532 - 13 Jan 2026
Cited by 1 | Viewed by 948
Abstract
This study develops a statistical and technological framework to analyze extreme rainfall in Colima, Mexico, by integrating historical precipitation records, probabilistic modeling, and spatial visualization. Using data from CONAGUA meteorological stations, we identify high-intensity rainfall events and model their recurrence using the Generalized [...] Read more.
This study develops a statistical and technological framework to analyze extreme rainfall in Colima, Mexico, by integrating historical precipitation records, probabilistic modeling, and spatial visualization. Using data from CONAGUA meteorological stations, we identify high-intensity rainfall events and model their recurrence using the Generalized Extreme Value (GEV) distribution to estimate key return periods. The results support flood-risk assessment and territorial planning in Colima. Spatial interpolation was performed in Python (version 3.13), and QGIS (version 3.38) produces exceedance maps that illustrate geographic variations in rainfall intensity across the state. These exceedance maps reveal a consistent spatial pattern, with the northern and western areas of Colima experiencing the highest frequencies of extreme events. Based on these results, the integration of real-time sensor technologies and satellite observations may improve flood monitoring and risk management frameworks. Full article
Show Figures

Figure 1

19 pages, 3374 KB  
Article
SpaceNet: A Multimodal Fusion Architecture for Sound Source Localization in Disaster Response
by Long Nguyen-Vu and Jonghoon Lee
Sensors 2026, 26(1), 168; https://doi.org/10.3390/s26010168 - 26 Dec 2025
Cited by 1 | Viewed by 990
Abstract
Sound source localization (SSL) has evolved from traditional signal-processing methods to sophisticated deep-learning architectures. However, applying these to distributed microphone arrays in adverse environments is complicated by high reverberation and potential sensor asynchrony, which can corrupt crucial Time-Difference-of-Arrival (TDoA) information. We introduce SpaceNet, [...] Read more.
Sound source localization (SSL) has evolved from traditional signal-processing methods to sophisticated deep-learning architectures. However, applying these to distributed microphone arrays in adverse environments is complicated by high reverberation and potential sensor asynchrony, which can corrupt crucial Time-Difference-of-Arrival (TDoA) information. We introduce SpaceNet, a multimodal deep-learning architecture designed to address such issues by explicitly fusing audio features with sensor geometry. SpaceNet features: (1) a dual-branch architecture with specialized spatial processing that decomposes microphone geometry into distances, azimuths, and elevations; and (2) a feature-normalization technique to ensure stable multimodal training. Evaluation on real-world datasets from disaster sites demonstrates that SpaceNet, when trained on ILD-only mel-spectra, achieves better accuracy compared to our baseline model (CHAWA) and identical models trained on full mel-spectrograms. This approach also reduces computational overhead by a factor of 24. Our findings suggest that for distributed arrays in adverse environments, time-invariant ILD cues are a more effective and efficient feature for localization than complex temporal features corrupted by reverberation and synchronization errors. Full article
Show Figures

Figure 1

30 pages, 4546 KB  
Article
TCN-LSTM-AM Short-Term Photovoltaic Power Forecasting Model Based on Improved Feature Selection and APO
by Ning Ye, Chaoyang Zhi, Yongchao Yu, Sen Lin and Fengxian Liu
Sensors 2025, 25(24), 7607; https://doi.org/10.3390/s25247607 - 15 Dec 2025
Cited by 7 | Viewed by 1311
Abstract
The inherent volatility and intermittency of solar power generation pose significant challenges to the stability of power systems. Consequently, high-precision power forecasting is critical for mitigating these impacts and ensuring reliable operation. This paper proposes a framework for photovoltaic (PV) power forecasting that [...] Read more.
The inherent volatility and intermittency of solar power generation pose significant challenges to the stability of power systems. Consequently, high-precision power forecasting is critical for mitigating these impacts and ensuring reliable operation. This paper proposes a framework for photovoltaic (PV) power forecasting that integrates refined feature engineering with deep learning models in a two-stage approach. In the feature engineering stage, a KNN-PCC-SHAP method is constructed. This method is initiated with the KNN algorithm, which is used to identify anomalous samples and perform data interpolation. PCC is then used to screen linearly correlated features. Finally, the SHAP value is used to quantitatively analyze the nonlinear contributions and interaction effects of each feature, thereby forming an optimal feature subset with higher information density. In the modeling stage, a TCN-LSTM-AM combined forecasting model is constructed to collaboratively capture the local details, long-term dependencies, and key timing features of the PV power sequence. The APO algorithm is utilized for the adaptive optimization of the crucial configuration parameters within the model. Experiments based on real PV power plants and public data show that the framework outperforms multiple comparison models in terms of key indicators such as RMSE (2.1098 kW), MAE (1.1073 kW), and R2 (0.9775), verifying that the deep integration of refined feature engineering and deep learning models is an effective way to improve the accuracy of PV power prediction. Full article
Show Figures

Figure 1

29 pages, 16069 KB  
Article
Dynamic Severity Assessment of Partial Discharge in HV Bushings Based on the Evolution Characteristics of Dense Clusters in PRPD Patterns
by Xiang Gao, Zhiyu Li, Zuoming Xu, Pengbo Yin, Xiongjie Xie, Xiaochen Yang and Baoquan Wan
Sensors 2025, 25(24), 7537; https://doi.org/10.3390/s25247537 - 11 Dec 2025
Cited by 2 | Viewed by 1140
Abstract
High-voltage bushings are critical insulation components, yet conventional PRPD-based severity assessment methods that rely on global pattern morphologies such as “rabbit ears” and “tortoise shell” remain coarse, lack local sensitivity, and fail to track continuous degradation. This paper proposes a dynamic severity assessment [...] Read more.
High-voltage bushings are critical insulation components, yet conventional PRPD-based severity assessment methods that rely on global pattern morphologies such as “rabbit ears” and “tortoise shell” remain coarse, lack local sensitivity, and fail to track continuous degradation. This paper proposes a dynamic severity assessment method that shifts the focus from global contours to dense partial discharge (PD) clusters, defined as high-density aggregations of PD pulses in specific phase–magnitude regions of PRPD patterns. Each dense cluster is treated as the statistical projection of a physical discharge channel, and the evolution of its number, intensity, location, and shape provides a fine-scale description of defect development. A multi-level relative density and morphological image processing algorithm is used to extract dense clusters directly from PRPD histograms, followed by a 20-dimensional feature set and a five-index system describing discharge activity, development speed, complexity, instability, and evolution trend. A fuzzy comprehensive evaluation model further converts these indices into three severity levels with confidence measures. Long-term degradation tests on defective bushings demonstrate that the proposed method captures key turning points from dispersed multi-cluster patterns to a single dominant cluster and yields a stable, stage-consistent severity evaluation, offering a more sensitive and physically interpretable tool for condition monitoring and early warning of HV bushings. The method achieved a high evaluation confidence (average 60.1%), which rose to 100% at the critical failure stage. It successfully identified three distinct degradation stages (stable, accelerated, and critical) across the 49 test intervals. A quantitative comparison demonstrated significant advantages: 8.3% improvement in early warning (4 windows earlier than IEC 60270), 50.6% higher monotonicity, 125.2% better stability, and 45.9% wider dynamic range, while maintaining physical interpretability and requiring no training data. Full article
Show Figures

Figure 1

21 pages, 335 KB  
Review
AI-Driven Motion Capture Data Recovery: A Comprehensive Review and Future Outlook
by Ahood Almaleh, Gary Ushaw and Rich Davison
Sensors 2025, 25(24), 7525; https://doi.org/10.3390/s25247525 - 11 Dec 2025
Viewed by 1819
Abstract
This paper presents a comprehensive review of motion capture (MoCap) data recovery techniques, with a particular focus on the suitability of artificial intelligence (AI) for addressing missing or corrupted motion data. Existing approaches are classified into three categories: non-data-driven, data-driven (AI-based), and hybrid [...] Read more.
This paper presents a comprehensive review of motion capture (MoCap) data recovery techniques, with a particular focus on the suitability of artificial intelligence (AI) for addressing missing or corrupted motion data. Existing approaches are classified into three categories: non-data-driven, data-driven (AI-based), and hybrid methods. Within the AI domain, frameworks such as generative adversarial networks (GANs), transformers, and graph neural networks (GNNs) demonstrate strong capabilities in modeling complex spatial–temporal dependencies and achieving accurate motion reconstruction. Compared with traditional methods, AI techniques offer greater adaptability and precision, though they remain limited by high computational costs and dependence on large, high-quality datasets. Hybrid approaches that combine AI models with physics-based or statistical algorithms provide a balance between efficiency, interpretability, and robustness. The review also examines benchmark datasets, including CMU MoCap and Human3.6M, while highlighting the growing role of synthetic and augmented data in improving AI model generalization. Despite notable progress, the absence of standardized evaluation protocols and diverse real-world datasets continues to hinder generalization. Emerging trends point toward real-time AI-driven recovery, multimodal data fusion, and unified performance benchmarks. By integrating traditional, AI-based, and hybrid approaches into a coherent taxonomy, this review provides a unique contribution to the literature. Unlike prior surveys focused on prediction, denoising, pose estimation, or generative modeling, it treats MoCap recovery as a standalone problem. It further synthesizes comparative insights across datasets, evaluation metrics, movement representations, and common failure cases, offering a comprehensive foundation for advancing MoCap recovery research. Full article
Show Figures

Figure 1

25 pages, 3204 KB  
Article
A Classified Branch–CapNet: A Multi-Modal Model with Classified Branches for the Capacity Prediction of Li–Ion Battery Cathodes
by Junghee Kim, Jaehyeok Yang and Daewon Chung
Mathematics 2025, 13(22), 3730; https://doi.org/10.3390/math13223730 - 20 Nov 2025
Viewed by 911
Abstract
Machine learning has emerged as a promising tool to accelerate the screening of lithium–ion battery electrode materials. Gravimetric capacity, a critical performance indicator governing electrode energy density, is intrinsically related to lithium insertion and extraction mechanisms, requiring sophisticated embedding approaches that capture the [...] Read more.
Machine learning has emerged as a promising tool to accelerate the screening of lithium–ion battery electrode materials. Gravimetric capacity, a critical performance indicator governing electrode energy density, is intrinsically related to lithium insertion and extraction mechanisms, requiring sophisticated embedding approaches that capture the structural characteristics of cathode materials. The cathode material dataset from the Materials Project database comprises heterogeneous data modalities: numerical features representing chemical properties and categorical features encoding structural characteristics. Naive integration of these disparate data types may introduce semantic gaps from statistical distributional discrepancies, potentially degrading predictive performance and limiting model generalization. To address these limitations, this study proposes a Classified Branch–CapNet model that individually embeds four distinct types of categorical structural data into separate classified branches along with numerical data for independent learning, subsequently integrating them through a late fusion strategy. This approach minimizes interference between heterogeneous data modalities while capturing structure–property relationships with enhanced precision. The proposed model achieved superior performance with a mean absolute error of 2.441 mAh/g, demonstrating substantial improvements of 56.2%, 71.2%, 73.9%, and 51.1% over conventional deep neural networks, recurrent neural networks, long short-term memory architectures, and the encoder-only Transformer, respectively. Furthermore, it achieved the lowest root mean square error of 15.236 mAh/g and the highest coefficient of determination of 0.961, confirming its superior predictive accuracy and generalization capability compared with all benchmark models. Our model therefore demonstrates significant potential to accelerate the efficient screening and discovery of high-performance battery electrode materials. Full article
Show Figures

Figure 1

30 pages, 8790 KB  
Article
An Adaptive Framework for Remaining Useful Life Prediction Integrating Attention Mechanism and Deep Reinforcement Learning
by Yanhui Bai, Jiajia Du, Honghui Li, Xintao Bao, Linjun Li, Chun Zhang, Jiahe Yan, Renliang Wang and Yi Xu
Sensors 2025, 25(20), 6354; https://doi.org/10.3390/s25206354 - 14 Oct 2025
Cited by 2 | Viewed by 1953
Abstract
The prediction of Remaining Useful Life (RUL) constitutes a vital aspect of Prognostics and Health Management (PHM), providing capabilities for the assessment of mechanical component health status and prediction of failure instances. Recent studies on feature extraction, time-series modeling, and multi-task learning have [...] Read more.
The prediction of Remaining Useful Life (RUL) constitutes a vital aspect of Prognostics and Health Management (PHM), providing capabilities for the assessment of mechanical component health status and prediction of failure instances. Recent studies on feature extraction, time-series modeling, and multi-task learning have shown remarkable advancements. However, most deep learning (DL) techniques predominantly focus on unimodal data or static feature extraction techniques, resulting in a lack of RUL prediction methods that can effectively capture the individual differences among heterogeneous sensors and failure modes under complex operational conditions. To overcome these limitations, an adaptive RUL prediction framework named ADAPT-RULNet is proposed for mechanical components, integrating the feature extraction capabilities of attention-enhanced deep learning (DL) and the decision-making abilities of deep reinforcement learning (DRL) to achieve end-to-end optimization from raw data to accurate RUL prediction. Initially, Functional Alignment Resampling (FAR) is employed to generate high-quality functional signals; then, attention-enhanced Dynamic Time Warping (DTW) is leveraged to obtain individual degradation stages. Subsequently, an attention-enhanced of hybrid multi-scale RUL prediction network is constructed to extract both local and global features from multi-format data. Furthermore, the network achieves optimal feature representation by adaptively fusing multi-source features through Bayesian methods. Finally, we innovatively introduce a Deep Deterministic Policy Gradient (DDPG) strategy from DRL to adaptively optimize key parameters in the construction of individual degradation stages and achieve a global balance between model complexity and prediction accuracy. The proposed model was evaluated on aircraft engines and railway freight car wheels. The results indicate that it achieves a lower average Root Mean Square Error (RMSE) and higher accuracy in comparison with current approaches. Moreover, the method shows strong potential for improving prediction accuracy and robustness in varied industrial applications. Full article
Show Figures

Figure 1

32 pages, 3383 KB  
Article
DLG–IDS: Dynamic Graph and LLM–Semantic Enhanced Spatiotemporal GNN for Lightweight Intrusion Detection in Industrial Control Systems
by Junyi Liu, Jiarong Wang, Tian Yan, Fazhi Qi and Gang Chen
Electronics 2025, 14(19), 3952; https://doi.org/10.3390/electronics14193952 - 7 Oct 2025
Cited by 3 | Viewed by 2180
Abstract
Industrial control systems (ICSs) face escalating security challenges due to evolving cyber threats and the inherent limitations of traditional intrusion detection methods, which fail to adequately model spatiotemporal dependencies or interpret complex protocol semantics. To address these gaps, this paper proposes DLG–IDS—a lightweight [...] Read more.
Industrial control systems (ICSs) face escalating security challenges due to evolving cyber threats and the inherent limitations of traditional intrusion detection methods, which fail to adequately model spatiotemporal dependencies or interpret complex protocol semantics. To address these gaps, this paper proposes DLG–IDS—a lightweight intrusion detection framework that innovatively integrates dynamic graph construction for capturing real–time device interactions and logical control relationships from traffic, LLM–driven semantic enhancement to extract fine–grained embeddings from graphs, and a spatio–temporal graph neural network (STGNN) optimized via sparse attention and local window Transformers to minimize computational overhead. Evaluations on SWaT and SBFF datasets demonstrate the framework’s superiority, achieving a state–of–the–art accuracy of 0.986 while reducing latency by 53.2% compared to baseline models. Ablation studies further validate the critical contributions of semantic fusion, sparse topology modeling, and localized temporal attention. The proposed solution establishes a robust, real–time detection mechanism tailored for resource–constrained industrial environments, effectively balancing high accuracy with operational efficiency. Full article
Show Figures

Figure 1

16 pages, 7297 KB  
Article
Attention-Based Multi-Agent RL for Multi-Machine Tending Using Mobile Robots
by Abdalwhab Bakheet Mohamed Abdalwhab, Giovanni Beltrame, Samira Ebrahimi Kahou and David St-Onge
AI 2025, 6(10), 252; https://doi.org/10.3390/ai6100252 - 1 Oct 2025
Cited by 3 | Viewed by 2680
Abstract
Robotics can help address the growing worker shortage challenge of the manufacturing industry. As such, machine tending is a task collaborative robots can tackle that can also greatly boost productivity. Nevertheless, existing robotics systems deployed in that sector rely on a fixed single-arm [...] Read more.
Robotics can help address the growing worker shortage challenge of the manufacturing industry. As such, machine tending is a task collaborative robots can tackle that can also greatly boost productivity. Nevertheless, existing robotics systems deployed in that sector rely on a fixed single-arm setup, whereas mobile robots can provide more flexibility and scalability. We introduce a multi-agent multi-machine-tending learning framework using mobile robots based on multi-agent reinforcement learning (MARL) techniques, with the design of a suitable observation and reward. Moreover, we integrate an attention-based encoding mechanism into the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm to boost its performance for machine-tending scenarios. Our model (AB-MAPPO) outperforms MAPPO in this new challenging scenario in terms of task success, safety, and resource utilization. Furthermore, we provided an extensive ablation study to support our design decisions. Full article
Show Figures

Figure 1

21 pages, 527 KB  
Article
Block-CITE: A Blockchain-Based Crowdsourcing Interactive Trust Evaluation
by Jiaxing Li, Lin Jiang, Haoxian Liang, Tao Peng, Shaowei Wang and Huanchun Wei
AI 2025, 6(10), 245; https://doi.org/10.3390/ai6100245 - 1 Oct 2025
Cited by 2 | Viewed by 1283
Abstract
Industrial trademark examination enables users to apply for and manage their trademarks efficiently, promoting industrial and commercial economic development. However, there still exist many challenges, e.g., how to customize a blockchain-based crowdsourcing method for interactive trust evaluation, how to decentralize the functionalities of [...] Read more.
Industrial trademark examination enables users to apply for and manage their trademarks efficiently, promoting industrial and commercial economic development. However, there still exist many challenges, e.g., how to customize a blockchain-based crowdsourcing method for interactive trust evaluation, how to decentralize the functionalities of a centralized entity to nodes in a blockchain network instead of removing the entity directly, how to design a protocol for the method and prove its security, etc. In order to overcome these challenges, in this paper, we propose the Blockchain-based Crowdsourcing Interactive Trust Evaluation (Block-CITE for short) method to improve the efficiency and security of the current industrial trademark management schemes. Specifically, Block-CITE adopts a dual-blockchain structure and a crowdsourcing technique to record operations and store relevant data in a decentralized way. Furthermore, Block-CITE customizes a protocol for blockchain-based crowdsourced industrial trademark examination and algorithms of smart contracts to run the protocol automatically. In addition, Block-CITE analyzes the threat model and proves the security of the protocol. Security analysis shows that Block-CITE is able to defend against the malicious entities and attacks in the blockchain network. Experimental analysis shows that Block-CITE has a higher transaction throughput and lower network latency and storage overhead than the baseline methods. Full article
Show Figures

Figure 1

27 pages, 4744 KB  
Article
Intelligent Soft Sensor for Spindle Convective Heat Transfer Coefficient Under Varying Operating Conditions Using Improved Grey Wolf Optimization Algorithm
by Jinxiang Pian and Gen Li
Sensors 2025, 25(18), 5806; https://doi.org/10.3390/s25185806 - 17 Sep 2025
Cited by 1 | Viewed by 983
Abstract
The thermal deformation of high-precision CNC machine tools has long been a significant barrier to improving machining accuracy. Accurately characterizing the thermal properties of the spindle, especially the convective heat transfer coefficients (CHTC), is essential for precise thermal analysis. However, due to the [...] Read more.
The thermal deformation of high-precision CNC machine tools has long been a significant barrier to improving machining accuracy. Accurately characterizing the thermal properties of the spindle, especially the convective heat transfer coefficients (CHTC), is essential for precise thermal analysis. However, due to the lack of dedicated instruments for directly measuring the CHTC, thermal analysis of the spindle faces substantial challenges. This study presents an innovative approach that combines multi-sensor data with intelligent optimization algorithms to address this issue. A distributed temperature monitoring network is constructed to capture real-time thermal field data across the spindle. At the same time, an improved Grey Wolf Optimization (IGWO) algorithm is employed to dynamically and accurately identify the CHTC. The proposed algorithm introduces an adaptive weight adjustment mechanism, which overcomes the limitations of traditional optimization methods in dynamic operating conditions. Experimental results show that the proposed method significantly outperforms conventional approaches in terms of temperature prediction accuracy across a broad operating range. This research provides a novel technical solution for machine tool thermal error compensation and establishes a scalable intelligent indirect measurement framework, even in the absence of specialized measurement instruments. Full article
Show Figures

Figure 1

27 pages, 30539 KB  
Article
Priori Knowledge Makes Low-Light Image Enhancement More Reasonable
by Zefei Chen, Yongjie Lin, Jianmin Xu, Kai Lu and Zihao Huang
Sensors 2025, 25(17), 5521; https://doi.org/10.3390/s25175521 - 4 Sep 2025
Viewed by 1852
Abstract
This paper presents a priori knowledge-based low-light image enhancement framework, termed Priori DCE (Priori Deep Curve Estimation). The priori knowledge consists of two key aspects: (1) enhancing a low-light image is an ill-posed task, as the brightness of the enhanced image corresponding to [...] Read more.
This paper presents a priori knowledge-based low-light image enhancement framework, termed Priori DCE (Priori Deep Curve Estimation). The priori knowledge consists of two key aspects: (1) enhancing a low-light image is an ill-posed task, as the brightness of the enhanced image corresponding to a low-light image is uncertain. To resolve this issue, we incorporate priori channels into the model to guide the brightness of the enhanced image; (2) during the enhancement of a low-light image, the brightness of pixels may increase or decrease. This paper explores the probability of a pixel’s brightness increasing/decreasing as its prior enhancement/suppression probability. Intuitively, pixels with higher brightness should have a higher priori suppression probability, while pixels with lower brightness should have a higher priori enhancement probability. Inspired by this, we propose an enhancement function that adaptively adjusts the priori enhancement probability based on variations in pixel brightness. In addition, we propose the Global-Attention Block (GA Block). The GA Block ensures that, during the low-light image enhancement process, each pixel in the enhanced image is computed based on all the pixels in the low-light image. This approach facilitates interactions between all pixels in the enhanced image, thereby achieving visual balance. The experimental results on the LOLv2-Synthetic dataset demonstrate that Priori DCE has a significant advantage. Specifically, compared to the SOTA Retinexformer, the Priori DCE improves the PSNR index and SSIM index from 25.67 and 92.82 to 29.49 and 93.6, respectively, while the NIQE index decreases from 3.94 to 3.91. Full article
Show Figures

Figure 1

17 pages, 624 KB  
Article
Predicting Out-of-Stock Risk Under Delivery Schedules Using Neural Networks
by Lu Xu
Electronics 2025, 14(15), 3012; https://doi.org/10.3390/electronics14153012 - 29 Jul 2025
Cited by 6 | Viewed by 1828
Abstract
In retail logistics, one typical task is to arrange a delivery schedule that guides the intake of inventory from the distribution center to stores. It is essential to accurately predict the out-of-stock (OOS) outcome for various delivery schedules to identify the optimal patterns [...] Read more.
In retail logistics, one typical task is to arrange a delivery schedule that guides the intake of inventory from the distribution center to stores. It is essential to accurately predict the out-of-stock (OOS) outcome for various delivery schedules to identify the optimal patterns for minimizing the OOS ratio. This paper investigates the feasibility of utilizing a neural network to accurately predict the out-of-stock (OOS) risk under each delivery pattern. Due to the zero-inflated distribution of the target values, it is necessary to evaluate two prediction accuracies simultaneously: the accuracy on data with a positive ground truth OOS rate and the accuracy on data with a zero ground truth OOS rate. In this paper, I examine how a selection of features associated with delivery schedules and the choice of activation function at the output layer, would impact the accuracy of the model. Full article
Show Figures

Figure 1

18 pages, 7391 KB  
Article
Reliable QoE Prediction in IMVCAs Using an LMM-Based Agent
by Michael Sidorov, Tamir Berger, Jonathan Sterenson, Raz Birman and Ofer Hadar
Sensors 2025, 25(14), 4450; https://doi.org/10.3390/s25144450 - 17 Jul 2025
Cited by 1 | Viewed by 1256
Abstract
Face-to-face interaction is one of the most natural forms of human communication. Unsurprisingly, Video Conferencing (VC) Applications have experienced a significant rise in demand over the past decade. With the widespread availability of cellular devices equipped with high-resolution cameras, Instant Messaging Video Call [...] Read more.
Face-to-face interaction is one of the most natural forms of human communication. Unsurprisingly, Video Conferencing (VC) Applications have experienced a significant rise in demand over the past decade. With the widespread availability of cellular devices equipped with high-resolution cameras, Instant Messaging Video Call Applications (IMVCAs) now constitute a substantial portion of VC communications. Given the multitude of IMVCA options, maintaining a high Quality of Experience (QoE) is critical. While content providers can measure QoE directly through end-to-end connections, Internet Service Providers (ISPs) must infer QoE indirectly from network traffic—a non-trivial task, especially when most traffic is encrypted. In this paper, we analyze a large dataset collected from WhatsApp IMVCA, comprising over 25,000 s of VC sessions. We apply four Machine Learning (ML) algorithms and a Large Multimodal Model (LMM)-based agent, achieving mean errors of 4.61%, 5.36%, and 13.24% for three popular QoE metrics: BRISQUE, PIQE, and FPS, respectively. Full article
Show Figures

Figure 1

38 pages, 3698 KB  
Review
Enhancing Autonomous Truck Navigation in Underground Mines: A Review of 3D Object Detection Systems, Challenges, and Future Trends
by Ellen Essien and Samuel Frimpong
Drones 2025, 9(6), 433; https://doi.org/10.3390/drones9060433 - 14 Jun 2025
Cited by 10 | Viewed by 6105
Abstract
Integrating autonomous haulage systems into underground mining has revolutionized safety and operational efficiency. However, deploying 3D detection systems for autonomous truck navigation in such an environment faces persistent challenges due to dust, occlusion, complex terrains, and low visibility. This affects their reliability and [...] Read more.
Integrating autonomous haulage systems into underground mining has revolutionized safety and operational efficiency. However, deploying 3D detection systems for autonomous truck navigation in such an environment faces persistent challenges due to dust, occlusion, complex terrains, and low visibility. This affects their reliability and real-time processing. While existing reviews have discussed object detection techniques and sensor-based systems, providing valuable insights into their applications, only a few have addressed the unique underground challenges that affect 3D detection models. This review synthesizes the current advancements in 3D object detection models for underground autonomous truck navigation. It assesses deep learning algorithms, fusion techniques, multi-modal sensor suites, and limited datasets in an underground detection system. This study uses systematic database searches with selection criteria for relevance to underground perception. The findings of this work show that the mid-level fusion method for combining different sensor suites enhances robust detection. Though YOLO (You Only Look Once)-based detection models provide superior real-time performance, challenges persist in small object detection, computational trade-offs, and data scarcity. This paper concludes by identifying research gaps and proposing future directions for a more scalable and resilient underground perception system. The main novelty is its review of underground 3D detection systems in autonomous trucks. Full article
Show Figures

Figure 1

30 pages, 5391 KB  
Article
Dual-Resource Scheduling with Improved Forensic-Based Investigation Algorithm in Smart Manufacturing
by Yuhang Zeng, Ping Lou, Jianmin Hu, Chuannian Fan, Quan Liu and Jiwei Hu
Mathematics 2025, 13(9), 1432; https://doi.org/10.3390/math13091432 - 27 Apr 2025
Viewed by 1377
Abstract
With increasing labor costs and rapidly dynamic changes in the market demand, as well as realizing the refined management of production, more and more attention is being given to considering workers, not just machines, in the process of flexible job shop scheduling. Hence, [...] Read more.
With increasing labor costs and rapidly dynamic changes in the market demand, as well as realizing the refined management of production, more and more attention is being given to considering workers, not just machines, in the process of flexible job shop scheduling. Hence, a new dual-resource flexible job shop scheduling problem (DRFJSP) is put forward in this paper, considering workers with flexible working time arrangements and machines with versatile functions in scheduling production, as well as a multi-objective mathematical model for formalizing the DRFJSP and tackling the complexity of scheduling in human-centric manufacturing environments. In addition, a two-stage approach based on a forensic-based investigation (TSFBI) is proposed to solve the problem. In the first stage, an improved multi-objective FBI algorithm is used to obtain the Pareto front solutions of this model, in which a hybrid real and integer encoding–decoding method is used for exploring the solution space and a fast non-dominated sorting method for improving efficiency. In the second stage, a multi-criteria decision analysis method based on an analytic hierarchy process (AHP) is used to select the optimal solution from the Pareto front solutions. Finally, experiments validated the TSFBI algorithm, showing its potential for smart manufacturing. Full article
Show Figures

Figure 1

21 pages, 9224 KB  
Article
A Multi-Scale Fusion Convolutional Network for Time-Series Silicon Prediction in Blast Furnaces
by Qiancheng Hao, Wenjing Liu, Wenze Gao and Xianpeng Wang
Mathematics 2025, 13(8), 1347; https://doi.org/10.3390/math13081347 - 20 Apr 2025
Cited by 2 | Viewed by 1454
Abstract
In steel production, the blast furnace is a critical element. In this process, precisely controlling the temperature of the molten iron is indispensable for attaining efficient operations and high-grade products. This temperature is often indirectly reflected by the silicon content in the hot [...] Read more.
In steel production, the blast furnace is a critical element. In this process, precisely controlling the temperature of the molten iron is indispensable for attaining efficient operations and high-grade products. This temperature is often indirectly reflected by the silicon content in the hot metal. However, due to the dynamic nature and inherent delays of the ironmaking process, real-time prediction of silicon content remains a significant challenge, and traditional methods often suffer from insufficient prediction accuracy. This study presents a novel Multi-Scale Fusion Convolutional Neural Network (MSF-CNN) to accurately predict the silicon content during the blast furnace smelting process, addressing the limitations of existing data-driven approaches. The proposed MSF-CNN model extracts temporal features at two distinct scales. The first scale utilizes a Convolutional Block Attention Module, which captures local temporal dependencies by focusing on the most relevant features across adjacent time steps. The second scale employs a Multi-Head Self-Attention mechanism to model long-term temporal dependencies, overcoming the inherent delay issues in the blast furnace process. By combining these two scales, the model effectively captures both short-term and long-term temporal dependencies, thereby enhancing prediction accuracy and real-time applicability. Validation using real blast furnace data demonstrates that MSF-CNN outperforms recurrent neural network models such as Long Short-Term Memory (LSTM) and the Gated Recurrent Unit (GRU). Compared with LSTM and the GRU, MSF-CNN reduces the Root Mean Square Error (RMSE) by approximately 22% and 21%, respectively, and improves the Hit Rate (HR) by over 3.5% and 4%, highlighting its superiority in capturing complex temporal dependencies. These results indicate that the MSF-CNN adapts better to the blast furnace’s dynamic variations and inherent delays, achieving significant improvements in prediction precision and robustness compared to state-of-the-art recurrent models. Full article
Show Figures

Figure 1

25 pages, 8768 KB  
Article
Towards More Accurate Industrial Anomaly Detection: A Component-Level Feature-Enhancement Approach
by Xiaodong Wang, Zhiyao Xie, Fei Yan, Jiayu Wang, Jiangtao Fan, Zhiqiang Zeng, Junwen Lu, Hangqi Zhang and Nianfeng Zeng
Electronics 2025, 14(8), 1613; https://doi.org/10.3390/electronics14081613 - 16 Apr 2025
Cited by 5 | Viewed by 4831
Abstract
Industrial visual inspection plays a crucial role in intelligent manufacturing. However, existing anomaly-detection methods based on unsupervised learning paradigms often struggle with issues such as overlooking minor defects and blurring component edges in confidence maps. To address these challenges, this paper proposes an [...] Read more.
Industrial visual inspection plays a crucial role in intelligent manufacturing. However, existing anomaly-detection methods based on unsupervised learning paradigms often struggle with issues such as overlooking minor defects and blurring component edges in confidence maps. To address these challenges, this paper proposes an industrial anomaly-detection method based on component-level feature enhancement. This method introduces a component-level feature-enhancement module, which optimizes feature matching by calculating the structural similarity between global coarse-grained confidence features and local fine-grained confidence features, thereby generating enhanced feature maps to improve the model’s detection accuracy for minor defects and local anomalies. Additionally, we propose a region-segmentation method based on multi-layer piecewise thresholds, which effectively distinguishes between foreground and background in confidence maps, circumvents background interference and ensures the integrity of structural information of foreground components. Experimental results demonstrate that the proposed method surpasses comparative methods in both logical and structural defect detection tasks, showing significant advantages, especially in fine-grained anomaly detection, with stronger robustness and accuracy. Full article
Show Figures

Figure 1

23 pages, 3749 KB  
Article
Proposed Long Short-Term Memory Model Utilizing Multiple Strands for Enhanced Forecasting and Classification of Sensory Measurements
by Sotirios Kontogiannis, George Kokkonis and Christos Pikridas
Mathematics 2025, 13(8), 1263; https://doi.org/10.3390/math13081263 - 11 Apr 2025
Cited by 3 | Viewed by 1465
Abstract
This paper presents a new deep learning model called the stranded Long Short-Term Memory. The model utilizes arbitrary LSTM recurrent neural networks of variable cell depths organized in classes. The proposed model can adapt to classifying emergencies at different intervals or provide measurement [...] Read more.
This paper presents a new deep learning model called the stranded Long Short-Term Memory. The model utilizes arbitrary LSTM recurrent neural networks of variable cell depths organized in classes. The proposed model can adapt to classifying emergencies at different intervals or provide measurement predictions using class-annotated or time-shifted series of sensory data inputs. In order to outperform the ordinary LSTM model’s classifications or forecasts by minimizing losses, stranded LSTM maintains three different weight-based strategies that can be arbitrarily selected prior to model training, as follows: least loss, weighted least loss, and fuzzy least loss in the LSTM model selection and inference process. The model has been tested against LSTM models for forecasting and classification, using a time series of temperature and humidity measurements taken from meteorological stations and class-annotated temperature measurements from Industrial compressors accordingly. From the experimental classification results, the stranded LSTM model outperformed 0.9–2.3% of the LSTM models carrying dual-stacked LSTM cells in terms of accuracy. Regarding the forecasting experimental results, the forecast aggregation weighted and fuzzy least loss strategies performed 5–7% better, with less loss, using the selected LSTM model strands supported by the model’s least loss strategy. Full article
Show Figures

Figure 1

21 pages, 20129 KB  
Article
UMAP-Based All-MLP Marine Diesel Engine Fault Detection Method
by Shengli Dong, Jilong Liu, Bing Han, Shengzheng Wang, Hong Zeng and Meng Zhang
Electronics 2025, 14(7), 1293; https://doi.org/10.3390/electronics14071293 - 25 Mar 2025
Cited by 4 | Viewed by 1635
Abstract
This study presents an innovative approach for marine diesel engine fault detection, integrating unsupervised learning through Uniform Manifold Approximation and Projection (UMAP) dimensionality reduction with time series prediction, offering significant improvements over existing methods. Unlike traditional model-based or expert-driven approaches, which struggle with [...] Read more.
This study presents an innovative approach for marine diesel engine fault detection, integrating unsupervised learning through Uniform Manifold Approximation and Projection (UMAP) dimensionality reduction with time series prediction, offering significant improvements over existing methods. Unlike traditional model-based or expert-driven approaches, which struggle with complex nonlinear systems, or supervised data-driven methods limited by scarce labeled fault data, our unsupervised method establishes a normal operational baseline without requiring fault labels, enhancing applicability across diverse conditions. Leveraging UMAP’s nonlinear dimensionality reduction, the proposed method outperforms conventional linear techniques (e.g., PCA) by amplifying subtle system anomalies, enabling earlier detection of state transitions—up to two batches before deviations appear in traditional performance indicators (Ps)—thus improving fault detection sensitivity. To address nonlinear relationships in UMAP-reduced dimensions, the proposed TimeMixer-FI model enhances the TimeMixer architecture with MLP-Mixer layers. The TimeMixer-FI model demonstrates consistent improvements over the original TimeMixer across various sequence lengths, achieving an MSE reduction of 69.1% (from 0.0544 to 0.0168) and an MAE reduction of 46.3% (from 0.1023 to 0.0549) at an input sequence length of 60 time steps, thereby enhancing the reliability of the time series prediction baseline. Experimental results validate that this approach significantly enhances both the sensitivity and accuracy of early fault detection, providing a more robust and efficient solution for predictive maintenance in marine diesel engines. Full article
Show Figures

Figure 1

21 pages, 2803 KB  
Article
Flexible Capacitated Vehicle Routing Problem Solution Method Based on Memory Pointer Network
by Enliang Wang, Yue Cai and Zhixin Sun
Mathematics 2025, 13(7), 1061; https://doi.org/10.3390/math13071061 - 25 Mar 2025
Cited by 1 | Viewed by 1943
Abstract
In real-world logistics scenarios, the complexities often surpass what traditional Capacitated Vehicle Routing Problem (CVRP) models can effectively address. For instance, when there is an excess of goods and limited vehicles, traditional CVRP models frequently fail to yield feasible solutions. Additionally, the time [...] Read more.
In real-world logistics scenarios, the complexities often surpass what traditional Capacitated Vehicle Routing Problem (CVRP) models can effectively address. For instance, when there is an excess of goods and limited vehicles, traditional CVRP models frequently fail to yield feasible solutions. Additionally, the time sensitivity of goods and the large scale of vehicles and goods in practical logistics scenarios present significant challenges for efficient problem-solving. This underscores the urgent need to develop a novel CVRP model that is better suited for logistics scenarios and enhances the scalability of CVRP. To address these limitations, we propose a flexible CVRP model, referred to as Flexible CVRP, which modifies the optimization objectives and constraints. This allows CVRP to provide a sensible solution even when no feasible solution exists in the traditional sense. To tackle the challenges posed by large-scale problems, we leverage the Memory Pointer Network (MemPtrN). This approach enables the modeling of solution strategies, offering strong generalization capabilities and mitigating the explosive growth in complexity to some extent. Compared to commonly used heuristic algorithms, our method achieves superior solution quality for large-scale problems. Specifically, when addressing large-scale scenarios, the MemPtrN outperforms Google’s OR-Tools solver, heuristic algorithms, enhanced evolutionary algorithms, and other reinforcement learning methods in terms of both solution speed and quality. Full article
Show Figures

Figure 1

20 pages, 12008 KB  
Article
Artificial Intelligence-Based Fault Diagnosis for Steam Traps Using Statistical Time Series Features and a Transformer Encoder-Decoder Model
by Chul Kim, Kwangjae Cho and Inwhee Joe
Electronics 2025, 14(5), 1010; https://doi.org/10.3390/electronics14051010 - 3 Mar 2025
Cited by 15 | Viewed by 3761
Abstract
Steam traps are essential for industrial systems, ensuring steam quality and energy efficiency by removing condensate and preventing steam leakage. However, their failure results in energy loss, operational disruptions, and increased greenhouse gas emissions. This paper proposes a novel predictive maintenance system for [...] Read more.
Steam traps are essential for industrial systems, ensuring steam quality and energy efficiency by removing condensate and preventing steam leakage. However, their failure results in energy loss, operational disruptions, and increased greenhouse gas emissions. This paper proposes a novel predictive maintenance system for steam traps that integrates statistical time series features and transformer encoder–decoder models for fault diagnosis and visualization. The proposed system combines IoT sensor data, operational parameters, open data (e.g., weather information and public holiday calendars), machine learning, and two-dimensional diagnostic projection to improve reliability and interpretability. Experiments were conducted in two industrial plants: an aluminum processing plant and a food manufacturing plant, and the system achieved superior defect detection accuracy and diagnostic reliability compared to existing methods. The transformer-based model outperformed traditional methods, including random forest, gradient boosting, and variational autoencoder, in classification and clustering. The system also demonstrated an average 6.92% reduction in thermal energy across both sites, highlighting its potential to improve energy efficiency and reduce carbon emissions. This research highlights the transformative impact of AI-based predictive maintenance technologies in industrial operations and provides a framework for sustainable manufacturing practices. Full article
Show Figures

Figure 1

21 pages, 2600 KB  
Article
A Particle Swarm Optimization-Based Ensemble Broad Learning System for Intelligent Fault Diagnosis in Safety-Critical Energy Systems with High-Dimensional Small Samples
by Jiasheng Yan, Yang Sui and Tao Dai
Mathematics 2025, 13(5), 797; https://doi.org/10.3390/math13050797 - 27 Feb 2025
Cited by 3 | Viewed by 1369
Abstract
Intelligent fault diagnosis (IFD) plays a crucial role in reducing maintenance costs and enhancing the reliability of safety-critical energy systems (SCESs). In recent years, deep learning-based IFD methods have achieved high fault diagnosis accuracy extracting implicit higher-order correlations between features. However, the excessive [...] Read more.
Intelligent fault diagnosis (IFD) plays a crucial role in reducing maintenance costs and enhancing the reliability of safety-critical energy systems (SCESs). In recent years, deep learning-based IFD methods have achieved high fault diagnosis accuracy extracting implicit higher-order correlations between features. However, the excessive long training time of deep learning models conflicts with the requirements of real-time analysis for IFD, hindering their further application in practical industrial environments. To address the aforementioned challenge, this paper proposes an innovative IFD method for SCES that combines the particle swarm optimization (PSO) algorithm and the ensemble broad learning system (EBLS). Specifically, the broad learning system (BLS), known for its low time complexity and high classification accuracy, is adopted as an alternative to deep learning for fault diagnosis in SCES. Furthermore, EBLS is designed to enhance model stability and classification accuracy with high-dimensional small samples by incorporating the random forest (RF) algorithm and an ensemble strategy into the traditional BLS framework. In order to reduce the computational cost of the EBLS, which is constrained by the selection of its hyperparameters, the PSO algorithm is employed to optimize the hyperparameters of the EBLS. Finally, the model is validated through simulated data from a complex nuclear power plant (NPP). Numerical experiments reveal that the proposed method significantly improved the diagnostic efficiency while maintaining high accuracy. In summary, the proposed approach shows great promise for boosting the capabilities of the IFD models for SCES. Full article
Show Figures

Figure 1

Back to TopTop