Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (2,515)

Search Parameters:
Keywords = K-nearest neighbors (KNNs)

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
29 pages, 2953 KB  
Article
A Flexible Framework for the Spatial Extension of Hyperspectral Classification Maps Using Multispectral Data
by Hideki Tsubomatsu, Satoru Yamamoto and Hideyuki Tonooka
Appl. Sci. 2026, 16(17), 8589; https://doi.org/10.3390/app16178589 (registering DOI) - 28 Aug 2026
Abstract
Hyperspectral (HS) sensors offer high spectral discrimination but generally limited spatial coverage, whereas multispectral (MS) sensors provide broad coverage with lower spectral detail. HS-MS complementary mapping aims to reduce coverage gaps in HS sensors by training classifiers within HS-MS overlap regions and applying [...] Read more.
Hyperspectral (HS) sensors offer high spectral discrimination but generally limited spatial coverage, whereas multispectral (MS) sensors provide broad coverage with lower spectral detail. HS-MS complementary mapping aims to reduce coverage gaps in HS sensors by training classifiers within HS-MS overlap regions and applying them to surrounding MS-only areas. However, comprehensive comparisons across classifiers remain scarce in this complementary mapping setting. Furthermore, lightweight pixel-wise classifiers often produce spatially inconsistent predictions, whereas spatial–spectral deep learning models demand substantial computational resources. In this study, we propose a flexible HS-MS complementary mapping framework by systematically evaluating 13 classifiers across mineral and land-use/land-cover (LULC) mapping tasks and introducing class-adaptive uncertainty revocation (CAUR), a lightweight, classifier-independent post-processing module. Performance was evaluated via spatial holdout cross-validation within the HS-MS overlap area. When computational resources are sufficient, 3D convolutional neural networks (3D-CNN) achieve the highest accuracy. Conversely, lightweight models such as random forest (RF) and k-nearest neighbors (kNN) provide computationally efficient and robust baselines. Applying CAUR consistently improves spatial consistency and overall classification accuracy without model retraining, with the largest improvements observed for coarser-resolution HS reference data. These findings provide practical design guidelines for constructing efficient HS-MS complementary mapping pipelines tailored to application demands and sensor characteristics. Full article
25 pages, 8158 KB  
Article
Diaphragm-Wall Settlement Prediction and Relative Anomaly Screening for Deep Excavations Using Multi-Model Comparison and Intelligent Optimization
by Yuhang Xu, Xinying Ai, Jian Fang, Dihua Yu, Wei Wang, Jianchao Zhang and Peiyu Zhong
Buildings 2026, 16(17), 3441; https://doi.org/10.3390/buildings16173441 - 28 Aug 2026
Abstract
Deep excavations are high-risk geotechnical activities, and accurate prediction of diaphragm-wall settlement is important for construction monitoring and deformation control. This study investigates cumulative vertical settlement at 23 diaphragm-wall monitoring points from the deep excavation of Tianjin Goldin Finance 117 in Tianjin, China. [...] Read more.
Deep excavations are high-risk geotechnical activities, and accurate prediction of diaphragm-wall settlement is important for construction monitoring and deformation control. This study investigates cumulative vertical settlement at 23 diaphragm-wall monitoring points from the deep excavation of Tianjin Goldin Finance 117 in Tianjin, China. Seven prediction models—a naïve persistence model, autoregressive integrated moving average (ARIMA), K-nearest neighbors (KNN), multilayer perceptron (MLP), gated recurrent unit (GRU), Transformer, and XGBoost—were evaluated using a unified five-fold rolling-origin expanding-window validation scheme. GRU achieved the best overall baseline performance, with a mean R2 of 0.9172, a mean absolute error (MAE) of 0.1009 mm, a root mean square error (RMSE) of 0.1341 mm, and a mean absolute percentage error (MAPE) of 0.5875%. GRU was subsequently optimized using the crow search algorithm (CSA), the genetic algorithm (GA), and the whale optimization algorithm (WOA). GRU-WOA achieved the best numerical performance, with a mean R2 of 0.9289 and an RMSE of 0.1215 mm. Relative anomaly levels were further identified from predicted settlement-change rates to characterize temporal concentration and spatial clustering of settlement-change activity. The proposed framework can support priority inspection and targeted monitoring, although the resulting anomaly levels represent project-relative statistical deviations rather than code-based engineering risk classes. Full article
(This article belongs to the Section Construction Management, and Computers & Digitization)
Show Figures

Figure 1

31 pages, 872 KB  
Article
Toward an Acoustic Characterization of Street Cries: A Machine Learning-Based Approach with Parsimonious Feature Selection
by Agosto de la Gala-Ureña, Julio-Alejandro Romero-González, M. Florencia Assaneo, José M. Álvarez-Alvarado, Diana-Margarita Córdova-Esparza, Ricardo Chaparro-Sánchez, Juan Terven and Juvenal Rodríguez-Reséndiz
Symmetry 2026, 18(9), 1429; https://doi.org/10.3390/sym18091429 - 26 Aug 2026
Viewed by 80
Abstract
Street cries are vocal expressions used by hawkers to advertise products or services in public spaces. Although they may reflect adaptation to noisy urban environments, their acoustic characteristics remain poorly understood. This study evaluated whether street cries (SC) differ acoustically from the hawkers’ [...] Read more.
Street cries are vocal expressions used by hawkers to advertise products or services in public spaces. Although they may reflect adaptation to noisy urban environments, their acoustic characteristics remain poorly understood. This study evaluated whether street cries (SC) differ acoustically from the hawkers’ normal speaking voices (NV) and whether these differences can be captured using compact feature subsets. Sixteen acoustic features related to fundamental frequency, formant structure, spectral properties, and voice quality were extracted. Redundant features were removed based on Kendall’s τ correlations, after which four classifiers were evaluated: support vector machine (SVM), random forest (RF), k-nearest neighbors (KNN), and Gaussian naive Bayes (GNB). Binary particle swarm optimization (BPSO) was then used for feature selection, with α controlling the trade-off between cross-validated balanced accuracy and subset size. In repeated participant-grouped cross-validation, SVM achieved the highest mean balanced accuracy with the 13-feature vector (0.887 ± 0.056). On the held-out test set, BPSO-selected subsets yielded balanced accuracies of 0.783–0.900. The best SVM and GNB configurations both achieved 0.900 using six and four features, respectively. The features f0 and F4 appeared in all selected subsets. These results indicate that SC and NV can be differentiated using parsimonious, interpretable acoustic feature sets. Full article
(This article belongs to the Special Issue Symmetry in Data Analysis and Optimization)
Show Figures

Figure 1

15 pages, 2170 KB  
Article
Identification of Cadmium Contamination in Rice Using Near-Infrared Reflectance Spectroscopy and Machine Learning
by Xuexue Miao, Ying Miao, Ni Li, Yang Liu and Weiping Wang
Foods 2026, 15(17), 3000; https://doi.org/10.3390/foods15173000 - 26 Aug 2026
Viewed by 139
Abstract
Routine monitoring of cadmium (Cd) contamination in rice is essential for public health protection and agricultural trade security. Conventional chemical detection methods are environmentally unfriendly, labor-intensive, and slow. This study presents a rapid, accurate classification approach based on near-infrared reflectance spectroscopy (NIRS) for [...] Read more.
Routine monitoring of cadmium (Cd) contamination in rice is essential for public health protection and agricultural trade security. Conventional chemical detection methods are environmentally unfriendly, labor-intensive, and slow. This study presents a rapid, accurate classification approach based on near-infrared reflectance spectroscopy (NIRS) for discriminating Cd-contaminated rice from uncontaminated rice. Five spectral preprocessing methods and three variable selection algorithms were systematically evaluated for their influence on model performance. Classification models were developed using partial least squares discriminant analysis (PLS-DA), K-nearest neighbors (KNN), and support vector machines (SVM). Second derivative (2D) preprocessing yielded the greatest performance gains, raising KNN and SVM test-set accuracy from 73% and 88% to 93% and 91%, respectively. Among the variable selection strategies, the successive projections algorithm (SPA) proved most effective. Under optimized conditions, PLS-DA achieved the best overall performance, attaining 92% accuracy, 89% specificity, and 95% sensitivity on the test set. These results demonstrate the strong potential of NIRS coupled with machine learning for rapid, large-scale Cd surveillance in rice, providing robust technical support for grain quality monitoring and low-cadmium variety breeding programs. Full article
(This article belongs to the Section Food Toxicology)
Show Figures

Figure 1

23 pages, 2818 KB  
Article
A Hybrid Analytical Approach for Voltage Stability Assessment in Microgrids Using Machine Learning
by Muhammad Jamshed Abbass and Robert Lis
Energies 2026, 19(17), 3983; https://doi.org/10.3390/en19173983 - 25 Aug 2026
Viewed by 122
Abstract
The complexity of voltage stability assessment in modern smart grids has increased significantly with the growing penetration of renewable energy sources and the dynamic nature of load variations. Although standard analytical methods are accurate, they are computationally expensive and unsuitable for real-time applications. [...] Read more.
The complexity of voltage stability assessment in modern smart grids has increased significantly with the growing penetration of renewable energy sources and the dynamic nature of load variations. Although standard analytical methods are accurate, they are computationally expensive and unsuitable for real-time applications. This paper proposes a hybrid analytical–machine learning framework for efficient voltage stability assessment and classification. The proposed approach consists of two stages. First, a power flow analysis is performed to compute the Fast Voltage Stability Index (FVSI) and quantify the proximity of the system operating conditions to voltage instability. Then, the FVSI values are converted into binary stability labels to formulate a supervised classification problem. In the second stage, the Extreme Gradient Boosting (XGBoost) algorithm is employed to learn the relationship between system operating variables and the corresponding stability states. The performance of the proposed method is evaluated on the IEEE 30-bus system and compared with that of conventional machine learning and deep learning models, such as Support Vector Machines (SVM), K-Nearest Neighbors (KNN), and Deep Neural Networks (DNNs). The simulation results show that the XGBoost-based framework outperforms the benchmark models in terms of classification accuracy, robustness, and computational efficiency. The proposed method provides a fast, reliable, and interpretable solution for real-time voltage stability monitoring. Therefore, it is suitable for modern smart grid applications. Full article
Show Figures

Figure 1

15 pages, 3148 KB  
Article
A Data-Driven EWMA-KNN Run-to-Run Controller for Drift-Dominant Processes with Application to Chemical Mechanical Planarization
by Ming-Cheng Hsu and Yaw-Jen Chang
Processes 2026, 14(17), 2714; https://doi.org/10.3390/pr14172714 - 25 Aug 2026
Viewed by 200
Abstract
This paper presents a data-driven run-to-run (R2R) controller for manufacturing processes subject to process drift. The proposed approach combines the exponentially weighted moving average (EWMA) method with the K-nearest neighbors (KNN) algorithm to determine process recipe adjustments. Control actions are derived entirely from [...] Read more.
This paper presents a data-driven run-to-run (R2R) controller for manufacturing processes subject to process drift. The proposed approach combines the exponentially weighted moving average (EWMA) method with the K-nearest neighbors (KNN) algorithm to determine process recipe adjustments. Control actions are derived entirely from historical process output data. In the hybrid controller, the EWMA estimator recursively updates the accumulated process drift using historical process errors and generates the corresponding recipe compensation. The KNN-based controller, in turn, identifies the K nearest neighbors in the historical feature database based on the current process error and determines the compensation action from the associated error–compensation relationships. The proposed controller was evaluated through simulations of a chemical mechanical planarization (CMP) process, with removal rate as the control objective. Under linear process drift with random white-noise disturbances, the proposed controller maintained the removal rate close to the target value, with a maximum overshoot of 4.40%, and satisfied the settling criterion from the beginning of the control process. Its performance was superior to that of the conventional EWMA controller and the standalone KNN controller. The EWMA controller exhibited several oscillations during the initial runs, with a maximum overshoot of 15.17%. Although the KNN controller satisfied the settling criterion from the beginning of the control process and produced a relatively small maximum overshoot of 3.10%, it did not consistently maintain the removal rate near the target value. Under nonlinear process drift with random disturbances, the proposed controller also maintained the process output near the target value with satisfactory stability, provided that the process drift remained within a bounded range. The controller also has a simple and intuitive implementation, which may facilitate practical industrial application. Full article
(This article belongs to the Section Process Control, Modeling and Optimization)
Show Figures

Figure 1

50 pages, 16998 KB  
Article
Multi-Strategy Improved Golden Sine Optimization Algorithm for Global Optimization and Corporate Bankruptcy Forecasting
by Yan Xu and Zhechun Li
Symmetry 2026, 18(9), 1412; https://doi.org/10.3390/sym18091412 - 22 Aug 2026
Viewed by 127
Abstract
With the increasing complexity of engineering optimization and intelligent decision-making problems, traditional metaheuristic algorithms often suffer from premature convergence, loss of population diversity, and insufficient adaptability to complex fitness landscapes. To address these issues, this paper proposes a Multi-strategy Symmetry-Aware Improved Golden Sine [...] Read more.
With the increasing complexity of engineering optimization and intelligent decision-making problems, traditional metaheuristic algorithms often suffer from premature convergence, loss of population diversity, and insufficient adaptability to complex fitness landscapes. To address these issues, this paper proposes a Multi-strategy Symmetry-Aware Improved Golden Sine Algorithm (MIGoldSA). The proposed algorithm introduces a symmetry-guided multi-strategy framework in which multiple complementary search operators are organized in a structurally balanced manner. Specifically, a strategy pool consisting of the original golden sine update rule, three differential evolution mutation strategies, and an elite-based quadratic interpolation local search operator is constructed. An adaptive strategy selection mechanism is further developed to dynamically regulate the selection probabilities of different strategies according to their historical success rates, forming a dynamic probabilistic symmetry that balances global exploration and local exploitation throughout the optimization process. The numerical performance of the resulting method is assessed using the CEC2014, 30-dimensional CEC2017, and 20-dimensional CEC2022 test collections. Comparative and statistical findings confirm that MIGoldSA generally delivers more accurate final solutions, more consistent outcomes across independent trials, and stronger convergence behavior than established algorithms and recently developed competitors. Its applicability is further examined in corporate insolvency forecasting by employing MIGoldSA to determine the hyperparameter configuration of a K-nearest neighbors classifier. Tests conducted on the Wieslaw financial database show that the resulting MIGoldSA-KNN system outperforms the selected reference models in classification accuracy, Matthews correlation coefficient, F1-score, and recall. These findings suggest that the proposed symmetry-inspired architecture offers an effective means of coordinating diversified search and intensive refinement, thereby providing a valuable computational approach for challenging global optimization and financial classification tasks. Full article
(This article belongs to the Special Issue Symmetry in Mathematical Optimization Algorithm and Its Applications)
Show Figures

Figure 1

18 pages, 6448 KB  
Article
Training a Model to Predict Asymbiotic Germination of Orchid Seeds on the Basis of Subfamily, Seed Morphology and Niche Profile
by Spyridon Oikonomidis, Anush Nersesyan, Hripsik Kosyan, Sonya Vardanyan and Costas A. Thanos
Plants 2026, 15(17), 2551; https://doi.org/10.3390/plants15172551 - 22 Aug 2026
Viewed by 208
Abstract
Although asymbiotic orchid seed germination was first achieved in vitro in 1922, the prediction of germination requirements under in vitro conditions still remains complicated. To address this, we developed a machine learning framework to classify the ex situ asymbiotic germination potential of wild [...] Read more.
Although asymbiotic orchid seed germination was first achieved in vitro in 1922, the prediction of germination requirements under in vitro conditions still remains complicated. To address this, we developed a machine learning framework to classify the ex situ asymbiotic germination potential of wild orchids into four discrete groups: Low (0–30%), Mid (31–50%), High (51–80%), and Max (81–100%). Models were trained on a dataset of 203 species, utilizing seed morphometrics—specifically, the embryo-to-testa (E:S) length ratio—alongside core ecological traits (subfamily, growth habit, habitat, and climate zone), as well as chemical scarification duration as a proxy of seed permeability. Validation leveraged novel germination and trait data from 26 taxa from Greece (17) and Armenia (9), published here for the first time. To mitigate class imbalance and prevent algorithmic bias toward highly germinating species, we applied inverse frequency weighting during training. Iterative testing of six algorithms revealed that the “Step 4” feature matrix (excluding climate zone and pretreatment duration) yielded the optimal predictive balance. K-Nearest Neighbor (KNN) and Support Vector Machine (SVM) emerged as the superior models, achieving overall accuracies of 44.4% and 61.1%, respectively, with both achieving 100% accuracy for low-germinating species. Finally, we synthesized a novel database compiling new seed morphometrics from Armenia (17 taxa), Greece (52 taxa), and the data from the literature (479 taxa). After filtering previously utilized species, we generated a prediction pool of 361 orchid taxa. Applying our Step 5 KNN and SVM models to forecast their germination behavior revealed distinct variations linked to ecological profiles. This high-accuracy framework, particularly for low-germinability groups, offers a powerful screening tool for ex situ conservation planning. The final trained models are compiled in the publicly available R (v. 4.6.0) package OrchidGermClass. Full article
(This article belongs to the Special Issue Orchid Diversity in Mediterranean-Type Climate Regions in the World)
Show Figures

Figure 1

45 pages, 5616 KB  
Article
Subject-Specific BCI Frameworks for Motor Imagery Classification Based on BSS-Free and BSS-Equipped Pipelines
by Nerita Ramsoonder, Rito Clifford Maswanganyi and Philani Khumalo
Big Data Cogn. Comput. 2026, 10(8), 280; https://doi.org/10.3390/bdcc10080280 - 20 Aug 2026
Viewed by 278
Abstract
The development of Motor Imagery (MI) Brain–Computer Interfaces (BCIs) is systematically constrained by low signal-to-noise ratios (SNRs), signal non-stationarity, and acute data scarcity. While complex Blind Source Separation (BSS) methods optimize signal clarity, their computational overhead introduces propagation delays that challenge real-time constraints. [...] Read more.
The development of Motor Imagery (MI) Brain–Computer Interfaces (BCIs) is systematically constrained by low signal-to-noise ratios (SNRs), signal non-stationarity, and acute data scarcity. While complex Blind Source Separation (BSS) methods optimize signal clarity, their computational overhead introduces propagation delays that challenge real-time constraints. This study addresses this engineering trade-off by introducing a localized architectural framework to evaluate whether a lightweight pipeline operating without BSS (No-BSS) is sufficiently efficient for real-time control when compared against two BSS-equipped pipelines utilizing Independent Component Analysis (ICA) and Empirical Mode Decomposition (EMD). Validated across the BCI Competition IV Dataset 2A and the PhysioNet MI dataset, all three pipelines share an identical processing chain designed to maximize efficiency. To mitigate low SNRs, an Adaptive Laplacian spatial filter isolates neural intent across target sensorimotor electrodes (C3, C4, and Cz). Data scarcity is countered via a Gaussian noise injection data augmentation strategy, while session-to-session variability is addressed during feature extraction using Wavelet Packet Decomposition (WPD) paired with a Fisher Score criterion to dynamically isolate subject-specific time-frequency nodes. Redundant features are subsequently eliminated using a Genetic Algorithm (GA) before classification. Experimental evaluation reveals a distinct performance stratification: while the ICA (92.80%) and EMD (92.69%) pipelines yield the highest average accuracy for the PhysioNet dataset by isolating non-stationary and physiological noise, the No-BSS baseline (90.28%) remains the superior framework for the BCI Dataset 2A. Across all pipelines across both datasets, a stable classification hierarchy emerges wherein the Support Vector Machine (SVM) leads performance due to its maximum-margin decision boundary, followed by k-Nearest Neighbors (kNN), a modified EEGNet, and Decision Trees. The No-BSS baseline achieves classification accuracies highly competitive with its BSS counterparts while entirely bypassing their algorithmic overhead. Given the strict latency constraints of live BCI control loops, these findings establish the optimized No-BSS pipeline as a highly viable alternative for low-latency, real-time implementations. Full article
Show Figures

Figure 1

20 pages, 4961 KB  
Article
Machine Learning-Driven Prediction of Optical Absorption in Composition-Dependent Truncated Pyramidal GaN/AlxGa1−xN Quantum Dots
by Tesnim Brahim, Adel Bouazra, Beriham Ibrahim Basha and Fatma Aouaini
Mathematics 2026, 14(16), 2938; https://doi.org/10.3390/math14162938 - 13 Aug 2026
Viewed by 183
Abstract
This study presents a comparative machine-learning investigation for predicting the optical absorption coefficient of truncated pyramidal GaN/AlxGa1−xN quantum dots. The physical dataset is generated by solving the three-dimensional Schrödinger equation using a coordinate-transformation method combined with the finite-difference [...] Read more.
This study presents a comparative machine-learning investigation for predicting the optical absorption coefficient of truncated pyramidal GaN/AlxGa1−xN quantum dots. The physical dataset is generated by solving the three-dimensional Schrödinger equation using a coordinate-transformation method combined with the finite-difference method (FDM). The coordinate transformation maps the sloping boundaries of the truncated pyramidal geometry onto a regular computational domain, enabling an accurate representation of the quantum-dot shape and facilitating its numerical treatment using the FDM. The absorption coefficient is then calculated as a function of photon energy for different alloy compositions. Using photon energy and alloy composition as input features, Artificial Neural Network (ANN), Random Forest (RFR), Decision Tree (DT), and k-Nearest Neighbor (KNN) models are developed and evaluated. A second-degree polynomial regression model is also considered as a classical baseline. Under the point-wise random 80/20 split, all models show excellent agreement with the numerical results, with R2 values close to unity. KNN generally provides the lowest prediction errors across most alloy compositions, whereas ANN achieves slightly lower MSE and RMSE values at x=0.5. Furthermore, leave-one-composition-out validation identifies ANN as the most effective model for predicting unseen compositions, achieving a mean R2 of 0.848 and an NRMSE of 7.19%. These findings demonstrate that KNN is particularly effective for local interpolation within the sampled domain, while ANN provides stronger composition-wise generalization. The proposed framework offers an efficient surrogate for computationally demanding numerical simulations of the optical properties of quantum nanostructures. Full article
(This article belongs to the Section E4: Mathematical Physics)
Show Figures

Figure 1

29 pages, 716 KB  
Article
Threat Actor Attribution Applying a Tactics–Techniques–Procedures Approach: An Empirical Investigation
by Shaheen Hussain and Krassie Petrova
Future Internet 2026, 18(8), 433; https://doi.org/10.3390/fi18080433 - 13 Aug 2026
Viewed by 453
Abstract
The increasing frequency and growing impact of cyberattacks have led organizations to adopt proactive defense approaches to cybersecurity risk mitigation, especially in the case of advanced persistent threats (APTs). The correct identification of the specific malicious actors behind a cyberattack is important for [...] Read more.
The increasing frequency and growing impact of cyberattacks have led organizations to adopt proactive defense approaches to cybersecurity risk mitigation, especially in the case of advanced persistent threats (APTs). The correct identification of the specific malicious actors behind a cyberattack is important for the success of incident response and for the investigative work of the security operations center (SOC) team. This research explores the capabilities and limitations of a machine learning (ML) approach to identifying malicious actors and the threats they pose (threat actor attribution) based on the tactics, techniques, and procedures (TTP) observed in specific cybersecurity incidents and on the incident context (the geographical location and industry affiliation of the victims targeted in the attack). A large language model (LLM) was used to extract TTPs from the MITRE ATT&CK database of cybersecurity incidents. The experiments included modeling threat actor attribution using five ML algorithms: k-nearest neighbors (KNN), decision tree (DT), random forest (RF), support vector machine (SVM), and naïve Bayes (NB), with different methods applied for feature selection and weighting. The results indicated that model accuracy and other performance metrics were significantly improved when the input dataset included both TTP and contextual features. The KNN and SVM models produced the best performance results; the highest classification accuracy achieved was 93.19%. The outcomes of this study may be applied by cybersecurity professionals to identify malicious actors, estimate the number and types of data points that are required to adequately attribute a cyberattack to an actor, and improve the accuracy of the classification by weighting the input dataset features. Full article
(This article belongs to the Special Issue Machine Learning and Internet of Things in Industry 4.0—2nd Edition)
Show Figures

Figure 1

23 pages, 5747 KB  
Article
Pilot Study Employing a Machine Learning Approach as a Potential Method for Predicting Parkinson’s Disease Using Voice as a Digital Biomarker and the SHAP Approach for Feature Engineering
by Mehdi Rashidi, Syed Adil Hussain Shah, Chiara Coppola, Andrea Buccoliero, Serena Arima, Angela Lupo, Filomena My, Marta Lorenzo, Marcello Donzella and Michele Maffia
Bioengineering 2026, 13(8), 917; https://doi.org/10.3390/bioengineering13080917 - 13 Aug 2026
Viewed by 412
Abstract
Introduction: Voice-based digital biomarkers have emerged as a promising approach for distinguishing individuals with neurodegenerative disorders, particularly Parkinson’s disease (PD), from healthy subjects (HS). With the increasing availability of smartphone and web-based recording tools, voice data can be collected efficiently in both [...] Read more.
Introduction: Voice-based digital biomarkers have emerged as a promising approach for distinguishing individuals with neurodegenerative disorders, particularly Parkinson’s disease (PD), from healthy subjects (HS). With the increasing availability of smartphone and web-based recording tools, voice data can be collected efficiently in both clinical and remote settings. However, further validation is required before such approaches can be translated into routine clinical practice. Methods: This study used a cross-sectional analysis at the recording level, treating repeated recordings from the same participant as separate observations collected at Vito Fazzi Hospital in Lecce, Italy. Speech recordings from individuals with Parkinson’s disease (PD) and healthy controls were collected using the dedicated Talia smartphone and web application. Sustained vowel phonation (/a/) was analyzed as the primary speech task. Following data acquisition, feature extraction was performed as a crucial step in the speech analysis pipeline, as the quality and relevance of the extracted features directly influence the ability of machine learning models to discriminate between Parkinson’s disease (PD) patients and healthy controls. To capture various aspects of speech impairment associated with PD, a comprehensive set of acoustic features was extracted, including long-term features (pitch, jitter, and shimmer), nonlinear descriptors such as Recurrence Period Density Entropy (RPDE), and short-term feature based on Mel-Frequency Cepstral Coefficients (MFCCs). These features were subsequently used to develop and evaluate machine learning models for the classification of Parkinson’s disease and healthy subjects. Feature selection was performed using SHAP to identify the most informative vocal biomarkers. Model performance was assessed using five independent random train–test splits (70% training and 30% testing), supported by an internal five-fold cross-validation procedure within the training data. Multiple machine learning models were developed and evaluated, including Random Forest, Logistic Regression, Support Vector Machine, Naive Bayes, K-Nearest Neighbors, Decision Tree, Artificial Neural Network, and Gradient Boosting. Results: The evaluated models demonstrated strong recording-level classification performance. Artificial Neural Networks (ANN) and K-Nearest Neighbors (KNN) achieved the highest accuracy scores (0.9545 and 0.9494, respectively), along with superior recall (up to 0.9500), precision (up to 0.9551), and F1-score (up to 0.9525). Both models also exhibited excellent discriminative ability, with ROC-AUC values reaching 0.9882 (ANN) and 0.9893 (KNN). In contrast, Naive Bayes and Decision Tree showed comparatively lower performance across all metrics. Log-loss analysis further confirmed the robustness of ANN and KNN, which achieved the lowest values (0.2552 and 0.2510, respectively), indicating well-calibrated predictions. Overall, the findings highlight the consistency and generalizability of ANN and KNN across cross-validation splits. Conclusions: This study demonstrates that machine learning models, particularly ANN and KNN, can effectively differentiate Parkinson’s disease from healthy conditions using voice recordings. The integration of explainable AI for feature selection enhances model transparency and clinical relevance. However, the reported performance estimates were obtained from a recording-level analysis and should be interpreted as preliminary findings. Further studies involving larger cohorts and participant-level validation strategies are required to determine the generalizability and clinical applicability of these approaches. Full article
(This article belongs to the Special Issue AI and Data Analysis in Neurological Disease Management)
Show Figures

Figure 1

30 pages, 8144 KB  
Article
Benchmarking RF, KNN, MLP, and CNN for FFT-Based PV Arc Fault Detection: Scaling Choice, Temporal Cross-Validation, and Latency Trade-Offs Toward Edge Deployment
by Michel Braulio de Oliveira, Filipe Ramos, José Cesar de Souza Almeida Neto, Fábio Jesus Moreira Almeida and Bruno Luis Soares Lima
Energies 2026, 19(16), 3787; https://doi.org/10.3390/en19163787 - 12 Aug 2026
Viewed by 182
Abstract
Ensuring the safety and reliability of photovoltaic (PV) installations requires accurate electrical arc fault detection. This work presents a computational arc fault detection framework that combines fixed-length windowing, Fast Fourier Transform (FFT)-based features, and supervised machine learning classifiers. Data were acquired using an [...] Read more.
Ensuring the safety and reliability of photovoltaic (PV) installations requires accurate electrical arc fault detection. This work presents a computational arc fault detection framework that combines fixed-length windowing, Fast Fourier Transform (FFT)-based features, and supervised machine learning classifiers. Data were acquired using an Arc Fault Circuit Interrupter (AFCI) test bench developed based on IEC 63027. Current and voltage signals were partitioned into 200-sample windows, DC-offset corrected, and Hann-windowed signals. Each window generated 204 statistical and spectral attributes used to train and evaluate Random Forest (RF), K-Nearest Neighbors (KNN), Multilayer Perceptron (MLP), and Convolutional Neural Network (CNN) models. Hyperparameters were tuned by grid search with TimeSeriesSplit cross-validation, comparing min–max normalization and Z–Score standardization. On a 15% hold-out test set, CNN with Z–Score achieved F1 = 0.9982 and recall = 0.9975, followed by MLP (F1 = 0.9957) and RF (F1 = 0.9821). Amortized per-window inference latencies were ≈0.0035 ms for RF, ≈0.0016 ms for MLP with Z–Score, and ≈0.032 ms for CNN. These classifier-stage timings indicate computational compatibility with edge-oriented implementation but do not constitute an end-to-end IEC 63027 AFCI compliance assessment. The framework targets integration into PV inverters at Mackenzie Presbyterian University’s solar plant. Full article
Show Figures

Graphical abstract

25 pages, 1476 KB  
Article
Food Production Index Forecasting for Sustainable Food Systems in Türkiye: A Machine Learning-Based Approach
by Ferhan Balci Torun, Mehmet Kayakuş, Onder Kabas, Georgiana Moiceanu and Mariana-Gabriela Munteanu
Foods 2026, 15(16), 2814; https://doi.org/10.3390/foods15162814 - 12 Aug 2026
Viewed by 410
Abstract
Sustainable food systems are increasingly challenged by climate change, resource constraints, market volatility, and growing food demand, making accurate forecasting of food production essential for food security and long-term sustainability. Despite the growing use of machine learning in agricultural forecasting, studies directly modeling [...] Read more.
Sustainable food systems are increasingly challenged by climate change, resource constraints, market volatility, and growing food demand, making accurate forecasting of food production essential for food security and long-term sustainability. Despite the growing use of machine learning in agricultural forecasting, studies directly modeling the Food Production Index (FPI) within a sustainable food systems framework remain limited, particularly in emerging economies. This study addresses this gap by forecasting Türkiye’s Food Production Index using agricultural, macroeconomic, and trade-related indicators covering the period 1962–2023. Seven predictive approaches, including Multiple Linear Regression (MLR), Bayesian Ridge Regression, Support Vector Regression (SVR), Random Forest, Gradient Boosting, Artificial Neural Networks (ANNs), and K-Nearest Neighbors (KNN), were comparatively evaluated using R2, RMSE, and MAE metrics. The results demonstrate that Bayesian Ridge Regression (R2 = 0.968) and MLR (R2 = 0.918) significantly outperform more complex machine learning algorithms, indicating that model–data compatibility is more critical than algorithmic complexity in long-term food production forecasting. The findings reveal that economic growth, agricultural inputs, and structural transformation processes play a decisive role in shaping food production dynamics. By integrating machine learning with sustainability-oriented food system analysis, this study provides a robust evidence base for supporting food security strategies, resource-efficient agricultural planning, and resilient food system governance. The proposed framework offers macro-level decision-support insights for policymakers engaged in long-term food system planning, strategic risk monitoring, and evidence-based policy evaluation. Full article
Show Figures

Figure 1

25 pages, 15267 KB  
Article
KSR-Huber: A Robust Method for Wind Vector Retrieval from Doppler Wind Lidar Observations
by Yuefeng Zhao, Zhongyue Zhang, Xueting Liu and Nannan Hu
Remote Sens. 2026, 18(16), 2698; https://doi.org/10.3390/rs18162698 - 11 Aug 2026
Viewed by 284
Abstract
Three-dimensional wind vector retrieval from Coherent Doppler Wind Lidar (CDWL) in Velocity–Azimuth Display (VAD) mode is susceptible to anomalous radial velocity observations induced by low signal-to-noise ratios, clutter echoes, and spectral estimation errors, which degrade inversion accuracy. To address this issue, a robust [...] Read more.
Three-dimensional wind vector retrieval from Coherent Doppler Wind Lidar (CDWL) in Velocity–Azimuth Display (VAD) mode is susceptible to anomalous radial velocity observations induced by low signal-to-noise ratios, clutter echoes, and spectral estimation errors, which degrade inversion accuracy. To address this issue, a robust retrieval method, termed KSR-Huber, is proposed by integrating K-nearest-neighbor (KNN)-based local statistical priors with Huber iterative reweighted least squares (IRLS). The method employs KNN-based local consistency and adaptive Sigmoid weighting, together with Huber residual reweighting within the IRLS framework, to suppress anomalous observations while preserving valid data. Simulations across diverse scenarios, conducted under controlled numerical experiments with varying observation redundancies and outlier contamination levels, show that the proposed method consistently outperforms existing approaches, including DSWF, KNN-COOKS, and airSWF, particularly in terms of robustness under controlled noise and outlier conditions. Real lidar observations further demonstrate the practical applicability of the method, while comprehensive validation against independent reference measurements is left for future work. Full article
(This article belongs to the Section Atmospheric Remote Sensing)
Show Figures

Figure 1

Back to TopTop