Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (1,076)

Search Parameters:
Keywords = K-Fold Cross-Validation

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
23 pages, 4312 KB  
Article
Machine Learning-Based Prediction of 48-Hour Extubation Success in Mechanically Ventilated Children: A Single-Center Retrospective Cohort Study
by Ferhat Sarı and Aynur Aliyeva
Children 2026, 13(9), 1166; https://doi.org/10.3390/children13091166 (registering DOI) - 29 Aug 2026
Abstract
Background: Accurate assessment of extubation readiness in mechanically ventilated children remains difficult because successful sustained breathing depends on the interaction of respiratory, metabolic, inflammatory, neurological, and cardiovascular factors. This study developed and internally validated machine-learning models for predicting 48 h extubation success using [...] Read more.
Background: Accurate assessment of extubation readiness in mechanically ventilated children remains difficult because successful sustained breathing depends on the interaction of respiratory, metabolic, inflammatory, neurological, and cardiovascular factors. This study developed and internally validated machine-learning models for predicting 48 h extubation success using routinely available pre-extubation data. Methods: Of 1097 total PICU admissions, 341 mechanically ventilated children treated between 2021 and 2026 constituted the final analytic cohort. The primary outcome was survival without reintubation during the first 48 h after planned extubation. Demographic, clinical, laboratory, blood gas, illness severity, and ventilator variables were evaluated. Logistic Regression, Random Forest, Gradient Boosting, and Support Vector Machine models were assessed using stratified 5-fold cross-validation. Unsupervised k-means clustering was performed to identify physiological phenotypes. Results: Extubation was successful in 298 children (87.4%) and failed in 43 (12.6%). Failure was associated with higher oxygenation index, lactate, procalcitonin, C-reactive protein, PaCO2, pSOFA, PEEP, and rapid shallow breathing index, together with lower arterial pH, bicarbonate, ionized calcium, hemoglobin, albumin, sodium, and Glasgow Coma Scale scores. Random Forest yielded the numerically highest discrimination, with an AUC of 0.983 (95% CI, 0.971–0.993), a sensitivity of 0.980, a specificity of 0.814, an accuracy of 0.959, and a Brier score of 0.032. Arterial pH, the oxygenation index, bicarbonate, ionized calcium, procalcitonin, lactate, and PaCO2 showed the highest Random Forest Gini importance scores. Clustering identified a low-severity phenotype (n = 302; 97% success; 0% mortality) and a high-severity phenotype (n = 39; 10% success; 100% mortality). Conclusions: Multidimensional machine-learning models predicted 48 h pediatric extubation success with strong internal discrimination. Prospective multicenter external validation is required before clinical implementation. Full article
Show Figures

Figure 1

40 pages, 9689 KB  
Article
FTU-Seek: Foundation Model-Guided Hard-Negative Learning for Sparse Functional Tissue Unit Segmentation
by Zonghao Liu, Lei Su, Jiguang Yu, Xuqing Geng, Louis Shuo Wang, Jianmin Wang and Jingfeng Liu
Biomedicines 2026, 14(9), 1935; https://doi.org/10.3390/biomedicines14091935 - 28 Aug 2026
Abstract
Background/Objectives: Functional tissue units (FTUs), including tertiary lymphoid structures (TLSs), blood vessels, and glands, encode localized immune, vascular, and epithelial organization in histopathology. Accurate quantification of these structures is important for studying tissue architecture and disease-associated tissue organization. However, FTUs are frequently [...] Read more.
Background/Objectives: Functional tissue units (FTUs), including tertiary lymphoid structures (TLSs), blood vessels, and glands, encode localized immune, vascular, and epithelial organization in histopathology. Accurate quantification of these structures is important for studying tissue architecture and disease-associated tissue organization. However, FTUs are frequently sparse, heterogeneous, and surrounded by large amounts of morphologically similar background tissue, making automated segmentation in whole-slide images (WSIs) challenging. We therefore developed FTU-Seek, a pathology foundation model-guided framework that treats morphology-aware negative-patch selection as a key component of sparse FTU segmentation. Methods: FTU-Seek uses frozen multi-depth features from the UNI pathology foundation model to train a patch-level classifier that distinguishes FTU-containing from FTU-absent tissue. Target-absent patches are subsequently ranked according to their predicted target-containing probabilities, and the highest-scoring hard negatives are selected through a static TopK strategy to construct compact segmentation training sets. The framework was evaluated using five-fold cross-validation and internal test cohorts across TLS, blood-vessel, and gland segmentation tasks, with an additional independent 30-WSI held-out cohort for TLS. Positive-only, all-tissue, random-negative, and matched random TopK sampling strategies served as comparators. Segmentation-derived phenotypes were further explored in external TCGA cohorts. Results: The patch-level classifiers achieved mean validation AUCs of 95.92%, 90.11%, and 98.16% for TLS, blood vessel, and gland classification, respectively. For TLS segmentation, the pre-specified Top1000 configuration retained 27.6% of the all-tissue training workload and achieved a slide-level Dice of 76.69 ± 11.89% on the independent 30-WSI held-out cohort. Compared with matched random Top1000 sampling, it improved Dice by 3.72 percentage points (95% CI, 2.05–5.38). Blood vessel and gland segmentation achieved performance approaching all-tissue training while reducing the retained training workload by approximately one-half and one-third, respectively. Compared with matched random sampling, classifier-guided hard-negative selection produced the greatest improvements for sparse and morphologically ambiguous FTUs. Exploratory TCGA analyses further showed associations of TLS phenotypes with overall survival, vascular phenotypes with overall survival and microvascular invasion, and glandular phenotypes with clinicopathological characteristics. Conclusions: FTU-Seek demonstrates that pathology foundation models can support sparse FTU segmentation not only through feature representation but also through morphology-aware construction of segmentation training sets. By prioritizing informative hard negatives, the framework reduces redundant segmentation-training workload while maintaining competitive segmentation performance and supporting quantitative tissue phenotyping from routine histopathology. Full article
(This article belongs to the Special Issue Human Stem Cells in Disease Modelling and Treatment (2nd Edition))
Show Figures

Figure 1

37 pages, 9998 KB  
Article
AOA-Constrained, Calibrated Random-Forest Prospectivity Mapping for Au-Ag in Nevada Great Basin: Spatially Independent Validation, Uncertainty, and Decision-Focused Top-k Targets
by Alisher Saduov, Geroy Zholtayev, Kuanysh Togizov, Zamzagul Umarbekova, Nurbakyt Zhumabay and Nursat Amangeldi
Minerals 2026, 16(9), 886; https://doi.org/10.3390/min16090886 (registering DOI) - 28 Aug 2026
Abstract
Mineral prospectivity mapping (MPM) increasingly relies on machine learning classifiers, but model performance may be overstated when spatial dependence, extrapolation beyond the training domain, and poorly calibrated prediction scores are not adequately addressed. This study presents a reproducible workflow for regional Au-Ag prospectivity [...] Read more.
Mineral prospectivity mapping (MPM) increasingly relies on machine learning classifiers, but model performance may be overstated when spatial dependence, extrapolation beyond the training domain, and poorly calibrated prediction scores are not adequately addressed. This study presents a reproducible workflow for regional Au-Ag prospectivity mapping in Nevada that integrates positive–unlabeled learning with Random Forest models and spatial GroupKFold cross-validation. Predictions were restricted to the Area of Applicability (AOA), defined using the 0.95 quantile of a training-derived feature-space dissimilarity index, resulting in areal coverage of 54.331% for Ag and 73.118% for Au. Discrimination based on pooled out-of-fold (OOF) predictions was strong for both commodities (Ag: AP = 0.745, ROC-AUC = 0.908; Au: AP = 0.777, ROC-AUC = 0.912). Post hoc OOF calibration with isotonic regression reduced Brier scores and improved probability reliability within the supported prediction domain. Area-based validation showed that the highest-ranked 10% of the AOA captured 85.6% of the Ag and 84.2% of the Au pooled-OOF positives, corresponding to targeting efficiencies of 8.6-fold and 8.4-fold relative to random spatial selection. Ensemble dispersion was used to map predictive uncertainty and to identify areas where additional data may be most useful for reducing uncertainty. SHAP importance and partial-dependence diagnostics consistently highlighted proximity to Quaternary faults, elevated slip and dilation tendency, and potential-field edges associated with intrusive or volcanic contacts, while surface heat flow acted mainly as a permissive control. The resulting AOA-masked prospectivity maps, uncertainty layers, and top-k targeting products provide a transparent regional framework for prioritizing follow-up exploration and allocating reconnaissance effort. Full article
(This article belongs to the Section Mineral Exploration Methods and Applications)
24 pages, 10814 KB  
Article
A Spiking Neural Network for Non-Invasive Glucose Estimation on Wearable Bioimpedance Biosensors, with a Multiplication-Free Neuromorphic Path
by Matheus Willian Sprotte and Pedro Bertemes Filho
Biosensors 2026, 16(9), 469; https://doi.org/10.3390/bios16090469 - 27 Aug 2026
Abstract
Wearable glucose monitoring demands low-power local processing, but conventional neural networks rely on energy-intensive multiply–accumulate (MAC) operations that limit battery life. This study shows that a Spiking Neural Network (SNN), built on a regression-adapted Leaky Integrate-and-Fire (LIF) neuron, can estimate blood glucose from [...] Read more.
Wearable glucose monitoring demands low-power local processing, but conventional neural networks rely on energy-intensive multiply–accumulate (MAC) operations that limit battery life. This study shows that a Spiking Neural Network (SNN), built on a regression-adapted Leaky Integrate-and-Fire (LIF) neuron, can estimate blood glucose from multi-frequency bioimpedance and auxiliary biosignals with clinically auditable accuracy at low computational and memory cost. Using data from 98 patients (717 measurements, eGluco3 device, Azambuja Hospital, Brusque, Brazil) evaluated by 5-fold walk-forward cross-validation under ISO 15197:2013, three main findings emerge. First, a new calibration method—the Patient Fingerprint, built from each patient’s first K sensor readings—outperforms conventional one-hot patient encoding (14.2 ± 2.6 mg/dL vs. 15.4 ± 3.3 mg/dL mean absolute error) and, unlike one-hot, requires only these K readings rather than the patient’s presence in the training set; a leave-patients-out analysis confirms that the fingerprint captures individual physiology and that unseen-patient accuracy improves with calibration depth but remains clinically insufficient (MAE 94.865.9 mg/dL from K=3 to K=5), positioning clinical-grade cross-patient generalization on a larger cohort as the primary scaling axis. Second, the direct-injection fingerprint model reaches 100% of the samples within Consensus Error Grid Zones A+B across all validation folds (the rate-coding variant reaches 98.8%, just below the 99% Criterion B threshold), without requiring any demographic or clinical metadata; sensor history alone renders such records redundant; and Criterion A, however, stays below the 95% normative threshold, so the results support clinical safety rather than formal certification. Third, replacing the analog input encoding with a multiplication-free rate-coding scheme removes all first-layer MAC operations at a cost of 2.7 mg/dL additional error; because the additional microticks raise the total operation count, this defines a design lever whose energy payoff is specific to neuromorphic hardware rather than a net saving on conventional microcontrollers. Together, these results demonstrate that SNNs offer a clinically auditable, self-calibrating, and memory-efficient path to continuous glucose estimation on embedded wearable devices. Full article
(This article belongs to the Special Issue Bioimpedance-Based Biosensors)
Show Figures

Graphical abstract

22 pages, 21454 KB  
Article
The Lettuce Nutritional Diagnosis Model of ResNet Improved by Integrating the MSA Mechanism
by Shiwei He, Iftikhar Hussain Shah, Zhengheng Shen, Weihang Zhang, Zhou Shen, Enqi Zhang, Shubo Wang, Qingliang Niu and Liying Chang
Horticulturae 2026, 12(9), 1063; https://doi.org/10.3390/horticulturae12091063 - 25 Aug 2026
Viewed by 195
Abstract
The sustainable production of leafy vegetables in Mediterranean and East Asian regions is increasingly constrained by water scarcity and nutrient imbalance in soil–plant systems, making timely and accurate nutrient diagnosis essential for precision fertilization. Conventional tissue analysis of nitrogen (N), phosphorus (P), and [...] Read more.
The sustainable production of leafy vegetables in Mediterranean and East Asian regions is increasingly constrained by water scarcity and nutrient imbalance in soil–plant systems, making timely and accurate nutrient diagnosis essential for precision fertilization. Conventional tissue analysis of nitrogen (N), phosphorus (P), and potassium (K) in lettuce is destructive, costly, and time-consuming, while existing non-destructive approaches based on traditional machine learning or deep learning still suffer from limited accuracy and poor generalization. To address these limitations, this study proposes ResNet-SA, a residual convolutional network assisted by a multi-head self-attention (MSA) mechanism, for the rapid and non-destructive estimation of leaf N, P, and K contents from top-view RGB images of lettuce trays under soilless cultivation. Two fusion strategies between the MSA mechanism and the ResNet50 trunk were evaluated, namely lateral side-connection of the attention block (ResNet50_SA_R series) and replacement of a trunk stage (ResNet50_SA_E series), each with three insertion depths. In the fixed-split evaluation, ResNet50_SA_RV1 achieved the best performance, with a test-set R2 of 0.92 versus 0.81 for the ResNet50 baseline (absolute R2 gains of 0.11 and 0.08 for RV1 and RV3, respectively). Grouped five-fold cross-validation by sampling batch tentatively verified the stability of this improvement (ResNet50_SA_RV1: R2 = 0.92 ± 0.03; RMSE = 7.16 ± 1.77 mg/g; MAE = 4.06 ± 1.58 mg/g, macro-averaged across N, P, and K), with significantly lower prediction error than both the ResNet50 baseline (ΔMAE = −2.25 mg/g, 95% CI: −2.88 to −1.62, Holm-adjusted p < 0.001) and four conventional CNN architectures (R2 range: 0.49–0.80). These results demonstrate that integrating the MSA mechanism into ResNet provides a reliable, non-destructive tool for lettuce nutrient diagnosis, offering practical support for precision fertilization and sustainable greenhouse production. Full article
Show Figures

Graphical abstract

28 pages, 2404 KB  
Article
HGSM-YOLO: A Small-Lesion-Oriented Lightweight YOLO11n Framework for Citrus Leaf Disease Detection
by Rui Zheng, Jing Zhao, Xinwei Wang and Feng Wang
Sensors 2026, 26(17), 5345; https://doi.org/10.3390/s26175345 - 24 Aug 2026
Viewed by 252
Abstract
Accurate and rapid detection of citrus leaf diseases is important for early diagnosis, precision orchard management, and the reduction of economic losses in citrus production. Automatic detection remains difficult because early lesions are often small and irregular. Several disease categories also share similar [...] Read more.
Accurate and rapid detection of citrus leaf diseases is important for early diagnosis, precision orchard management, and the reduction of economic losses in citrus production. Automatic detection remains difficult because early lesions are often small and irregular. Several disease categories also share similar visual appearances, and localization is easily affected by veins, shadows, and cluttered backgrounds. To address these task-specific challenges, we propose HGSM-YOLO, where HGSM denotes the coordinated use of heterogeneous convolution, a GSConv-based slim neck, and multi-scale dilated local attention. The framework is built on YOLO11n because its 2.59 M-parameter and 6.4 GFLOP design provides a stringent compact baseline for edge-oriented improvement. The method follows a hierarchical design: C3k2-HetConv preserves lesion edges and local morphology in the backbone; the GSConv-based slim neck reduces part of the feature fusion cost; and an MSDA module in the high-resolution P3 branch enhances the context of small lesions. Following model selection on the validation split, the final locked models were evaluated once on the held-out test split, with HGSM-YOLO reaching 77.5% precision, 66.8% recall, 71.7% F1-score, 70.8% mAP@0.5, and 44.2% mAP@0.5:0.95, compared with 68.5%, 61.5%, 64.8%, 66.0%, and 40.2% for YOLO11n. A stratified outer five-fold cross-validation further yields 71.0% ± 1.4% mAP@0.5 and 44.4% ± 1.1% mAP@0.5:0.95 for HGSM-YOLO, versus 65.9% ± 1.1% and 40.2% ± 0.9% for YOLO11n. On the independent 1871-image citrus-leaf-disease-2 dataset, retraining under the same protocol gives 94.4% mAP@0.5 for HGSM-YOLO versus 92.2% for YOLO11n and 93.1% for the public Roboflow YOLOv11 reference model. The complete HGSM-YOLO architecture uses 7.2 GFLOPs, 2.82 M parameters, and runs at 90.9 FPS on the RTX 4090, compared with 6.4 GFLOPs, 2.59 M parameters, and 110.1 FPS for the baseline. Thus, the contribution provides a recall- and localization-oriented accuracy–efficiency trade-off rather than universal superiority in every individual metric. Full article
(This article belongs to the Section Smart Agriculture)
Show Figures

Figure 1

33 pages, 14775 KB  
Article
Mutation-Aware Machine Learning Framework for Predicting Binding Affinity of Nirmatrelvir Analogs Targeting Coronavirus Main Proteases
by Md Saidur Rahman, Md Mehedi Hasan and Shahidul M. Islam
Molecules 2026, 31(17), 2949; https://doi.org/10.3390/molecules31172949 - 22 Aug 2026
Viewed by 194
Abstract
The emergence of resistance-associated mutations in coronavirus main protease (Mpro) poses a significant challenge to the development of broad-spectrum antiviral therapeutics. In this study, we improved and accelerated a mutation-aware machine learning (ML) framework to predict the binding score of Nirmatrelvir analogue ligands [...] Read more.
The emergence of resistance-associated mutations in coronavirus main protease (Mpro) poses a significant challenge to the development of broad-spectrum antiviral therapeutics. In this study, we improved and accelerated a mutation-aware machine learning (ML) framework to predict the binding score of Nirmatrelvir analogue ligands against wild-type and mutant MERS-CoV Mpro. A library of 15,889 Nirmatrelvir derivatives generated through systematic scaffold modification was docked against the wild-type and five variants of the Mpro, producing a total of 95,334 structural and docking score datasets of these protein–ligand complexes. During the ML model development phase, ligand effects were learned from RDKit molecular descriptors and graph-based representations, and the mutation-induced effects were captured through delta-encoded physicochemical properties (hydrophobicity, charge, aromaticity, and polarity) of the active-site residues. Among the evaluated models, the CatBoost regressor tree-based algorithm achieved the lowest mean absolute error (MAE) value of 0.23 Kcal/mol and an R2 of 0.87. Further improvement was achieved by creating a weighted ensemble model combining the CatBoost regressor, XGBoost and LightGBM regressor, resulting in a prediction accuracy with a MAE of 0.19 Kcal/mol and an R2 of 0.90 relative to docking scores. Model robustness was further evaluated through random-, ligand group- and scaffold group- K-fold cross-validation along with their Y-randomization. Moreover, the models were also tested with a new set of 1000 structurally diverse compounds. SHAP analysis was conducted, which identified 20 molecular descriptors critical for accurate predictions. The ensemble model accurately predicted the binding affinities of Nirmatrelvir and its four analogues (E1–E4), reproducing the experimental pIC50 trend and correctly identifying the most potent inhibitors. The ensemble model also showed consistent performance across all MERS-CoV Mpro variants, S147Y, S142G, L144A, S142G/S147Y, and S142G/L144A/S147Y, demonstrating its potential for rapidly discovering mutation-resistant antiviral drugs. Full article
(This article belongs to the Special Issue Computational Approaches for Drug and Protein Design)
Show Figures

Figure 1

24 pages, 2621 KB  
Article
Interpretable Prediction of Geopolymer Concrete Compressive Strength Using DBO–CatBoost and SHAP Analysis
by Nima Saeedi, Zahra Mohammadipour Novin, Amirreza Shirini, Sina Samadi Gharehveran, Siamak Pedrammehr and Mohammad Fotouhi
Buildings 2026, 16(16), 3326; https://doi.org/10.3390/buildings16163326 - 21 Aug 2026
Viewed by 290
Abstract
The construction sector faces a critical need to minimize its carbon footprint, which is currently stimulating the development of geopolymer concrete using recycled coarse aggregates as an eco-friendly material compared with Portland cement. Accurate prediction of the compressive strength of this eco-efficient concrete [...] Read more.
The construction sector faces a critical need to minimize its carbon footprint, which is currently stimulating the development of geopolymer concrete using recycled coarse aggregates as an eco-friendly material compared with Portland cement. Accurate prediction of the compressive strength of this eco-efficient concrete is complex, however, as a result of the complex, non-linear interactions between many of the mix-design and curing parameters. Although modern scientific literature and engineering practices have increasingly adopted machine learning (ML) for concrete strength prediction, a significant scientific gap remains. Most existing studies rely on “black-box” models that lack sufficient interpretability and frequently overlook the severe risk of data leakage during validation, limiting their practical engineering application. To address this gap, this study proposes a robust, data-leakage-aware framework driven by a rigorous nested GroupKFold cross-validation strategy. By grouping concrete samples by their unique Mix_ID, this approach ensures genuine generalization to entirely unseen mixtures. Within this reliable validation scheme, the CatBoost algorithm is utilized for compressive-strength prediction, with the Dung Beetle Optimizer (DBO) serving as an effective tool for hyperparameter tuning. The evaluation results across multiple random seeds show that the DBO–CatBoost model significantly outperforms the default CatBoost, rigorously tuned baseline models (Support Vector Regression and Random Forest), and a comparative metaheuristic benchmark (PSO–CatBoost). It achieves the most stable distribution of errors and excellent predictive accuracy (Test R2=0.9995±0.0002, RMSE = 0.3828±0.0909). In addition, the model predictions were demystified using the methods of SHapley Additive exPlanations (SHAP) and partial dependence plots (PDPs). The interpretability analysis revealed strong statistical associations, showing that Curing Time and Coarse Aggregate are the most prominent predictive features and the strongest pairwise interaction between each other; the NaOH molar concentration is the most important second-level influence on optimization of strength. Overall, the framework provides a robust data-driven screening tool that can assist in preliminary mix-design evaluation. By reducing the reliance on extensive empirical “trial and error” approaches, this predictive model supports more efficient material usage and facilitates preliminary optimization of low-carbon concrete formulations. Theoretically, this study advances the fundamental science of geopolymer materials by explicitly quantifying the complex, non-linear interactions between alkaline activators, curing conditions, and recycled aggregates. This provides a robust data-driven theoretical foundation for designing and optimizing next-generation eco-friendly concrete products and structures. Full article
Show Figures

Figure 1

39 pages, 7462 KB  
Article
A Robust Model Evaluation Process for Early-Stage Cooling Load Prediction of Buildings
by Yaren Aydın, Ümit Işıkdağ, Sinan Melih Nigdeli, Gebrail Bekdaş, Wook-Won Kim and Zong Woo Geem
Processes 2026, 14(16), 2667; https://doi.org/10.3390/pr14162667 - 20 Aug 2026
Viewed by 286
Abstract
In the construction industry, a large portion of energy is spent on heating and cooling, which both increases costs and contributes to resource depletion. The aim of the study was to provide and evaluate a robust ML model evaluation process for early design [...] Read more.
In the construction industry, a large portion of energy is spent on heating and cooling, which both increases costs and contributes to resource depletion. The aim of the study was to provide and evaluate a robust ML model evaluation process for early design stage cooling load prediction of buildings. For this purpose, 18 different machine learning models were evaluated using a Nested Cross-Validation approach consisting of 50 outer fold and 50 inner Optuna trials, along with hyperparameter optimization. To avoid model selection being dependent on small decimal differences, paired model comparisons, effect sizes, Holm-corrected statistical tests, and the 1-SE economy rule were applied over the same outer folds. As a result of the analysis, Categorical Boosting (CatBoost) was selected as the final model, and within the Nested-CV framework, R2 = 0.8275 ± 0.0072, RMSE = 1.6792 ± 0.0210 kWh, MAE = 1.4356 ± 0.0233 kWh, and MAPE = 0.0521 ± 0.0009 were obtained. Model interpretability analyses showed that the variables Ambient Temperature, Solar Radiation, and Heat Reflective Treatment had the highest permutation importance values. Residual analyses revealed that the model exhibited low systematic bias, but the residual variance was dependent on the estimate value, and the residuals deviated from a normal distribution. This study provides a framework that evaluates not only the prediction performance but also model selection, generalization stability, interpretability, and residual behavior together. The findings demonstrate that CatBoost is a strong option for cooling load prediction in this simulation-based dataset. However, validation of the obtained results with real building data and different climatic conditions is considered an important requirement for future studies in terms of evaluating the external validity of the model. Full article
Show Figures

Figure 1

24 pages, 515 KB  
Article
Reliable Machine Learning Screening of Adsorption Energies Is Better Assessed with Formula-Grouped Cross-Validation
by Wenjie Wu, Mingling Yang, Ping Cheng and Yangning Wang
Catalysts 2026, 16(8), 742; https://doi.org/10.3390/catal16080742 - 20 Aug 2026
Viewed by 168
Abstract
Machine learning (ML) models trained on bulk-crystal descriptors are increasingly used to prescreen catalysts by predicting adsorption energies, yet reported performances often rely on random K-fold cross-validation that permits the same bulk formula to appear in both training and test sets. We [...] Read more.
Machine learning (ML) models trained on bulk-crystal descriptors are increasingly used to prescreen catalysts by predicting adsorption energies, yet reported performances often rely on random K-fold cross-validation that permits the same bulk formula to appear in both training and test sets. We construct a reproducible benchmark that fuses 936 CatApp DFT adsorption energies with bulk descriptors from the Materials Project for H*, O*, and OH* on metal and alloy surfaces. We compare random K-fold cross-validation with GroupKFold grouped by parsed formula, the latter mimicking the realistic task of predicting adsorption on entirely new catalyst compositions. Under formula-grouped evaluation, random CV materially overestimates apparent generalization performance, with the largest and most robust effects for H* and OH* (protocol-inflation gaps up to approximately 0.8). The H* and OH* results are based on only 20 and 31 unique formulas, so their GroupKFold Spearman point estimates should be read as directional evidence rather than quantitative estimates. O* shows a smaller and statistically fragile protocol-inflation signal and, even where composition-plus-bulk features improve Random Forest and Ridge, the usable signal is best described as a very coarse pre-filter within a limited domain. Bulk descriptors are adsorbate-dependent: they improve O* prediction for Random Forest and Ridge, but degrade H* and OH*—a qualitative, directional observation given the small formula counts—whose binding is poorly captured by bulk crystal descriptors, consistent with the established view that it is governed by surface-localized electronic structure. These results outline a realistic performance boundary for bulk-to-surface ML in this benchmark: O* can be very coarsely prioritized from bulk descriptors within a limited domain, whereas H* and OH* are unlikely to be quantitatively predicted from bulk descriptors alone and would benefit from surface-aware models. We therefore recommend that bulk-to-surface adsorption-energy benchmarks report formula-grouped cross-validation alongside random cross-validation as a more robust and transparent practice. Full article
(This article belongs to the Section Electrocatalysis)
Show Figures

Figure 1

35 pages, 6931 KB  
Article
A Prediction Model for Operator Diagnosis Level Integrating SACADA Database and Machine Learning in a Main Control Room of Nuclear Power Plants
by Huan Xiao, Jianjun Jiang, Wenming Chen and Zetian Tao
Appl. Sci. 2026, 16(16), 8264; https://doi.org/10.3390/app16168264 - 19 Aug 2026
Viewed by 156
Abstract
Operator diagnosis level in a main control room (MCR) of Nuclear Power Plants (NPPs) is a core factor in preventing human errors and ensuring the safe operation of NPPs. Due to the high uncertainty of human behaviors and the scarcity of relevant data, [...] Read more.
Operator diagnosis level in a main control room (MCR) of Nuclear Power Plants (NPPs) is a core factor in preventing human errors and ensuring the safe operation of NPPs. Due to the high uncertainty of human behaviors and the scarcity of relevant data, traditional analysis methods mainly rely on empirical judgment, which suffer from insufficient dynamics and poor engineering adaptability. To address the issues, this paper conducts a study on an AI prediction model for operator diagnosis level in a MCR of NPPs based on the SACADA database and machine learning technology. The model adopts a probabilistic neural network (PNN) as the main architecture, and proposes a hybrid method of network search considering density distribution combined with K-fold cross-validation, which breaks the traditional mode of a single smoothing factor adapting to an entire dataset. The analysis results show that the performance of the hybrid method proposed in this paper outperforms network search + K-fold cross-validation and particle swarm optimization + K-fold cross-validation methods in terms of accuracy, precision, recall, and F1-score. The five-fold cross-validation verifies that the model has good stability and good generalization ability. Further, the model is compared with common AI models such as BP neural network and RBF neural network. The results demonstrate that the proposed model has advantages in core indicators including overall accuracy (0.9444), macro-precision (0.9783), macro-recall (0.9063), and macro-F1-score (0.9362), and can effectively solve the problems of insufficient recognition of minority-class samples, overfitting, and underfitting. This research achieves professional and in-depth application of the SACADA database for diagnosis level prediction, extends existing research on prediction tasks, and delivers valuable theoretical insights and practical application significance. Full article
Show Figures

Figure 1

19 pages, 2289 KB  
Article
Machine Learning Identifies High-Risk Suicide Profiles in a Population-Based Forensic Registry
by Alin Ionut Piraianu, Anisia-Luiza Culea-Florescu, Elena Stamate, Ana Fulga, Doriana Iancu, Octavian Stefan Patrascanu and Iuliu Fulga
Diagnostics 2026, 16(16), 2615; https://doi.org/10.3390/diagnostics16162615 - 18 Aug 2026
Viewed by 230
Abstract
Background: Suicide is a leading cause of preventable death, yet machine learning (ML) analyses of forensic (medico-legal) suicide data are scarce and, to our knowledge, absent for Romania. Population-based forensic registries offer exhaustive, autopsy-confirmed coverage that is structurally distinct from clinical or civil [...] Read more.
Background: Suicide is a leading cause of preventable death, yet machine learning (ML) analyses of forensic (medico-legal) suicide data are scarce and, to our knowledge, absent for Romania. Population-based forensic registries offer exhaustive, autopsy-confirmed coverage that is structurally distinct from clinical or civil death-registration data. We applied supervised and unsupervised ML to a complete regional medico-legal suicide registry to profile the method of death and to identify latent victim subgroups of preventive relevance. Methods: We analysed 395 consecutive suicide deaths (Galați and Brăila counties, ≈750,000 inhabitants; 2018–2024). Two supervised classifiers—L2-regularised logistic regression (LR) and random forest (RF, 200 trees)—were trained to discriminate hanging from other methods, using eleven sociodemographic and clinical predictors, and evaluated by 10-fold stratified cross-validation. Given severe class imbalance, the area under the ROC curve (AUC) was the primary metric. Model hyperparameters were fixed a priori, and no class-imbalance correction was applied; both decisions are pre-specified and justified in the Methods. Robustness was assessed by stratified non-parametric bootstrap confidence intervals for the odds ratios, a tipping-point sensitivity analysis for the undocumented clinical fields, and Ward-linkage hierarchical clustering as an independent partitioning check. Predictor importance was quantified by out-of-bag (OOB) permutation importance and Spearman correlations. Unsupervised structure was assessed by K-means clustering (k = 2–7), with the optimal solution selected by the average silhouette coefficient and the elbow (WCSS) criterion. Reporting followed TRIPOD+AI and STROBE. Results: The study population was predominantly male (87.1%) and rural (73.2%), with a mean age of 54.1 years; hanging accounted for 94.2% of deaths—far above the European average (≈50%). RF achieved AUC = 0.865 ± 0.181 and LR AUC = 0.847 ± 0.192, both within the “excellent” discrimination band; sensitivity was very high (0.995–0.997) and specificity was limited (0.233–0.367), an expected consequence of imbalance. Prior suicide attempts (OOB importance 0.959; Spearman ρ = −0.549, p < 0.001; OR = 0.496, 95% CI 0.25–0.72) and the presence of a suicide note (importance 0.575; ρ = −0.439, p < 0.001; OR = 0.549, 95% CI 0.34–0.76) were the dominant predictors. K-means identified two well-separated clusters (silhouette = 0.707), and the partition was reproduced exactly by Ward-linkage hierarchical clustering (adjusted Rand index = 1.000). Cluster 2 (n = 19; 4.8%) was a clinically distinct, younger subgroup (42.4 vs. 54.7 years) characterised by prior attempts (57.9% vs. 0%), suicide notes (68.4% vs. 0%), higher psychiatric comorbidity (52.6% vs. 30.9%) and lower hanging proportion (36.8% vs. 97.1%)—an exploratory, hypothesis-generating profile of recurrent suicidal behaviour with documented prior contact with the medical or medico-legal system. The principal findings were stable across all plausible degrees of clinical under-documentation in the tipping-point sensitivity analysis. Conclusions: ML applied to a complete forensic suicide registry reproduced known regional epidemiology and, beyond classical statistics, isolated an exploratory but clinically coherent high-risk subgroup of direct relevance to the audit of structured post-attempt follow-up. This is, to our knowledge, the first ML study of Romanian forensic suicide data and supports integrating ML into medico-legal research and into the regional targeting and audit of existing post-attempt follow-up provision. Full article
(This article belongs to the Section Forensic Diagnostics)
Show Figures

Figure 1

32 pages, 45242 KB  
Article
Automated Multimodal Sleep Staging Using DWT-Based Wavelet Decomposition and Explainable Machine Learning with Signal Sculpting Topographies
by Adnan Sami Sarker, Kazi Mahatir Mohammed Samir, Zunayed Khan Shakib, Md Kishor Morol and Tze Hui Liew
Diagnostics 2026, 16(16), 2609; https://doi.org/10.3390/diagnostics16162609 - 17 Aug 2026
Viewed by 314
Abstract
Objectives: Sleep staging from polysomnographic (PSG) recordings is clinically critical for diagnosing sleep-related disorders, yet manual scoring by certified technologists remains time-consuming, costly, and subject to inter-rater variability. Methods: This study presents an automated, explainable, and multimodal framework for five-class sleep [...] Read more.
Objectives: Sleep staging from polysomnographic (PSG) recordings is clinically critical for diagnosing sleep-related disorders, yet manual scoring by certified technologists remains time-consuming, costly, and subject to inter-rater variability. Methods: This study presents an automated, explainable, and multimodal framework for five-class sleep stage classification using simultaneously acquired electroencephalography (EEG), electrooculography (EOG), and electromyography (EMG) signals. A total of 1946 annotated 30 s epochs from 30 healthy adult recording sessions (Sleep-EDF Expanded and Sleep Cassette subset) were processed through a 37-dimensional multimodal feature extraction pipeline encompassing temporal amplitude statistics, frequency-domain spectral band powers, nonlinear entropy and complexity measures, and Daubechies-4 discrete wavelet transform (DWT) energy coefficients. Four classical machine learning classifiers -Random Forest (RF), Support Vector Machine with radial basis function kernel (SVM-RBF), Gradient Boosting (GB), and K-Nearest Neighbours (KNN, k = 7) were benchmarked under stratified five-fold cross-validation. Results: SVM-RBF achieved the highest macro-averaged F1-score of 0.7322 (Cohen’s kappa 0.6784, overall accuracy 75.18%). N3 deep slow-wave sleep achieved the highest per-class F1 of 0.879, while N1 light sleep was the most challenging (F1 = 0.668). SHapley Additive exPlanations (SHAP) and RF mean decrease in Gini impurity (MDGI) analysis jointly identified EMG root mean square amplitude (MDGI = 0.0805), gamma band power (0.0784), and permutation entropy (0.0434) as the three most discriminative features. As a novel methodological contribution, sixteen categories of signal sculpting visualisations were developed, translating abstract multivariate features into clinically interpretable graphical representations. Conclusions: The proposed framework achieves substantial kappa agreement approaching the lower bound of expert inter-rater reliability (0.76–0.82) while providing full model transparency, with direct implications for wearable sleep monitoring device design. Full article
(This article belongs to the Section Machine Learning and Artificial Intelligence in Diagnostics)
Show Figures

Graphical abstract

25 pages, 455 KB  
Article
Benchmarking Supervised Classifiers for Concurrent Multidomain Dropout-Intention Attributions in Higher Education: Evidence from a Colombian Public University
by Marieth Agnes Guillen-García, Osnamir Elias Bru-Cordero and Cristian David Correa-Álvarez
Big Data Cogn. Comput. 2026, 10(8), 274; https://doi.org/10.3390/bdcc10080274 - 16 Aug 2026
Viewed by 226
Abstract
Student retention analytics often treats withdrawal as a single outcome, although students may attribute dropout intention to personal, socioeconomic, and academic pressures simultaneously. We benchmarked nine supervised classifiers for identifying a concurrent three-domain attribution profile in a cross-sectional survey of 333 undergraduates at [...] Read more.
Student retention analytics often treats withdrawal as a single outcome, although students may attribute dropout intention to personal, socioeconomic, and academic pressures simultaneously. We benchmarked nine supervised classifiers for identifying a concurrent three-domain attribution profile in a cross-sectional survey of 333 undergraduates at a Colombian public university campus. The response came from a semi-structured weight-allocation item; an audit found that literal label matching altered 21 classifications because of spelling variants and decimal notation. Nine classifiers—logistic regression, decision tree, random forest, neural network, Gaussian Naïve Bayes, k-nearest neighbors, AdaBoost, gradient boosting, and XGBoost—were fitted using six pre-specified predictors. Models were compared by repeated nested stratified cross-validation (five outer folds, three repeats), inner tuning, fold-contained preprocessing, and training-only threshold selection. The concurrent profile occurred in 256 students (76.9%). Logistic regression achieved the highest mean held-out ROC AUC (0.674, 95% CI 0.644–0.704), closely followed by random forest (0.671, 0.641–0.702); their paired difference was nonsignificant after Holm adjustment. Logistic regression had the highest F1 score (0.788), whereas random forest had the highest balanced accuracy (0.617). AdaBoost did not retain its apparent single-holdout advantage. Housing and financial aid had the largest held-out permutation importance. The predictors provided moderate discrimination of a perceptual profile, not a validated prediction of future dropout. Outcome auditing and leakage-free validation materially changed the model ranking. Full article
Show Figures

Figure 1

20 pages, 1763 KB  
Article
Temperature-Adaptive Activation Energy for Maturity-Based Strength Prediction of Sustainable, SCM-Blended Self-Compacting Concrete
by Abdulaziz Aldawish, Sivakumar Kulasegaram, Ayman Almutlaqah and Abdullah Alshahrani
Materials 2026, 19(16), 3462; https://doi.org/10.3390/ma19163462 - 14 Aug 2026
Viewed by 277
Abstract
The maturity method (ASTM C1074) predicts in situ concrete strength from a recorded temperature history but assumes a constant apparent activation energy, contradicting the experimental evidence that the activation energy falls as hydration shifts from kinetics control to diffusion control—an effect that differs [...] Read more.
The maturity method (ASTM C1074) predicts in situ concrete strength from a recorded temperature history but assumes a constant apparent activation energy, contradicting the experimental evidence that the activation energy falls as hydration shifts from kinetics control to diffusion control—an effect that differs between binder chemistries when supplementary cementitious materials (SCMs) are used. This study develops a physics-based maturity model in which the apparent activation energy varies linearly with temperature, Q(T) = Q0 + βQ(TTref), coupling a variable-energy Arrhenius equivalent age to a hyperbolic strength–maturity relationship. The model was calibrated on 196 mean-strength observations (588 cube tests) from seven self-compacting concrete mixtures cured isothermally at 10, 20, 35 and 50 °C and tested at seven ages (1–90 days). All four SCM systems (fly ash, GGBS, silica fume and rice husk ash) returned a negative coefficient (−210 to −974), enclosing the temperature sensitivity implied by independent calorimetric measurements on Portland cement paste (≈−580 J/(mol·K)), whereas the ordinary Portland cement control returned a positive point estimate (+101) that is not statistically distinguishable from zero. The model achieved R2 = 0.929 (RMSE = 4.74 MPa), outperforming the constant-energy ASTM C1074 baseline in both accuracy and the Akaike Information Criterion while eliminating its systematic bias at the temperature extremes. Five-fold cross-validation confirms the out-of-sample accuracy (R2 = 0.901, RMSE = 5.59 MPa), and bootstrap analysis shows the negative coefficients of the fly ash, GGBS and rice husk ash systems to be statistically significant. External validation on 120 independent literature observations gave R2 = 0.881. Full article
Show Figures

Graphical abstract

Back to TopTop