Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (263)

Search Parameters:
Keywords = ensemble voting classification

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
26 pages, 642 KB  
Article
Predicting Student Stress Using Machine Learning Ensemble Models: A Multi-Criteria Comparison with Explainable Artificial Intelligence Analysis
by Daniel Cristóbal Andrade-Girón, William Joel Marin-Rodriguez, Marcelo Gumercindo Zuñiga-Rojas, Abrahan Cesar Neri-Ayala, Edgar Tito Susanibar-Ramírez and Miguel Angel Aguilar-Luna-Victoria
AI 2026, 7(7), 268; https://doi.org/10.3390/ai7070268 - 18 Jul 2026
Viewed by 329
Abstract
Student stress is a significant mental health issue in educational settings; therefore, developing reliable, calibrated, and interpretable predictive models can support the classification of observed stress levels. This study analyzed the public Student Stress Factors dataset, comprising 1100 records, 20 predictors, and one [...] Read more.
Student stress is a significant mental health issue in educational settings; therefore, developing reliable, calibrated, and interpretable predictive models can support the classification of observed stress levels. This study analyzed the public Student Stress Factors dataset, comprising 1100 records, 20 predictors, and one target variable, using a supervised machine learning pipeline designed to reduce information leakage. The pipeline included stratified data partitioning, encapsulated preprocessing, nested cross-validation restricted to the training data, and independent holdout evaluation. Nine ensemble and boosting algorithms for tabular data were compared: AdaBoost, Gradient Boosting, Random Forest, Extra Trees, Bagging, Voting, Stacking, XGBoost, and LightGBM. Model performance was assessed using key discrimination and calibration metrics, together with the nonparametric Friedman test for statistical comparison. Gradient Boosting achieved the best average performance in nested cross-validation, with an accuracy of 89.55 ± 3.16%, F1-weighted of 89.54 ± 3.17%, MCC of 0.845 ± 0.047, and ROC-AUC weighted of 98.59 ± 0.92%. XGBoost and LightGBM showed comparable performance. In the independent holdout set, the final calibrated model maintained robust predictive performance, achieving an accuracy of 0.8818, F1-weighted of 0.8818, MCC of 0.8237, and ROC-AUC weighted of 0.9861. Although the overall results indicate stable and high predictive performance, the Friedman test did not identify statistically significant differences among the algorithms, χ2 = 10.953, p = 0.204. Therefore, model selection should consider not only predictive accuracy but also computational efficiency, calibration, interpretability, and implementation feasibility. Despite the internal stability of the pipeline and satisfactory holdout performance, the public and cross-sectional nature of the dataset limits causal inference and model transferability. Consequently, external and prospective validation is required before integration into institutional early warning systems. Full article
Show Figures

Figure 1

23 pages, 1802 KB  
Article
Leakage-Aware Transfer Learning with Explainable AI and CPU-Efficient Deployment for Mango Leaf Disease Classification on the MangoLeafBD Benchmark
by Wirapong Chansanam, Avshalom Elmalech, Suparp Kanyacome, Prasert Luekhong, Natthakan Iam-On and Tossapon Boongoen
Appl. Sci. 2026, 16(14), 6989; https://doi.org/10.3390/app16146989 - 12 Jul 2026
Viewed by 383
Abstract
Background and Aim: Mango (Mangifera indica) is a globally important fruit crop whose productivity is repeatedly threatened by foliar diseases such as anthracnose, bacterial canker, powdery mildew, die-back, and sooty mould. Although deep learning has rapidly advanced automated leaf-disease diagnosis, recently reported accuracies [...] Read more.
Background and Aim: Mango (Mangifera indica) is a globally important fruit crop whose productivity is repeatedly threatened by foliar diseases such as anthracnose, bacterial canker, powdery mildew, die-back, and sooty mould. Although deep learning has rapidly advanced automated leaf-disease diagnosis, recently reported accuracies on the MangoLeafBD benchmark are approaching saturation, and many studies still rely on random data splits that may inflate performance through leakage of duplicate or near-duplicate images. This study aimed to develop and rigorously evaluate a leakage-aware, deployment-oriented deep learning framework for the complete eight-class MangoLeafBD task. Methods: Three modern transfer-learning backbones—EfficientNetB0, MobileNetV3Large, and ConvNeXtTiny—were fine-tuned using a two-stage training strategy on a group-aware data partition. Duplicate and near-duplicate images were detected with cleaned filename stems, average hashing (aHash), and difference hashing (dHash) and grouped before splitting to guarantee zero cross-partition overlap. The strongest models were combined through probability-level soft voting. Robustness was further assessed using 5-fold StratifiedGroupKFold cross-validation. Explainability was examined with Gradient-weighted Class Activation Mapping (Grad-CAM), deployment suitability was characterized through CPU latency benchmarking, and the framework was operationalized as a publicly accessible web-based diagnostic system. Results: EfficientNetB0 achieved the highest performance under the leakage-controlled protocol, reaching 99.50% accuracy and 99.50% weighted F1-score on the 599-image group-aware test set. The heterogeneous soft-voting ensemble matched EfficientNetB0 but did not exceed it, indicating that the ensemble gains reported in earlier studies may partly reflect optimistic split conditions. Five-fold grouped cross-validation confirmed stability, yielding a mean accuracy of 99.63% (SD = 0.20%) and a mean weighted F1 of 99.62% (SD = 0.20%). Grad-CAM visualizations showed that the model attended to biologically meaningful lesion regions across all eight classes, and CPU benchmarking produced a mean single-image latency of 25.3 ms (p95 = 27.6 ms) with batch throughput scaling up to 166.5 images per second. Conclusions: The proposed framework demonstrates that a compact, leakage-aware EfficientNetB0 model can match the accuracy of more complex hybrid CNN–transformer architectures while remaining interpretable and deployable on commodity CPU hardware. By coupling group-aware evaluation, explainable AI, latency benchmarking, and a publicly accessible web application, the study advances reproducible, deployment-ready precision agriculture research and offers a more rigorous benchmarking protocol for future mango leaf disease classification studies. Full article
(This article belongs to the Special Issue The Application of Deep Learning in Image Processing)
Show Figures

Figure 1

14 pages, 3460 KB  
Article
Pilot-Site Land Cover Mapping Using an Externally-Guided Clustering Framework: A Case Study from Ontario, Canada
by Sondos Omar, Reza Shahidi, Masoud Mahdianpari and Fariba Mohammadimanesh
Geomatics 2026, 6(4), 77; https://doi.org/10.3390/geomatics6040077 - 10 Jul 2026
Viewed by 219
Abstract
High-resolution land cover classification is critical for monitoring environmental change and managing natural resources. This study presents an unsupervised framework with externally guided feature prioritization that integrates Sentinel-1 synthetic aperture radar (SAR) and Sentinel-2 optical imagery at 10 m spatial resolution. A cloud-native [...] Read more.
High-resolution land cover classification is critical for monitoring environmental change and managing natural resources. This study presents an unsupervised framework with externally guided feature prioritization that integrates Sentinel-1 synthetic aperture radar (SAR) and Sentinel-2 optical imagery at 10 m spatial resolution. A cloud-native export protocol in Google Earth Engine (GEE) enables the generation of consistent, cloud-free, and snow-free seasonal composites across Ontario, Canada. A comprehensive feature engineering pipeline combines spectral indices, radar backscatter metrics, terrain derivatives from digital elevation models (DEMs), and temporal statistics to create a rich multi-sensor input space. Dimensionality reduction is performed using Sparse Principal Component Analysis (SparsePCA) and mutual-information-based feature selection. Clustering is conducted using three complementary algorithms: centroid-based K-means, density-based Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN), and reachability-based Ordering Points To Identify the Clustering Structure (OPTICS). Final land cover labels are assigned via a majority-voting ensemble, with prediction ties resolved deterministically using OPTICS. OPTICS is particularly effective for modeling heterogeneous landscapes due to its ability to detect clusters of varying density without requiring a global threshold. This study is designed as a pilot-site methodological demonstration using three representative 2 km × 2 km regions in Ontario, rather than a full provincial-scale land cover product. The resulting classification maps are validated against reference land cover data, demonstrating the effectiveness and potential scalability of the proposed external-label guided unsupervised mapping approach. Full article
Show Figures

Figure 1

42 pages, 1832 KB  
Article
Profiling Organizational AI Readiness in Thailand’s Logistics Industry Using TOE–UTAUT Features, Clustering Analysis, and Explainable Machine Learning
by Wipada Sriwichien, Warawut Narkbunnum and Kittipol Wisaeng
Information 2026, 17(7), 672; https://doi.org/10.3390/info17070672 - 10 Jul 2026
Viewed by 360
Abstract
Artificial intelligence (AI) adoption within logistics organizations remains uneven despite increasing digital transformation initiatives in emerging economies. This study investigates respondent-perceived organizational AI readiness profiles in Thailand’s logistics industry using an integrated analytical framework combining TOE–UTAUT predictors, clustering analysis, supervised machine learning, and [...] Read more.
Artificial intelligence (AI) adoption within logistics organizations remains uneven despite increasing digital transformation initiatives in emerging economies. This study investigates respondent-perceived organizational AI readiness profiles in Thailand’s logistics industry using an integrated analytical framework combining TOE–UTAUT predictors, clustering analysis, supervised machine learning, and explainable artificial intelligence techniques. Data were collected from 520 logistics and supply chain professionals in Thailand using a structured questionnaire. K-means clustering was applied to identify internally derived respondent-perceived AI readiness profiles, while Random Forest, Support Vector Machine (SVM), XGBoost, and LightGBM models were developed to classify readiness-profile membership. A weighted voting ensemble model was additionally employed to assess classification robustness and profile-differentiation stability across multiple learning algorithms. The findings identified three internally derived respondent-perceived AI readiness profiles representing relatively low, moderate, and advanced readiness patterns within the TOE–UTAUT feature space. Among the evaluated models, the SVM classifier achieved the strongest classification performance, obtaining the highest accuracy and AUC values. SHAP analysis indicated that Actual Use, Technological Factors, Facilitating Conditions, and Behavioral Intention exhibited the largest feature-attribution contributions within the readiness-profile classification framework. The study contributes to AI adoption research by integrating clustering-based segmentation, machine-learning classification, and explainable artificial intelligence into a unified readiness-profiling framework. The findings provide practical insights for managers and policymakers seeking to understand respondent-perceived organizational AI readiness patterns and support digital transformation initiatives within logistics professional contexts. Full article
Show Figures

Figure 1

44 pages, 6943 KB  
Article
HFW-NPO: A Dual a Paradigm Hybrid Filter–Wrapper Nomadic People Optimizer Framework for High-Dimensional Alzheimer’s Gene Expression Classification
by Almuntadher Mahmood Alwhelat and Rahib H. Abiyev
Electronics 2026, 15(13), 2970; https://doi.org/10.3390/electronics15132970 - 7 Jul 2026
Viewed by 381
Abstract
Alzheimer’s Disease (AD) necessitates high-resolution transcriptomic biomarkers for early detection, yet current computational methods are hampered by high-dimensional search space and publication bias regarding imbalanced datasets. We propose the Hybrid Filter–Wrapper Nomadic People Optimizer, a three-stage pipeline integrating a tri-criterion filter, an enhanced [...] Read more.
Alzheimer’s Disease (AD) necessitates high-resolution transcriptomic biomarkers for early detection, yet current computational methods are hampered by high-dimensional search space and publication bias regarding imbalanced datasets. We propose the Hybrid Filter–Wrapper Nomadic People Optimizer, a three-stage pipeline integrating a tri-criterion filter, an enhanced NPO wrapper with adaptive Lévy-scale anti-stagnation mechanism, and a five-member soft-voting ensemble. The system was evaluated using a dual-paradigm protocol; Scenario A (balance brain tissue; GEO dataset GSE 33000, GSE 132903, GSE122063) and Scenario B (imbalanced peripheral blood: GSE 63060 + GSE 636061). In scenario A, HFW-NPO outperformed 13 published methods, achieving balanced accuracy of 85.28%, 87.16%, and 96.67% while identifying compact panels of 29–32 probes per fold (observed range: 24–38). Scenario B, evaluated on a merged 478-samples peripheral blood cohort (GSE63060 + GSE 636061 imbalanced 1.48:1) with z-score batch harmonization and RSKF (5 × 10) cross-validation, achieved a balanced accuracy of 59.53% and MCI Recall of 63.50 ± 14.02%, providing the first reproducible baseline for this clinically challenging task, while acknowledging that 59.53% balanced accuracy does not yet reach clinically actionable levels. By providing transparent reporting across both balanced and severely imbalanced datasets, this study establishes a state-of-the-art, reproducible framework for AD biomarker discovery and provides a critical baseline for the challenging task of transcriptomic-based classification in peripheral blood samples. Result is currently scoped to Illumina HumanHT-12 microarray data, and cross-platform validation on RNA-seq cohorts is identified as a priority future extension. Full article
Show Figures

Graphical abstract

33 pages, 3549 KB  
Article
A Stability-Driven Framework for Automated Operational Crop Mapping Using Optical and Radar Satellite Image Time Series
by Maryam Choukri, Yacine Bouroubi, Jamal-Eddine Ouzemou, Abdelghani Chehbouni and Ahmed Laamrani
Remote Sens. 2026, 18(13), 2149; https://doi.org/10.3390/rs18132149 - 2 Jul 2026
Viewed by 295
Abstract
Operational crop mapping requires classifiers capable of robust generalization across years. While feature importance is routinely used for model optimization, its temporal stability has rarely been systematically investigated, creating a critical gap in deploying reliable monitoring systems. This study moves beyond identifying “most [...] Read more.
Operational crop mapping requires classifiers capable of robust generalization across years. While feature importance is routinely used for model optimization, its temporal stability has rarely been systematically investigated, creating a critical gap in deploying reliable monitoring systems. This study moves beyond identifying “most important” features to systematically evaluate and quantify their inter-annual stability for enabling automated classification. Using six agricultural years (2018, 2019, 2020, 2023, 2024 and 2025) of Sentinel-1 and Sentinel-2 data over Morocco, we extracted 156 multi-sensor features across 12 monthly composites and analyzed their importance stability through statistical metrics, clustering, and novel composite indices: the Reliability Index (RI) and Automatic Selection Score (AuSS). This framework automates feature selection by ranking features with RI and AuSS and then applying Pareto optimization to identify a minimal stable feature set—without requiring annual retraining or expert intervention. Our analysis confirms a fundamental tension: the most discriminative features (e.g., NDVI, VH, VV) are also the most volatile, while stable features (e.g., NDRE, MSI, NDMI) offer modest predictive power. Hierarchical clustering revealed four behavioral typologies (Dominant Stable, Performant Volatile, Stable Minor, and Noise), guiding strategic feature management. Crucially, a Pareto analysis demonstrated that a refined portfolio of 6 indices (VH, VV, NDVI, NDRE, GCVI, RVI) captures 57.2% of cumulative predictive importance, filtering out inter-annual noise while preserving discriminative signal. The Voting Ensemble leveraging this Stable Portfolio maintained consistent high accuracy (87.4% accuracy, 87.2% F1-score) with minimal performance degradation during temporal transfer, while models based on volatile top features exhibited significant drops. Entropy analysis confirmed that all features in the Stable Portfolio provide consistent informational certainty, indicating that stability-driven selection does not increase model uncertainty. We conclude that feature stability is not merely a diagnostic metric but a foundational criterion for operational design. We propose a practical, metrics-driven framework for constructing automated crop classification systems that are more resilient to inter-annual climate variability. Full article
Show Figures

Figure 1

27 pages, 11691 KB  
Article
GoldFormer: A Texture-Aware Vision Transformer-Based Algorithm for Detecting Near-Identical Images
by Zobeir Raisi
Algorithms 2026, 19(7), 530; https://doi.org/10.3390/a19070530 - 1 Jul 2026
Viewed by 355
Abstract
Distinguishing authentic gold products from high-quality counterfeits is a challenging fine-grained computer vision problem; counterfeit items are engineered to replicate surface texture, hallmark engravings, color, and geometry with remarkable fidelity, making visual discrimination unreliable even for trained professionals. In this paper, we address [...] Read more.
Distinguishing authentic gold products from high-quality counterfeits is a challenging fine-grained computer vision problem; counterfeit items are engineered to replicate surface texture, hallmark engravings, color, and geometry with remarkable fidelity, making visual discrimination unreliable even for trained professionals. In this paper, we address the problem of visual gold authentication from unconstrained smartphone imagery in three main contributions. First, we introduce GoldNet, a public benchmark dataset designed for this task, comprising 2127 real-world images of authentic and counterfeit gold items collected under diverse real-world conditions. Second, we evaluate fourteen classification architectures spanning classical handcrafted texture descriptors, convolutional neural networks (CNNs), and vision transformers under a rigorous transfer learning protocol, establishing the first comprehensive baseline for this problem. Third, we propose GoldFormer, a hybrid dual-stream algorithm that combines the local texture representations of ResNet-50 with the global contextual modeling capability of the Swin Transformer (Swin-T) through a newly designed Texture-Aware Attention Gate (TAAG) module. The TAAG dynamically modulates Swin feature dimensions using CNN-derived texture energy, providing improved discriminability and per-prediction interpretability without requiring post hoc attribution. Experimental results show that, under matched-resolution 5-fold cross-validation, the proposed GoldFormer attains the highest overall accuracy (95.02%, F1-score 0.9502) at roughly half the FLOPs of its higher-resolution setting, statistically tied with the strongest individual backbone (ViT-B/16, 94.31%; McNemar p=0.23) and on par with a training-free soft-voting ensemble (94.92%), while significantly improving on its own Swin-T backbone (93.65%) and adding built-in, attribution-free texture-gate interpretability. GoldFormer surpasses trained human-expert performance (89.80%) by approximately 5 percentage points. Full article
(This article belongs to the Section Algorithms for Multidisciplinary Applications)
Show Figures

Figure 1

20 pages, 777 KB  
Article
Machine Learning-Based Forecasting of Sukuk Index Movements: Evidence from Turkey and Malaysia
by Mehmet Ekmekcioglu and Kaya Tokmakcioglu
J. Risk Financial Manag. 2026, 19(7), 486; https://doi.org/10.3390/jrfm19070486 - 1 Jul 2026
Viewed by 358
Abstract
Islamic finance has rapidly expanded into a major component of global financial systems, positioning Sukuk as a core instrument for Sharia-compliant investment and capital raising. Despite growing academic attention to Sukuk pricing and valuation, to the best of our knowledge, no prior study [...] Read more.
Islamic finance has rapidly expanded into a major component of global financial systems, positioning Sukuk as a core instrument for Sharia-compliant investment and capital raising. Despite growing academic attention to Sukuk pricing and valuation, to the best of our knowledge, no prior study has systematically classified the directional movements of Sukuk index prices using machine learning techniques. This study addresses that gap by developing predictive classification models for the directional (upward or downward) movements of Sukuk indices and applying them to both Turkey and Malaysia. It represents one of the first systematic attempts to forecast Sukuk index direction and the first application of machine learning-based directional forecasting to the Turkish Sukuk market. Using a diverse set of financial and macroeconomic indicators, we employ advanced machine learning algorithms including extreme gradient boosting (XGBoost), categorical boosting (CatBoost), and support vector machines (SVMs), and further enhance prediction accuracy through ensemble methods such as majority voting, weighted averaging, and soft voting. Model performance is evaluated using accuracy, precision, recall, F1-score, and ROC-AUC. The findings indicate that SVM consistently delivers the strongest standalone performance across both markets, while ensemble methods generate substantial improvements in Malaysia. Overall, predictive performance is higher in Malaysia, which may be associated with its more stable and liquid Sukuk market environment compared to Turkey. Full article
(This article belongs to the Special Issue Islamic Financial Markets in Times of Global Uncertainty)
Show Figures

Figure 1

34 pages, 1813 KB  
Article
Large Language Models as Explainable AI Ensemble Aggregators for Business Review Sentiment Analysis: A Comparative Study with Classical Ensembles
by Konstantinos I. Roumeliotis, Dionisis Margaris, Dimitris Spiliotopoulos and Costas Vassilakis
Appl. Sci. 2026, 16(13), 6479; https://doi.org/10.3390/app16136479 - 29 Jun 2026
Viewed by 255
Abstract
Online business reviews encode rich customer sentiment that is critical for commercial decision making, yet accurately predicting star ratings from free text remains a challenging five-class classification problem. Classical ensemble methods—Soft Voting, Weighted Voting, and Stacking—aggregate complementary base-model outputs to improve predictive performance, [...] Read more.
Online business reviews encode rich customer sentiment that is critical for commercial decision making, yet accurately predicting star ratings from free text remains a challenging five-class classification problem. Classical ensemble methods—Soft Voting, Weighted Voting, and Stacking—aggregate complementary base-model outputs to improve predictive performance, but they produce opaque decisions that are unintelligible to business stakeholders. This paper proposes using a large language model (LLM), specifically unsloth/LLaMA-3.3-70B-Instruct, as an Explainable AI (XAI) ensemble aggregator: the LLM receives the predictions and confidence scores of four heterogeneous base models (Logistic Regression, Support Vector Machine, Naïve Bayes, and BERT-base-uncased) and reasons over them to produce both a final star-rating prediction and a natural-language explanation. We evaluate the full pipeline on 10,000-sample balanced and natural-distribution test sets derived from the Yelp Academic Dataset, with additional cross-lingual validation on Spanish Amazon Reviews. The LLM aggregator (LLAMA_AGG) achieves the highest macro-F1 on both pipelines (0.6800 on balanced; 0.6720 on natural) and the best ordinal calibration (QWK = 0.9111 on balanced; 0.9337 on natural), outperforming all classical aggregators and base models. A detailed Explainable AI analysis reveals that the LLM revises 28.07% of its standalone predictions after observing the ensemble outputs, improving the accuracy by +22.2 percentage points on the revised cases. The aggregator corrects severe polar bias in the standalone LLM (±0.35 recall improvement on mid-range star classes) and produces longer explanations when evidence is conflicted—a quantitative signal of deliberative reasoning. A formal human evaluation with two judges confirms high explanation faithfulness (4.47/5) and readability (4.82/5). Model scale ablation shows an 8B parameter variant achieves 90.8% agreement with the 70B model, enabling practical deployment. These findings demonstrate that Explainable AI can be achieved through LLM-based ensemble aggregation, establishing a principled approach for business-review sentiment analysis. Full article
(This article belongs to the Special Issue The Age of Transformers: Emerging Trends and Applications)
Show Figures

Figure 1

27 pages, 662 KB  
Article
LLM-Augmented Ensemble Reasoning for Adversarial-Aware Power Quality Monitoring in Smart Grids
by Mubarak Alanazi
Electronics 2026, 15(13), 2788; https://doi.org/10.3390/electronics15132788 - 24 Jun 2026
Viewed by 279
Abstract
Deep learning models for power quality (PQ) disturbance classification remain critically vulnerable to adversarial perturbations, with classification performance degrading severely under white-box attacks. Existing defenses address individual models in isolation and provide no mechanism for operators to assess whether the system is under [...] Read more.
Deep learning models for power quality (PQ) disturbance classification remain critically vulnerable to adversarial perturbations, with classification performance degrading severely under white-box attacks. Existing defenses address individual models in isolation and provide no mechanism for operators to assess whether the system is under attack or which classifier remains trustworthy. This paper proposes a two-stage framework that combines adversarial training with large language model (LLM) reasoning to improve both robustness and interpretability. In the first stage, four architecturally diverse classifiers, including a proposed Multi-Scale Temporal Attention Network (MSTAN), are evaluated under four adversarial attacks (FGSM, PGD, C&W, and UAP), and their failure patterns are recorded as structured vulnerability fingerprints. The ensemble is then retrained via adversarial training on mixed clean and perturbed signals. In the second stage, an LLM analyzes the ensemble predictions alongside the fingerprint knowledge base to perform attack detection, fingerprint-guided meta-classification, and operator-facing threat report generation. On a 17-class, 255,000-signal synthetic benchmark, adversarial training recovers FGSM and PGD accuracy from below 25% to the 53–78% range, with MSTAN achieving the highest post-training robustness (78.26% under FGSM, 65.41% under PGD). The LLM reasoning layer provides an additional 3.5–6.2 percentage point improvement over majority voting by selecting the most reliable ensemble member based on the inferred attack condition, and detects adversarial attacks with 87.6% overall accuracy. To our knowledge, this is the first integration of LLM-based ensemble reasoning into the PQ adversarial robustness pipeline and the first application of the C&W optimization attack to power quality signals. Full article
Show Figures

Figure 1

31 pages, 1630 KB  
Article
DWRF-MVC: A Novel Random Forest Optimization Framework Combining DBSCAN Clustering and Multi-Metric Weighted Voting
by Tianhe Liu, Yanliang Zhou and Jie Cheng
Electronics 2026, 15(12), 2674; https://doi.org/10.3390/electronics15122674 - 16 Jun 2026
Viewed by 173
Abstract
Although Random Forest (RF) is a widely adopted ensemble method for classification, preserving diversity among decision trees and reducing the negative impact of underperforming trees remain major challenges. This paper proposes DWRF-MVC, a novel RF optimization framework that integrates DBSCAN clustering with a [...] Read more.
Although Random Forest (RF) is a widely adopted ensemble method for classification, preserving diversity among decision trees and reducing the negative impact of underperforming trees remain major challenges. This paper proposes DWRF-MVC, a novel RF optimization framework that integrates DBSCAN clustering with a multi-metric performance-weighted voting mechanism. Using a dissimilarity metric derived from classic RF outcomes, the framework applies DBSCAN to cluster decision trees, selects top-performing representatives from each cluster, and retains noise points to maintain diversity. A weighted voting mechanism is then introduced to further improve model performance. Experimental results on 14 benchmark datasets show that DWRF-MVC achieves an average accuracy improvement of 7.45% over classical RF with a 91.79% reduction in the number of trees. Moreover, DWRF-MVC surpasses two cutting-edge methods by 1.82% and 3.83% in average accuracy, respectively. Full article
(This article belongs to the Section Artificial Intelligence)
Show Figures

Figure 1

25 pages, 4402 KB  
Article
Sleep Stage Classification During CPAP Therapy from CPAP-Airflow and Wearable Fingertip Signals
by Hsin-Yu Chen, Aatif Husain, Andrey V. Zinchuk, Henry K. Yaggi, Muneeb Ahsan, Cheng-Yao Chen, Shirah Pokusa and Hau-Tieng Wu
Sensors 2026, 26(12), 3720; https://doi.org/10.3390/s26123720 - 11 Jun 2026
Viewed by 483
Abstract
Background: Continuous Positive Airway Pressure (CPAP) therapy is the standard treatment for obstructive sleep apnea–hypopnea syndrome (OSAHS), and photoplethysmography (PPG) sensors are commonly used in wearable devices for home sleep apnea testing. The recorded airflow and PPG signals from both sensors capture rich [...] Read more.
Background: Continuous Positive Airway Pressure (CPAP) therapy is the standard treatment for obstructive sleep apnea–hypopnea syndrome (OSAHS), and photoplethysmography (PPG) sensors are commonly used in wearable devices for home sleep apnea testing. The recorded airflow and PPG signals from both sensors capture rich physiological patterns. We hypothesize that by combining information from these signals, we can efficiently estimate sleep dynamics of patients receiving CPAP treatment. Methods: The airflow signals were obtained from CPAP titration devices, denoted as CPAP-airflow, while the PPG signals were collected using the PranaQ TipTraQ (TTQ001), a fingertip-worn wearable device. We separately trained one-dimensional convolutional neural networks for CPAP-airflow and PPG signals and fused their outputs through probabilistic ensembling to predict sleep stages. The ensemble method is a late-fusion soft-voting scheme that computes a linearly weighted combination of synchronized softmax probability vectors from the modality-specific models. Results: For three-stage classification (Wake, REM, NREM), the PPG-based and CPAP-airflow-based models achieved overall Cohen’s kappa scores of 0.511 and 0.452, respectively, while the ensembled model improved the overall kappa to 0.587. The F1-score for the REM stage improved to 0.706 using the ensemble method, compared to 0.685 and 0.532 achieved by the individual models, respectively. In the four-stage classification (Wake, REM, Light, Deep) task, a deep sleep sensitivity of 0.596 was attained through the application of probabilistic ensembling. Conclusions: A fusion scheme of complementary information from the CPAP and PPG enhances the accuracy of sleep stage detection and hence enables more precise sleep monitoring, especially with an improved REM identification. Clinical implications include applying the proposed algorithm to improve in-home auto-CPAP titration by capturing REM-related respiratory instability and avoiding under-titration in REM-dominant OSAHS, better reflecting the patient’s true nocturnal respiratory needs. Full article
(This article belongs to the Special Issue Wearable Technologies and Sensors for Health Monitoring)
Show Figures

Figure 1

28 pages, 2692 KB  
Article
Explainable Ensemble Convolutional Neural Networks for Automated Post-Disaster Structural Damage Assessment
by Anıl Sezgin, Merve Açıkgenç Ulaş, Görkem Gök, Hakan Güler, Nuray Beyza Avcı, Betül Bektaş Ekici, Nihal Arda Akyıldız, Mustafa Ulaş and Aytuğ Boyacı
Appl. Sci. 2026, 16(11), 5682; https://doi.org/10.3390/app16115682 - 5 Jun 2026
Viewed by 314
Abstract
The recent seismic activity in southeastern Turkey in February 2023 again emphasized the critical need to promptly evaluate structural damage to assist in emergency response operations. This study introduces a comprehensive ensemble deep learning approach to structural damage classification following earthquake events, based [...] Read more.
The recent seismic activity in southeastern Turkey in February 2023 again emphasized the critical need to promptly evaluate structural damage to assist in emergency response operations. This study introduces a comprehensive ensemble deep learning approach to structural damage classification following earthquake events, based on a dataset containing 13,270 high-resolution images with 15 different damage classes. Six different state-of-the-art convolutional neural network models (VGG16, ResNet50, InceptionV3, DenseNet121, EfficientNetB0, and MobileNetV2) are combined using a weighted voting approach to handle extreme class imbalance using weighted categorical cross-entropy loss. An integrated explainability component is incorporated into the trained convolutional neural network models to highlight the image regions that contribute to the predicted damage class, thereby improving the interpretability of deep learning decisions in safety-critical post-disaster assessment scenarios. The performance evaluation results show that the ensemble model achieves a test accuracy of 93.77%, with an increase of 2.67% compared to the best performing model individually. Notably, the ensemble model improves performance in minority classes like collapsed buildings. The proposed framework can be used to provide a powerful approach to structural damage evaluation, balancing accuracy with interpretability, to assist structural engineers in post-earthquake evaluation procedures. Full article
Show Figures

Figure 1

12 pages, 2179 KB  
Article
Raman Spectroscopy of Protein–Polysaccharide Conjugates: A Comparative Study of Tree-Based Ensemble Models
by Svetlana A. Shevtsova, Samvel A. Grigoryan, Oksana A. Mayorova, Mariia S. Saveleva and Ekaterina S. Prikhozhdenko
Macromol 2026, 6(2), 37; https://doi.org/10.3390/macromol6020037 - 3 Jun 2026
Viewed by 574
Abstract
Proteins with additives, especially in small quantities, are of great interest as a subject of study. Machine learning approaches implemented on Raman spectroscopy data could provide an insight into the chemical structures of such mixtures or conjugates. Although decision tree models could be [...] Read more.
Proteins with additives, especially in small quantities, are of great interest as a subject of study. Machine learning approaches implemented on Raman spectroscopy data could provide an insight into the chemical structures of such mixtures or conjugates. Although decision tree models could be powerful in solving either classification or regression tasks and could provide accessible predictions, they are prone to overfitting. Ensemble models that implement several decision trees could overcome the determined problem. Five different model types are discussed: RandomForest, GradientBoosting, AdaBoost, Voting, and Stacking. Raman spectroscopy data of whey protein isolates (5 wt.%) with different amounts of hyaluronic acid (0, 0.1, 0.25, and 0.5 wt.%) were used as datasets. In order to generalize the results of the study, WPI samples from three different manufacturers were used. Optimization established that ensembles of 200 decision trees with a maximum depth of four were optimal. The Stacking algorithm, which used RandomForest, GradientBoosting, and AdaBoost as base models with either LogisticRegressor (classification task) or RidgeCV (regression task), was found to be the most efficient in finding differences between the whey protein isolate and its conjugates with hyaluronic acid: specificity of 68.7% and sensitivity of 95.4% (classification task); R2 = 0.764 with mean absolute error of 0.068 (regression task). According to the feature importance plots, the Raman bands that were most influential in predicting the results were 1003 cm−1 (phenylalanine, ring breath), 1125 cm−1 (rocking of NH3+), 1206 cm−1 (C–C stretching), 1240 cm−1 (amide III (β-sheet), N–H in-plane bend, C–N stretch), and 1399 cm−1 (aspartic and glutamic acids, C=O stretch of COO–). The findings of this study may contribute to the development of novel methods for quality control and analysis of complex multicomponent systems in various industrial settings. In particular, the ensemble approach can be adapted for monitoring in food processing or as a screening tool in pharmaceutical formulation development. Full article
Show Figures

Figure 1

22 pages, 715 KB  
Article
Benchmarking of Ensembles and Meta-Ensembles in the Multiclass Classification of Obesity-Status Classification: Predictive Performance, Calibration and Interpretability
by Daniel Andrade-Girón, William Marin-Rodriguez, Americo Peña, Elsa Oscuvilca-Tapia and Fredy Bermejo-Sanchez
Informatics 2026, 13(6), 80; https://doi.org/10.3390/informatics13060080 - 3 Jun 2026
Viewed by 460
Abstract
Obesity is a major public health concern because of its high prevalence and association with cardiometabolic comorbidities. This study compared nine ensemble and meta-ensemble learning models for multiclass obesity-status classification using the Obesity Dataset, comprising 1610 records, 14 predictors, and four body-weight status [...] Read more.
Obesity is a major public health concern because of its high prevalence and association with cardiometabolic comorbidities. This study compared nine ensemble and meta-ensemble learning models for multiclass obesity-status classification using the Obesity Dataset, comprising 1610 records, 14 predictors, and four body-weight status classes. To ensure a leakage-aware evaluation, all preprocessing and resampling steps were embedded within the validation workflow. Standardization, one-hot encoding, and RandomOverSampler were applied only within the training folds; SMOTE and no-resampling configurations were retained as configurable alternatives but were not used to generate the reported results. Model performance was assessed using complementary classification, discrimination, agreement, and calibration metrics, including accuracy, balanced accuracy, weighted F1-score, macro F1-score, weighted ROC-AUC, Matthews correlation coefficient, Brier score, and multiclass expected calibration error. Overall, the ensemble models achieved strong discriminative performance, with eight of nine classifiers exceeding 82% accuracy and obtaining weighted ROC-AUC values close to or above 94%. LightGBM showed the strongest mean metric-based profile, with an accuracy of 85.41 ± 2.85%, weighted F1-score of 85.25 ± 2.88%, weighted ROC-AUC of 95.58 ± 1.52%, and MCC of 0.779 ± 0.042. Random Forest and Stacking achieved comparable classification performance, although Stacking presented poorer calibration. The Friedman test detected significant global differences among classifiers, χ2 = 38.7733, p = 0.000005. However, the Nemenyi post hoc test indicated that Stacking, Random Forest, LightGBM, Voting, Gradient Boosting, and Extra Trees belonged to the same high-performance statistical group. Therefore, LightGBM was selected as the final model based on its practical balance of predictive performance, calibration behavior, stability, and implementation feasibility, rather than on unequivocal statistical superiority. On the independent holdout set, LightGBM maintained strong generalization, achieving accuracy = 0.8447, weighted F1-score = 0.8435, MCC = 0.7653, and weighted ROC-AUC = 0.9464. Calibration was moderate, with Brier score = 0.2575 and multiclass ECE = 0.1070, indicating that predicted probabilities should be interpreted cautiously when used to support threshold-based decisions. Full article
Show Figures

Figure 1

Back to TopTop