Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (330)

Search Parameters:
Keywords = stacking model fusion

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
32 pages, 10029 KB  
Article
Multiclass Defect Classification from Legacy Foundry Data: A Decision Support System for Reducing Manual Inspection Time
by Joachim Denker, Loui Al-Shrouf and Mohieddine Jelali
Processes 2026, 14(18), 2885; https://doi.org/10.3390/pr14182885 - 10 Sep 2026
Abstract
This paper presents a machine learning-based decision support system for multiclass defect detection, utilizing exclusively heterogeneous legacy process data to minimize manual inspection time in foundries. Validated on 51,377 products across 193 defect categories, the methodology resolves structural data inconsistencies through k-nearest neighbor [...] Read more.
This paper presents a machine learning-based decision support system for multiclass defect detection, utilizing exclusively heterogeneous legacy process data to minimize manual inspection time in foundries. Validated on 51,377 products across 193 defect categories, the methodology resolves structural data inconsistencies through k-nearest neighbor (kNN) imputation and piecewise winsorization. A multi-stage feature selection cascade, incorporating variance thresholding, correlation filtering, and Random Forest Feature Importance (RFFI), reduces the feature space from 139 to 78 process-critical variables. Following Synthetic Minority Over-sampling Technique (SMOTE)-based class balancing, five classifiers were benchmarked via 10-fold cross-validation and optimized using the macro-averaged F3-score to mathematically penalize undetected defects. Light Gradient Boosting Machine (LightGBM) and Random Forest (RF) provided superior predictive baselines. To enforce strict zero-defect constraints, an asymmetric risk function shifted decision boundaries, enabling the risk-calibrated LightGBM model to reduce manual inspection volume by 9.72% with zero defect escapes. For resolving conflicting predictions, multi-algorithm decision fusion was implemented. By statistically evaluating the joint probabilities of the base models’ post-calibration outputs, a Naive Bayes Stacking meta-classifier effectively neutralizes single-algorithm inductive biases. Ultimately, synthesizing these F3-optimized, risk-calibrated base models via meta-learning successfully isolated true defect-free components, maximizing the final inspection time reduction to 14.82% while strictly maintaining zero defect escapes. Full article
Show Figures

Figure 1

27 pages, 2966 KB  
Article
FusionDroid: A Lightweight Multimodal Android Malware Detection Framework for Mobile Security
by Arockia Xavier Annie Rayan and Ajai Ram
Symmetry 2026, 18(9), 1501; https://doi.org/10.3390/sym18091501 - 8 Sep 2026
Viewed by 168
Abstract
The rapid adoption of Android applications in mobile commerce has increased exposure to sophisticated malware capable of bypassing traditional security mechanisms through code obfuscation, dynamic code loading, and runtime-triggered malicious behaviors. Although static analysis offers efficient large-scale detection, it often fails to identify [...] Read more.
The rapid adoption of Android applications in mobile commerce has increased exposure to sophisticated malware capable of bypassing traditional security mechanisms through code obfuscation, dynamic code loading, and runtime-triggered malicious behaviors. Although static analysis offers efficient large-scale detection, it often fails to identify concealed runtime activities, while dynamic analysis provides richer behavioral evidence but suffers from limited execution coverage and high computational overhead. To address these complementary limitations, this paper proposes FusionDroid, a lightweight multimodal Android malware detection framework that integrates permission-based static features, permission co-occurrence graph representations, and engineered runtime behavioral features through probability-level ensemble fusion. The framework was evaluated using Android applications collected from the AndroZoo repository, comprising 24,055 valid applications for static analysis and a balanced paired benchmark of 1716 applications for multimodal evaluation. The experimental results show that complementary static and dynamic representations can improve Android malware detection, although the magnitude and nature of the improvement depend on class distribution and evaluation metric. The best-performing model, StackedFusion-LightGBM, achieved 93.31% accuracy, 96.86% precision, 89.53% recall, a 93.05% F1 score, 97.69% ROC–AUC, and 98.14% PR–AUC. A controlled five-fold evaluation on the common paired benchmark further showed an F1 score of 0.9175±0.0176 for the full Static+Graph+Dynamic configuration compared with 0.8781±0.0081 for the Static-only baseline. Paired statistical analysis further supported the improvement. These findings show that multimodal fusion can improve Android malware detection while preserving low-complexity manifest-derived representations. The results support FusionDroid as a staged, sandbox-assisted framework in which lightweight static analysis is complemented by runtime behavioral evidence when deeper inspection is required. Full article
(This article belongs to the Section A: Computer Science)
Show Figures

Figure 1

24 pages, 50905 KB  
Article
Anchor-Constrained Residual Stacking for Missing Data Reconstruction in Structural Health Monitoring
by Du Guo, Chunfeng Wan, Miaomiao Peng, Yixu Wang, Caiqian Yang, Changqing Miao and Songtao Xue
Buildings 2026, 16(17), 3563; https://doi.org/10.3390/buildings16173563 - 7 Sep 2026
Viewed by 188
Abstract
During long-term monitoring, data missing often happens due to environmental interference, sensor malfunction or transmission problems, which will threaten the effectiveness of structural health monitoring. This study proposes an anchor-constrained residual stacking method, which can improve data reconstruction performance relative to individual models [...] Read more.
During long-term monitoring, data missing often happens due to environmental interference, sensor malfunction or transmission problems, which will threaten the effectiveness of structural health monitoring. This study proposes an anchor-constrained residual stacking method, which can improve data reconstruction performance relative to individual models across diverse missing patterns while reducing the risk of performance degradation associated with unconstrained ensemble learning. The method adopts a two-level stacking framework in which heterogeneous base learners generate candidate reconstructed data and grouped out-of-fold predictions are used to select an anchor learner for different missing patterns. A residual meta-learner then learns complementary residual information relative to the anchor learner, while validation-gated fusion regulates residual correction and final fusion based on reserved validation data, reducing unreliable fusion contributions. The effectiveness of the proposed method is verified using monitoring data with different missing patterns from a steel stringer bridge under multiple structural states. It achieves a coefficient of determination (R2) of 0.9020, together with the lowest relative root mean square error (RRMSE) and mean absolute error (MAE) among the compared methods. Operational modal analysis further confirms that the method preserves the main dynamic characteristics of structural responses, supporting reliable reconstruction across diverse missing patterns and multiple structural states. Full article
Show Figures

Figure 1

20 pages, 5284 KB  
Article
Tri-Band Vis–NIR Spectroscopy with Color Residual-Variance Gated Attention Fusion for Rapid Assessment of Hongmeiren (Citrus reticulata) Soluble Solids Content
by Anan Tao, Longfei Ye, Chaoxu Yu, Liuye Cao, Tiantian Pan and Fei Liu
Foods 2026, 15(17), 3166; https://doi.org/10.3390/foods15173166 - 7 Sep 2026
Viewed by 139
Abstract
Rapid assessment of internal fruit quality is essential for fruit grading, postharvest management, and consumer-oriented quality evaluation. Among the quality attributes, soluble solids content (SSC) is a key indicator of citrus sweetness and maturity. Visible and near-infrared (Vis–NIR) spectroscopy provides an effective approach [...] Read more.
Rapid assessment of internal fruit quality is essential for fruit grading, postharvest management, and consumer-oriented quality evaluation. Among the quality attributes, soluble solids content (SSC) is a key indicator of citrus sweetness and maturity. Visible and near-infrared (Vis–NIR) spectroscopy provides an effective approach for rapid SSC detection in fruit. However, most existing studies rely on single full-spectrum models or simple band stacking strategies, which limits their ability to fully exploit complementary information among different spectral sub-bands. To address this limitation, a color residual-variance gated attention fusion network (CR-VGAFNet) is proposed for efficient SSC assessment in Hongmeiren. On the independent prediction set, CR-VGAFNet achieved a prediction correlation coefficient (RP) of 0.7744, a root mean square error of prediction (RMSEP) of 0.6530 °Brix, and a mean absolute percentage error of prediction (MAPEP) of 4.83%. These findings suggest the potential of the framework for multi-band spectral fusion. This study provides a new technical perspective for multi-band spectral fusion and rapid fruit quality assessment. Full article
Show Figures

Figure 1

12 pages, 699 KB  
Article
EDBERT: Predicting Emergency Department Disposition Using a BERT-Based Architecture
by Md Ali Hossain, Krishan Chavinda, Mitchell D. Woodbright, Isuru Senadheera, Srikandabala Kogul, Sam Freeman, Mark Putland, Hamed Akhlaghi, Damminda Alahakoon and Md Anisur Rahman
Algorithms 2026, 19(9), 729; https://doi.org/10.3390/a19090729 - 30 Aug 2026
Viewed by 222
Abstract
Objective: To develop and evaluate EDBERT (Emergency Department Bidirectional Encoder Representations from Transformers), a BERT-based architecture built around a multi-input fusion network that enriches free-text triage notes with the additional context of presenting complaints and patient age to predict Emergency Department (ED) disposition, [...] Read more.
Objective: To develop and evaluate EDBERT (Emergency Department Bidirectional Encoder Representations from Transformers), a BERT-based architecture built around a multi-input fusion network that enriches free-text triage notes with the additional context of presenting complaints and patient age to predict Emergency Department (ED) disposition, supporting early decision-making and reduced ED length of stay (LOS). Methods: A retrospective cohort of 570,143 ED presentations to the Royal Melbourne Hospital, Australia, was used. At the core of EDBERT, a fusion network integrates three complementary input sources—triage notes, presenting complaints, and patient age—through learnable weights, with the two free-text inputs encoded by customised eight-layer BERT encoder stacks pre-trained on triage text with Masked Language Modelling (MLM). The resulting architecture is also lightweight, containing 66 million parameters, substantially fewer than the 110 million of BERT-BASE. Eighty percent of the dataset was used for training, and twenty percent for testing. Performance was benchmarked against BERT-Base configurations of varying depth. Results: EDBERT achieved an accuracy of 83.26%, a macro-averaged F1-score of 81.99%, and an area under the receiver operating characteristic curve of 0.91, outperforming all single-input BERT-Base variants, including a depth-matched eight-layer model (82.08%), indicating that the gain derives from the fusion of complementary inputs combined with domain-adaptive pre-training rather than from model capacity. Conclusions: A fusion-based BERT architecture that supplements the triage narrative with presenting complaint and age context predicts ED disposition more accurately than larger single-input BERT models, while requiring substantially less memory and computation, making it practical for deployment in hospitals with limited computing resources. Full article
(This article belongs to the Special Issue Algorithms in Data Classification (4th Edition))
Show Figures

Figure 1

31 pages, 41740 KB  
Article
A Stacking-Fusion Feature Selection Framework for Cross-Year and Cross-Cultivar Leaf Hyperspectral Rice Yield Estimation
by Yu Wang, Huaqi Ji, Shaozhong Song, Chunyan Qi and Xu Yang
Agriculture 2026, 16(17), 1847; https://doi.org/10.3390/agriculture16171847 - 27 Aug 2026
Viewed by 180
Abstract
Single-criterion feature selection for hyperspectral crop yield estimation suffers from methodological bias and limited generalisation. This study proposes a stacking-fusion feature selection framework with a ridge-regression meta-learner (STACKING_FUSION) that transfers the stacked-generalisation concept to the feature-evaluation space, integrating the Pearson correlation coefficient (PCC), [...] Read more.
Single-criterion feature selection for hyperspectral crop yield estimation suffers from methodological bias and limited generalisation. This study proposes a stacking-fusion feature selection framework with a ridge-regression meta-learner (STACKING_FUSION) that transfers the stacked-generalisation concept to the feature-evaluation space, integrating the Pearson correlation coefficient (PCC), grey relational analysis (GRA), and variable importance in projection (VIP) and using the out-of-fold R2 as a meta-supervision signal for adaptive weighting. In field experiments (2024–2025, Gongzhuling, Jilin Province) on rice cultivars Jijing 830 and Jijing 855, leaf hyperspectral reflectance (400–2400 nm) was acquired under controlled indoor measurement conditions at the tillering, jointing, flowering, and milking stages; the study was thus conducted at the leaf scale rather than at the canopy scale of UAV or satellite remote sensing. Second-derivative spectra outperformed original and first-derivative spectra at most stages, and STACKING_FUSION with XGBoost achieved the highest accuracy (R2 = 0.948, RMSE = 0.018 kg m−2, ratio of performance to deviation (RPD) = 4.399). Joint interpretation using SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME), computed on a held-out test subset, identified flowering-stage indices as the leading predictors within the model and suggested a flowering–tillering cross-stage association. With the configuration fixed from the 2024 development dataset, the 2025 analysis showed that the framework was reusable as a locked feature-engineering and algorithmic configuration rather than as a directly portable fitted predictor: under strict zero-shot application it retained high predicted–measured correlations (Pearson r≈ 0.91–0.93) but showed a consistent negative bias together with additional scale and residual error, whereas recalibration on a target-domain calibration subset (70% of the 2025 samples) achieved RPD > 3.0 in both the cross-year and cross-cultivar evaluations. These results indicate that the reusable component is the locked feature-engineering and algorithmic setting rather than the fitted coefficients. Full article
(This article belongs to the Section Artificial Intelligence and Digital Agriculture)
Show Figures

Figure 1

24 pages, 40484 KB  
Article
BC-GECO2: A Coarse and Fine Aggregate Segmentation and Counting Method for Hydraulic Concrete with Dense Depth Feature Fusion and Edge Enhancement
by Jiandong Wu, Baijing Wu, Jianwei Deng, Long Ma, Shuhong Liu and Shufan Zhang
Infrastructures 2026, 11(9), 297; https://doi.org/10.3390/infrastructures11090297 - 25 Aug 2026
Viewed by 186
Abstract
To reduce aggregate gradation counting errors caused by over-segmentation and under-segmentation of stacked and clustered aggregates with mixed types and diverse spatial distributions in hydraulic concrete, this study proposes BC-GECO2, a coarse and fine aggregate segmentation and counting method. Firstly, a BAHiera feature [...] Read more.
To reduce aggregate gradation counting errors caused by over-segmentation and under-segmentation of stacked and clustered aggregates with mixed types and diverse spatial distributions in hydraulic concrete, this study proposes BC-GECO2, a coarse and fine aggregate segmentation and counting method. Firstly, a BAHiera feature extraction network is designed to extract multi-scale deep features through edge-aware attention. In addition, a DFG-Edge module is developed to enhance the boundary features of densely distributed aggregates by integrating wavelet transform with a gated fusion mechanism, thereby alleviating the loss of small aggregate features during downsampling. Secondly, a CSFM-GFFCA module is constructed, in which a dual-branch structure is employed to adaptively fuse adjacent-scale features, strengthen the edge responses of densely distributed small aggregates, and enhance cross-layer feature interaction. Finally, a joint optimization function combining Focal loss and counting loss is established to guide the model toward hard-to-classify pixels, especially boundary pixels, thereby improving segmentation integrity and counting accuracy. Experiments conducted on an aggregate dataset collected from practical construction sites show that, compared with the baseline GECO2 model, the proposed method improves the average segmentation IoU, Dice, and BIoU by 2.92%, 5.04%, and 2.83%, respectively, while reducing the average counting MAE and RMSE by 6.92 and 15.65, respectively. Moreover, BC-GECO2 exhibits superior robustness and generalization capability under different stacking densities and blurred-boundary scenarios, providing technical support for the intelligent development of rapid concrete gradation detection. Full article
(This article belongs to the Section Infrastructures Materials and Constructions)
Show Figures

Figure 1

26 pages, 94016 KB  
Article
LMF-CP: An Interpretable Multimodal Late-Fusion Framework for Compound Carcinogenicity Prediction
by Yingjie Zhu, Liujie He and Xinjie Liang
Int. J. Mol. Sci. 2026, 27(16), 7499; https://doi.org/10.3390/ijms27167499 - 21 Aug 2026
Viewed by 368
Abstract
Accurately predicting the carcinogenicity of compounds is of great significance for drug discovery, clinical drug safety, and chemical risk assessment. Traditional methods for assessing carcinogenicity rely on animal testing, which suffers from limitations such as time-consuming processes, high costs, significant interspecies differences, and [...] Read more.
Accurately predicting the carcinogenicity of compounds is of great significance for drug discovery, clinical drug safety, and chemical risk assessment. Traditional methods for assessing carcinogenicity rely on animal testing, which suffers from limitations such as time-consuming processes, high costs, significant interspecies differences, and low predictive throughput. In recent years, computational modeling-based prediction methods (such as Quantitative Structure–Activity Relationships, QSAR) have made some progress, but they still face challenges such as insufficient molecular feature information and poor model interpretability. To overcome these barriers, the multimodal deep learning framework LMF-CP (Late Multimodal Fusion of Carcinogenicity Prediction) is proposed to enhance the performance and interpretability of compound carcinogenicity prediction. First, to comprehensively characterize the structural and physicochemical properties of compounds, a multimodal representation system based on four molecular modalities is constructed, namely SMILES sequences, molecular fingerprints, molecular images, and molecular graph structures. Specifically, Text Convolutional Neural Network (TextCNN), Multi-Layer Perceptron (MLP), Visual Geometry Group Network (VGGNet), as well as Molecular Graph Attention Network (MGAT) are employed to process this information, respectively. Second, to integrate information from different molecular representations, a late-stage fusion strategy based on Lasso stacking is employed. On the test set, LMF-CP achieves an area under curve (AUC) of 0.828, an accuracy (ACC) of 0.782, an F1 score of 0.786, a sensitivity (SEN) of 0.786, and a specificity (SPE) of 0.779. In addition, this paper combines Shapley Additive Explanations (SHAP) analysis with Bemis–Murcko scaffold analysis to interpret the model results from two perspectives. Finally, a visual online platform for predicting the carcinogenicity of compounds is designed, providing a convenient tool for the rapid assessment of compound carcinogenicity and structural interpretation. Full article
(This article belongs to the Special Issue Computational Strategies in Toxicology)
Show Figures

Figure 1

27 pages, 16065 KB  
Article
Spatial Domain Mismatch Between Field Plots and GEDI Inflates Aboveground Biomass Model Accuracy in a Sudanian Savanna Woodland
by Ahmed M. M. Hasoba and Kornél Czimber
Remote Sens. 2026, 18(16), 2751; https://doi.org/10.3390/rs18162751 - 14 Aug 2026
Viewed by 631
Abstract
Accurate estimation of aboveground biomass (AGB) in dryland savanna woodlands is constrained by sparse field data, which has motivated widespread fusion of field plots with spaceborne LiDAR reference data from the Global Ecosystem Dynamics Investigation (GEDI). Here, we show that such fusion can [...] Read more.
Accurate estimation of aboveground biomass (AGB) in dryland savanna woodlands is constrained by sparse field data, which has motivated widespread fusion of field plots with spaceborne LiDAR reference data from the Global Ecosystem Dynamics Investigation (GEDI). Here, we show that such fusion can substantially inflate apparent model accuracy when the two reference sources sample different spatial domains. Using 44 field plots from the Abu-Gadaf Natural Reserved Forest (AGNRF), Sudan, and 56 GEDI L4A footprints drawn from a 50 km buffer surrounding the reserve, we trained Random Forest (RF), Gradient Boosting (GB) and Classification and Regression Tree (CART) models on Sentinel-1, Sentinel-2, SRTM and Dynamic World predictors and evaluated them under 10-fold, 2 km block spatial cross-validation. The merged dataset yielded apparently moderate performance (RF: RMSE = 9.40 Mg ha−1, R2 = 0.33). However, GEDI-derived AGB was 2.1 times higher than field-measured AGB (18.71 vs. 8.89 Mg ha−1; Kolmogorov–Smirnov D = 0.53, p < 0.001), and decomposing performance by source revealed that predictive skill within the field plot population was effectively absent (R2 = 0.001–0.023). A classifier trained to discriminate data source from the predictor stack alone achieved 85% accuracy against a 56% baseline, quantile calibration removing the inter-source level difference reduced pooled R2 from 0.33 to 0.13, and restricting GEDI footprints to within 20 km of the reserve reduced R2 to 0.008. Apparent accuracy therefore derived largely from between-source separation rather than from structural prediction of AGB. We conclude that spatial cross-validation does not detect population heterogeneity arising from multi-source reference fusion, and that source-stratified validation is necessary. The AGB maps presented are interpreted as relative spatial patterns rather than validated absolute estimates. Full article
(This article belongs to the Section Forest Remote Sensing)
Show Figures

Figure 1

20 pages, 12232 KB  
Article
Fatigue State Assessment Based on Soft Voting Using Surface Electromyography Signals
by Fangcao Zhang, Kunpeng Chen, Fei Guo and Hao Yan
Appl. Sci. 2026, 16(16), 7946; https://doi.org/10.3390/app16167946 - 10 Aug 2026
Viewed by 246
Abstract
To improve the accuracy of fatigue assessment in patients undergoing upper limb rehabilitation, this study proposes a lightweight fatigue detection algorithm based on ensemble learning and single-channel surface electromyography (sEMG) signals. Thirty healthy subjects without upper limb injuries or severe chronic diseases were [...] Read more.
To improve the accuracy of fatigue assessment in patients undergoing upper limb rehabilitation, this study proposes a lightweight fatigue detection algorithm based on ensemble learning and single-channel surface electromyography (sEMG) signals. Thirty healthy subjects without upper limb injuries or severe chronic diseases were recruited, and dynamic sEMG signals of the biceps brachii were collected during dumbbell bicep curls at a sampling frequency of 2048 Hz, yielding 6650 valid experimental samples. To address the class imbalance in the sEMG dataset, the SMOTETomek hybrid sampling algorithm was employed for data balancing. Three ensemble strategies—voting, stacking, and mean fusion—were integrated and combined with four different sets of base classifiers to construct a total of 12 fusion models, and the optimal model was identified through comparative screening. The experimental results demonstrated that, after SMOTETomek sample balancing, the LightGBM-LR-MLP soft voting ensemble model achieved the best overall performance, with the accuracy, recall, precision, and F1-score for both fatigue categories all exceeding 0.93. Multiple statistical analyses, including paired t-tests, effect sizes, and 95% confidence intervals, verified the reliability of the model selection and confirmed that the soft voting algorithm significantly outperformed the other comparison schemes. Machine learning can efficiently interpret dynamic biceps brachii sEMG signals; the proposed method improves the accuracy of fatigue recognition and provides an objective, quantitative fatigue reference index for upper limb rehabilitation training. It holds promise for assisting the dynamic adjustment of rehabilitation training intensity, thereby potentially reducing the risk of overtraining injuries and enhancing the safety of rehabilitation training. Full article
Show Figures

Figure 1

28 pages, 29047 KB  
Article
Integrating Multi-Season Sentinel-1/2 and Topographic Features to Improve Tree Species Diversity Estimation Accuracy
by Wendou Liu, Shaozhi Chen, Tianbao Huang, Ram P. Sharma, Dongyang Han, Jiang Liu, Pengfei Zheng and Xin Huang
Remote Sens. 2026, 18(16), 2651; https://doi.org/10.3390/rs18162651 - 7 Aug 2026
Viewed by 453
Abstract
Accurate estimation of forest tree species diversity at regional scales is essential for biodiversity monitoring, forest resource management, and ecological conservation. Because tree species differ in canopy spectral responses and phenological dynamics, multi-season remote sensing observations can provide critical information for characterizing species [...] Read more.
Accurate estimation of forest tree species diversity at regional scales is essential for biodiversity monitoring, forest resource management, and ecological conservation. Because tree species differ in canopy spectral responses and phenological dynamics, multi-season remote sensing observations can provide critical information for characterizing species composition and diversity patterns. However, the potential contribution of seasonal image features to improving remote-sensing-based tree species diversity estimation has often been insufficiently considered. In this study, the Yichun forest region in Heilongjiang Province, northeastern China, was selected as the study area. Sentinel-1, Sentinel-2, and topographic data were integrated to extract multi-seasonal spectral, vegetation index, texture, radar, and topographic features. The Boruta algorithm was used for feature selection, and random forest (RF), extreme gradient boosting (XGBoost), k-nearest neighbor (KNN), support vector regression (SVR), Bayesian regularized neural network (BRNN), and Stacking ensemble learning were developed to estimate and map Richness, Shannon, and Gini–Simpson indices. The results showed that: (1) Sentinel-2 optical features were the primary information source for tree species diversity estimation, topographic factors further improved model performance, and Sentinel-1 radar features mainly provided complementary structural information; (2) seasonal remote sensing features differed in their predictive ability, with Richness performing better in spring, while Shannon and Gini–Simpson achieved higher accuracy in winter. The four-season fusion scenario produced the highest accuracy for all three indices, with optimal R2 values of 0.51, 0.63, and 0.57, respectively; (3) the Stacking ensemble generally improved estimation accuracy and model stability, although the optimal model differed among diversity indices, with Stacking, SVR, and RF performing best for Richness, Shannon, and Gini–Simpson, respectively; and (4) summer Sentinel-2 NDVI, GNDVI, and NDWI contributed strongly to all three indices, elevation was particularly important for Richness, and winter vegetation indices and autumn red-edge bands and texture features were also informative for Shannon and Gini–Simpson. These findings indicate that integrating multi-seasonal remote sensing features and multi-source data using machine learning models can effectively improve forest tree species diversity estimation, providing technical support for regional forest biodiversity monitoring and precision forest management. Full article
Show Figures

Figure 1

22 pages, 1216 KB  
Article
Ensemble-Based Approach for Amazigh POS Tagging: Leveraging Multiple Models for Enhanced Performance in Low-Resource Language Processing
by Abdelouahed Moussaoui, Nor-Eddine Azalmad, Said Bahassine and Khalid Housni
Information 2026, 17(8), 753; https://doi.org/10.3390/info17080753 - 5 Aug 2026
Viewed by 834
Abstract
Part-of-Speech (POS) tagging is a foundational task in Natural Language Processing (NLP), yet it remains challenging for low-resource and morphologically rich languages such as Amazigh. This paper proposes a hybrid ensemble framework for Amazigh POS tagging that integrates three complementary models: a Bidirectional [...] Read more.
Part-of-Speech (POS) tagging is a foundational task in Natural Language Processing (NLP), yet it remains challenging for low-resource and morphologically rich languages such as Amazigh. This paper proposes a hybrid ensemble framework for Amazigh POS tagging that integrates three complementary models: a Bidirectional Long Short-Term Memory network (BiLSTM), a Conditional Random Field model (CRF), and a rule-based morphological analyzer (RBMA). Rather than treating prior results obtained on different corpora and tag inventories as directly comparable, the study evaluates all proposed components under a common 54-tag experimental setting based on the publicly available Amazigh Linguistic Dataset. Three ensemble strategies are examined: majority voting, validation-weighted voting, and logistic-regression stacking. An additional late-fusion ablation applies hard and soft RBMA constraints to CRF and Stacking outputs; hard masking degrades performance substantially, whereas soft masking is more robust but remains below unconstrained decoding. The best micro-level performance is obtained by the stacking ensemble, which reaches 98.51% Micro-F1/accuracy, whereas the boosting-like weighted ensemble obtains the strongest Macro-F1 among the ensemble variants, reaching 74.24%. These results show that hybrid ensemble methods can improve token-level accuracy in low-resource POS tagging, while also revealing a trade-off between frequent-tag accuracy and rare-tag robustness. The findings highlight the usefulness of combining neural, probabilistic, and rule-based information for Amazigh POS tagging, and point to class-balanced meta-learning and character/subword representations as important directions for improving rare and out-of-vocabulary categories. Full article
Show Figures

Graphical abstract

32 pages, 1860 KB  
Article
3MuViS, 3-Phase Multi-View Stacking: A Model Selection Algorithm for Multi-View Fusion
by Guillermo Villegas-Morales, Enrique Garcia-Ceja, Salvador Hinojosa, Jesús Arturo Pérez-Díaz and Mahdi Zareei
Mach. Learn. Knowl. Extr. 2026, 8(8), 227; https://doi.org/10.3390/make8080227 - 3 Aug 2026
Viewed by 447
Abstract
Multi-view learning has emerged as an effective paradigm for integrating heterogeneous data representations in complex classification problems, yet selecting suitable learners for Multi-View Stacking architectures remains computationally challenging and highly dependent on expert decisions. This work proposes 3MuViS, a three-phase methodology for automatic [...] Read more.
Multi-view learning has emerged as an effective paradigm for integrating heterogeneous data representations in complex classification problems, yet selecting suitable learners for Multi-View Stacking architectures remains computationally challenging and highly dependent on expert decisions. This work proposes 3MuViS, a three-phase methodology for automatic learner selection in Multi-View Stacking that optimizes both view-level learners and the meta-learner according to a target classification metric. The method evaluates candidate machine learning algorithms through cross-validation and constructs optimized stacking configurations. Experiments were conducted on six heterogeneous datasets spanning network intrusion detection, human activity recognition, handwritten digit classification, and transportation mode detection. Performance was evaluated over 25 iterations using the Matthews Correlation Coefficient (MCC), and statistical significance was assessed using Wilcoxon tests. The results show that 3MuViS consistently outperformed stochastic selection strategies and achieved superior or comparable performance to fixed models in five of the six evaluated datasets while frequently approaching the Brute Force upper-bound baseline at a substantially lower computational cost. The findings indicate that jointly optimizing view-level and meta-level learners improves both predictive performance and stability, demonstrating the potential of 3MuViS as a general and efficient framework for multi-view classification problems across diverse domains. Full article
Show Figures

Graphical abstract

29 pages, 13085 KB  
Article
PSA-DMCF-SVAE: A Lithium Battery SOH Prediction Framework with ProbSparse Self-Attention and Deep Multidimensional Features
by Weihao Sun, Gang Liu, Jiawei Chen, Yiyao Zhao, Yuting Cheng, Gang Xiao and Durga Prasad Bavirisetti
Energy Storage Appl. 2026, 3(3), 11; https://doi.org/10.3390/esa3030011 - 1 Aug 2026
Viewed by 293
Abstract
Accurate state of health (SOH) evaluation is essential for lithium-ion battery safety. Conventional prediction models suffer high computation overhead and insufficient multidimensional feature fusion, leading to poor generalization across working conditions. This work develops the PSA-DMCF-SVAE framework. Ten health factors are extracted from [...] Read more.
Accurate state of health (SOH) evaluation is essential for lithium-ion battery safety. Conventional prediction models suffer high computation overhead and insufficient multidimensional feature fusion, leading to poor generalization across working conditions. This work develops the PSA-DMCF-SVAE framework. Ten health factors are extracted from cycling data, and four high-correlation indicators are selected via Pearson screening. The method combines stacked ensemble learning and sparse variational autoencoder (SVAE) for multi-scale aging feature extraction, with an embedded ProbSparse self-attention module to adaptively weight base learners and cut computational complexity. Experiments on NASA battery data reveal over 40% lower prediction error and 42% less feature redundancy than classic stacking architectures. Cross-domain validation on an MIT fast-charging dataset confirms the model’s zero-shot generalization under high-rate, noisy discharge conditions, offering a practical solution for battery management systems. Full article
Show Figures

Figure 1

23 pages, 4133 KB  
Article
Integrating Local and Global Representation Learning for Pediatric Pneumonia Detection: A Hybrid CNN–Transformer Ensemble Framework
by Ece Meltem Yalçın, Hayriye Tanyıldız, Serpil Aslan, Mustafa Yıldız, Damla Ağaçkıran and Gül Fidan
Diagnostics 2026, 16(15), 2399; https://doi.org/10.3390/diagnostics16152399 - 30 Jul 2026
Viewed by 371
Abstract
Background/Objectives: Pneumonia remains a leading cause of childhood morbidity and mortality worldwide. Accurate interpretation of pediatric chest radiographs is challenging because of anatomical variability, subtle radiographic findings, and inter-observer variability. This study evaluates different CNN–Transformer ensemble strategies for pediatric pneumonia detection by [...] Read more.
Background/Objectives: Pneumonia remains a leading cause of childhood morbidity and mortality worldwide. Accurate interpretation of pediatric chest radiographs is challenging because of anatomical variability, subtle radiographic findings, and inter-observer variability. This study evaluates different CNN–Transformer ensemble strategies for pediatric pneumonia detection by combining complementary local and global image representations. Methods: Experiments were conducted on the publicly available Pediatric Pneumonia Chest X-ray dataset containing 5856 radiographs. EfficientNetV2-S was used to extract local features, whereas Swin Transformer-T modeled global anatomical relationships. Soft voting, weighted voting, and stacking were evaluated under a unified training protocol. Image preprocessing, data augmentation, and Weighted Random Sampling were applied to improve robustness and address class imbalance. Performance was assessed using an independent hold-out test set and five-fold cross-validation. Grad-CAM was used to interpret model predictions. Results: Ensemble learning improved classification performance compared with individual models while revealing different trade-offs among fusion strategies. The Soft Ensemble achieved the highest hold-out accuracy (96.96%) and F1-score (97.59%). The Hybrid CNN–Transformer Stacking model achieved the highest sensitivity (99.49%) and produced the fewest false-negative predictions (n = 2), while demonstrating the most consistent performance across five-fold cross-validation. Grad-CAM visualizations indicated that the CNN and Transformer models captured complementary radiographic information. Conclusions: The proposed framework demonstrates that different ensemble strategies offer distinct advantages. Soft voting provided the best overall hold-out performance, whereas stacking minimized false-negative predictions and achieved the highest sensitivity, indicating its potential for AI-assisted pediatric pneumonia screening. Full article
(This article belongs to the Special Issue Artificial Intelligence for Health and Medicine—2nd Edition)
Show Figures

Figure 1

Back to TopTop