Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (1,160)

Search Parameters:
Keywords = statistical regression algorithm

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
34 pages, 1671 KB  
Article
Toward More Resilient Agricultural Enterprises: Integrating Agricultural Environment Information into Explainable Financial Distress Prediction
by Dominika Gajdosikova
Agriculture 2026, 16(17), 1920; https://doi.org/10.3390/agriculture16171920 - 4 Sep 2026
Viewed by 181
Abstract
Financial vulnerability in agriculture is shaped by both enterprise-specific factors and the agricultural systems in which firms operate, challenging financial distress prediction models based exclusively on firm-level financial metrics. This study assesses the predictive performance of logistic regression (LR), extreme gradient boosting (XGBoost), [...] Read more.
Financial vulnerability in agriculture is shaped by both enterprise-specific factors and the agricultural systems in which firms operate, challenging financial distress prediction models based exclusively on firm-level financial metrics. This study assesses the predictive performance of logistic regression (LR), extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM), and category boosting (CatBoost), examines whether country-level agricultural environment characteristics improve prediction beyond traditional financial indicators, and uses Shapley additive explanations (SHAP) to analyze predictor contributions. The empirical analysis is based on 28,745 firm-year observations of agricultural enterprises in the Visegrad Group (V4) countries over 2019–2024. Adding agricultural environment variables produced a modest but consistent improvement in predictive performance, increasing AUC from approximately 0.854–0.855 in firm-level boosting models to 0.864–0.865 in integrated specifications. All three gradient boosting algorithms outperform conventional LR, although differences among XGBoost, LightGBM, and CatBoost remain statistically insignificant. SHAP analysis identifies liquidity, leverage, profitability, and firm size as the primary predictors while revealing nonlinear and threshold-dependent relationships not captured by conventional linear models. Integrating agricultural environment characteristics with explainable machine learning (ML) therefore supports more comprehensive, transparent, and context-aware early-warning systems for agricultural financial distress. Full article
Show Figures

Figure 1

23 pages, 11839 KB  
Article
A Terrain-Corrected Vegetation Index Strategy for Improving Leaf Area Index Estimation in Mountainous Areas
by Haier Liu, Guyue Hu, Chenghao Liu, Yakun Han, Siqi Li, Ronghao Yang, Junxiang Tan and Shaoda Li
Remote Sens. 2026, 18(17), 2977; https://doi.org/10.3390/rs18172977 - 2 Sep 2026
Viewed by 131
Abstract
Leaf Area Index (LAI) is an important biophysical parameter in studies of regional and global ecosystems. However, terrain-induced distortion of surface reflectance can reduce the reliability of vegetation indices (VIs) in characterizing the canopy structure, thereby introducing uncertainties in LAI retrieval. In this [...] Read more.
Leaf Area Index (LAI) is an important biophysical parameter in studies of regional and global ecosystems. However, terrain-induced distortion of surface reflectance can reduce the reliability of vegetation indices (VIs) in characterizing the canopy structure, thereby introducing uncertainties in LAI retrieval. In this study, an LAI retrieval method for mountainous areas based on the combination of terrain-corrected VIs and the random forest algorithm was proposed. Typical topographic correction models (Cosine+C, SCS+C, and Statistical–Empirical) were applied to normalize surface reflectance, and the terrain-corrected normalized difference vegetation index (NDVI) and modified soil-adjusted vegetation index (MSAVI) were constructed accordingly. Then, random forest regression was used for LAI retrieval, and the proposed method was validated through comparisons of this LAI with the field observations and original VI-based methods. The results showed that topographic correction effectively reduced the radiometric distortions induced by topography, and the NDVISCSC-based retrieval method performed well under various terrain conditions (with R2 and RMSE of 0.927 and 0.151, respectively). In addition, to investigate the effects of different terrain factors and illumination conditions on LAI retrieval, methods based on the original and terrain-corrected VIs were compared for surfaces with different slopes and aspects. The results revealed that the terrain-corrected VIs can improve the performance of LAI retrieval in areas with terrain-induced reflectance distortion. Finally, the optimal method successfully estimated LAI in the study area. Therefore, the proposed LAI retrieval method for mountainous areas is an effective tool for extracting surface biophysical parameters, and can provide a reliable approach for regional ecological monitoring and evaluation. Full article
(This article belongs to the Section Forest Remote Sensing)
Show Figures

Figure 1

23 pages, 3824 KB  
Article
Comparative Performance of Classical Statistical and Machine Learning Models for Melanoma Classification
by Diego Pagnoncelli, Gian Luca Viganò and Veronica Cimolin
Electronics 2026, 15(17), 3944; https://doi.org/10.3390/electronics15173944 - 2 Sep 2026
Viewed by 103
Abstract
Early and accurate classification of melanoma is essential for improving patient outcomes and supporting clinical decision-making. Although numerous predictive models have been proposed, comparisons between classical statistical approaches and modern machine learning algorithms are often limited by heterogeneous analytical workflows and inconsistent validation [...] Read more.
Early and accurate classification of melanoma is essential for improving patient outcomes and supporting clinical decision-making. Although numerous predictive models have been proposed, comparisons between classical statistical approaches and modern machine learning algorithms are often limited by heterogeneous analytical workflows and inconsistent validation strategies. This study aimed to compare the predictive performance of classical statistical and machine learning models for melanoma classification using a fully reproducible analytical framework. A retrospective observational study was conducted using the publicly available BCN20000 dermoscopic dataset from the ISIC Archive. After standardized data preprocessing, four routinely available clinical variables (age, sex, anatomical site and melanocytic status) were used to develop Logistic Regression, Generalized Additive Models, Random Forest and Extreme Gradient Boosting (XGBoost) classifiers. All models were trained and evaluated using the same stratified training/testing split, and their performance was assessed through discrimination, calibration and SHAP explainability analysis. Machine learning models, particularly XGBoost and Random Forest, achieved superior predictive performance compared with conventional statistical approaches, while patient age emerged as the most influential predictor of malignancy. The proposed framework provides a transparent and reproducible approach for objectively comparing predictive models and supports the development of accurate, interpretable, and reproducible clinical decision-support systems for melanoma classification. Full article
Show Figures

Figure 1

22 pages, 2194 KB  
Article
Airline Pricing Dynamics on Thin Routes from a Central European Regional Airport: A 30-Day Case Study
by Peter Hanák, Luboš Socha and Edina Jenčová
Sustainability 2026, 18(17), 8943; https://doi.org/10.3390/su18178943 - 1 Sep 2026
Viewed by 177
Abstract
Regional airports in post-EU accession Central Europe operate under demand conditions that differ substantially from those of major hubs, yet revenue management in these thin markets remains underexplored. This study examines whether temporal fare escalation observed on high-frequency routes also occurs for low-frequency [...] Read more.
Regional airports in post-EU accession Central Europe operate under demand conditions that differ substantially from those of major hubs, yet revenue management in these thin markets remains underexplored. This study examines whether temporal fare escalation observed on high-frequency routes also occurs for low-frequency regional services and whether carriers apply different pricing strategies on overlapping routes. Daily one-way fares were tracked on three routes from Košice International Airport, Slovakia, to the United Kingdom (Dublin, London Luton, and London Stansted) between 4 January and 2 February 2025, yielding 1738 observations across direct and connecting itineraries via Liverpool. Statistical analyses included paired t-tests, OLS regression with Newey–West standard errors and collection-session-clustered robustness checks, bootstrap inference for coefficient-of-variation ratios (10,000 resamples), Holm–Bonferroni correction, and log-transformation robustness checks. Four of six route–flight type combinations exhibited significant temporal fare escalation after multiple-testing correction. Wizz Air’s Luton service showed substantially greater fare variability than Ryanair’s Stansted service (coefficient of variation of 77.5% vs. 31.3%; bootstrap 95% CI for the ratio [1.29, 10.32]), despite similar mean fares. Connecting fares via Liverpool averaged 171% above direct Dublin fares, while daily prices on the London routes were strongly correlated (r = 0.816). These results reveal carrier-specific algorithmic pricing differences in thin regional markets with implications for consumer decision-making and for the affordability dimension of regional connectivity. Full article
Show Figures

Figure 1

30 pages, 54034 KB  
Article
GWBASE: An Algorithm for Screening Groundwater–Baseflow Coupling Using Paired USGS Well and Streamflow Records
by Xueyi Li, Norman L. Jones, Gustavious P. Williams, Amin Aghababaei, Riley C. Hales, Eniola Webster-Esho, Ryan van der Heijden, T. Prabhakar Clement and Donna M. Rizzo
Hydrology 2026, 13(9), 233; https://doi.org/10.3390/hydrology13090233 - 30 Aug 2026
Viewed by 273
Abstract
Groundwater discharge contributes to stream baseflow, but coupling strength varies among catchments and transferable quantification methods are limited. We present GWBASE, an open-source Python algorithm that pairs U.S. Geological Survey (USGS) wells and gages within hydrographic catchments and ranks coupling at each gage [...] Read more.
Groundwater discharge contributes to stream baseflow, but coupling strength varies among catchments and transferable quantification methods are limited. We present GWBASE, an open-source Python algorithm that pairs U.S. Geological Survey (USGS) wells and gages within hydrographic catchments and ranks coupling at each gage by linear regression and mutual information (MI, capturing nonlinear and lagged dependence) on monthly ΔWTE–ΔQ records from baseflow-dominated months. We apply it to the Great Salt Lake Basin (Utah; ∼93,000km2 with 8752 USGS wells, 1906–2025). GWBASE ranked four terminal-gage catchments using a seasonally corrected, within-well regression as the primary estimate; three of the four show statistically significant groundwater–baseflow coupling. The ranking depends on the metric: absolute magnitude (cfs per foot) is dominated by the large Bear River catchment, whereas size-normalized sensitivity and coupling tightness identify the smaller Little Cottonwood Creek. The basin-scale aggregate (∼4 cfs per foot of basin-averaged decline) is dominated by Bear River and, once catchment-level uncertainty is propagated, is not distinguishable from zero. Because national groundwater records are predominantly intermittent, GWBASE resolves seasonal-to-interannual storage coupling rather than event-scale exchange. It is best used for screening and ranking catchment-scale coupling rather than yielding a single basin-scale coefficient. Full article
Show Figures

Figure 1

21 pages, 5387 KB  
Article
Double-Diode Modeling and Simulation of PV Cell Performance: Statistical Analysis and Machine-Learning Validation
by Nowrin Jannat, Saleha Nasrin Mishu, Prithwiraj Biswas Pallab, Md. Atik Hasan Nishat, Md. Firoz Ahmed and M. Hasnat Kabir
Lights 2026, 2(3), 7; https://doi.org/10.3390/lights2030007 - 29 Aug 2026
Viewed by 336
Abstract
Accurate modeling of photovoltaic (PV) cell behavior under varying operational conditions is essential for optimizing energy yield and system reliability. This study presents an extended simulation-based methodology for analyzing monocrystalline silicon PV cells using a double-diode model (DDM) with a physics-based, temperature- and [...] Read more.
Accurate modeling of photovoltaic (PV) cell behavior under varying operational conditions is essential for optimizing energy yield and system reliability. This study presents an extended simulation-based methodology for analyzing monocrystalline silicon PV cells using a double-diode model (DDM) with a physics-based, temperature- and irradiance-dependent parameterization. Building on a SPICE-equivalent circuit formulation, the governing implicit DDM equation is solved numerically to regenerate every current–voltage (I–V) and power–voltage (P–V) curve, and all circuit, block and flow diagrams are redrawn as vector-quality figures. Beyond the deterministic analysis, the manuscript introduces two extensions: (i) a quantitative statistical analysis of the influence of temperature (T), irradiance (G) and series resistance (Rs) on open-circuit voltage, short-circuit current, maximum power and fill factor, using linear/log-linear regression, a multiple linear regression model and a Pearson correlation analysis; and (ii) a machine-learning (ML) validation study in which a random-forest surrogate model is trained on a 600-point physics-consistent synthetic dataset spanning the full (T, G, Rs) operating envelope and evaluated with a held-out test split and 5-fold cross-validation. The surrogate reproduces the DDM outputs with cross-validated coefficients of determination above 0.98 for maximum power, open-circuit voltage, short-circuit current and fill factor, confirming that the DDM response surface is smooth, learnable and suitable for fast surrogate-based design optimization and maximum-power-point-tracking (MPPT) algorithm testing. Simulated outputs at standard test conditions (25 °C, 1000 W/m2, AM 1.5) are compared against manufacturer datasheet values, and residual errors are analyzed and attributed to specific modeling assumptions. Full article
Show Figures

Figure 1

30 pages, 5276 KB  
Article
Novel Heterogeneous Dynamic Fusion Model Based on a Data-Mechanism Dual-Driven Framework for a State-of-Health Prediction of Lithium Batteries in Autonomous Underwater Vehicle Applications
by Yongxun Liu, Zijun Wang, Yibo Shen, Feng Zhao and Bin Wang
World Electr. Veh. J. 2026, 17(9), 455; https://doi.org/10.3390/wevj17090455 - 28 Aug 2026
Viewed by 209
Abstract
Accurate state-of-health (SOH) prediction of lithium batteries is critical for guaranteeing the long endurance and system safety of autonomous underwater vehicles (AUVs) in marine operations. However, owing to complex underwater-operation conditions, data acquisition in AUVs is typically restricted to rest stages in the [...] Read more.
Accurate state-of-health (SOH) prediction of lithium batteries is critical for guaranteeing the long endurance and system safety of autonomous underwater vehicles (AUVs) in marine operations. However, owing to complex underwater-operation conditions, data acquisition in AUVs is typically restricted to rest stages in the communication period, which would be characterized by incomplete data with high-frequency sensor noise. As a result, existing battery SOH prediction approaches would struggle to ensure estimation accuracy and model robustness for within-cell degradation trajectories in AUV applications. This paper proposes a novel heterogeneous dynamic fusion model based on a data-mechanism dual-driven (DMDD) framework for the SOH prediction of lithium batteries in AUV applications, innovatively utilizing features extracted from the rest stage after discharge. At first, a dual-filter strategy based on the interquartile range interception and the Savitzky–Golay algorithms is designed to effectively eliminate transient spikes and high-frequency artifacts of raw data. Furthermore, a two-stage feature-screening architecture is developed, which can not only filter out statistical redundancies but also elucidate the electrochemical mechanisms between extracted features and battery degradation. Moreover, a heterogeneous fusion model comprising random forest, support vector regression, and gated recurrent unit networks is constructed. On this basis, an adaptive dynamic fusion strategy based on the K-nearest neighbor and the minimum-variance unbiased estimation (MVUE) is proposed, which enables locally optimal credit assignments tailored to the specific characteristics of different aging stages. Experimental validations comprehensively demonstrate the superior performance of the heterogeneous dynamic fusion model based on the DMDD framework across the entire battery lifecycle. Specifically, the proposed heterogeneous dynamic fusion model can achieve a coefficient of determination (R2) over 0.99907, while the MAE and the RMSE can be maintained within 0.33958% and 0.56211%, respectively, showing satisfactory accuracy for battery SOH prediction in AUV applications. Full article
(This article belongs to the Section Storage Systems)
Show Figures

Figure 1

25 pages, 2017 KB  
Article
An Explainable Machine Learning Framework for Adaptive Multi-Mode CORDIC Iteration Optimization and Hardware-Efficient Computation
by Ratheesh Sudheerbabu, Lekshmi Chandrika Reghunath, Cristian Randieri, Brunella Botte and Alfredo Milani
Mathematics 2026, 14(17), 3096; https://doi.org/10.3390/math14173096 - 28 Aug 2026
Viewed by 177
Abstract
The Coordinate Rotation Digital Computer (CORDIC) algorithm is widely employed in digital signal processing and hardware accelerators because it computes a broad range of elementary functions using iterative shift-and-add operations. Conventional CORDIC implementations, however, execute a fixed number of iterations irrespective of the [...] Read more.
The Coordinate Rotation Digital Computer (CORDIC) algorithm is widely employed in digital signal processing and hardware accelerators because it computes a broad range of elementary functions using iterative shift-and-add operations. Conventional CORDIC implementations, however, execute a fixed number of iterations irrespective of the input characteristics or the precision required, resulting in unnecessary computational overhead and increased execution latency. This work presents an explainable machine learning framework for adaptive iteration optimization in a multi-mode CORDIC architecture supporting circular, hyperbolic, and linear operating modes. A unified prediction framework for calculating the optimal number of iterations is made possible by the developing a generic feature representation to describe the numerical behavior of CORDIC computations across various modes. We systematically evaluated eight regression models, including Linear Regression, Decision Tree, Random Forest, Extra Trees, Support Vector Regression, Multi-Layer Perceptron, and Extreme Gradient Boosting (XGBoost) and LightGBM. Among the models evaluated, the Decision Tree achieved the best performance on an independent test set of 2305 samples from 461 previously unseen input groups, with a MAE of 0.9160 iterations, RMSE of 1.9671, and R2 of 0.6076. Predictions were within one and two iterations of the reference value for 80.26% and 90.07% of the test samples, respectively. Since prediction accuracy alone does not guarantee that the required numerical tolerance will be satisfied, the predicted iteration count was further evaluated using the actual CORDIC error, followed by a safety-correction procedure. The safety-corrected approach achieved 100% tolerance satisfaction on the independent test set, reducing the mean number of iterations from 20 to 11.739, corresponding to a 41.31% reduction in iterations. Model behavior was further interpreted using feature importance analysis, permutation importance, and feature ablation studies to examine the contribution of individual features to iteration prediction. Statistical robustness is established using bootstrap confidence intervals, the Friedman test, and Holm-corrected Wilcoxon signed-rank tests. Full article
Show Figures

Figure 1

18 pages, 2488 KB  
Article
Data-Driven Workforce Optimization in Weaving Manufacturing Systems: An Integrated Machine Learning and Queueing Framework
by Bilge Berkhan Kastacı
Processes 2026, 14(17), 2745; https://doi.org/10.3390/pr14172745 - 27 Aug 2026
Viewed by 256
Abstract
This study proposes a predictive–prescriptive framework for workforce planning and queueing-based capacity assessment in weaving manufacturing systems by integrating statistical analysis, machine learning, and queueing theory. The dataset comprised 12 months of hall-level observations combining four operational indicators recorded by the LoomData system—warp [...] Read more.
This study proposes a predictive–prescriptive framework for workforce planning and queueing-based capacity assessment in weaving manufacturing systems by integrating statistical analysis, machine learning, and queueing theory. The dataset comprised 12 months of hall-level observations combining four operational indicators recorded by the LoomData system—warp and weft breakage frequencies and their corresponding repair durations—with workforce requirements obtained from the company’s operational planning records. Statistical analyses identified operational differences among production halls and time periods. Workforce demand was modeled using multiple linear regression and tree-based machine learning algorithms, including M5P, REPTree, Random Forest, and Random Tree. Model performance was evaluated through cross-validation and independent testing, where multiple linear regression demonstrated the highest predictive accuracy and stability (R2 > 0.93). The predicted workforce estimates were subsequently converted into integer staffing levels and incorporated into an M/M/c queueing model to analytically assess system utilization, expected waiting times, expected queue lengths, and capacity pressure. The results revealed hall-specific differences in queueing performance and capacity utilization. These findings showed that predictive model performance should be complemented by a queueing-based capacity assessment to support workforce planning. By integrating predictive modeling with analytical queueing assessment, the proposed framework bridges machine learning and operations research while providing practical decision support for data-driven workforce planning in manufacturing environments. Full article
Show Figures

Figure 1

19 pages, 2529 KB  
Article
An Explainable Machine Learning Framework for Predicting Hearing Aid Satisfaction: Integrating the HATASS Instrument and Clinical Insights
by Seyma Arslanbas, Tahir Cetin Akinci, Ümit Can Çetinkaya and Sengul Terlemez
Bioengineering 2026, 13(9), 985; https://doi.org/10.3390/bioengineering13090985 - 26 Aug 2026
Viewed by 212
Abstract
Hearing aid technology adaptation and user satisfaction are influenced by multiple interacting demographic, clinical, and behavioral factors, making reliable prediction of outcomes challenging with conventional statistical approaches alone. This study proposes an explainable machine learning framework to investigate the multidimensional determinants of hearing [...] Read more.
Hearing aid technology adaptation and user satisfaction are influenced by multiple interacting demographic, clinical, and behavioral factors, making reliable prediction of outcomes challenging with conventional statistical approaches alone. This study proposes an explainable machine learning framework to investigate the multidimensional determinants of hearing aid satisfaction by integrating demographic characteristics, hearing aid-related variables, and patient-reported outcomes obtained from the Hearing Aid Technology Adaptation and Satisfaction Scale (HATASS). Five regression algorithms—Linear Regression, Decision Tree Regression (DTR), Random Forest Regression (RFR), Support Vector Regression, and Gradient Boosting Regression (GBR)—were comparatively evaluated using a five-fold cross-validation strategy. Predictive performance was assessed using the root mean square error (RMSE), mean absolute error (MAE), and the coefficient of determination (R2), while model interpretability was investigated through cross-validated out-of-bag permutation feature importance analysis. Among the evaluated algorithms, Random Forest Regression achieved the most consistent predictive performance, yielding the lowest average RMSE (14.225) and the highest average R2 (0.146) under the adopted validation framework. Although the overall predictive performance remained modest, the explainability analysis consistently identified age as the most influential predictor, followed by onset year, education level, hearing aid usage duration, and daily hearing aid use. In contrast, gender, battery type, tinnitus, and vertigo contributed comparatively less to model predictions. These findings indicate that hearing aid adaptation and satisfaction arise from complex nonlinear interactions among demographic, behavioral, clinical, and device-related characteristics rather than isolated linear associations. The proposed framework provides an interpretable, internally validated analytical approach for investigating hearing aid technology adaptation and satisfaction and establishes a foundation for future studies that integrate comprehensive audiological measurements, longitudinal follow-up data, and independent external validation to support the development of more transparent and personalized hearing healthcare systems. Full article
(This article belongs to the Section Biosignal Processing)
Show Figures

Graphical abstract

29 pages, 4607 KB  
Article
Machine Learning-Based Classification of Glycemic Status Using Routine Laboratory Data: A Comparative Study of Statistical and Ensemble Models
by Argyrios Ginoudis, Dimitra Pardali, Eleni Vagdatli, Evgenia Lymperaki and Dimitrios Galiatsatos
BioMedInformatics 2026, 6(5), 63; https://doi.org/10.3390/biomedinformatics6050063 - 25 Aug 2026
Viewed by 227
Abstract
Early identification of individuals with abnormal glucose metabolism is essential for timely intervention and prevention of diabetes-related complications. Routine laboratory testing generates large amounts of clinical data that may support automated glycemic classification through machine learning approaches. This study aimed to develop and [...] Read more.
Early identification of individuals with abnormal glucose metabolism is essential for timely intervention and prevention of diabetes-related complications. Routine laboratory testing generates large amounts of clinical data that may support automated glycemic classification through machine learning approaches. This study aimed to develop and evaluate a machine learning framework for the classification of HbA1c-defined glycemic status using routinely available clinical laboratory features. A retrospective dataset of 1434 individuals with available glycemic measurements was analyzed. Participants were categorized into HbA1c-defined normoglycemic, prediabetic-range, or diabetic-range groups. Three concurrent classification tasks were examined: HbA1c-defined dysglycemia classification, diabetic-range HbA1c classification, and multiclass HbA1c-defined glycemic-status classification. Demographic, biochemical, and hematological variables were used as predictors. Data preprocessing included missing-value handling, feature filtering, and outlier treatment. Several supervised learning algorithms were evaluated, including Logistic Regression, Random Forest, Gradient Boosting, Support Vector Machine, and Multinomial Logistic Regression. Model performance was assessed using train–test validation and cross-validation with accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve. For dysglycemia, Gradient Boosting achieved the highest AUC (0.848), while Random Forest achieved the highest accuracy (0.801) and sensitivity (0.908). For diabetic-range HbA1c, Random Forest achieved the highest AUC (0.864), whereas SVM achieved the highest accuracy (0.794). In multiclass classification, Random Forest achieved the highest accuracy (0.610), while Gradient Boosting achieved the highest macro-AUC (0.796) and macro-F1 score (0.603). Pairwise comparisons showed no statistically significant superiority of any classifier after Holm correction. Clinical-baseline and ablation analyses demonstrated that fasting glucose accounted for a substantial proportion of discrimination, with only modest incremental value from additional laboratory variables. These findings support cautious interpretation of routine laboratory-based classification models pending further validation and clinical-utility assessment. Full article
(This article belongs to the Section Applied Biomedical Data Science)
Show Figures

Figure 1

20 pages, 551 KB  
Article
Socioenvironmental Vulnerability Profiles and Health Expenditure in Mexican Households Using Data Science
by Héctor Alejandro Acuña-Cid, Eduardo Ahumada-Tello, Cristina Almeida-Perales, Mónica Judith Chávez-Soto, Pablo Gerardo Guerrero-Herrera and José Eduardo Briceño-Muro
Big Data Cogn. Comput. 2026, 10(9), 284; https://doi.org/10.3390/bdcc10090284 - 25 Aug 2026
Viewed by 187
Abstract
This study aimed to identify socioenvironmental vulnerability profiles among Mexican households and analyze their association with health expenditure. Data from 86,102 households included in the 2024 National Household Income and Expenditure Survey were analyzed. Socioenvironmental profiles were constructed using housing, basic services, sanitation, [...] Read more.
This study aimed to identify socioenvironmental vulnerability profiles among Mexican households and analyze their association with health expenditure. Data from 86,102 households included in the 2024 National Household Income and Expenditure Survey were analyzed. Socioenvironmental profiles were constructed using housing, basic services, sanitation, household energy, socioeconomic stratum, and overcrowding through factor analysis of mixed data and k-means. Internal validation, stability analyses, algorithm comparisons, and sensitivity analyses supported a three-profile solution representing low, intermediate, and high vulnerability. Health expenditure was examined using survey-weighted descriptive estimates, exploratory nonparametric comparisons, and a survey-adjusted two-part model. The intermediate vulnerability profile showed higher odds of reporting health expenditure than the low vulnerability profile (OR = 1.129, 95% CI: 1.048 to 1.217, p = 0.002), whereas the high vulnerability profile showed no significant difference. Among households with positive expenditure, high vulnerability was associated with lower logarithmic expenditure (β=0.341, 95% CI: 0.490 to 0.192, p < 0.001), whereas the intermediate profile was not significantly different in the main model. The association for high vulnerability remained significant in sensitivity analyses, while results for the intermediate profile were more sensitive to model specification and income adjustment. Despite several statistically significant associations, effect sizes in the exploratory comparisons were small and the regression models explained a limited proportion of the variability in health expenditure. The findings therefore indicate modest associations between socioenvironmental vulnerability and health expenditure, with household economic resources and other unmeasured health-related factors likely contributing to the observed differences. Full article
(This article belongs to the Section Data Mining and Machine Learning)
Show Figures

Figure 1

27 pages, 2719 KB  
Article
Driver Behavior Classification on Secondary Roads Using Machine Learning Models
by Albert Jose Potams, Raymond Ghandour, Zaher Al Barakeh and Karim Youssef
Technologies 2026, 14(9), 524; https://doi.org/10.3390/technologies14090524 - 25 Aug 2026
Viewed by 271
Abstract
Most existing driver behavior classification technologies have focused on highways and other primary road infrastructures, despite secondary roads accounting for a disproportionately large number of traffic fatalities worldwide. Compared with highways, secondary roads present greater variability in road geometry, infrastructure quality, and traffic [...] Read more.
Most existing driver behavior classification technologies have focused on highways and other primary road infrastructures, despite secondary roads accounting for a disproportionately large number of traffic fatalities worldwide. Compared with highways, secondary roads present greater variability in road geometry, infrastructure quality, and traffic interactions, making driver behavior recognition considerably more challenging. This paper investigates the classification of driver behavior on secondary roads using machine learning techniques. Naturalistic driving data obtained from the publicly available UAH-DriveSet dataset were analyzed using two complementary feature groups describing lane detection and traffic status. Four supervised machine learning algorithms, namely, Logistic Regression (LR), gradient boosting (GB), Random Forest (RF), and Artificial Neural Networks (ANNs), were evaluated to classify driving behavior into three categories: Normal, Aggressive, and Drowsy. The extracted features were first analyzed through statistical profiling and exploratory feature analysis before training and evaluating the classification models. The experimental results show that gradient boosting consistently achieved the highest performance for both feature groups, attaining an overall classification accuracy of approximately 67% while providing balanced precision, recall, and F1-scores across all behavioral classes. Logistic regression and random forest produced competitive but lower performance, whereas the Artificial Neural Network yielded the lowest classification accuracy. The obtained results demonstrate the effectiveness of ensemble learning methods for driver behavior recognition under secondary-road conditions and highlight their potential for integration into intelligent driver monitoring and Advanced Driver Assistance Systems (ADASs). By enabling earlier identification of aggressive and drowsy driving behaviors on secondary roads, the proposed approach could support timely driver warnings and safety interventions, potentially reducing accident risk. Furthermore, the findings provide a benchmark for future machine learning models designed for real-world secondary-road environments, where driving conditions are more variable and challenging than on highways. Full article
(This article belongs to the Special Issue Advanced Intelligent Driving Technology)
Show Figures

Figure 1

16 pages, 1493 KB  
Systematic Review
Reference Ranges for Fetal Ventricular Global Longitudinal Strain (GLS) Using Bidimensional Speckle-Tracking Echocardiography: A Systematic Review
by Danielle Bittencourt Sodré Barmpas, Maria de Fátima Monteiro Pereira Leite, Saint Clair Gomes Junior, Karla G. Camacho, Maria Virginia M. Peixoto, Heron Werner and Renato Augusto Moreira de Sá
J. Clin. Med. 2026, 15(17), 6536; https://doi.org/10.3390/jcm15176536 - 24 Aug 2026
Viewed by 185
Abstract
Background/Objectives: The primary objective was to assess reference intervals for fetal Global Longitudinal Strain (GLS) using bidimensional speckle-tracking echocardiography (2D-STE), including only prospective studies specifically designed for this purpose. An additional objective was to evaluate studies’ methodological quality and reproducibility. Methods: This is [...] Read more.
Background/Objectives: The primary objective was to assess reference intervals for fetal Global Longitudinal Strain (GLS) using bidimensional speckle-tracking echocardiography (2D-STE), including only prospective studies specifically designed for this purpose. An additional objective was to evaluate studies’ methodological quality and reproducibility. Methods: This is a systematic review registered at PROSPERO (CRD420251038889). Five electronic databases (Web of Science, Scopus, MEDLINE/PubMed, EMBASE and LILACS) were searched, from inception to May 2025. Prospective studies specifically designed to establish 2D-STE GLS reference intervals in low-risk singleton pregnancies with normal fetuses were included. Data were independently extracted by two reviewers. Risk of bias was assessed using an adapted tool, including study design and statistical and reporting methods. Results: After the initial identification of 187 records, nine studies published between 2012 and 2025 were included. There was marked heterogeneity among the studies. Four articles achieved high-quality scores (>70%) and three of them reported similar left ventricular (LV) GLS at 24 weeks (−22%). Right ventricular absolute GLS values were slightly lower than LV numbers. Regression models for both ventricles showed GLS absolute values decreased with gestation across studies. The 2D-STE algorithm (endocardial versus myocardial) was the main source of discrepancy between studies. Conclusions: High-quality prospective studies show a consistent pattern of biventricular GLS variation with gestational age. However, technical heterogeneity, lack of standardization, operator subjectivity and vendor-specific algorithm differences currently limit the applicability of the method. Multicentric studies with large sample sizes, standardized protocols and artificial intelligence-assisted tools are needed to consolidate this technique. Full article
(This article belongs to the Special Issue Challenges and Opportunities in Prenatal Diagnosis)
Show Figures

Graphical abstract

42 pages, 2533 KB  
Article
Governance-Centered AI Framework for Public-Sector Budgetary Decision-Making and Financial Risk Management
by Hasan A. Hashim
Electronics 2026, 15(17), 3786; https://doi.org/10.3390/electronics15173786 - 24 Aug 2026
Viewed by 251
Abstract
The increasing adoption of artificial intelligence (AI) in public-sector financial management has raised significant concerns regarding interpretability, accountability, governance alignment, and institutional transparency. Existing AI-based fiscal analytical approaches frequently emphasize predictive capability while providing limited integration with formal governance structures and public-sector oversight [...] Read more.
The increasing adoption of artificial intelligence (AI) in public-sector financial management has raised significant concerns regarding interpretability, accountability, governance alignment, and institutional transparency. Existing AI-based fiscal analytical approaches frequently emphasize predictive capability while providing limited integration with formal governance structures and public-sector oversight requirements. This study proposes a governance-centered AI consultancy framework that embeds AI-assisted fiscal analysis directly within institutional budgeting, accountability, and governance-oriented decision-support workflows. Rather than treating AI as an isolated predictive or automation technology, the proposed framework operationalizes analytical intelligence within governance-aware consultancy structures emphasizing interpretability, auditability, traceability, and institutional usability. The framework was evaluated using authentic longitudinal public-sector fiscal records obtained from the official Ministry of Finance budget performance reports of Saudi Arabia for fiscal year 2023. The experimental evaluation incorporated temporal fiscal monitoring, robustness analysis under heterogeneous budgetary conditions, and comparative assessment against conventional descriptive budgetary analysis and standalone AI-based fiscal analytical procedures. The experiments utilized quarterly governmental fiscal indicators including revenues, expenditures, deficit progression, debt accumulation, expenditure volatility, and oil and non-oil revenue behavior across multiple reporting intervals. The findings demonstrate that the proposed governance-centered framework preserves strong temporal analytical consistency (81.9%) while achieving an algorithmically computed interpretability-support score of 4.6/5, a Governance Alignment Index of 0.94, and an operational Decision Usability Index of 4.7/5 relative to the conventional statistical baseline (Logistic Regression) and standalone AI-based analytical approaches. Improvements over conventional descriptive budgetary analysis are reported separately through the governance-oriented institutional comparison. Additional validation studies showed that the Decision Usability Index and Governance Alignment Index provided the strongest predictive contributions, while Traceability Index and Temporal Support Index exhibited the strongest construct-level statistical validity evidence. Robustness analysis further showed stable governance-aware analytical behavior across heterogeneous fiscal conditions involving expenditure volatility, debt progression, and changing revenue structures. Additional statistical validation demonstrated empirical support for the Traceability Index and Temporal Support Index, while other governance metrics exhibited weaker evidence and should be interpreted primarily as governance-support indicators rather than primary predictive drivers. The findings suggest that governance-aware analytical operationalization can provide measurable value beyond standalone AI models for formula-based operational fiscal-risk categorization when supported by reproducible governance-oriented analytical procedures. Because the supervised labels represent deterministic operational fiscal-risk categories rather than independently verified fiscal anomalies, the reported results should not be interpreted as direct validation of real-world fiscal anomaly detection or financial misconduct identification. Full article
Show Figures

Figure 1

Back to TopTop