Machine Learning for Predictive Analytics: Models, Applications, and Challenges

A Special Issue of Information (ISSN 2078-2489) belonging to the section "Artificial Intelligence".

Deadline for manuscript submissions: 31 January 2027 | Viewed by 14834

Editors

School of Technology and Maritime Industries, Southampton Solent University, Southampton SO140YN, UK
Interests: AI; machine learning; machine vision; LLM; data visualization
Special Issues, Collections and Topics in MDPI journals

E-Mail Website
Guest Editor
Department of Science and Engineering, Southampton Solent University, Southampton SO14 0YN, UK
Interests: affective computing; investigating multimodal data; hybrid DNNs; applications of AI; data science; computer vision; time-series and financial market analysis; FinTech
Special Issues, Collections and Topics in MDPI journals

E-Mail Website
Guest Editor
Division of Information and Communication Engineering, Kitami Institute of Technology, Kitami 090-8507, Hokkaido, Japan
Interests: computer vision; control theory and application; artificial intelligence

Special Issue Information

Dear Colleagues,

The MDPI Information journal invites submissions to a Special Issue on "Machine Learning for Predictive Analytics: Models, Applications, and Challenges".

Machine learning (ML) continues to revolutionize predictive capabilities across scientific and industrial domains. While achieving remarkable success, ML-based prediction systems face persistent challenges in interpretability, generalization, computational efficiency, and ethical implementation. This Special Issue seeks to advance the field by publishing innovative research that bridges theoretical developments with practical solutions across the predictive analytics pipeline.

Contributions are invited across (but are not limited to) the following themes:

  1. Model Development and Innovation
  • Novel architectures (transformers, graph neural networks, neurosymbolic systems);
  • Time-series, spatial–temporal, and multimodal forecasting;
  • Uncertainty quantification and confidence calibration;
  • Federated and distributed learning approaches.
  1. Domain-Specific Applications
  • Healthcare: clinical outcome prediction and medical imaging analytics;
  • Cybersecurity: threat detection and adversarial attack forecasting;
  • Engineering: predictive maintenance and structural health monitoring;
  • Climate Science: extreme weather modeling and carbon emission prediction;
  • Finance (FinTech): algorithmic trading, fraud detection systems, and credit risk assessment;
  • Education (EdTech): learning outcome prediction, adaptive learning systems, student performance analytics, and educational resource optimization;
  • Smart Cities: traffic flow optimization;
  • Agriculture: precision farming and crop yield forecasting;
  • Social Good: poverty mapping, disaster response optimization, and computer forensic analytics.
  1. Critical Challenges and Solutions
  • Explainable AI (XAI) for high-consequence decisions;
  • Bias detection and fairness-aware modeling;
  • Edge deployment and resource-efficient inference;
  • Hybrid modeling;
  • Data scarcity solutions.
  1. Evaluation and Reproducibility
  • Benchmark datasets and metrics;
  • Reproducibility frameworks;
  • Real-world validation studies.

We welcome original research and reviews that demonstrate rigorous methodology with clear practical implications. Interdisciplinary contributions connecting ML theory with domain expertise are particularly encouraged. Join us in shaping the future of predictive analytics—submit your work to advance methodologies, tools, and applications that empower equitable and sustainable decision-making.

Dr. Raza Hasan
Dr. Bacha Rehman
Prof. Dr. Wei Xie
Guest Editors

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Information is an international peer-reviewed open access monthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 1800 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • explainable AI
  • FinTech
  • cybersecurity
  • EduTech
  • healthcare analytics
  • forensic AI
  • hybrid deep learning
  • multimodal data fusion
  • predictive modeling
  • ethical AI
  • algorithmic fairness
  • adaptive learning systems
  • financial forecasting
  • threat intelligence

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (12 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

Jump to: Review

21 pages, 2531 KB  
Article
A Reproducible Framework for Monitoring and Forecasting Regional Morbidity in Kazakhstan Using Harmonized Annual Official Statistics
by Zhanar Oralbekova, Zhaniya Karabayeva, Marzhan Turarova, Natalya Demidchik, Akmaral Oralbekova, Lyailya Kurmangaziyeva and Balbupe Utenova
Information 2026, 17(9), 900; https://doi.org/10.3390/info17090900 - 15 Sep 2026
Viewed by 148
Abstract
This study evaluates a reproducible framework for forecasting and retrospective monitoring of regional morbidity in Kazakhstan using harmonized annual official statistics. Ministry of Healthcare compendia were combined with demographic and living-standards tables from the Bureau of National Statistics. The panel comprises 1344 observations [...] Read more.
This study evaluates a reproducible framework for forecasting and retrospective monitoring of regional morbidity in Kazakhstan using harmonized annual official statistics. Ministry of Healthcare compendia were combined with demographic and living-standards tables from the Bureau of National Statistics. The panel comprises 1344 observations for 2012–2025 across 16 stable territories and six disease classes. Forty-two primary forecasting configurations and two sensitivity configurations were evaluated using expanding-window one-year-ahead validation, with selection based on 2017–2024 forecasts. Results for 2025 are exploratory because the outcome was inspected during model development. The selected specification combines the last-observation forecast with an Extra Trees residual correction, core lagged predictors, and a shrinkage weight selected within earlier temporal folds. Pooled mean absolute error (MAE) was 163.65 vs. 166.87 for the last-observation forecast, a 1.93% reduction; 2025 values were 108.05 and 110.69. The exact year-block test yielded a one-sided p-value of 0.066 and a familywise-adjusted p-value of 0.387. The hybrid reduced MAE in four of six disease classes. Grouped Shapley additive explanations (SHAP) decomposed the complete forecast. Monitoring used absolute error relative to preceding MAE; the categories are retrospective aids and were not validated against clinical outcomes or interventions. Full article
Show Figures

Figure 1

24 pages, 643 KB  
Article
Long-Term Fairness-Aware Recommendation via Adaptive Fairness Metric Selection
by Jijun Yu and Minghua Xiong
Information 2026, 17(9), 834; https://doi.org/10.3390/info17090834 - 28 Aug 2026
Viewed by 349
Abstract
Fairness in recommendation systems has drawn growing attention due to rising societal and regulatory concerns over algorithmic bias. Existing fairness-aware approaches typically mitigate bias by either removing sensitive attributes via representation learning or leveraging causal-path interventions (e.g., counterfactual or specific-path debiasing) to distinguish [...] Read more.
Fairness in recommendation systems has drawn growing attention due to rising societal and regulatory concerns over algorithmic bias. Existing fairness-aware approaches typically mitigate bias by either removing sensitive attributes via representation learning or leveraging causal-path interventions (e.g., counterfactual or specific-path debiasing) to distinguish genuine causal effects from confounder-induced correlations between sensitive attributes and user preferences. However, when it comes to evaluation, most prior work adopts both Demographic Parity (DP) and Equal Opportunity (EO) as simultaneous criteria, yet overlooks their inherent tension and the causal nature of the sensitive attribute. Specifically, if a sensitive attribute genuinely drives preference variation, enforcing DP forces equal exposure across groups, contradicting natural interest diversity and severely hurting accuracy; conversely, for spurious correlations, relying solely on EO fails to remove confounder-introduced bias. More importantly, these metrics are typically computed in a static, one-shot manner, ignoring that recommendation is an iterative process where even minor initial disparities can be amplified over time through feedback loops, eventually leading to substantial long-term unfairness. Nevertheless, existing studies rarely address such dynamic, long-term fairness implications, leaving a critical gap in both evaluation and optimization. To resolve this, we propose Long-term Fairness-aware Recommendation via Adaptive Fairness Metric Selection (LFR-via-AFMS). Our framework first learns the causal structure to identify whether the sensitive attribute has a genuine causal effect or merely a spurious association with user preferences. Based on this diagnosis, it adaptively selects the most appropriate fairness criterion: Equal Opportunity for true causality, which allows legitimate group differences in preference, and Demographic Parity for spurious correlations, which eliminates unjustified disparities entirely. The adaptively chosen metric is then integrated into an actor–critic reinforcement learning reward to optimize long-term fairness without sacrificing accuracy. Extensive experiments on Alibaba and MovieLens datasets, with five independent runs and statistical significance testing, demonstrate that the proposed method achieves a superior fairness-accuracy trade-off compared with state-of-the-art baselines, and the adaptive metric selection proves indispensable for maintaining both equity and recommendation quality. We validate the causal diagnosis module through simulation studies with known ground truth and sensitivity analyses confirming robustness across threshold choices. Full article
Show Figures

Graphical abstract

32 pages, 3698 KB  
Article
Spatial Predictive Patterns of Cause-Specific Mortality: Evidence from East Africa
by Sally Sonia Simmons, John Elvis Hagan, Jr., Imanol L. Nieto-González and Thomas Schack
Information 2026, 17(8), 804; https://doi.org/10.3390/info17080804 - 20 Aug 2026
Viewed by 280
Abstract
(1) Background: Whether spatial predictive patterns in non-communicable disease mortality persist after accounting for socio-demographic development and biomarkers remains understudied in East Africa. (2) Methods: This study used heterogeneous graph transformer (HGT) models and other techniques to model spatial patterns in cause- and [...] Read more.
(1) Background: Whether spatial predictive patterns in non-communicable disease mortality persist after accounting for socio-demographic development and biomarkers remains understudied in East Africa. (2) Methods: This study used heterogeneous graph transformer (HGT) models and other techniques to model spatial patterns in cause- and sex/age-specific mortality (hypertensive heart disease [HHD], ischaemic heart disease [IHD], stroke, and diabetes), incorporating risk factors and socio-demographic development (SDI), using data from the Global Burden of Disease (GBD) study, 1990–2023, across Burundi, Kenya, Rwanda, Tanzania, and Uganda. (3) Results: HGT achieved higher performance than OLS spatial lag benchmarks (R2 0.948–0.970 vs. 0.194–0.376). Spatial predictive patterns were disease-specific. Stroke was the only disease with consistent positive spatial structure (SDI-only: 0.645%, 95% CI [0.380, 0.907]), with spatial structure strengthening after 2015. HHD exhibited severe and stable degradation (Risk-only: −137.892%, 95% CI [−181.908, −96.380]), driven by the interaction between metabolic risk covariates and geographic adjacency. Diabetes showed consistently severe degradation (SDI + Risk: −201.941%, 95% CI [−257.349, −150.082]). IHD patterns were weak and unstable. Sex disaggregation revealed stronger stroke spatial signals, indicating latent sex-specific patterns masked by aggregation. GBD measurement uncertainty contributed less than 0.025% of result variance, with model randomness dominating. (4) Conclusions: Spatial predictive patterns in NCD mortality in East Africa are disease-specific. Stroke shows emerging cross-border spatial structure after 2015, while HHD and diabetes reflect country-specific determinants. Sex-disaggregated graph construction reveals latent spatial heterogeneity invisible to aggregate models, supporting disease-specific, sex-stratified regional health strategies. Full article
Show Figures

Figure 1

30 pages, 3364 KB  
Article
A Multi-Attribute Predictive Analysis Model for University Student Sentiment Public Opinion Based on Big Data
by Baoguo Chen and Yongsheng Hao
Information 2026, 17(8), 792; https://doi.org/10.3390/info17080792 - 18 Aug 2026
Viewed by 302
Abstract
With social media as the main channel for college students to express emotions, sentiment public opinion analysis in big data environments poses three core challenges to campus sentiment monitoring and psychological counseling: severe data noise interference, insufficient multi-attribute feature extraction, and the trade-off [...] Read more.
With social media as the main channel for college students to express emotions, sentiment public opinion analysis in big data environments poses three core challenges to campus sentiment monitoring and psychological counseling: severe data noise interference, insufficient multi-attribute feature extraction, and the trade-off between recognition accuracy and inference efficiency. This paper proposes a university student public opinion prediction model integrating multi-attribute decision-making and BERT–Mamba. First, an anti-interference matching filter cleans raw data by filtering out advertisements and irrelevant comments to improve data quality. Second, a multi-attribute decision object model extracts quantifiable attributes covering media sources, themes, and temporal dimensions. Third, BERT generates textual sentiment representations, and a three-stage deep feature extraction architecture with Mamba balances accuracy and efficiency. Finally, multi-attribute features and sentiment representations are fused for dynamic public opinion prediction. Validated using the ChnSentiCorp Chinese sentiment analysis benchmark dataset and university student Weibo public opinion corpus, the model achieves 97.44% average sentiment recognition accuracy. It provides technical support for universities to understand student sentiment trends and address negative public opinions, with practical value for enhancing campus public opinion monitoring and assisting mental health counseling. Full article
Show Figures

Figure 1

45 pages, 10654 KB  
Article
Persistent Highway–Rail Grade Crossing Incidents: A Spatial Analytics and Explainable Machine-Learning Framework
by Raj Bridgelall
Information 2026, 17(8), 718; https://doi.org/10.3390/info17080718 - 23 Jul 2026
Viewed by 535
Abstract
Highway–rail grade crossing (HRGC) incidents in the United States declined substantially for several decades before stabilizing in recent years. Understanding this persistence is important because future safety improvements may depend on identifying locations where incident occurrence remains resistant to further reduction. This study [...] Read more.
Highway–rail grade crossing (HRGC) incidents in the United States declined substantially for several decades before stabilizing in recent years. Understanding this persistence is important because future safety improvements may depend on identifying locations where incident occurrence remains resistant to further reduction. This study developed an integrated framework to characterize persistent HRGC incident environments using 50 years (1976–2025) of Federal Railroad Administration incident records. Trend, structural-break, variance, and stationarity tests were first applied to determine whether the historical decline transitioned into a distinct persistence regime. A county-level persistence index (PI) was then developed to quantify the combined effects of incident burden and resistance to decline during the plateau period. Distributional analysis characterized the statistical behavior of the PI, while global and local Moran’s I statistics evaluated its spatial organization. Explainable machine learning methods were subsequently used to identify incident characteristics associated with elevated persistence. The results identified a statistically significant regime change around 2010. Prior to 2010, incidents exhibited a strong declining trend, whereas the subsequent period displayed a statistically significant but substantially weaker decline, lower variance, and behavior consistent with a persistence regime characterized by a markedly attenuated rate of improvement. The PI followed a strongly right-skewed distribution that was best represented by a bounded heavy-tailed unit log-logistic model, indicating that persistence is concentrated within a relatively small subset of counties. Spatial analysis revealed significant positive spatial autocorrelation (Moran’s I = 0.180, p = 0.001) and geographically coherent clusters concentrated primarily in the southeastern United States and several major freight-oriented regions. Explainable machine learning models identified train-operating characteristics, warning device contexts, movement patterns, and temporal conditions as key attributes associated with high-persistence counties. The findings demonstrate that the post-2010 incident plateau is sustained disproportionately by a limited number of geographically concentrated environments and provide a framework for supporting more targeted safety interventions. Full article
Show Figures

Graphical abstract

20 pages, 1374 KB  
Article
Dynamic Cost Prediction for State Grid Engineering Projects Based on Multi-Source Business Data Fusion and Data-Driven Methods
by Weiqiong Wang, Qidong Xu, Tianyu Zhao and Fang Fang
Information 2026, 17(7), 691; https://doi.org/10.3390/info17070691 - 16 Jul 2026
Viewed by 436
Abstract
Accurate dynamic cost prediction is essential for budget optimization and risk mitigation in State Grid projects. However, traditional models and even recent deep learning approaches fall short, as they treat cost drivers independently, adopt simplistic concatenation that destroys sourcewise structure, or fail to [...] Read more.
Accurate dynamic cost prediction is essential for budget optimization and risk mitigation in State Grid projects. However, traditional models and even recent deep learning approaches fall short, as they treat cost drivers independently, adopt simplistic concatenation that destroys sourcewise structure, or fail to handle irregularly sampled and partially missing multi-source data. This paper proposes a novel data-driven framework that integrates multi-source business data through a hierarchical tensor fusion mechanism and a hybrid spatiotemporal architecture. The problem is formalized as multivariate time-series prediction with irregular sampling and missing modalities. The framework comprises three synergistic innovations: a differentiable low-rank CANDECOMP/PARAFAC (CP) decomposition layer with adaptive attention weights that preserves cross-source structure while enabling compact dimensionality reduction; a spatiotemporal attention-based bidirectional gated recurrent unit (Bi-GRU) that captures long-range temporal dependencies; and a graph convolutional network (GCN) that explicitly learns interrelations among cost drivers, a capability absent in most existing forecasting methods. The entire system is trained end to end with a customized loss combining mean squared error, quantile loss, and temporal consistency regularization. Extensive experiments on three State Grid substation projects demonstrate that the proposed method outperforms state-of-the-art baselines by 12.7–18.4% in MAPE and maintains robust performance with up to 40% of data missing. These results confirm that explicitly modeling both temporal evolution and driver interdependencies within a unified fusion framework is the key to reliable cost forecasting in large-scale infrastructure projects. Full article
Show Figures

Figure 1

30 pages, 1675 KB  
Article
Predicting Academic Award Recognition Across Disciplines Using Publication-Based Bibliometric Indices and SHAP-Driven Explainability
by Muhammad Shaban Qabil, Hafiza Zarafshan Mukhtiar, Ghulam Mustafa, Muhammad Tanvir Afzal, Isabel De la Torre Díez, Elizabeth Caro Montero and Mirtha Silvana Garat de Marin
Information 2026, 17(6), 515; https://doi.org/10.3390/info17060515 - 22 May 2026
Cited by 1 | Viewed by 711
Abstract
Researcher evaluation underpins critical academic decisions, yet traditional bibliometric indicators lack predictive capability and cross-domain generalizability, while most predictive approaches offer limited interpretability and narrow domain validation. This study proposes a SHAP interpretable, multi-domain supervised learning framework for predicting academic award recognition using [...] Read more.
Researcher evaluation underpins critical academic decisions, yet traditional bibliometric indicators lack predictive capability and cross-domain generalizability, while most predictive approaches offer limited interpretability and narrow domain validation. This study proposes a SHAP interpretable, multi-domain supervised learning framework for predicting academic award recognition using thirty two publication count-based bibliometric indices. A balanced dataset was constructed across four disciplines, namely Computer Science, Neuroscience, Mathematics, and Civil Engineering, comprising verified awardees from recognized professional societies and matched non-awardee researchers. Eight classifiers were evaluated under stratified five fold cross validation, assessed via accuracy, precision, recall, F1-score, and ROC AUC. The framework achieved domain-specific F1-scores of 0.70 in Computer Science, 0.73 in Neuroscience, 0.72 in Civil Engineering, and 0.78 in Mathematics, with SVM and XGBoost demonstrating the strongest cross-domain robustness across disciplines. SHAP analysis consistently identified normalized h index, h2 family, q2 index, and g index as dominant cross-domain predictors, while domain-specific indicators, including Rm and w indices in Neuroscience and P index in Civil Engineering, reflected disciplinary recognition patterns. By unifying publication-based feature engineering, multi-domain classification, and SHAP explainability within a single reproducible pipeline, this framework offers a scalable, transparent, and evidence-based tool for institutional researcher evaluation. Full article
Show Figures

Figure 1

20 pages, 2397 KB  
Article
Towards Sustainable AI: Benchmarking Energy Efficiency of Deep Neural Networks for Resource-Constrained Edge Devices
by Rohail Qamar, Raheela Asif and Syed Muslim Jameel
Information 2026, 17(4), 380; https://doi.org/10.3390/info17040380 - 17 Apr 2026
Cited by 1 | Viewed by 1483
Abstract
Deep learning models represent one of the most advanced and effective approaches in predictive modeling. Their hierarchical architectures enable the extraction of complex, non-linear feature relationships and the identification of latent patterns within data, making them highly suitable for tasks involving high-dimensional or [...] Read more.
Deep learning models represent one of the most advanced and effective approaches in predictive modeling. Their hierarchical architectures enable the extraction of complex, non-linear feature relationships and the identification of latent patterns within data, making them highly suitable for tasks involving high-dimensional or unstructured inputs. However, these models are computationally demanding, requiring significant processing resources and time. Furthermore, their predictive performance is largely contingent upon the availability of large-scale datasets. In this study, a Deep Green Framework is employed for the prediction of two computer vision tasks. CIFAR-10 and CIFAR-00 have been taken for image classification. Fifteen convolutional neural network (CNN) variants categorized into light-weight and heavy-weight are trained for the prediction of these two datasets. Based on energy footprint, time, memory usage, Top-1 accuracy, Top-3 accuracy, model size, and model parameters. The study highlights that MobileNetV3-Small produces the best outcomes when compared to other trained models having low task latency and higher efficiency, making it highly suitable for edge environments where resources are scarce. Full article
Show Figures

Graphical abstract

33 pages, 3715 KB  
Article
Enhancing Multi-Level Spatio-Temporal Forecasting of Adjudicated Crime Occurrence Trends in Indonesia
by Firman Arifman, Teddy Mantoro and Media Anugerah Ayu
Information 2026, 17(4), 331; https://doi.org/10.3390/info17040331 - 1 Apr 2026
Viewed by 1187
Abstract
Indonesia faces persistent challenges in crime forecasting and judicial resource management, compounded by chronic underreporting and inconsistent spatial resolution in official crime statistics. In this study, a multi-level spatio-temporal machine learning framework is developed and applied to 95,666 adjudicated crime records from the [...] Read more.
Indonesia faces persistent challenges in crime forecasting and judicial resource management, compounded by chronic underreporting and inconsistent spatial resolution in official crime statistics. In this study, a multi-level spatio-temporal machine learning framework is developed and applied to 95,666 adjudicated crime records from the Supreme Court of Indonesia spanning January 2023 to June 2024. Following the CRISP-DM methodology, a hybrid STL-XGBoost v. 3.2.0 model is trained on a chronological split to forecast daily judicial caseloads, achieving an R2 of 0.8070, MAE of 16.52, and sMAPE of 9.76% on the held-out test set. DBSCAN spatial clustering, parameterized via k-distance plot analysis (ϵ=0.3, minPts = 3) and validated through Jaccard Similarity Index sensitivity analysis, identifies 29 distinct adjudicated crime hubs concentrated along Java and Sumatra’s urban and transit corridors. Comparative analysis of reported versus adjudicated crime data reveals systematic judicial funnel attrition ranging from 199.12% in Riau to 2436.02% in Papua, establishing that adjudicated crime records provide a reliable indicator of judicial workload rather than a comprehensive measure of social deviance. Key limitations, including the 18-month observation window that may not capture long-term policy shifts and the use of city centroids as spatial proxies that introduces a degree of ecological fallacy, are acknowledged. The framework offers a scalable, interpretable decision support tool for evidence-based judicial resource planning across national, provincial, and city scales in Indonesia. Full article
Show Figures

Figure 1

32 pages, 2526 KB  
Article
HSE-GNN-CP: Spatiotemporal Teleconnection Modeling and Conformalized Uncertainty Quantification for Global Crop Yield Forecasting
by Salman Mahmood, Raza Hasan and Shakeel Ahmad
Information 2026, 17(2), 141; https://doi.org/10.3390/info17020141 - 1 Feb 2026
Cited by 2 | Viewed by 1423
Abstract
Global food security faces escalating threats from climate variability and resource constraints. Accurate crop yield forecasting is essential; however, existing methods frequently overlook complex spatial dependencies driven by climate teleconnections, such as the ENSO, and lacks rigorous uncertainty quantification. This paper presents HSE-GNN-CP, [...] Read more.
Global food security faces escalating threats from climate variability and resource constraints. Accurate crop yield forecasting is essential; however, existing methods frequently overlook complex spatial dependencies driven by climate teleconnections, such as the ENSO, and lacks rigorous uncertainty quantification. This paper presents HSE-GNN-CP, a novel framework integrating heterogeneous stacked ensembles, graph neural networks (GNNs), and conformal prediction (CP). Domain-specific features are engineered, including growing degree days and climate suitability scores, and explicitly model spatial patterns via rainfall correlation graphs. The ensemble combines random forest and gradient boosting learners with bootstrap aggregation, while GNNs encode inter-regional climate dependencies. Conformalized quantile regression ensures statistically valid prediction intervals. Evaluated on a global dataset spanning 15 countries and six major crops from 1990 to 2023, the framework achieves an R2 of 0.9594 and an RMSE of 4882 hg/ha. Crucially, it delivers calibrated 80% prediction intervals with 80.72% empirical coverage, significantly outperforming uncalibrated baselines at 40.03%. SHAP analysis identifies crop type and rainfall as dominant predictors, while the integrated drought classifier achieves perfect accuracy. These contributions advance agricultural AI by merging robust ensemble learning with explicit teleconnection modeling and trustworthy uncertainty quantification. Full article
Show Figures

Graphical abstract

21 pages, 1753 KB  
Article
A Personality-Informed Candidate Recommendation Framework for Recruitment Using MBTI Typology
by Hamza Wazir Khan, Mian Usman Sattar, Samreen Noor and Muna I. Alyousef
Information 2025, 16(10), 863; https://doi.org/10.3390/info16100863 - 5 Oct 2025
Cited by 2 | Viewed by 5514
Abstract
In many developing regions, recruitment still relies heavily on traditional methods that often ignore the importance of aligning a candidate’s personality with the job role. This mismatch can lead to poor performance, dissatisfaction, and high turnover. To address this, the study presents a [...] Read more.
In many developing regions, recruitment still relies heavily on traditional methods that often ignore the importance of aligning a candidate’s personality with the job role. This mismatch can lead to poor performance, dissatisfaction, and high turnover. To address this, the study presents a personality-aware recommendation system that combines the Myers–Briggs Type Indicator (MBTI) with machine learning to support smarter hiring decisions. The system is tailored for the South Asian job market and includes two main components: a web-based MBTI assessment for applicants and a dashboard for HR professionals powered by a XGBoost classifier. This model was trained on a dataset correlating applicant profiles and the flagged preferences of MBTI with the job. Experience and the number of skills, education level, and encoded MBTI types were the key features, and the SMOTE method was employed to balance the dataset. The model attained an accuracy of 74.30%, having balanced precision and recall measures. It was also discriminative, the ROC AUC was 0.84, and the precision–recall AUC was 0.85. One example of utilizing the Software Developer position in real life demonstrated the success of the system to filter and rank candidates at the same time according to both technical and personality-specific criteria. Overall, this study emphasizes the worth of combining insights from psychological profiling with machine learning in order to develop a more holistically, fair, and efficient hiring process. Full article
Show Figures

Figure 1

Review

Jump to: Research

27 pages, 1252 KB  
Review
Beyond Occam’s Razor: Double Descent and the Potential Paradigm Shift Toward Over-Parameterized Personalization in Higher Education
by Chong Ho Yu and Han Nee Chong
Information 2026, 17(7), 696; https://doi.org/10.3390/info17070696 - 17 Jul 2026
Viewed by 742
Abstract
This paper examines how the emergence of over-parameterized artificial intelligence models and the phenomenon of double descent challenge the classical assumption that simpler models generalize better. Traditional predictive analytics relied on parsimonious models grounded in the bias-variance trade-off, where increasing complexity was expected [...] Read more.
This paper examines how the emergence of over-parameterized artificial intelligence models and the phenomenon of double descent challenge the classical assumption that simpler models generalize better. Traditional predictive analytics relied on parsimonious models grounded in the bias-variance trade-off, where increasing complexity was expected to produce overfitting. However, recent advances in deep learning demonstrate that highly over-parameterized models can achieve superior generalization after surpassing the interpolation threshold. This paradigm shift has enabled systems such as AlphaFold, Aurora, Delphi-2M, and recommenders to model complex, high-dimensional relationships through contextual attention rather than global feature selection. The paper argues that higher education analytics remains largely reductionist, relying on limited variables such as GPA, demographics, and course completion rates to identify “at-risk” students. While interpretable, these approaches often fail to capture the dynamic and multidimensional nature of student success. In response, this study proposes a transition toward over-parameterized personalization, where students’ academic and behavioral histories are modeled as longitudinal high-dimensional sequences. Drawing parallels to commercial recommendation systems such as Amazon, Netflix, and YouTube, the paper explores how higher education can move from generalized early-warning systems toward adaptive “n-of-1” interventions. Importantly, the paper is conceptual rather than empirical: it develops a research agenda and a set of testable propositions, and it identifies the evaluation designs—temporally valid prediction protocols and causal intervention studies—by which the promise of over-parameterized personalization in higher education should be assessed before any claim of superiority can be made. Full article
Show Figures

Graphical abstract

Back to TopTop