3.1. Results of Feature Selection
To gain an initial understanding of the relationships between input variables and the target outcomes, a correlation analysis was first conducted. This step provides insight into the strength and direction of association between individual features and the dependent variables (
G3_level,
G3_grade, and
G3_pass/fail), offering a basis for subsequent feature selection. By identifying variables with stronger associations, this preliminary analysis supports the selection of the most informative predictors and helps reduce dimensionality before applying ML models.
Table 3 presents the correlation coefficients between the selected input features with the strongest associations and the outcome variables.
As shown in
Table 3, prior academic indicators, particularly
G1,
G2, and
failures, exhibit the strongest correlations with all outcome variables. In contrast, demographic and behavioral attributes exhibit weaker associations, suggesting a more limited direct influence on academic performance. These findings further justify the application of feature selection techniques to identify the most informative predictors.
The application of the CFS method yielded compact and interpretable feature subsets that are largely consistent with theoretical expectations in educational research, supporting its suitability for comparative analysis across multiple predictive tasks. For the ordinal performance levels (G3_level), the method selected mother’s education (medu), the number of prior course failures (failures), and the second-period grade (G2). For the final grade categorization (G3_grade), the selected subset comprised the home address type (address), the number of prior course failures (failures), the first- and second-period grades (G2, G1), the attendance in preschool (nursery), and relationship status (romantic). Finally, for the binary outcome (G3_pass/fail), the retained features were the student’s sex (sex), the home address type (address), the number of prior course failures (failures), the number of school absences (absences), and the first- and second-period grades (G1, G2).
These results demonstrate the ability of the CFS method to identify both academic and socio-demographic variables that are strongly associated with student achievement. In particular, the consistent selection of failures G1 and G2 across multiple outcomes underscores the dominant predictive role of prior academic performance, a finding consistent with established evidence in educational analytics. At the same time, the inclusion of contextual attributes such as address and nursery for the four-level grade outcome suggests that environmental and early-life factors may exert additional influence on student performance. Overall, CFS proved effective in reducing the feature space to manageable, interpretable subsets, emphasizing the most relevant predictors while mitigating redundancy, an important advantage when working with educational datasets containing many potentially overlapping attributes.
The CorrEval method produced slightly broader feature sets by ranking individual variables according to their pairwise linear association with the target outcome. For G3_level, the selected attributes were the student’s age (age), the number of prior course failures (failures), the peer social engagement (goout), and the first- and second-period grades (G1, G2). For G3_grade, the method identified mother’s education (medu), the number of prior course failures (failures), the student’s age (age), the peer social engagement (goout), and the first- and second-period grades (G1, G2). Finally, for G3_pass/fail, the retained features included mother’s education (medu), the number of prior course failures (failures), the extra educational support (schoolsup), and the first- and second-period grades (G1, G2).
These results further highlight the central importance of academic indicators, as G1, G2 and failures were consistently selected across all outcome definitions. In addition, socio-demographic and behavioral variables emerged as relevant predictors. For example, medu was selected for both the four-level and binary outcomes, highlighting the role of parental education in shaping academic success. Similarly, goout reflecting students’ social activity pattern was associated with both ordinal and four-level performance measures, suggesting that lifestyle factors may have measurable links to achievement. The inclusion of schoolsup for the binary outcome further points to the potential impact of supplementary academic support on course completion. Nevertheless, because this method relies solely on pairwise correlations, it does not explicitly address redundancy among highly correlated predictors. Consequently, closely related variables such as G1 and G2 were frequently selected together. Moreover, nonlinear associations are only captured to the extent that they manifest as monotonic relationships.
When applied to the same prediction tasks, the InfoGain method identified a broader set of informative features. For G3_level, the most informative attributes were the first- and second-period grades (G1, G2), the number of prior course failures (failures), mother’s occupation (mjob), mother’s education (medu), aspiration for higher education (higher), father’s occupation (fjob), and the extra educational support (schoolsup). For G3_grade, InfoGain selected the first- and second-period grades (G1, G2), the number of prior course failures (failures), mother’s occupation (mjob), father’s occupation (fjob), and the extra educational support (schoolsup). For G3_pass/fail, the identified features were the first- and second-period grades (G1, G2), the number of prior course failures (failures), mother’s occupation (mjob), aspiration for higher education (higher), the number of school absences (absences), father’s occupation (fjob), and the extra educational support (schoolsup).
Across all three outcomes, InfoGain consistently emphasized prior academic performance and failure history, confirming their central predictive role. In addition, the method highlighted several socio-economic and aspirational variables, including parental occupation, educational aspirations, and access to supplementary support. This pattern illustrates the capacity of this approach to capture broader contextual influences on academic outcomes by quantifying uncertainty reduction, rather than relying solely on linear associations. However, because it evaluates features independently, it does not account for redundancy or inter-feature interactions. As a result, highly correlated predictors were consistently selected together, leading to slightly larger feature sets compared to CFS.
Figure 2 presents a heatmap illustrating the selection frequency (low, medium, high) of each feature across the three outcome definitions (
G3_level,
G3_grade,
G3_pass/fail) and the applied filter-based techniques (CFS, CorrEval, InfoGain). The results indicate that the three procedures provide complementary perspectives. Academic indicators (
G1,
G2, and
failures) were consistently identified as key predictors across all methods and outcomes, underscoring their dominant role in modeling student achievement. In contrast, InfoGain highlighted additional socio-economic and support-related factors (
mjob,
fjob,
higher,
schoolsup) that were less emphasized by correlation-based techniques. Overall, these results demonstrate that combining information-theoretic and correlation-based approaches provides a more comprehensive view of the academic, behavioral, and contextual drivers of student performance.
Beyond reporting which variables/features were selected by each method, it is important to interpret why certain socio-demographic attributes emerge or disappear across outcome granularities. The results indicate that the relevance of contextual variables varies according to how academic achievement is defined. For example, parental education and occupation (e.g., medu, mjob, fjob) are selected more often when predicting detailed grade levels (G3_level, G3_grade), where distinguishing between different performance categories requires considering broader socio-economic influences. In contrast, the binary outcome (G3_pass/fail) is mainly driven by prior academic indicators such as G1, G2, and failures, which directly reflect whether a student meets the minimum passing requirement. This suggests that socio-demographic factors are more relevant for explaining differences across performance levels, while recent academic history plays a stronger role in determining simple course completion. Behavioral variables such as goout or schoolsup also appear selectively, depending on whether the outcome captures detailed performance variation or only overall success. Overall, these patterns indicate that feature selection not only reduces dimensionality, but also helps reveal how different definitions of achievement highlight different aspects of students’ academic and social context.
3.2. Classification Accuracy Under Different Feature Selection Methods
This section examines in detail how each combination of feature-selection strategy and ML algorithm performs across three outcomes defined at different levels of granularity derived from the final grade
G3 (
G3_level,
G3_grade,
G3_pass/fail). Accuracy scores are macro-averaged across ten-fold stratified cross-validation and compared under four feature-selection conditions: no filtering, CFS, CorrEval, and InfoGain. The ML algorithms employed in this study, as described in
Section 2.4, include Naive Bayes, Logistic Regression, Multilayer Perceptron (MLP), Support Vector Machine trained via SMO, k-Nearest Neighbors (IBk,
k = 5), Bagging, J48 decision trees, Random Forest, and Random Tree. Model evaluation was intentionally designed to go beyond overall accuracy by incorporating class-wise precision, recall, F1-score, AUC, confusion matrices, statistical testing, and sensitivity analyses.
Table 4 and
Figure 3 summarize and compare classification accuracies obtained by the nine ML algorithms under four feature selection settings: no feature selection, CFS, CorrEval, and InfoGain for
G3_level. For each classifier, the configuration yielding the highest accuracy is highlighted in
Table 4. Although
Figure 3 uses line plots for visualization, the categories along the horizontal axis represent nominal variables without inherent ordering. The plots are intended to facilitate comparison of performance across configurations, and no sequential relationship between categories is implied.
Overall, the results demonstrate that feature selection substantially improves predictive performance for most classifiers, particularly for algorithms that are sensitive to feature dimensionality and noise. Performance without feature selection was generally moderate, with most algorithms achieving accuracy levels between 78% and 83%. Notably, IBk performed poorly (47.61%), illustrating the vulnerability of distance-based techniques to irrelevant attributes. Feature selection markedly improved performance. Logistic Regression rose from 83.70% to 88.97% with CFS or CorrEval, while MLP jumped from 79.69% to 88.97% under CFS. These gains suggest that dimensionality reduction sharpens the decision boundaries for both linear and neural models.
Ensemble methods achieved the best results. Bagging with CorrEval reached 90.47%, while Random Forest with InfoGain attained 90.72%, surpassing benchmarks reported in similar educational data mining studies (i.e., Wahdan et al. [
53]). Decision tree models remained competitive. J48 achieved 89.97% with CorrEval, confirming that interpretable classifiers can rival black-box ensembles when paired with effective feature filtering. IBk benefited most dramatically, improving by over 40 percentage points when coupled with CFS (from 47.61% to 87.96%), underscoring that irrelevant dimensions disproportionately distort distance computations. Such a finding corroborates the claim by Hall and Holmes [
54] that distance metrics suffer from redundant dimensions. The results indicate that combining ensemble models with correlation-based filtering yields highly reliable classifiers, while simpler learners can also achieve strong performance when noise is reduced. The findings further show that both the feature selection technique and the classification algorithm have a significant influence on model performance. Ensemble methods and decision trees, especially Bagging and Random Forest, consistently achieved high accuracy, whereas algorithms such as IBk were more sensitive to feature reduction strategies.
To complement the accuracy-based evaluation, class-wise precision, recall, F1-scores, AUC values and confusion matrix (
Table 5) were computed for the best-performing configuration (Random Forest with InfoGain) on the
G3_level prediction task. The model demonstrates strong and balanced performance across all classes, with precision, recall, and F1-scores ranging from 0.87 to 0.92. The macro-averaged F1-score (≈0.90) indicates consistent predictive efficiency across the three levels. The confusion matrix reveals that most misclassifications occur between adjacent categories, particularly for medium-performing students, suggesting that borderline academic profiles are more difficult to distinguish. In contrast, the extreme categories (low and high performance) exhibit higher recall (0.91 and 0.92, respectively), indicating more reliable identification of these groups. The high macro-averaged AUC (0.95) further indicates strong discriminative ability across classes. Overall, these results confirm that Random Forest with InfoGain provides a robust and well-balanced classification of student performance levels.
Table 6 and
Figure 4 summarize the classification accuracy obtained by the nine ML algorithms under the four feature selection settings for the four-level categorization of the final grade (
G3_grade). The best-performing configuration for each classifier is highlighted in
Table 6.
The four-category grade prediction task proved more challenging, as expected, with accuracies dropping relative to the three-class case. Performance without feature selection declined for most models. Logistic Regression (74.18%), MLP (72.68%), and SVM (74.43%) struggled to distinguish between the four grade bands, reflecting the greater complexity of this task. Feature selection mitigated this difficulty. Logistic Regression improved to 85.46% with CFS, while MLP increased to 83.70%. Naive Bayes also showed a modest improvement, rising from 78.94% to 81.20% under InfoGain.
Ensemble and tree-based methods produced the strongest and most consistent performance across feature selection settings. Bagging maintained stable performance at approximately 87% across all configurations, supporting Dietterich’s [
55] findings on variance reduction through bootstrap aggregation. Random Forest achieved a maximum accuracy of 84.71% with CFS, while J48 reached 86.90% under CorrEval. Random Tree also showed substantial improvement, increasing from 63.65% without filtering to 81.95% with CFS [
56].
Instance-based learning was particularly sensitive to irrelevant features. IBk improved from a near-random accuracy of 40.85% to 77.44% when CorrEval filtering was applied. In contrast, SVM showed only moderate improvement, increasing from 74.43% to 78.94%, suggesting that kernel-based models may require more extensive hyperparameter tuning to fully benefit from feature reduction.
Taken together, the results highlight the importance of feature selection in more granular grade prediction tasks. CFS and CorrEval provided the most consistent improvements across classifiers, while ensemble and tree-based methods maintained strong overall performance. Although the four-level prediction task is inherently more challenging, appropriate feature selection helps mitigate the loss in accuracy.
Table 7 provides class-wise evaluation metrics and confusion matrix for the best-performing configuration (Bagging with CorrEval) on
G3_grade. The classifier shows stable performance across the four grade categories, with precision, recall, and F1-scores ranging from 0.82 to 0.92. The macro-averaged F1-score (≈0.86) indicates balanced performance despite the increased complexity of the four-level prediction task. The confusion matrix shows that most instances are correctly classified, particularly in the Fail and Excellent categories, which achieve the highest recall (0.92 and 0.87, respectively). In contrast, misclassifications mainly occur between adjacent grade levels, especially between Satisfactory and Good, reflecting the difficulty of distinguishing borderline cases. The high macro-averaged AUC (0.95) further indicates strong discriminative ability across classes. Overall, these results suggest that Bagging with CorrEval provides a robust and well-balanced multiclass prediction of final grade outcomes.
Table 8 and
Figure 5 present the classification accuracy of the nine ML algorithms across the four feature selection settings for
G3_pass/fail. The best-performing configuration for each classifier is highlighted in
Table 8.
As expected, binary classification produced the strongest results overall, with high accuracy observed even without feature selection for most ML algorithms. Logistic Regression achieved 89.97% without filtering and improved to 90.47% with both CFS and CorrEval, demonstrating both strong performance and stability. MLP also maintained high accuracy across all configurations (87.71–89.22%), reflecting its robustness to different feature subsets. Random Forest reached 90.72%, further confirming that the pass/fail distinction is relatively easier to model.
Across most classifiers, accuracy was high and generally improved after applying feature selection techniques [
2]. Bagging with InfoGain and J48 with CorrEval both achieved 92.48%, the highest values across all tasks, suggesting strong performance of ensemble and tree-based methods for binary outcomes. Random Forest, another ensemble method, reached 91.22% with CorrEval and remained above 88% across all configurations, confirming its effectiveness. Probabilistic and margin-based classifiers also benefited: Naive Bayes peaked at 87.96% with InfoGain, while SVM improved from 86.96% to 88.97% with CFS. IBk showed the largest gain, increasing from 65.91% without feature selection to 88.72% with CFS, highlighting its sensitivity to irrelevant or noisy features [
37].
Overall, feature selection improved or preserved performance across most classifiers. Ensemble methods (Bagging and Random Forest) achieved the highest accuracy, followed closely by Logistic Regression and MLP. While binary tasks are inherently simpler, ensemble classifiers provided an additional advantage, reaching accuracies above 92%. These findings also emphasize the importance of feature selection, particularly for algorithms such as IBk, which are more sensitive to high-dimensional data.
Table 9 presents class-wise evaluation metrics and the confusion matrix for the best-performing configuration (Bagging with InfoGain) on the
G3_pass/fail prediction task.
The model achieves strong predictive performance, with a macro-averaged F1-score of approximately 0.92 and an AUC of 0.96, indicating excellent discrimination between the two classes. The confusion matrix shows that most instances are correctly classified, with particularly high recall for the Fail class (0.95), correctly identifying 127 out of 133 students. Only a small number of failing students are misclassified as passing, reducing the risk of overlooking at-risk individuals. For the Pass class, the model also performs well, with high precision (0.98) and recall (0.92). Most errors occur when passing students are predicted as failing, slightly increasing false positives while maintaining strong detection of failing cases. Overall, the high F1-scores and AUC values indicate a well-balanced model for distinguishing between passing and failing students.
Figure 6 summarizes the highest classification accuracy achieved by each ML algorithm across the three outcome definitions (
G3_level,
G3_grade, and
G3_pass/fail), considering all feature selection configurations.
The comparison of results across the three outcome variables (G3_level, G3_grade, and G3_pass/fail) reveals several important trends related to algorithm performance, feature selection, and predictive complexity. Ensemble methods such as Bagging and Random Forest consistently achieved the highest accuracy in both binary and multiclass tasks. Bagging performed best in the binary setting and also showed strong results in multiclass scenarios, especially when combined with CorrEval for G3_level and G3_grade. These findings highlight the interaction between outcome granularity, feature selection, and classifier choice in student performance prediction.
Feature selection proved beneficial across all tasks, particularly for algorithms sensitive to high-dimensional data, such as IBk and SVM (SMO). IBk, which performed poorly without feature selection, showed substantial accuracy gains when combined with CFS and CorrEval, highlighting its reliance on a reduced and relevant feature space. Similarly, SVM showed more stable and improved results after feature selection, suggesting that its decision boundaries benefit from noise reduction and more informative attributes.
Logistic Regression and MLP classifiers exhibited robust and stable performance across all representations of G3. Logistic Regression showed notable improvement after feature selection for both G3_level and G3_grade. MLP followed closely, demonstrating consistent adaptability, especially with CFS and CorrEval, although it slightly trailed Logistic Regression in some cases.
Naive Bayes, although not the top-performing model, demonstrated consistent moderate accuracy. It benefited from feature selection, particularly InfoGain and CorrEval, achieving its best results in the G3_pass/fail and G3_level tasks. These findings suggest that, despite its simplicity, Naive Bayes remains a useful baseline for educational data classification when combined with appropriate preprocessing.
In terms of outcome complexity, G3_pass/fail, as a binary task, achieved the highest accuracy across all classifiers, indicating that the simpler decision boundary is easier for models to learn. In contrast, G3_grade, with four distinct categories, posed greater challenges and resulted in lower performance, even for top-performing algorithms. G3_level, with three ordinal classes, showed intermediate performance levels between the binary and four-class settings. Overall, prediction was easiest for the binary pass/fail task (>92%), moderately difficult for the three-level classification (≈90%), and most challenging for the four-grade task (≈87%).
Among the feature selection methods, CorrEval consistently led to the highest or near-highest accuracy, particularly when combined with Bagging and J48. InfoGain also produced strong results, although its impact was less pronounced in some configurations, while remaining particularly effective in binary tasks. CFS often resulted in noticeable improvements, especially for algorithms that do not perform internal feature weighting or selection [
13].
To assess whether performance differences among classifiers were statistically significant, a Friedman test was applied to fold-level cross-validation accuracies under the InfoGain setting. The test compares the performance of multiple classifiers across repeated cross-validation folds. For all three target variables (
G3_level,
G3_grade, and
G3_pass/fail), the test indicated statistically significant differences (
χ2 = 38.26,
p < 0.001;
χ2 = 29.86,
p < 0.001; and
χ2 = 25.46,
p < 0.01, respectively), indicating that at least one classifier differed significantly in performance.
Table 10 presents the average classifier ranks derived from this analysis. Logistic Regression and Random Forest achieved the best overall ranks across the three targets, indicating stable performance. Ensemble techniques generally achieved higher ranks than single-tree models, while Naive Bayes consistently ranked lowest.
In addition to individual classifier comparisons, the results are summarized using two radar charts that provide an aggregated view of model performance across outcome granularities.
Figure 7 contrasts ensemble (Bagging, Random Forest) and non-ensemble models (all other models), highlighting differences in average classification accuracy across the three outcomes.
Figure 8 extends this analysis by grouping classifiers into four model families: ensemble (Bagging, Random Forest), tree-based (J48, Random Tree), linear (Logistic Regression), and instance-based (IBk), offering a higher-level comparison of learning paradigms. Together, these visualizations complement the classifier-level results by emphasizing overall performance trends and robustness patterns across outcome definitions.
As shown in the figures, ensemble models generally achieve the highest average accuracy across all outcome definitions (G3_level, G3_grade, and G3_pass/fail), demonstrating strong robustness to changes in outcome granularity. In contrast, linear and instance-based models show greater variability, particularly for more complex outcome formulations, while tree-based approaches exhibit more stable intermediate performance.
Key features remain stable across methods. Prior academic performance (
G1,
G2) and failure history (
failures) appear in all top-performing models, reaffirming their central role in predicting achievement and aligning with the academic-momentum theory of Sticca et al. [
57]. The consistent presence of socio-supportive features (
schoolsup,
higher,
address,
nursery) further highlights that environmental and aspirational factors may also influence performance trajectories. From a practical perspective, ensemble models such as Bagging and Random Forest can serve as high-accuracy components of early warning systems, while interpretable models like J48 provide transparent decision rules for policy applications.
Overall, the findings indicate that ensemble classifiers provide a robust modeling approach for educational datasets, particularly when combined with attribute evaluation techniques such as CorrEval and InfoGain. While binary outcomes are easier to predict, with accuracies exceeding 90%, multiclass outcomes, such as G3_grade and G3_level, require more careful feature selection and model design to achieve comparable performance. The consistent performance of Bagging and J48 across all grade representations further supports their suitability for educational data mining applications.
Beyond predictive performance, the use of socio-demographic attributes in educational models raises important considerations regarding fairness and potential bias. Variables such as gender, parental education, and family background can enhance predictive robustness, but they also reflect structural socio-economic conditions that are largely beyond students’ control. The observed distributional patterns indicate that lower academic outcomes are more prevalent among students from less advantaged backgrounds, reflecting well-documented educational inequalities rather than differences in individual ability. If used without careful interpretation, such models may inadvertently reproduce or amplify these inequalities. In this study, socio-demographic attributes were included solely for comparative purposes, to examine their predictive contribution across outcome granularities, and not for automated decision-making. It should also be noted that the present study does not include a formal fairness evaluation across demographic groups. The results also show that prior academic performance remains the dominant predictor, particularly in binary settings, while socio-demographic variables play a secondary and context-dependent role. Future applications should include fairness-aware evaluation and bias mitigation strategies to ensure equitable outcomes. When used responsibly, predictive models can support early identification of at-risk students, complementing, rather than replacing, human judgment in educational decision-making.
3.4. Early Prediction Experiments Excluding Prior Grades
To examine the early intervention potential of the proposed framework, additional experiments were conducted after excluding the prior grade variables G1 and G2 from the feature set. As these variables represent intermediate assessments close to the final outcome (G3), their inclusion may reduce the practical applicability of predictive models for early identification of at-risk students. Removing these variables enables the evaluation of models based solely on demographic, behavioral, and contextual attributes available at earlier stages of the academic process. To ensure consistency, the same preprocessing procedures, feature selection methods (CFS, CorrEval, and InfoGain), and ML algorithms were applied. Model evaluation was again performed using stratified ten-fold cross-validation.
Considering feature selection for the G3_level prediction task, all methods consistently identified failures as a highly informative attribute, underscoring the strong influence of prior academic difficulties on future performance. The CFS method selected a compact subset consisting of failures, medu, and higher, suggesting that parental education and students’ intention to pursue higher education are important indicators of academic achievement. CorrEval identified failures, age and goout, indicating that maturity and social activity patterns may also contribute to variation in performance levels. InfoGain produced a slightly larger feature set, including failures, mjob, medu, higher, fjob, and schoolsup, highlighting the potential influence of family background and additional academic support in explaining variations in student outcomes.
For the G3_grade outcome, the selected features varied more substantially across the methods, reflecting the increased complexity of predicting four distinct grade categories. CFS emphasized contextual and social factors such as address, nursery, and romantic, along with failures, suggesting that environmental and social conditions may influence finer distinctions in academic performance. In contrast, CorrEval highlighted demographic and behavioral attributes including medu, age, goout, and failures, while InfoGain prioritized family and support-related variables such as mjob, fjob, schoolsup, higher, and failures.
For the G3_pass/fail outcome, the CFS method retained failures, famsize, reason, schoolsup, internet, and goout, indicating the relevance of family background, school motivation, academic support, and student engagement. CorrEval selected failures, age, goout, higher, and paid, emphasizing behavioral patterns and educational aspirations. InfoGain produced a more compact subset consisting of failures, goout, and age, suggesting that these variables provide the strongest predictive contribution when intermediate grades are unavailable. Notably, failures and goout were consistently identified across methods, highlighting the importance of accumulated academic difficulty and behavioral engagement for early-stage prediction.
Table 12 summarizes the classification accuracy obtained by the nine ML algorithms across different feature selection methods (without feature selection, CFS, CorrEval, and InfoGain) for
G3_level after excluding prior grades (
G1 and
G2) from the feature set. The best-performing configuration for each classifier appears highlighted.
The overall predictive performance decreases substantially compared to experiments that include prior academic indicators. Accuracy across all algorithms and configurations ranges between approximately 41% and 52%, highlighting the strong influence of previous grades on student achievement prediction. Despite this reduction, feature selection consistently improves model performance compared to using the full feature set, with the highest accuracies typically achieved under InfoGain. Ensemble methods demonstrate relatively stable performance under this setting, with Bagging reaching the highest accuracy of 52.41%. In contrast, instance-based learning (IBk) and Random Tree demonstrate lower predictive performance, although they also benefit from feature filtering.
Table 13 presents class-wise evaluation metrics and the confusion matrix for the best-performing configuration (Bagging with InfoGain) on the
G3_level prediction task without
G1 and
G2. Overall performance decreases compared to models that include prior academic information, with a macro-averaged F1-score of approximately 0.70 and an AUC of 0.77, indicating moderate discriminative capability. The model performs best for the High-performance category, achieving the highest precision, recall, and F1-score. In contrast, the Low and Medium categories show lower and more balanced scores, reflecting greater difficulty in distinguishing these levels without prior grade information. The confusion matrix indicates that most errors occur between adjacent categories, particularly between Low and Medium and between Medium and High, suggesting increased ambiguity in predicting borderline cases.
Table 14 presents the classification accuracy of the evaluated ML algorithms for the
G3_grade task after excluding prior grades (
G1 and
G2), with the best-performing feature selection configuration highlighted. Overall, performance declines substantially compared to models that include prior academic information, with accuracy ranging from approximately 39% to 49% across classifiers. Despite this reduction, feature selection consistently improves results, with InfoGain and CorrEval generally yielding the highest accuracy. The best performance is achieved by Logistic Regression and Bagging with InfoGain, both reaching 48.61%. In contrast, IBk and Random Tree remain among the weaker performers. These findings further confirm the dominant predictive role of prior grades.
Table 15 presents class-wise evaluation metrics and the confusion matrix for the best-performing configuration (Bagging with InfoGain) on the
G3_grade task without
G1 and
G2. Overall performance decreases compared to models that include prior academic indicators, with a macro-averaged F1-score of 0.56 and an AUC of 0.73, indicating moderate predictive ability. The Excellent category shows the strongest performance, suggesting that top-performing students can still be identified relatively reliably. In contrast, the Fail, Satisfactory, and Good categories exhibit lower and similar F1-scores, reflecting greater difficulty in distinguishing intermediate levels. The confusion matrix indicates that most errors occur between adjacent categories, particularly between Satisfactory and Good, and between Good and Excellent, highlighting increased ambiguity in predicting more precise grade distinctions without prior grade information.
Table 16 summarizes the classification accuracy of the nine ML algorithms for the
G3_pass/fail outcome without prior grades. Despite the absence of these predictors, the models retain meaningful performance, particularly when feature selection is applied. Several classifiers achieve accuracy levels around or above 70%, especially with InfoGain and CorrEval. Logistic Regression, SVM, IBk, and tree-based models demonstrate competitive results, while ensemble methods remain relatively stable.
Table 17 presents class-wise evaluation metrics and the confusion matrix for the best-performing configuration (Bagging with InfoGain) on the
G3_pass/fail task without
G1 and
G2. Overall performance decreases compared to models that incorporate prior academic information, with a macro-averaged F1-score of 0.71 and an AUC of 0.78, indicating moderate discriminative ability. The model performs better for the Pass class, correctly identifying most passing students. In contrast, the Fail class shows lower precision but relatively high recall, indicating that most failing students are detected, although some passing students are misclassified. The confusion matrix indicates that most errors occur when passing students are predicted as failing, reflecting increased uncertainty in predicting student outcomes without prior grade information.
The results across the three grade granularities indicate that, in the absence of prior grades (G1 and G2), model performance declines substantially. Nevertheless, for binary risk identification, demographic, behavioral, and school-related variables still provide meaningful predictive signals. This finding suggests that the proposed framework retains practical relevance for early intervention, particularly for identifying students at risk of failing before intermediate grades become available.
Although prior academic indicators such as G1, G2, and previous failures significantly improve predictive accuracy, their proximity to the final outcome limits their usefulness for early intervention. The aim of this study was not to develop models based solely on pre-course or demographic variables, but to examine how feature selection and ML methods behave across different grade granularities. The experiments excluding prior grades therefore provide an exploratory assessment of model behavior without near-outcome academic indicators. The results highlight the distinction between models optimized for predictive accuracy and those designed for early warning systems, which must rely on information available at earlier stages of the academic process.