Student dropout prediction is a key application of Educational Data Mining for supporting early intervention in higher education. However, previous studies have primarily focused on improving predictive accuracy, while the effects of data balancing on decision bias and algorithmic fairness remain underexplored. This study proposes a comprehensive evaluation framework that integrates predictive performance, decision bias, and algorithmic fairness to assess the impact of the Synthetic Minority Over-sampling Technique (SMOTE) on Extreme Gradient Boosting (XGBoost) for student dropout prediction. Experiments were conducted using the publicly available Predict Students Dropout and Academic Success dataset containing 4,424 student records. After excluding the Enrolled class, the dataset was transformed into a binary classification problem consisting of 2,209 Graduate (60.9%) and 1,421 Dropout (39.1%) instances. Two models were compared: a baseline XGBoost classifier and an XGBoost classifier trained with SMOTE. Predictive performance was evaluated using Accuracy, Precision, Recall, F1-score, and ROC-AUC, while decision bias and algorithmic fairness were assessed using the False Negative Rate (FNR), Statistical Parity Difference (SPD), Disparate Impact (DI), Equal Opportunity Difference (EOD), and Average Odds Difference (AOD). The baseline model achieved higher Accuracy (93.11% vs. 92.29%), Precision (91.49% vs. 89.86%), F1-score (91.17% vs. 90.18%), and a lower FNR (0.0915 vs. 0.0951), whereas both models produced comparable ROC-AUC values (0.972). McNemar's test indicated that the difference in predictive performance was not statistically significant (p = 0.264). Although SMOTE did not improve predictive performance, it produced modest reductions in Statistical Parity Difference (0.2310–0.2218), Equal Opportunity Difference (0.0298–0.0233), and Average Odds Difference (0.0295–0.0235), indicating a slight improvement in fairness metrics while maintaining comparable discrimination capability. These findings highlight the trade-off between predictive performance and algorithmic fairness and demonstrate that evaluating predictive performance together with decision bias and fairness provides a more comprehensive assessment of educational machine learning models, supporting the development of responsible AI-based educational decision-support systems.
Copyrights © 2026