Outcome-Based Education (OBE) implementation requires continuous evaluation of graduate learning outcomes (CPL). While machine learning models are widely used for predicting academic performance, most studies overlook fundamental issues like data leakage and class imbalance. This study evaluates the impact of data leakage audits on the performance of four models (Logistic Regression, Random Forest, XGBoost, and Deep Neural Network) in predicting CPL. The novelty lies in applying a structured audit framework prior to model comparison. Experiments utilized 27,262 academic records with stratified cross-validation. Audit results proved that proxy features caused unrealistic performance (Accuracy 0.999). After removing leaked features, Logistic Regression achieved the highest discrimination stability (AUC 0.772), while DNN recorded the highest F1-score. Wilcoxon tests confirmed no statistically significant performance difference among the models (α=0.05). In conclusion, on leakage-free OBE data, simple linear models remain highly competitive and suitable as the foundation for early warning systems.
Copyrights © 2026