Claim Missing Document
Check
Articles

Found 1 Documents
Search

Student Dropout Prediction Using XGBoost and Explainable AI (SHAP): A Case Study of Portuguese Student Dataset Chandra Kesuma; Vembria Rose Handayani
International Journal of Informatics, Economics, Management and Science Vol. 5 No. 2 (2026): IJIEMS (August 2026)
Publisher : Sekolah Tinggi Manajemen Informatika dan Komputer Jayakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.52362/ijiems.v5i2.2649

Abstract

Student dropout is one of the major challenges in higher education, as it can affect students’ academic continuity as well as the efficiency of institutional resource management. The availability of academic, administrative, demographic, and socioeconomic data provides opportunities to develop machine learning models capable of identifying students based on their dropout status. However, models with high predictive performance often have limited interpretability, making them difficult to use as an informative basis for higher education decision-makers. This study aims to develop a dropout status classification model using Extreme Gradient Boosting (XGBoost) and to explain the contribution of features to the model’s decisions using Explainable AI through the SHapley Additive exPlanations (SHAP) method. The dataset consists of 4,424 students from higher education institutions in Portugal, with 36 predictor variables. The original three-class target, consisting of Dropout, Enrolled, and Graduate, was transformed into a binary classification problem, with Dropout as the positive class and Graduate and Enrolled as the Non-Dropout class. The study compares Logistic Regression, Random Forest, Support Vector Machine (SVM), and XGBoost using accuracy, precision, recall, F1-score, and ROC-AUC. The results of stratified 5-fold cross-validation show that XGBoost achieves the best overall performance, with an accuracy of 88.07%, recall of 77.90%, F1-score of 80.76%, and ROC-AUC of 92.64%. On the testing data, XGBoost achieves an accuracy of 88.14%, precision of 83.03%, recall of 79.23%, F1-score of 81.08%, and ROC-AUC of 93.40%. SHAP analysis indicates that Curricular units 2nd sem (approved) is the feature with the greatest contribution to the model’s decisions, followed by Tuition fees up to date, Curricular units 1st sem (approved), Course, and Age at enrollment. Additional experiments removing first- and second-semester academic features resulted in a substantial decrease in model performance. These findings indicate that the model has strong predictive capability but is more appropriately used to predict dropout status based on students’ academic trajectories rather than being claimed as an early warning system from the beginning of their studies.