Student dropout is a persistent challenge in higher education, and predictive models can support early identification of students who may require academic or financial intervention. This study develops an explainable multiclass machine learning approach to predict three academic outcomes—Dropout, Enrolled, and Graduate—using the public Predict Students' Dropout and Academic Success dataset containing 4,424 student records. Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost) were compared using a stratified 80:20 hold-out design. XGBoost hyperparameters were optimized through randomized search with five-fold stratified cross-validation, and SHapley Additive exPlanations (SHAP) were used to interpret global and class-specific predictions. Random Forest achieved the highest overall accuracy of 77.18%, whereas the optimized XGBoost model produced the highest macro recall of 69.58% and macro F1-score of 70.08%. XGBoost improved recall for the minority Enrolled class to 46.54%, compared with 38.36% for Random Forest and 33.33% for Logistic Regression. SHAP analysis identified the number of curricular units approved in the second and first semesters, tuition-fee status, course, second-semester grade, and age at enrollment among the most influential predictors. Low academic progression and unpaid tuition status contributed strongly toward Dropout predictions, while stronger academic progression shifted predictions toward Graduate. These findings show that explainability complements predictive performance by revealing actionable patterns behind multiclass student-outcome predictions.Student dropout is a persistent challenge in higher education, and predictive models can support early identification of students who may require academic or financial intervention. This study develops an explainable multiclass machine learning approach to predict three academic outcomes—Dropout, Enrolled, and Graduate—using the public Predict Students' Dropout and Academic Success dataset containing 4,424 student records. Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost) were compared using a stratified 80:20 hold-out design. XGBoost hyperparameters were optimized through randomized search with five-fold stratified cross-validation, and SHapley Additive exPlanations (SHAP) were used to interpret global and class-specific predictions. Random Forest achieved the highest overall accuracy of 77.18%, whereas the optimized XGBoost model produced the highest macro recall of 69.58% and macro F1-score of 70.08%. XGBoost improved recall for the minority Enrolled class to 46.54%, compared with 38.36% for Random Forest and 33.33% for Logistic Regression. SHAP analysis identified the number of curricular units approved in the second and first semesters, tuition-fee status, course, second-semester grade, and age at enrollment among the most influential predictors. Low academic progression and unpaid tuition status contributed strongly toward Dropout predictions, while stronger academic progression shifted predictions toward Graduate. These findings show that explainability complements predictive performance by revealing actionable patterns behind multiclass student-outcome predictions.
Copyrights © 2026