This Author published in this journals
All Journal Telematika
Eka Rahmawati
Sistem Informasi, Fakultas Teknik dan Informatika, Universitas Bina Sarana Informatika, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Interpretable Machine Learning for Early Detection of Academically At-Risk Students Using Behavioral, Socio-Demographic and Learning-Related Factors Vadlya Maarif; Eka Rahmawati; Candra Agustina
Telematika Vol 19, No 2: August (2026)
Publisher : Universitas Amikom Purwokerto

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35671/telematika.v19i2.3318

Abstract

The increasing adoption of learning analytics in higher education has encouraged the development of machine learning models for the early detection of academically at-risk students. However, many predictive models emphasize accuracy while providing limited interpretability for educators and academic advisors. This study proposes an interpretable machine learning framework for predicting academic risk using behavioral, socio-demographic, and learning-related student variables. A publicly available synthetic Kaggle dataset consisting of 500 student records was used as a controlled dataset for methodological validation. The data were preprocessed through missing-value handling, standardization, and one-hot encoding before being divided into training and testing sets. Several classifiers were evaluated, including Naive Bayes, Support Vector Machine, Random Forest, Logistic Regression, and XGBoost. Logistic Regression was employed as an interpretable baseline model, while XGBoost was used as a comparative ensemble classifier. Model performance was evaluated using accuracy, precision, recall, F1-score, AUC, and confusion matrix analysis. The results show that Logistic Regression achieved the highest accuracy and F1-score, with an accuracy of 0.8600 and an F1-score of 0.7812. XGBoost achieved the highest AUC value of 0.9278, followed closely by Logistic Regression with an AUC of 0.9246. Random Forest and XGBoost produced the lowest false negative values, indicating their ability to identify at-risk students more effectively. SHAP-based explainability revealed that weekly study hours, assignment completion, class attendance, and sleep duration were the most influential predictors. These findings suggest that interpretable machine learning can support academic early-warning systems by providing transparent prediction results. Since the dataset is synthetic, the findings should be interpreted as methodological validation rather than direct generalization to real student populations.