Building of Informatics, Technology and Science
Vol 8 No 1 (2026): June 2026

Komparasi Decision Tree, Random Forest, dan XGBoost untuk Deteksi Dini Risiko Akademik, Kehadiran, dan Karakter Siswa

M. Hafidhatul Fathoni (Universitas Sains dan Teknologi Indonesia, Pekanbaru)
Junadhi Junadhi (Universitas Sains dan Teknologi Indonesia, Pekanbaru)
Susanti Susanti (Universitas Sains dan Teknologi Indonesia, Pekanbaru)
Hadi Asnal (Universitas Sains dan Teknologi Indonesia, Pekanbaru)



Article Info

Publish Date
30 Jun 2026

Abstract

Schools routinely collect academic, attendance, and character data, but their use remains limited to administrative recording. As a result, changes in student conditions are not always transformed into early risk signals, such as academic decline, repeated absenteeism, or accumulated character-related violations. The problem addressed in this study is how to evaluate an early detection model that can capture these three risk dimensions simultaneously using imbalanced and temporally ordered monthly longitudinal data. This study compares Decision Tree, Random Forest, and XGBoost for detecting Academic Risk, Attendance Risk, and Character Risk in the following month. The dataset consists of longitudinal student records from SMKS YUM Pesantren Teknologi Riau in 2025, covering 206 students across 12 months. The raw dataset contained 2,472 rows and was reduced to 2,266 valid rows after feature engineering and time-shifting target construction. Model evaluation used rolling-origin temporal evaluation, majority baseline, temporal persistence baseline, and imbalance handling strategies: no balancing, class_weight, and partial SMOTE on training data. Macro F1 was used as the primary metric because the class distribution was imbalanced. The best models were Random Forest with class_weight for Academic Risk, with Macro F1 of 0.5005 and High-Risk recall of 0.6185; Decision Tree with class_weight for Attendance Risk, with Macro F1 of 0.4586 and High-Risk recall of 0.5339; and Random Forest with class_weight for Character Risk, with Macro F1 of 0.4761 and High-Risk recall of 0.5698. These findings indicate that machine learning models can provide early High-Risk signals, but predictions should be verified by teachers or homeroom teachers.

Copyrights © 2026






Journal Info

Abbrev

bits

Publisher

Subject

Computer Science & IT

Description

Building of Informatics, Technology and Science (BITS) is an open access media in publishing scientific articles that contain the results of research in information technology and computers. Paper that enters this journal will be checked for plagiarism and peer-rewiew first to maintain its quality. ...