Sinkron : Jurnal dan Penelitian Teknik Informatika
Vol. 10 No. 3 (2026): Article Research July 2026

Classifying Student Academic Achievement from Limited Categorical Institutional Records: A Comparative Study of Naive Bayes, K-Nearest Neighbor, and Decision Tree

Relita Buaton (Information System, Faculty of Computer Science, STMIK Kaputama, Binjai, Indonesia)
I Gusti Prahmana (Information System, Faculty of Computer Science, STMIK Kaputama, Binjai, Indonesia)
Siti Nur Azizah (Information System, Faculty of Computer Science, STMIK Kaputama, Binjai, Indonesia)
Elisiya Putri (Information System, Faculty of Computer Science, STMIK Kaputama, Binjai, Indonesia)
Windy Indah Sary Sinaga (Information System, Faculty of Computer Science, STMIK Kaputama, Binjai, Indonesia)



Article Info

Publish Date
05 Jul 2026

Abstract

Student academic achievement prediction is an important application in Educational Data Mining (EDM) that supports proactive academic decision-making. This study investigates a specific and underexplored condition in the literature the classification of student academic achievement when the only available predictors are categorical institutional background attributes  without behavioral, attendance, or course-level data. This condition reflects data infrastructure limitations commonly found in Indonesian private higher education institutions. Three widely used classification algorithms Naive Bayes (BernoulliNB), K-Nearest Neighbor (KNN), and Decision Tree (CART) are compared against a majority class baseline through a five-stage preprocessing pipeline encompassing label normalization, cohort feature extraction, KNN k-value sensitivity analysis, and reporting of balanced accuracy and macro F1-score for fair evaluation under mild class imbalance. Results show that Decision Tree (depth=5) achieved the highest balanced accuracy (57.77%) and macro F1-score (57.51%), while Naive Bayes demonstrated the best generalization stability based on 10-fold cross-validation (60.07% ± 6.02%). All three models substantially outperformed the majority class baseline on balanced accuracy (+5–8 percentage points) and macro F1-score (+19–21 percentage points). Feature importance analysis identified IPS prior major background (15.6%) and the 2020 cohort (14.4%) as the most discriminative features. These findings provide evidence based algorithm selection guidance for data-constrained institutions and establish a reproducible performance benchmark for the categorical attributes only classification condition.

Copyrights © 2026






Journal Info

Abbrev

sinkron

Publisher

Subject

Computer Science & IT

Description

Scope of SinkrOns Scientific Discussion 1. Machine Learning 2. Cryptography 3. Steganography 4. Digital Image Processing 5. Networking 6. Security 7. Algorithm and Programming 8. Computer Vision 9. Troubleshooting 10. Internet and E-Commerce 11. Artificial Intelligence 12. Data Mining 13. Artificial ...