Timely graduation is a key indicator of student success and institutional effectiveness in higher education. However, clustering student academic records containing mixed data types (numerical and categorical) using the conventional K-Means algorithm often leads to distance bias and reduced clustering quality due to the curse of dimensionality. This study proposes an optimized K-Means++ approach integrated with One-Hot Encoding and Principal Component Analysis (PCA) to improve clustering performance. The model was evaluated using 200 graduate records from STMIK El Rahma Yogyakarta. The results show that reducing the dataset to two principal components significantly enhances cluster quality. Validation metrics indicate that the Silhouette Score increased from 0.3275 to 0.4979, the Davies–Bouldin Index decreased from 1.405 to 0.871, and the Calinski–Harabasz Index improved from 70.448 to 168.035. The optimized model identified two distinct groups: Academically Stable Students (157 students) and At-Risk Working Students (43 students), the latter predominantly consisting of part-time employed students. These findings provide valuable insights for developing data-driven Academic Early Warning Systems (EWS) that enable higher education institutions to identify students at risk of delayed graduation and implement targeted intervention strategies.
Copyrights © 2026