Claim Missing Document
Check
Articles

Found 2 Documents
Search
Journal : building of informatics technology and science

Perbandingan Kinerja Random forest dan SVM Pada Klasifikasi Tingkat Kekumuhan Permukiman Menggunakan SMOTE Nurika Dwi Wahyuni; Fadhilah Syafria; Novi Yanti; Surya Agustian
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.10101

Abstract

Classifying slum levels is essential for a structured, data-driven analysis of settlement conditions. This study compares the performance of Random forest and Support vector machine (SVM) in classifying slum levels in Pekanbaru City across two scenarios with and without SMOTE using slum indicator scoring data. Its contributions include analyzing SMOTE's impact on model performance and evaluating the top 10 features against the full feature set. The dataset comprises 992 RT-level records from Disperkim Pekanbaru City (2020, 2021, and 2023) featuring 16 slum indicator scores based on PUPR Ministerial Regulation No. 14/2018, categorized into three classes: Non-Slum, Low Slum, and Moderate Slum. Following the KDD process (selection, preprocessing, transformation, data mining, evaluation, and analysis), the data was split 80:20 using stratified sampling and evaluated based on accuracy, precision, recall, F1-score, and confusion matrix. Results show that the Linear SVM without SMOTE achieved perfect evaluation metrics (1.0000); however, this is interpreted cautiously as the class labels derive from strict regulatory scoring rules, making class boundaries inherently linear. Random forest saw its F1-score rise from 0.9660 to 0.9700 after SMOTE, while the most significant improvement occurred in SVM RBF, jumping from 0.9214 to 0.9779. Testing the top 10 features led to a decreased F1-score across models, indicating that utilizing all 16 features remains optimal for this dataset.
Comparative Study of Agglomerative Hierarchical Clustering and K-Means for Student Academic Stress Grouping Irfan Arifin; Iwan Iskandar; Elvia Budianita; Novi Yanti; Fitri Insani
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.10265

Abstract

Academic stress is a common problem experienced by college students due to high academic demands, parental expectations, and social pressures during their college years. The high levels of academic stress experienced by students underscore the need for a data-driven approach to more accurately identify and map students’ stress levels. This research aims to compare the performance of the Agglomerative Hierarchical Clustering (AHC) and K-Means methods in clustering students’ academic stress levels and to determine which method produces the best clustering quality. Data were obtained from the distribution of the Perception of Academic Stress Scale (PAS) questionnaire, consisting of 18 statement items, with 361 valid respondents from the Informatics Engineering Program at UIN SUSKA Riau, class of 2022–2025. The selection of the best linkage method in AHC was performed using the Cophentic Correlation Coefficient (CCC), where Ward Linkage was selected with the highest CCC value of 0.8180. Comparative evaluation was conducted using the Silhouette Coefficient, Davies-Bouldin Index, and Calinski-Harabasz Index for variations in the number of clusters from K=2 to K=7. The test results showed that AHC Ward Linkage with K=2 was the best configuration with a Silhouette Coefficient of 0.4407 and a Davies-Bouldin Index of 0.8373, outperforming K-Means, which only excelled in the Calinski-Harabasz Index with a value of 419.7405 The clustering resulted in two clusters: High Stress with 244 students (67.6%) and Low Stress with 117 students (32.4%). The 2023 and 2024 cohorts had the highest proportions of high stress at 90.4% and 90.6%, respectively. This research contributes empirical evidence comparing hierarchy-based and partition-based clustering methods for academic stress data, while also demonstrating the use of the Cophenetic Correlation Coefficient as an objective basis for linkage method selection in AHC. It is hoped that the results of this study can serve as a basis for the institution in designing targeted mental health intervention programs for students.