Claim Missing Document
Check
Articles

Found 2 Documents
Search

Prediksi Tingkat Pengetahuan Mahasiswa Menggunakan Logistic Regression dan Random Forest Triyasri, Novita; Andini, Rekha Apriliana; Chandra, Aurea Ivana; Saprianti, Assyifa; Yusiandra, Erick
Journal of Data Science Methods and Applications Vol. 2 No. 1 (2026)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

particularly for predicting students’ knowledge levels more accurately. This study aims to compare the performance of Logistic Regression and Random Forest algorithms in predicting students’ knowledge levels using the User Knowledge Modeling dataset. The dataset consists of 258 instances with five numerical attributes, namely STG, SCG, STR, LPR, and PEG, and one target variable UNS representing students’ knowledge levels. The reserch stages include data selection, preprocessing, data normalization, train-test splitting, and handling class imbalance using the SMOTE method. Model performance is evaluated using accuracy, precision, recall, and F1-score metrics. The results show that Logistic Regression outperforms Random Forest, achieving higher accuracy and F1-score values. These findings indicate that the relationships among variables in the dataset tend to be linear. Therefore, Logistic Regression is considered more suitable for predicting students’ knowledge levels in this study.
Komparasi Model Machine Learning dalam Memprediksi Penyakit Jantung dengan Pengoptimalisasian Hyperparameter Tunning Permata, Maharani Aulia; Saprianti, Assyifa; Chandra, Aurea Ivana; Yuni, Sundari Putri; Salsabila, Aghitsna; Ningrum, Margareta Oktavia
Journal of Data Science Methods and Applications Vol. 2 No. 1 (2026)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Heart disease remains one of the leading causes of death worldwide, making early detection efforts essential to minimize more serious risks. This study compares four machine learning algorithms—Random Forest, CatBoost, LightGBM, and XGBoost—to determine which model is most effective in predicting heart disease risk. The dataset used was sourced from Kaggle, comprising a total of 918 data points and 12 clinical features related to cardiovascular conditions. The research process included data pre-processing, class balancing using SMOTE, data partitioning, model training with hyperparameter tuning, and evaluation using various performance metrics. The results showed that Random Forest had the highest discriminatory ability with an AUC value of 0.9385. CatBoost, on the other hand, showed the most stable performance with an accuracy of 0.91 after tuning, and had balanced precision and recall in both classes. LightGBM and XGBoost also provided competitive results, although they were still slightly below the two best models. Overall, this study shows that ensemble methods such as Random Forest and CatBoost have great potential for use as decision support in detecting heart disease earlier and more accurately