Claim Missing Document
Check
Articles

Found 2 Documents
Search

Analisis Klasifikasi Multikelas Obesitas Menggunakan Algoritma Decision Tree Classifier, Random Forest Classifier, dan Support Vector Classifier (SVC) Septa, Oon; Triyasri, Novita; Permata, Maharani Aulia; Salsabila, Aghitsna; Firdani, Fahri
Journal of Data Science Methods and Applications Vol. 2 No. 1 (2026)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Obesity remains a significant global health challenge, making early classification and detection essential to minimize the risk of more serious degenerative diseases. This study compares three machine learning algorithms—Random Forest, Decision Tree, and Support Vector Classifier (SVC) to determine which model is most effective in predicting weight status categories and obesity risks. The dataset used includes physical features such as age, height, weight, and Body Mass Index (BMI). The research process encompasses data preprocessing stages, feature correlation analysis, data splitting, model training, and evaluation using various performance metrics.The research results indicate that Random Forest demonstrates the highest discriminative ability with an AUC of 0.99, showing perfect accuracy in distinguishing between obesity categories. Decision Tree provides identical results in terms of accuracy at 95.45%, but with a slightly lower AUC value of 0.96. Meanwhile, SVC yields competitive results with an accuracy of 81.82%, although its performance remains below the two tree-based models. Overall, this study demonstrates that ensemble methods such as Random Forest hold great potential for use as decision support systems in detecting and classifying obesity status more accurately and reliably.
Komparasi Model Machine Learning dalam Memprediksi Penyakit Jantung dengan Pengoptimalisasian Hyperparameter Tunning Permata, Maharani Aulia; Saprianti, Assyifa; Chandra, Aurea Ivana; Yuni, Sundari Putri; Salsabila, Aghitsna; Ningrum, Margareta Oktavia
Journal of Data Science Methods and Applications Vol. 2 No. 1 (2026)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Heart disease remains one of the leading causes of death worldwide, making early detection efforts essential to minimize more serious risks. This study compares four machine learning algorithms—Random Forest, CatBoost, LightGBM, and XGBoost—to determine which model is most effective in predicting heart disease risk. The dataset used was sourced from Kaggle, comprising a total of 918 data points and 12 clinical features related to cardiovascular conditions. The research process included data pre-processing, class balancing using SMOTE, data partitioning, model training with hyperparameter tuning, and evaluation using various performance metrics. The results showed that Random Forest had the highest discriminatory ability with an AUC value of 0.9385. CatBoost, on the other hand, showed the most stable performance with an accuracy of 0.91 after tuning, and had balanced precision and recall in both classes. LightGBM and XGBoost also provided competitive results, although they were still slightly below the two best models. Overall, this study shows that ensemble methods such as Random Forest and CatBoost have great potential for use as decision support in detecting heart disease earlier and more accurately