Claim Missing Document
Check
Articles

Peningkatan Sensitivitas Model Boosting untuk Deteksi Diabetes Menggunakan SMOTE pada Imbalanced Dataset Setyawan Wibisono; Eko Nur Wahyudi; Imam Husni Al Amin
Dinamik Vol 31 No 2 (2026)
Publisher : Universitas Stikubank

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35315/dinamik.v31i2.10606

Abstract

This study aims to analyze the effect of SMOTE on the sensitivity of boosting models for diabetes detection using the BRFSS 2015 dataset. The dataset consists of 253,680 instances with 21 features and a binary target, namely diabetes and non-diabetes. The primary issue in the dataset is class imbalance, causing the models to be more biased toward recognizing the non-diabetes class. The algorithms employed in this study include AdaBoost, XGBoost, and Gradient Boosting, evaluated under two scenarios: without SMOTE and with SMOTE. Model performance was assessed using accuracy, precision, recall, F1-score, ROC-AUC, confusion matrix, and 10-fold cross validation. The results demonstrate that SMOTE improves recall across all models. The most significant improvement occurred in AdaBoost, where recall increased from 0.016551 to 0.711840. The cross-validation results also showed that AdaBoost + SMOTE achieved a recall value of 0.721384. Although accuracy and precision decreased, AdaBoost + SMOTE became the most sensitive model for detecting diabetes. Therefore, this model has potential to be utilized as an early diabetes screening support tool.
Ensemble Learning untuk Klasifikasi Penyakit Kardiovaskuler dengan Optimasi Hyperparameter menggunakan Grid Search Eko Nur Wahyudi; Setyawan Wibisono
Dinamik Vol 31 No 2 (2026)
Publisher : Universitas Stikubank

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35315/dinamik.v31i2.10608

Abstract

This study aims to optimize ensemble learning models for cardiovascular disease classification based on patients’ clinical data using the Grid Search method. The dataset used consists of 70,000 patient records with clinical attributes such as age, gender, height, weight, blood pressure, cholesterol, glucose, smoking habits, alcohol consumption, physical activity, and cardiovascular disease status. After preprocessing, 68,606 records were utilized, with a relatively balanced class distribution. The algorithms employed in this study include LightGBM, AdaBoost, and Gradient Boosting. Model evaluation was conducted using accuracy, precision, recall, F1-score, ROC-AUC, confusion matrix, and 10-fold cross validation. The results indicate that Gradient Boosting achieved the best performance with a ROC-AUC score of 0.804400 on the testing data and 0.801757 on cross validation. This model also produced the highest recall and F1-score values. Therefore, Gradient Boosting optimized with Grid Search is considered suitable for cardiovascular disease classification based on clinical data.