Lukman Azhari
Universitas Muhammadiyah Tangerang, Tangerang

Published : 2 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 2 Documents
Search

Pendekatan Hybrid K-Means SMOTE dan Logistic Regression Untuk Deteksi Dini Diabetes Mellitus Pada Imbalanced Data Abdus Salam; Lukman Azhari; Ri Sabti Septarini; Nofitri Heriyani
Bulletin of Computer Science Research Vol. 5 No. 3 (2025): April 2025
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bulletincsr.v5i3.502

Abstract

The increasing global prevalence of Diabetes Mellitus necessitates more accurate early detection efforts, particularly through machine learning-based approaches. However, one of the main challenges in medical classification lies in data imbalance, where the number of diabetic cases is significantly lower than that of non-diabetic ones. This study aims to develop a hybrid model by integrating Logistic Regression and K-Means SMOTE to enhance the sensitivity of early detection for Diabetes Mellitus, especially toward the minority class. Logistic Regression is chosen for its computational efficiency and interpretability, while K-Means SMOTE plays a role in balancing class distribution by generating synthetic samples in a structured manner based on clusters of minority class data. The dataset used consists of 2,000 records with 9 health-related features, obtained from the Kaggle platform. Evaluation results indicate that the model utilizing K-Means SMOTE achieves the best performance, with an accuracy of 82.00%, an F1-score of 72.73% for the Diabetes class, and the highest ROC-AUC score of 87.48%. Compared to models without oversampling and with standard SMOTE, this approach improves model generalization and sensitivity to positive cases. These findings have practical implications for the development of fairer and more effective machine learning-based early detection systems, particularly for implementation in healthcare facilities with limited resources.
Model Prediksi Penyakit Jantung dengan Penanganan Outlier Menggunakan Interquartile Range dan Extreme Gradient Boosting Lukman Azhari; Novi Wulandari; Feru Adiningrat; Allan Desi Alexander
Journal of Information System Research (JOSH) Vol 6 No 2 (2025): January 2025
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/josh.v6i2.6390

Abstract

Heart disease remains one of the leading causes of death worldwide, with increasing prevalence rates, including in Indonesia. Delayed detection and diagnosis are the main challenges in treating this disease, as most cases are only identified after patients experience serious symptoms or heart attacks. Medical data often containing outliers and noise adds to the complexity of developing accurate predictive models. This study aims to develop a heart disease prediction model using a combination of the Interquartile Range (IQR) method for outlier handling and the Extreme Gradient Boosting (XGBoost) algorithm for predictive modeling. The IQR method is applied at the pre-processing stage to identify and eliminate outliers robustly without reducing data integrity, while XGBoost is used to build an efficient prediction model through an ensemble learning approach. The results showed significant improvements in model performance, with accuracy increasing from 75.41% to 89.47% and AUC-ROC from 0.8615 to 0.9450. The model demonstrates balanced predictive capabilities with precision of 95.24% and recall of 80.00% for cases without disease, and precision of 86.11% and recall of 96.88% for cases with disease. The developed model makes significant contributions by improving data quality through robust outlier handling using the IQR method, building a more accurate prediction model by leveraging the advantages of the XGBoost algorithm in the ensemble learning approach.