Ahmad Fatoni Dwi Putra
Department of Computer Science, Universitas Qamarul Huda Badaruddin Bagu, Central Lombok, Indonesia

Published : 2 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 2 Documents
Search

Enhancing Early Diabetes Detection Using Tree-Based Machine Learning Algorithms with SMOTEENN Balancing Syahrani Lonang; Ahmad Fatoni Dwi Putra; Asno Azzawagama Firdaus; Fahmi Syuhada; Yuan Sa'adati
Mobile and Forensics Vol. 8 No. 1 (2026)
Publisher : Universitas Ahmad Dahlan

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.12928/mf.v8i1.14495

Abstract

Diabetes continues to be a critical global health issue, demanding accurate predictive systems to enable preventive interventions. Traditional diagnostic tests lack efficiency for large-scale early screening, which has led to growing interest in artificial intelligence solutions. This research proposed an effective methodology for diabetes classification based on tree-based algorithms enhanced with SMOTEENN balancing. The study employed the Kaggle Diabetes Prediction Dataset with 100,000 instances and eight medical and demographic features. Preprocessing steps included handling missing and duplicate values, encoding categorical variables, and scaling numerical attributes with Min-Max normalization. To address severe class imbalance, SMOTEENN was adopted, producing a cleaner and more balanced dataset. Model evaluation was performed using Stratified 5-Fold cross-validation on six classifiers: Decision Tree, Random Forest, Gradient Boosting, AdaBoost, XGBoost, and CatBoost. Experimental results indicated significant gains after balancing, with ensemble methods outperforming single-tree baselines. Random Forest delivered the best overall performance (98.93% accuracy, 98.96% F1-score, 99.16% recall, 99.94% AUC), followed by CatBoost and XGBoost with comparable results above 99% AUC. While Decision Tree benefited most from SMOTEENN in relative terms, it remained less competitive. Analysis of the importance of the analysis revealed HbA1c level and blood glucose level as dominant predictors, validating clinically meaningful learning. These findings suggest that integrating hybrid resampling with ensemble tree classifiers provides reliable and general predictions for diabetes risk. The approach holds promise for deployment in healthcare decision support systems.
Hybrid Feature Selection for Effective Heart Disease Detection: A Multi-Algorithm Machine Learning Approach Syahrani Lonang; Ahmad Fatoni Dwi Putra; Fahmi Syuhada; Asno Azzawagama Firdaus; Alya Masitha
Scientific Journal of Informatics Vol. 13 No. 1: February 2026
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v13i1.38815

Abstract

Purpose: This research aims to develop an effective early detection model for heart disease with data balancing and hybrid feature selection. The study seeks to enhance predictive accuracy and minimize errors, providing a robust model for clinical decision support systems. Methods: The study used the Heart Failure Prediction dataset derived from Kaggle. A novel hybrid framework was implemented, integrating SMOTEENN (Synthetic Minority Over-sampling Technique + Edited Nearest Neighbors) for data balancing and a Hybrid Feature Selection (HFS) method combining Chi-square and Backward Elimination. Eight machine learning algorithms, including Logistic Regression, Naïve Bayes, Decision Tree, K Nearest Neighbor, Random Forest, Gradient Boosting, Support Vector Machine, and XGBoost. Performance was assessed based on accuracy, precision, recall, f1-score, specificity, AUC Score, fallout and miss rate. Result: The proposed framework significantly improved classification performance across all algorithms. The Random Forest model emerged as the optimal classifier, achieving an accuracy of 99.44%, AUC Score of 99.98%, and a specific reduction in miss rate to 0.92% (from 10.03% baseline). The HFS method successfully reduced the feature space by 54%, identifying 'ExerciseAngina', 'FastingBS', 'ST_Slope', 'ChestPainType', and 'Sex' as the most critical predictors. The model outperformed standard approaches and recent state-of-the-art benchmarks by over 10% in accuracy. Novelty: This study introduces a synergistic integration of SMOTEENN with hybrid feature selection. The combination significantly improves model performance in early heart disease detection.