Vanya Dwi Nabila
Universitas Sriwijaya, Palembang

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Comparative Evaluation of Machine Learning Models with Explainability and Fairness for Diabetes Risk Prediction Vanya Dwi Nabila; Bayu Wijaya Putra; M Rudi Sanjaya; Dwi Rosa Indah
Building of Informatics, Technology and Science (BITS) Vol 8 No 2 (2026): September 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i2.10763

Abstract

Diabetes mellitus remains a major global health challenge, making early risk prediction essential for timely intervention and prevention. This study proposes a trustworthy machine learning framework for diabetes risk prediction by integrating predictive performance, explainability, and fairness evaluation. Six classification algorithms Logistic Regression, Decision Tree, Random Forest, XGBoost, LightGBM, and CatBoost were comparatively evaluated, with LightGBM selected as the best baseline model. Hyperparameter optimization was subsequently performed using Optuna with macro F1-score as the optimization objective to better address the imbalanced multiclass nature of the dataset. Model performance was assessed using Accuracy, Balanced Accuracy, Precision, Recall, F1-score, and Receiver Operating Characteristic–Area Under the Curve (ROC-AUC). Model explainability was analyzed using SHapley Additive exPlanations (SHAP), while fairness was evaluated using Fairlearn based on gender and race through Demographic Parity Difference and Equalized Odds Difference.Experimental results show that hyperparameter optimization increased Balanced Accuracy from 0.3755 to 0.4891 and macro F1-score from 0.3797 to 0.4333, indicating improved recognition of minority classes. Although overall Accuracy decreased from 0.8357 to 0.7023, this trade-off reflects a more balanced classification across diabetes categories, which is preferable for imbalanced clinical datasets where identifying minority cases is essential for early risk detection. SHAP analysis identified Body Mass Index (BMI), Age, Physical Health Days, Mental Health Days, and Income Level as the most influential predictors. Fairness evaluation further demonstrated low demographic disparities across gender and race. These findings demonstrate that integrating predictive performance, explainability, and fairness enables the development of a more transparent, equitable, and clinically reliable diabetes risk prediction framework.