Victoria Zevini Sabo
Federal University Wukari, Nigeria

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Optimizing Chronic Kidney Disease Prediction Via Ensemble Learning On Imbalanced Multi-Feature Clinical Data Emmanuel John Anagu; Gani Timothy Abe; Victoria Zevini Sabo; Sunday Jatau Lamiri
Brilliance: Research of Artificial Intelligence Vol. 6 No. 2 (2026): Brilliance: Research of Artificial Intelligence, Article Research May 2026
Publisher : Yayasan Cita Cendekiawan Al Khwarizmi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47709/brilliance.v6i2.8210

Abstract

Chronic Kidney Disease (CKD) remains a critical global health burden, characterized by its asymptomatic progression in early stages and high risk of culminating in end-stage renal disease, yet timely detection remains elusive within conventional diagnostic frameworks. This study addresses this gap by comparatively evaluating four machine learning classifiers Random Forest, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), and Logistic Regression for early CKD prediction using a multi-feature, class-imbalanced clinical dataset. A dataset comprising 1,659 patient records and 54 clinical, demographic, and laboratory features was sourced from the Kaggle repository, preprocessed through feature elimination (reducing features to 40), standardization, and Random Oversampling to correct class imbalance. An 80-20 train-test split was applied prior to model training and hyperparameter tuning. Classification performance was assessed using accuracy, precision, recall, F1-score, and Area Under the Receiver Operating Characteristic Curve (AUC-ROC). Random Forest achieved the highest accuracy of 99.67%, an AUC of 1.00, and near-perfect precision and recall, substantially outperforming KNN (95.26%), SVM (80.00%), and Logistic Regression (78.53%). These findings confirm the superiority of ensemble bagging methods over distance-based and linear classifiers in managing high-dimensional, imbalanced medical datasets. The study contributes to the growing body of evidence supporting machine learning integration into CKD screening pathways, while underscoring the critical role of class-balancing strategies in preventing diagnostic bias. The deployed Streamlit application further demonstrates a viable pathway toward accessible clinical decision-support tools for CKD early detection.