Mirza Afif Pradivta
STIKOM Tunas Bangsa, Pematangsiantar

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Klasifikasi Penyakit Ginjal Kronis pada Data Tidak Seimbang Menggunakan K-Nearest Neighbor Berbasis Seleksi Fitur Mutual Information dan GridSearchCV Mirza Afif Pradivta; Solikhun Solikhun; Timbo Faritcan P Siallagan
Bulletin of Artificial Intelligence Vol 5 No 1 (2026): April 2026
Publisher : Graha Mitra Edukasi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.62866/buai.v5i1.227

Abstract

Chronic Kidney Disease (CKD) is a progressive disease characterized by a gradual decline in kidney function and requires early detection to reduce the risk of severe complications. Machine learning has been widely applied to support CKD classification based on clinical attributes; however, medical datasets often contain missing values, a combination of numerical and categorical features, and class imbalance. This study aims to evaluate the performance of the K-Nearest Neighbor (KNN) algorithm for CKD classification using Mutual Information feature selection and GridSearchCV. The dataset consisted of 400 samples, including 250 CKD cases and 150 non-CKD cases. The proposed methodology included data cleaning, missing value imputation, categorical feature encoding, numerical feature normalization using MinMaxScaler, feature selection using SelectKBest with Mutual Information, and hyperparameter tuning using GridSearchCV. Model performance was evaluated using hold-out testing and 10-fold cross-validation. The hold-out evaluation showed that the KNN model with GridSearchCV achieved 100.00% accuracy, precision, recall, F1-score, and AUC on the test set. To ensure that this result was not dependent on a single train-test split, additional evaluation was conducted using 10-fold cross-validation. The cross-validation results yielded an average accuracy of 99.25% for the KNN model with GridSearchCV, indicating consistent performance across different data partitions. Meanwhile, the KNN model with Mutual Information feature selection and GridSearchCV achieved 98.75% accuracy, 100.00% recall, and a 99.01% F1-score, demonstrating competitive performance while using a more compact feature subset. The findings indicate that the application of GridSearchCV improved the performance of the KNN model on the dataset used, while Mutual Information contributed to selecting relevant features, enabling the model to maintain strong classification performance with a reduced number of features