Claim Missing Document
Check
Articles

Found 1 Documents
Search

Comparative Analysis of K-Nearest Neighbor and Fuzzy K-Nearest Neighbor for Diabetes Classification Using the Pima Indians Diabetes Dataset Ulfasari Rafflesia; Siska Dwi Kumala; Ratna Widayati; Aisyah Nooravieta S.; Oon Septa
Journal of Innovation in Applied Natural Science Vol. 2 No. 2 (2026): Journal of Innovation in Applied Natural Science
Publisher : CV Media Inti Teknologi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.58723/jinas.v2i2.223

Abstract

Background of study: Diabetes mellitus remains one of the most common chronic diseases worldwide; thus, early and accurate prediction is critical for appropriate intervention. Aims: This study examines the performance of the standard K-Nearest Neighbor (KNN) method and its fuzzy counterpart, Fuzzy K-Nearest Neighbor (Fuzzy KNN), in classifying diabetes status and compares their behavior across different neighborhood sizes (k). Methods: The Pima Indians Diabetes Dataset (768 cases, 8 clinical variables) was used. Zero values in physiologically implausible attributes (glucose, blood pressure, skin thickness, insulin, and BMI) were treated as missing values based on the physiological interpretation of these variables and imputed using class-wise medians. The dataset was then split into 80% training and 20% testing subsets using stratified sampling to preserve class proportions. Five-fold cross-validation on the training set was used to determine the optimal number of neighbors (k), after which both models were evaluated on the held-out test set using accuracy, precision, recall, and F1-score. Result: The optimal k varied between models: k=15 for KNN and k=3 for Fuzzy KNN. Cross-validation accuracy across the tested k range (3–21) remained relatively stable for both methods (approximately 0.80–0.83), indicating that performance was not highly sensitive to k selection. On the test set, Fuzzy KNN outperformed conventional KNN across all evaluation metrics (accuracy: 0.779 vs. 0.773; precision: 0.679 vs. 0.673; recall: 0.704 vs. 0.685; F1-score: 0.691 vs. 0.679). Conclusion: For the Pima Indians Diabetes Dataset, Fuzzy KNN achieved slightly better performance than conventional KNN across all evaluated metrics on the held-out test set. These findings indicate that distance-based membership weighting may provide a modest advantage over crisp majority voting for diabetes classification in this dataset.