Journal of Innovation in Applied Natural Science
Vol. 2 No. 2 (2026): Journal of Innovation in Applied Natural Science

Comparative Analysis of K-Nearest Neighbor and Fuzzy K-Nearest Neighbor for Diabetes Classification Using the Pima Indians Diabetes Dataset

Ulfasari Rafflesia (Universitas Bengkulu)
Siska Dwi Kumala (Universitas Bengkulu)
Ratna Widayati (Universitas Bengkulu)
Aisyah Nooravieta S. (Universitas Bengkulu)
Oon Septa (Universitas Bengkulu)



Article Info

Publish Date
08 Sep 2026

Abstract

Background of study: Diabetes mellitus remains one of the most common chronic diseases worldwide; thus, early and accurate prediction is critical for appropriate intervention. Aims: This study examines the performance of the standard K-Nearest Neighbor (KNN) method and its fuzzy counterpart, Fuzzy K-Nearest Neighbor (Fuzzy KNN), in classifying diabetes status and compares their behavior across different neighborhood sizes (k). Methods: The Pima Indians Diabetes Dataset (768 cases, 8 clinical variables) was used. Zero values in physiologically implausible attributes (glucose, blood pressure, skin thickness, insulin, and BMI) were treated as missing values based on the physiological interpretation of these variables and imputed using class-wise medians. The dataset was then split into 80% training and 20% testing subsets using stratified sampling to preserve class proportions. Five-fold cross-validation on the training set was used to determine the optimal number of neighbors (k), after which both models were evaluated on the held-out test set using accuracy, precision, recall, and F1-score. Result: The optimal k varied between models: k=15 for KNN and k=3 for Fuzzy KNN. Cross-validation accuracy across the tested k range (3–21) remained relatively stable for both methods (approximately 0.80–0.83), indicating that performance was not highly sensitive to k selection. On the test set, Fuzzy KNN outperformed conventional KNN across all evaluation metrics (accuracy: 0.779 vs. 0.773; precision: 0.679 vs. 0.673; recall: 0.704 vs. 0.685; F1-score: 0.691 vs. 0.679). Conclusion: For the Pima Indians Diabetes Dataset, Fuzzy KNN achieved slightly better performance than conventional KNN across all evaluated metrics on the held-out test set. These findings indicate that distance-based membership weighting may provide a modest advantage over crisp majority voting for diabetes classification in this dataset.

Copyrights © 2026






Journal Info

Abbrev

jinas

Publisher

Subject

Aerospace Engineering Agriculture, Biological Sciences & Forestry Astronomy Biochemistry, Genetics & Molecular Biology Chemical Engineering, Chemistry & Bioengineering Chemistry Energy Environmental Science Immunology & microbiology Materials Science & Nanotechnology Mathematics Physics

Description

Journal of Integrated Natural Sciences (JINAS) is an academic peer-reviewed journal that publishes high-quality research in the field of natural sciences, encompassing both fundamental and applied scientific studies. The journal aims to provide a scholarly platform for researchers, academics, and ...