Claim Missing Document
Check
Articles

Found 1 Documents
Search

GAUSSIAN NAIVE BAYES FOR EARLY DIABETES PREDICTION: A COMPREHENSIVE EVALUATION OF CLASSIFICATION PERFORMANCE ACROSS VARYING TRAINING TEST PROPORTIONS Fahrur Rozi; Bian Dwi Pamungkas; Vertika Panggayuh; Feraldy Satria Putra
JoEICT (Jurnal of Education And ICT) Vol 8, No 2 (2024)
Publisher : STKIP PGRI TULUNGAGUNG

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29100/joeict.v8i2.9282

Abstract

Diabetes mellitus is a chronic metabolic disorder whose global prevalence continues to rise at an alarming rate, imposing substantial burdens on public health systems worldwide. Early and accurate prediction of diabetes risk is critical to enable timely clinical intervention and to mitigate life-threatening complications such as cardiovascular disease, nephropathy, and retinopathy. Machine learning algorithms have emerged as promising tools for automated diabetes risk classification; however, the comparative stability of probabilistic classifiers across varying data-partitioning strategies remains insufficiently studied. This study presents a systematic evaluation of the Gaussian Naive Bayes (GNB) algorithm for binary diabetes prediction using a publicly available dataset sourced from Kaggle (n = 768 instances; 9 clinical and demographic features). The experimental protocol includes a standard 75:25 training–test split and a sensitivity analysis spanning ten training–test ratio configurations (10%–100%). Under the canonical split, the GNB model attained an overall accuracy of 78.89%, with precision of 64.29%, recall of 67.74%, and F1-score of 66.00%. Cross-partition sensitivity analysis demonstrated that classification accuracy remained relatively stable in the range of 71–76% across all ratio configurations, with the 50% training proportion yielding the most balanced performance (accuracy = 76.30%, F1-score = 67.20%). These findings confirm that GNB constitutes a computationally efficient and interpretable baseline for diabetes screening, while simultaneously revealing limitations in sensitivity that motivate the integration of feature selection and ensemble learning strategies in future research. The novelty of this work lies in its structured cross-partition stability analysis, which provides empirically grounded guidance for dataset splitting decisions in small-scale clinical prediction tasks.