Journal of Education and ICT
Vol 8, No 2 (2024)

GAUSSIAN NAIVE BAYES FOR EARLY DIABETES PREDICTION: A COMPREHENSIVE EVALUATION OF CLASSIFICATION PERFORMANCE ACROSS VARYING TRAINING TEST PROPORTIONS

Fahrur Rozi (Universitas Bhinneka PGRI)
Bian Dwi Pamungkas (Universitas Bhinneka PGRI)
Vertika Panggayuh (Universitas Bhinneka PGRI)
Feraldy Satria Putra (Universitas Bhinneka PGRI)



Article Info

Publish Date
01 Dec 2024

Abstract

Diabetes mellitus is a chronic metabolic disorder whose global prevalence continues to rise at an alarming rate, imposing substantial burdens on public health systems worldwide. Early and accurate prediction of diabetes risk is critical to enable timely clinical intervention and to mitigate life-threatening complications such as cardiovascular disease, nephropathy, and retinopathy. Machine learning algorithms have emerged as promising tools for automated diabetes risk classification; however, the comparative stability of probabilistic classifiers across varying data-partitioning strategies remains insufficiently studied. This study presents a systematic evaluation of the Gaussian Naive Bayes (GNB) algorithm for binary diabetes prediction using a publicly available dataset sourced from Kaggle (n = 768 instances; 9 clinical and demographic features). The experimental protocol includes a standard 75:25 training–test split and a sensitivity analysis spanning ten training–test ratio configurations (10%–100%). Under the canonical split, the GNB model attained an overall accuracy of 78.89%, with precision of 64.29%, recall of 67.74%, and F1-score of 66.00%. Cross-partition sensitivity analysis demonstrated that classification accuracy remained relatively stable in the range of 71–76% across all ratio configurations, with the 50% training proportion yielding the most balanced performance (accuracy = 76.30%, F1-score = 67.20%). These findings confirm that GNB constitutes a computationally efficient and interpretable baseline for diabetes screening, while simultaneously revealing limitations in sensitivity that motivate the integration of feature selection and ensemble learning strategies in future research. The novelty of this work lies in its structured cross-partition stability analysis, which provides empirically grounded guidance for dataset splitting decisions in small-scale clinical prediction tasks.

Copyrights © 2024






Journal Info

Abbrev

joeict

Publisher

Subject

Computer Science & IT Control & Systems Engineering Decision Sciences, Operations Research & Management

Description

This journal encompasses original research articles, review articles, and short communications, including: Pendidikan Teknologi Informasi Information System Artificial Intelligence AI & Expert systems Database Systems Computing Languages & Algorithms Computer Networks & Communications Computer ...