Stroke is a leading cause of death and permanent disability in Indonesia, where prevalence rose from 8.3 per 1,000 population in 2013 to 10.9 per 1,000 in 2018, creating an urgent need for accurate early detection. This study develops and compares stroke risk prediction models based on Naive Bayes (NB) and K-Nearest Neighbor (KNN), each integrated with the Synthetic Minority Over-sampling Technique (SMOTE) to handle class imbalance. The dataset consists of 512 medical records from a hospital in Batam collected between 2020 and 2023, covering 12 demographic, clinical, and lifestyle predictors. SMOTE was applied only to the training data of an 80 to 20 split, yielding a balanced distribution of 258 samples per class. The KNN parameter K was tuned over eight odd values and 7 was found optimal. Evaluation on 103 test records shows that KNN consistently outperforms Naive Bayes, with accuracy of 91.3 against 81.6 percent, F1-Score of 89.2 against 79.9 percent, and AUC of 0.934 against 0.862. SMOTE raised recall of the Stroke class by 18.8 percentage points. Blood glucose, age, and systolic blood pressure were the three strongest predictors. The model is proposed as an Early Warning System module for hospital information systems.
Copyrights © 2026