cover
Contact Name
Muhammad Nur Faiz
Contact Email
faiz@pnc.ac.id
Phone
+6282324039994
Journal Mail Official
jinita.ejournal@pnc.ac.id
Editorial Address
Department of Informatics Engineering Politeknik Negeri Cilacap Jln. Dr.Soetomo No.01 Sidakaya, Cilacap, Indonesia
Location
Kab. cilacap,
Jawa tengah
INDONESIA
Journal of Innovation Information Technology and Application (JINITA)
ISSN : 27160858     EISSN : 27159248     DOI : https://doi.org/10.35970/jinita.v2i01.119
Software Engineering, Mobile Technology and Applications, Robotics, Database System, Information Engineering, Interactive Multimedia, Computer Networking, Information System, Computer Architecture, Embedded System, Computer Security, Digital Forensic Human-Computer Interaction, Virtual/Augmented Reality, Intelligent System, IT Governance, Computer Vision, Distributed Computing System, Mobile Processing, Next Network Generation, Natural Language Processing, Business Process, Cognitive Systems, Networking Technology, and Pattern Recognition
Articles 191 Documents
Machine Learning-Based Diabetes Mellitus Classification Using Multi-Dataset Evaluation and Class Imbalance Resampling Wijiyanto; Agustinus Eko Setiawan; Ferly Ardhy; Ritzkal; Ummi Athiyah
Journal of Innovation Information Technology and Application (JINITA) Vol 8 No 1 (2026): JINITA, June 2026
Publisher : Politeknik Negeri Cilacap

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35970/jinita.v8i1.3315

Abstract

Diabetes mellitus (DM) remains a major global health challenge due to its increasing prevalence and long-term complications, emphasizing the need for accurate early prediction systems. This study proposes a machine learning-based framework for DM classification using a multi-dataset setting while addressing class imbalance issues. Two independent datasets from Iraq and Germany were employed to evaluate model robustness across different population characteristics. The experimental workflow consisted of data preprocessing, stratified train-test splitting, imbalance handling using Synthetic Minority Over-sampling Technique (SMOTE) and SMOTE-Tomek, 10-fold cross-validation, and hyperparameter optimization via GridSearchCV. Four classification algorithms were compared, namely Logistic Regression (LR), K-Nearest Neighbors (KNN), Random Forest (RF), and Support Vector Machine (SVM). Experimental results demonstrate that data distribution significantly affects classification performance. Under imbalanced conditions, RF achieved the best performance on the Iraqi dataset with an accuracy of 0.98 and an AUC of 1.00, while KNN and RF reached perfect accuracy (1.00) on the German dataset. After applying SMOTE, all models showed more stable performance, particularly in recall, which reached 1.00, indicating effective minority-class detection. In contrast, SMOTE-Tomek produced only marginal additional improvements. The findings suggest that no single classifier is universally optimal for DM prediction. Instead, model effectiveness depends on dataset characteristics and preprocessing strategies. From a practical perspective, the combination of RF and SMOTE shows strong potential for early diabetes screening and clinical decision-support systems. Further validation using larger and more heterogeneous external datasets is recommended.