Claim Missing Document
Check
Articles

Found 2 Documents
Search

Improving Classification Performance of Imbalanced Data Using SMOTE: empirical studies Ulfasari Rafflesia; Dedi Rosadi
Riemann: Research of Mathematics and Mathematics Education Vol. 8 No. 1 (2026): EDISI APRIL
Publisher : Program Studi Pendidikan Matematika Universitas Katolik Santo Agustinus Hippo

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.38114/riemann.v8i1.199

Abstract

Data balancing methods in multi-class settings continue to evolve as the importance of balanced data conditions for classification analysis grows. However, limited studies have provided comprehensive empirical comparisons across both binary and multi-class imbalanced datasets. Data imbalance can affect model predictions, particularly by leading to inaccurate identification of minority classes. Therefore, this study aims to evaluate the effectiveness of the Synthetic Minority Over-sampling Technique (SMOTE) in improving classification performance. Three benchmark datasets from the UCI Machine Learning Repository—Breast Cancer, Ecoli, and Glass—were selected to represent imbalanced classification problems in both binary and multi-class settings. The proposed framework addresses class imbalance during data preprocessing using SMOTE. Each dataset is first divided into training and testing subsets. SMOTE is applied only to the training data to address class imbalance, while the test data is kept unchanged for evaluation. Then, the classification process is applied to the original (imbalanced) data and to the balanced data generated by SMOTE. The classifiers used in this study are SVM, a decision tree, and AdaBoost. The classification results are evaluated based on accuracy, sensitivity, and F1-score. The results show that the decision tree and AdaBoost improve classification performance under imbalanced data conditions. In particular, AdaBoost achieves the best overall performance in terms of prediction accuracy and class balance, demonstrating the effectiveness of combining SMOTE with ensemble methods for handling imbalanced datasets.
Comparative Analysis of K-Nearest Neighbor and Fuzzy K-Nearest Neighbor for Diabetes Classification Using the Pima Indians Diabetes Dataset Ulfasari Rafflesia; Siska Dwi Kumala; Ratna Widayati; Aisyah Nooravieta S.; Oon Septa
Journal of Innovation in Applied Natural Science Vol. 2 No. 2 (2026): Journal of Innovation in Applied Natural Science
Publisher : CV Media Inti Teknologi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.58723/jinas.v2i2.223

Abstract

Background of study: Diabetes mellitus remains one of the most common chronic diseases worldwide; thus, early and accurate prediction is critical for appropriate intervention. Aims: This study examines the performance of the standard K-Nearest Neighbor (KNN) method and its fuzzy counterpart, Fuzzy K-Nearest Neighbor (Fuzzy KNN), in classifying diabetes status and compares their behavior across different neighborhood sizes (k). Methods: The Pima Indians Diabetes Dataset (768 cases, 8 clinical variables) was used. Zero values in physiologically implausible attributes (glucose, blood pressure, skin thickness, insulin, and BMI) were treated as missing values based on the physiological interpretation of these variables and imputed using class-wise medians. The dataset was then split into 80% training and 20% testing subsets using stratified sampling to preserve class proportions. Five-fold cross-validation on the training set was used to determine the optimal number of neighbors (k), after which both models were evaluated on the held-out test set using accuracy, precision, recall, and F1-score. Result: The optimal k varied between models: k=15 for KNN and k=3 for Fuzzy KNN. Cross-validation accuracy across the tested k range (3–21) remained relatively stable for both methods (approximately 0.80–0.83), indicating that performance was not highly sensitive to k selection. On the test set, Fuzzy KNN outperformed conventional KNN across all evaluation metrics (accuracy: 0.779 vs. 0.773; precision: 0.679 vs. 0.673; recall: 0.704 vs. 0.685; F1-score: 0.691 vs. 0.679). Conclusion: For the Pima Indians Diabetes Dataset, Fuzzy KNN achieved slightly better performance than conventional KNN across all evaluated metrics on the held-out test set. These findings indicate that distance-based membership weighting may provide a modest advantage over crisp majority voting for diabetes classification in this dataset.