Claim Missing Document
Check
Articles

Performance Comparison Of K-Nearest Neighbors And Decision Tree Algorithms With Random Oversampling For Imbalanced Heart Disease Classification Yustianisa, Dita; Wajidi, Farid; Firgiawan, Wawan; Sholeha, Adinda Gama
Jurnal Teknik Informatika (Jutif) Vol. 7 No. 3 (2026): JUTIF Volume 7, Number 3, June 2026
Publisher : Informatika, Universitas Jenderal Soedirman

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.52436/1.jutif.2026.7.3.5626

Abstract

Heart disease remains one of the leading causes of mortality worldwide, including in Indonesia, where delayed detection continues to be a serious challenge. In medical data mining, class imbalance often degrades classification performance by reducing sensitivity toward minority class cases. This study aims to compare the performance of the K-Nearest Neighbors (KNN) and Decision Tree algorithms for heart disease classification and to evaluate the effectiveness of random oversampling in handling imbalanced data. This research uses a heart disease dataset consisting of 10,000 medical records obtained from Kaggle. Data preprocessing includes categorical transformation, missing value imputation using KNN Imputer, and Min–Max normalization. Random oversampling is applied to increase minority class representation. Model evaluation is performed using stratified 10-fold cross-validation with accuracy, precision, recall, F1-score, and Receiver Operating Characteristic–Area Under the Curve (ROC–AUC) as performance metrics. Experimental results show that after random oversampling, the KNN model achieves the best performance with an accuracy of 94%, precision of 96%, recall of 90%, F1-score of 92%, and ROC–AUC of 90.2%. In comparison, the Decision Tree model records an accuracy of 80%, precision of 78%, recall of 81%, F1-score of 79%, and ROC–AUC of 81.5%. These findings demonstrate that random oversampling significantly improves minority class detection, particularly for KNN. This study contributes to Informatics by providing empirical evidence that simple and efficient data mining strategies can effectively address class imbalance in large-scale medical datasets, supporting the development of accurate, interpretable, and accessible AI-based diagnostic systems for early heart disease detection.
Comparison of SVM and Naive Bayes in Public Sentiment Analysis on Budget Efficiency Yusmita; Farid Wajidi; Muh.Rafli Rasyid
Jurnal Sistem Cerdas Vol. 8 No. 3 (2025)
Publisher : APIC

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.37396/jsc.v8i3.576

Abstract

Kebijakan efisiensi anggaran melalui Instruksi Presiden Nomor 1 Tahun 2025 memicu beragam respons publik di media sosial, khususnya X. Penelitian ini mengklasifikasikan sentimen publik menggunakan algoritma Naïve Bayes dan SVM dengan 6.596 twit setelah tahap praproses, menggunakan pelabelan Lexicon InSet, dan ekstraksi fitur TF-IDF. Hasilnya menunjukkan bahwa SVM-LinearSVC mencapai akurasi tertinggi sebesar 94%, sementara Naïve Bayes mencapai 86% tetapi lebih cepat dalam pelatihan dan prediksi. Temuan ini menegaskan bahwa algoritma pembelajaran mesin efektif untuk memetakan opini publik terkait kebijakan, sekaligus menjadi referensi penelitian analisis sentimen berbahasa Indonesia.
Perbandingan Algoritma Support Vector Machine dan Random Forest untuk Analisis Sentimen terkait Makan Bergizi Gratis (MBG) Saputra, Alfian; Zulkarnaim, Nuralamsah; Wajidi, Farid
JURNAL FASILKOM Vol. 16 No. 2 (2026): Jurnal FASILKOM (teknologi inFormASi dan ILmu KOMputer)
Publisher : Unversitas Muhammadiyah Riau

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.37859/jf.v16i2.11363

Abstract

Advances in digital technology have made social media the primary platform for the public to voice their opinions on government policies, including the Free Nutritious Meals (MBG) program. This study aims to compare the performance of the Support Vector Machine (SVM) and Random Forest algorithms in conducting sentiment analysis of public opinion on the X platform regarding this policy, in order to provide an objective overview for policymakers. A total of 2,164 tweets were collected via crawling using Tweet Harvest from May to August 2025. The methodological steps included preprocessing, BERT-based automatic labeling, Eliminate the neutral class to focus the analysis on binary classification (positive and negative), TF-IDF feature extraction, and the application of SMOTE to address class imbalance in the dataset. Model optimization was performed using Grid Search with a 5-Fold Cross Validation testing scheme and an 80:20 data split. The results of the study indicate that the majority of public responses to the MBG program were positive. Based on the final evaluation, the SVM with a linear kernel proved superior with an accuracy of 78.24% and a macro average F1-score of 0.76, outperforming the Random Forest (300 estimators), which achieved an accuracy of 75.00% and an F1-score of 0.73. The application of SMOTE has proven to be crucial in improving the negative class recall, enabling the model to identify minority sentiments more accurately. This study concludes that SVMs demonstrate more stable and objective generalization capabilities in mapping public opinion following data balancing.