Betharya Tampubolon
Universitas HKBP Nommensen Pematangsiantar, Pematangsiantar

Published : 2 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 2 Documents
Search

Komparasi Algoritma Naive Bayes, Random Forest, dan Decision Tree untuk Prediksi Penyakit Stroke Menggunakan Orange Data Mining Wulan Liviana Simbolon; David Ofel Gihon Purba; Ningsih Purba; Saudurma Sidabutar; Betharya Tampubolon; Jaya Tata Hardinata
Jurnal Ilmu Komputer, Teknologi Dan Informasi Vol 4 No 2 (2026): Juli
Publisher : CV. Graha Mitra Edukasi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.62866/jurikti.v4i2.359

Abstract

Stroke is one of the leading causes of death and long-term disability worldwide, making accurate prediction methods essential to support early detection and clinical decision-making. The problem addressed in this study is the lack of evidence regarding which classification algorithm provides the best performance for predicting stroke using the Healthcare Stroke Dataset. This study aims to compare the performance of the Naive Bayes, Random Forest, and Decision Tree algorithms using Orange Data Mining to identify the most effective predictive model. A quantitative approach with a comparative experimental design was employed. The dataset used in this research was the Healthcare Stroke Dataset obtained from Kaggle, consisting of 5,110 records with 12 attributes. The research process included data preprocessing using the Impute widget, feature selection using the Rank widget, classification model development, and model evaluation through 10-fold cross-validation. Performance was assessed using Accuracy, Area Under the Curve (AUC), Precision, Recall, F1-Score, Matthews Correlation Coefficient (MCC), Confusion Matrix, and Receiver Operating Characteristic (ROC) analysis. The results indicate that the Decision Tree algorithm achieved the highest accuracy of 95.1%, followed by Random Forest with 94.8%, while Naive Bayes achieved 92.4%. However, Naive Bayes obtained the highest AUC value of 0.804, demonstrating superior class discrimination capability on an imbalanced dataset. These findings suggest that algorithm selection should not rely solely on accuracy but also consider the model's ability to distinguish between classes consistently. This study contributes to providing recommendations for selecting appropriate classification algorithms to support the development of machine learning-based early stroke prediction systems
Analisis Perbandingan Algoritma Random Forest dan Support Vector Machine pada Prediksi Customer Churn Menggunakan IBM Telco Customer Churn Dataset Saudurma Seven Septiana Sidabutar; Ningsih Septi Uli Purba; Wulan Liviana Simbolon; Betharya Tampubolon; Jaya Tata Hardinata
Jurnal Ilmu Komputer, Teknologi Dan Informasi Vol 4 No 2 (2026): Juli
Publisher : CV. Graha Mitra Edukasi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.62866/jurikti.v4i2.377

Abstract

Customer churn is a condition in which customers decide to discontinue the services provided by a company. In the telecommunications industry, a high customer churn rate can reduce company revenue and customer loyalty. Therefore, an accurate prediction method is needed to identify customers who are likely to churn so that preventive strategies can be implemented at an early stage. This study aims to compare the performance of the Random Forest and Support Vector Machine (SVM) algorithms in predicting customer churn using the IBM Telco Customer Churn Dataset. The research stages include data collection, data preprocessing, model development using Orange Data Mining, model evaluation through Test and Score, Confusion matrix, and Receiver operating characteristic (ROC), as well as data visualization using Microsoft Power BI. The results indicate that both algorithms are capable of classifying customer data; however, the Random Forest algorithm achieves better performance than the Support Vector Machine based on the evaluation metrics obtained. Furthermore, data visualization using Microsoft Power BI provides a clearer understanding of customer characteristics and supports the interpretation of the research findings. Therefore, Random Forest is recommended as a more effective algorithm for customer churn prediction in the telecommunications sector.