Claim Missing Document
Check
Articles

Found 1 Documents
Search

Comparative Analysis of Machine Learning Algorithms for Optimal Decision-Making in Data-Driven Applications Fatyanosa, Tirana Noor; Brata, Gede Indra Adi; Reansyah, Javier Aahmes; Athaya, Haikal Thoriq; Adam, Muhammad Herdi; Aranda, Achmad Fauzi
Journal of Information Technology and Computer Science Vol. 10 No. 1: April 2025
Publisher : Faculty of Computer Science (FILKOM) Brawijaya University

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.25126/jitecs.2025101921

Abstract

This research examines the effectiveness of machine learning algorithms, including K-Nearest Neighbors (KNN), Naive Bayes, Support Vector Machine, Decision Tree, Logistic Regression, and Gradient Boosting, applied to Kaggle competition datasets. It also investigates essential data preprocessing techniques, such as Standard Scaler, Power Transform, and Simple Imputer, to enhance model performance. The paper provides a comprehensive analysis of best practices in model selection and data preparation to achieve optimal results across different domains. From our experimental results on different datasets, we found that combining Logistic Regression with XGBoost and using ensemble methods like XGBoost and CatBoost achieved the highest scores in several Kaggle competitions. Moreover, using preprocessing such as missing value imputation, data normalization, and feature engineering significantly improved the performance of all models. Our findings suggest that the selection of appropriate machine learning algorithms and data preprocessing techniques is crucial for achieving optimal results in data-driven decision-making. By employing these methodologies, we achieved commendable results across the utilized datasets. For instance, the Predict Failure Keep It Dry dataset yielded a Kaggle score of 0.5901, while the Smoker Status using Bio Signals dataset achieved a score of 0.8675. The Bank Churn Dataset resulted in a Kaggle score of 0.89046, the Spaceship Titanic dataset scored 0.80617, and the health of horses dataset attained a score of 0.76212. These results highlight the effectiveness of the proposed combination of machine learning algorithms and preprocessing techniques in enhancing model performance on real-world datasets.