Claim Missing Document
Check
Articles

Found 2 Documents
Search

Analisis Perbandingan Model Machine Learning Tree-Based dan Non-Tree-Based untuk Tugas Klasifikasi Hilmi, Fadhilah; Taqiyassar, Kenzie; Pratama, Naufal Romero Putra; Kusuma, Satrio Condro; Nurwachid, Hafiz Rizky; Fatyanosa, Tirana Noor
Jurnal Teknologi Informasi dan Ilmu Komputer Vol 12 No 4: Agustus 2025
Publisher : Fakultas Ilmu Komputer, Universitas Brawijaya

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.25126/jtiik.124

Abstract

Penelitian ini membahas perbandingan performa model machine learning berbasis pohon keputusan (Tree-Based) dan non-pohon keputusan (Non-Tree-Based) dalam tugas klasifikasi. Model Tree-based yang diuji meliputi LightGBM, CatBoost, XGBoost, dan Random Forest, sedangkan model Non-tree-based meliputi SVM, KNN, dan GaussianNB. Evaluasi dilakukan pada tiga dataset berbeda, yaitu Spaceship Titanic, Horse Health, dan Keep It Dry. Metrik yang digunakan untuk mengevaluasi performa model adalah AUC-ROC, akurasi, dan F1-score Micro. Hasil penelitian menunjukkan bahwa model berbasis pohon keputusan seperti CatBoost dan LightGBM umumnya memberikan performa yang lebih baik dibandingkan dengan model non-pohon keputusan. CatBoost khususnya menunjukkan hasil terbaik dalam hal akurasi, AUC-ROC, dan F1-score Micro di sebagian besar dataset yang diuji. Selain itu, penelitian ini juga menyoroti pentingnya pemilihan model yang tepat berdasarkan karakteristik dataset yang digunakan. Faktor-faktor seperti kompleksitas data, jumlah fitur, dan distribusi kelas sangat mempengaruhi hasil akhir dari setiap model yang diterapkan. Dengan demikian, temuan ini dapat membantu praktisi machine learning dalam memilih model yang paling sesuai untuk tugas klasifikasi tertentu.   Abstract This study discusses the performance comparison of tree-based and non-tree-based machine learning models for classification tasks. The Tree-based models tested include LightGBM, CatBoost, XGBoost, and Random Forest, while the Non-tree-based models include SVM, KNN, and GaussianNB. The evaluation was conducted on three different datasets, namely Spaceship Titanic, Horse Health, and Keep It Dry. The metrics used to evaluate model performance are AUC-ROC, accuracy, and F1-score Micro. The results show that tree-based models such as CatBoost and LightGBM generally provide better performance compared to non-tree-based models. CatBoost, in particular, showed the best results in terms of accuracy, AUC-ROC, and F1-score Micro in most of the datasets tested. Additionally, this study highlights the importance of selecting the appropriate model based on the characteristics of the datasets used. Factors such as data complexity, number of features, and class distribution significantly affect the final results of each applied model. Thus, these findings can assist machine learning practitioners in choosing the most suitable model for specific classification tasks.
A Comparative Study: Can Deep Learning Outperform Tree-Based Models in Tabular Data Classification? Fatyanosa, Tirana Noor; Hilmi, Fadhilah; Taqiyassar, Kenzie; Pratama, Naufal Romero Putra; Satrio Condro Kusuma; Hafiz Rizky Nurwachid
Journal of Information Technology and Computer Science Vol. 10 No. 1: April 2025
Publisher : Faculty of Computer Science (FILKOM) Brawijaya University

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.25126/jitecs.2025101922

Abstract

The rapid growth of data in various domains has heightened the need for accurate and efficient predictive models, particularly for tabular data. While deep learning has revolutionized fields like computer vision and natural language processing, its effectiveness in tabular data classification remains a topic of debate. This study conducts a comprehensive comparison between deep learning models (NODE, SAINT, TabNet, Tab Transformers) and tree-based models (Random Forest, XGBoost, LightGBM, CatBoost) to determine whether deep learning can outperform traditional methods in this context. The results indicate that tree-based models, particularly LightGBM and CatBoost, consistently achieve the highest accuracy and F1 scores, coupled with efficient execution times, making them more suitable for real-world applications that require quick and accurate predictions. In contrast, deep learning models show varied performance. Although SAINT sometimes achieves comparable accuracy, its processing time makes it less practical. The findings suggest that despite the potential of deep learning, tree-based models remain superior for tabular data classification tasks, particularly when considering a balance of accuracy, speed, and robustness. This study contributes to the ongoing discussion on the role of deep learning in tabular data and highlights the conditions under which traditional models may still be preferred.