Ilham Putra Ariatama
Department of Informatics, Sepuluh Nopember Institute of Technology (ITS), Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Comparative Evaluation of Machine Learning and Ensemble Learning with Synthetic Minority Oversampling and Random Undersampling for Loan Approval Prediction Across Multiple Datasets Ilham Putra Ariatama; Wahyu Fajar Setiawan; Afif Amirullah; Ratih Nur Esti Anggraini
Jurnal Teknik Informatika (Jutif) Vol. 7 No. 4 (2026): JUTIF Volume 7, Number 4, August 2026
Publisher : Informatika, Universitas Jenderal Soedirman

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.52436/1.jutif.2026.7.4.5854

Abstract

Loan approval prediction remains a critical yet challenging task in financial risk management, particularly due to class imbalance and variability across lending datasets. This study presents a comparative evaluation of eleven machine learning and ensemble learning algorithms for loan approval prediction, evaluated across three publicly available datasets with distinct characteristics to ensure generalizability of findings. The algorithms evaluated include Logistic Regression, Decision Tree, Random Forest, Soft Voting, Hard Voting, AdaBoost, Gradient Boosting, XGBoost, Bagging, Pasting, and Stacking. To address class imbalance, three data balancing strategies are applied: original imbalanced data, Synthetic Minority Oversampling Technique, and random undersampling. Model performance is assessed using accuracy, precision, recall, and F1-score under 5-fold cross-validation with a 75:25 stratified train-test split. Experimental results demonstrate that ensemble methods consistently outperform single classifiers across all datasets. Stacking achieves the highest accuracy on original data across all three datasets (88.33%, 76.00%, and 93.16%, respectively), while XGBoost  demonstrates competitive and stable performance across multiple data balancing conditions. Results also reveal that Synthetic Minority Oversampling does not universally improve model performance, as certain dataset characteristics lead to accuracy degradation under oversampling conditions. These findings contribute to the field of computer science by establishing a multi-dataset benchmark for ensemble learning in automated credit scoring and by providing empirical evidence on the interaction between data balancing strategies and classifier architectures, offering a practical framework for selecting robust machine learning pipelines in financial decision support systems.