Non-Performing Loans (NPL) are a fundamental indicator of a financial institution's asset health, reflecting loans that fail to meet interest or principal payment obligations as agreed. A high NPL ratio negatively impacts a bank's financial performance, such as decreased profitability as measured by Return on Assets (ROA) and decreased liquidity. Bank Indonesia sets an NPL tolerance limit of 5% of total credit provided by banking financial institutions. Therefore, a predictive model is needed that can detect the possibility of customers experiencing NPLs early. This study aims to identify relevant factors in predicting NPLs and create an NPL prediction model based on these factors. The contribution of this study lies in combining the results of three feature selection techniques: Chi-Square, Mutual Information, and Random Forest feature importance, using the average score eliminated by the Recursive Feature Elimination technique. Several ensemble algorithms, namely Random Forest, XGBoost, Gradient Boosting, and LightGBM, were explored to produce the best-performing model. Then, hyperparameter tuning was performed on the best model. The Random Forest model produced the best performance, with 92.17% accuracy, 78.1% precision, 98.1% recall, and 95.5% AUC. Hyperparameter tuning was shown to improve recall, thus improving the model's ability to measure how much positive data (Current class) was successfully predicted by the model. The results of this study can assist management in making credit decisions. Thus, it is hoped that it can help reduce the number of NPL cases.
Copyrights © 2026