Creditworthiness is critical to financial system stability, yet conventional methods struggle with imbalanced data and mixed numeric-categorical features. This study develops the CreditRF-OptSHAP pipeline, integrating Random Forest, SMOTENC, GridSearchCV, and SHAP to improve credit classification performance and interpretability. The dataset used, Statlog German Credit Data (UCI), comprises 1,000 instances, 20 mixed features, and a 70:30 imbalance ratio. The pipeline comprises five phases: data exploration; preprocessing via label encoding and StandardScaler; SMOTENC-based training-set balancing; hyperparameter optimization via GridSearchCV with Stratified 10-Fold Cross Validation; and model interpretation via SHAP TreeExplainer so that each feature's contribution to prediction is explained, making the model no longer a black box. The significance of this recall gain over RF Default was validated using McNemar's Test to rule out statistical coincidence. The main model (RF+SMOTENC+GridSearchCV) achieved a recall of 0.5833 for the bad credit class and an AUC-ROC of 0.7638, up from 0.5000 for RF Default. SHAP analysis identified checking account status as the most dominant feature (importance 0.1239). The McNemar test confirmed a statistically significant difference (p=0.041), confirming its validity. The CreditRF-OptSHAP pipeline yields a model that is more sensitive to non-performing loans, transparent, and statistically validated, thus supporting accountable credit decisions in financial institutions.
Copyrights © 2026