Predicting creditworthiness is crucial for mitigating loan risks, but missing data and extreme outliers often compromise model accuracy. This study evaluates the impact of data imputation quality on financial credit predictions using the K-Nearest Neighbor algorithm. Utilizing a Finance Loan Approval dataset of 614 records, this research compared four imputation methods: Mean, Median, Linear Interpolation, and KNN Imputer. The results demonstrate that outlier-robust imputation methods significantly enhance predictive integrity. Median and Linear Interpolation achieved the highest accuracy of 85.37% and an F1-Score of 90.32% by maintaining the spatial distances essential for the algorithm. Conversely, Mean imputation produced the lowest accuracy of 83.71% due to distortions from extreme income anomalies. Furthermore, the evaluation exposed algorithmic bias driven by class imbalance. The model achieved a near-perfect Recall of 99% but struggled with a lower Precision of 83%, indicating a high rate of false positives. In conclusion, while robust imputation strategies restore data structure and optimize accuracy, they cannot resolve class imbalance biases. Future scoring systems must integrate outlier-resistant imputation with advanced resampling techniques to improve precision and mitigate financial risks.
Copyrights © 2026