Missing data represent a common challenge in statistical modeling and can substantially reduce the performance of classification algorithms. This study examines the impact of missing values on the performance of the XGBoost model by considering different proportions of missingness (50% and 75%) and various combinations of affected variables under the Missing Completely at Random (MCAR) and Missing at Random (MAR) mechanisms. Two imputation methods, MissForest and Multiple imputation by chained equations (MICE), were compared with a baseline model without imputation. The analysis of variance revealed that the interaction between imputation method, missing data proportion, and variable combinations had a significant effect on accuracy and specificity, while sensitivity remained relatively stable across scenarios. Tukey tests confirmed that MissForest consistently outperformed the other approaches, producing the highest accuracy and specificity, especially at 50% missingness with three variables affected. Moreover, the evaluation of categorical distributions before and after imputation indicated that MissForest better preserved category balance compared to MICE. These findings highlight that the performance of imputation methods strongly depends on the characteristics of missing data. Overall, MissForest demonstrated clear superiority in handling missing categorical data, maintaining distributional integrity while enhancing the classification performance of the XGBoost model. This study advances statistical learning by giving empirical evidence and practical tips for choosing strong imputation strategies in categorical datasets. This improves the reliability of predictive modeling when data is missing.