The rapid growth of the e-commerce industry has intensified competition and increased the risk of customer churn, defined as the discontinuation of service usage. This study aims to develop an accurate churn prediction model using four tree-based classifiers—Random Forest, XGBoost, LightGBM, and CatBoost—and to evaluate feature contributions using SHAP (Shapley Additive Explanations). The study employs the E-commerce Customer Churn dataset from Kaggle, consisting of 5,630 observations and 20 behavioral features. All models demonstrate strong performance, achieving accuracy above 95%. XGBoost outperforms the other models, with an accuracy of 97.42% and an F1-score of 0.918, indicating a good balance between precision and recall. SHAP analysis identifies Tenure as the most influential feature, followed by Complain, NumberOfAddress, CashbackAmount, and MaritalStatus_Married. Lower tenure increases churn likelihood, while higher complaint levels significantly raise churn risk. Feature effects vary across individuals due to interactions among variables. These findings highlight the importance of interpretability in supporting accurate and data-driven customer retention strategies in e-commerce.
Copyrights © 2026