Journal of Computing Theories and Applications
Vol. 2 No. 1 (2024): JCTA 2(1) 2024

Effects of Data Resampling on Predicting Customer Churn via a Comparative Tree-based Random Forest and XGBoost

Rita Erhovwo Ako (Federal University of Petroleum Resources)
Fidelis Obukohwo Aghware (University of Delta Agbor)
Margaret Dumebi Okpor (Delta State University of Science and Technology Ozoro)
Maureen Ifeanyi Akazue (Delta State University Abraka)
Rume Elizabeth Yoro (Dennis Osadebay University Asaba)
Arnold Adimabua Ojugo (Federal University of Petroleum Resources Effurun)
De Rosal Ignatius Moses Setiadi (Dian Nuswantoro University)
Chris Chukwufunaya Odiakaose (Dennis Osadebay University Anwai-Asaba)
Reuben Akporube Abere (Federal University of Petroleum Resources Effurun)
Frances Uche Emordi (Dennis Osadebay University Asaba)
Victor Ochuko Geteloma (Federal University of Petroleum Resources Effurun)
Patrick Ogholuwarami Ejeh (Dennis Osadebay University Anwai-Asaba)



Article Info

Publish Date
27 Jun 2024

Abstract

Customer attrition has become the focus of many businesses today – since the online market space has continued to proffer customers, various choices and alternatives to goods, services, and products for their monies. Businesses must seek to improve value, meet customers' teething demands/needs, enhance their strategies toward customer retention, and better monetize. The study compares the effects of data resampling schemes on predicting customer churn for both Random Forest (RF) and XGBoost ensembles. Data resampling schemes used include: (a) default mode, (b) random-under-sampling RUS, (c) synthetic minority oversampling technique (SMOTE), and (d) SMOTE-edited nearest neighbor (SMOTEEN). Both tree-based ensembles were constructed and trained to assess how well they performed with the chi-square feature selection mode. The result shows that RF achieved F1 0.9898, Accuracy 0.9973, Precision 0.9457, and Recall 0.9698 for the default, RUS, SMOTE, and SMOTEEN resampling, respectively. Xgboost outperformed Random Forest with F1 0.9945, Accuracy 0.9984, Precision 0.9616, and Recall 0.9890 for the default, RUS, SMOTE, and SMOTEEN, respectively. Studies support that the use of SMOTEEN resampling outperforms other schemes; while, it attributed XGBoost enhanced performance to hyper-parameter tuning of its decision trees. Retention strategies of recency-frequency-monetization were used and have been found to curb churn and improve monetization policies that will place business managers ahead of the curve of churning by customers.

Copyrights © 2024






Journal Info

Abbrev

jcta

Publisher

Subject

Computer Science & IT Decision Sciences, Operations Research & Management

Description

Journal of Computing Theories and Applications (JCTA) is a refereed, international journal that covers all aspects of foundations, theories and the practical applications of computer science. FREE OF CHARGE for submission and publication. All accepted articles will be published online and accessed ...