Building of Informatics, Technology and Science
Vol 8 No 2 (2026): September 2026

Comparative Customer Segmentation Pipelines for E-Commerce Using K-Means-KNN and UMAP-K-Means-XGBoost

Dzidan Aditya Gumilang (Universitas Sriwijaya, Palembang)
Endang Lestari Ruskan (Universitas Sriwijaya, Palembang)
Ardina Ariani (Universitas Sriwijaya, Palembang)
Ken Dhita Tania (Universitas Sriwijaya, Palembang)
Ahmad Rifai (Universitas Sriwijaya, Palembang)



Article Info

Publish Date
08 Sep 2026

Abstract

The rapid expansion of e-commerce has generated massive volumes of customer data that remain underutilized for supporting Customer Relationship Management (CRM) strategies. Conventional customer segmentation approaches commonly employ a pipeline consisting of K-Means clustering followed by K-Nearest Neighbors (KNN) classification. However, this approach exhibits limitations in handling high-dimensional data and maintaining classification performance on large-scale datasets. This study presents a comparative analysis of two customer segmentation pipelines: the conventional K-Means-KNN pipeline and the proposed Uniform Manifold Approximation and Projection (UMAP)-K-Means-XGBoost pipeline. The experiments were conducted using the E-Commerce Shopper Behavior & Lifestyle dataset, comprising approximately one million customer records and eight selected features representing transactional, psychographic, and financial behavioral characteristics. Clustering performance was evaluated using the Silhouette Score, Davies-Bouldin Index, and Calinski-Harabasz Index, while classification performance was assessed using accuracy, precision, recall, and F1-score. Experimental results demonstrate that incorporating UMAP improves cluster separability by preserving the intrinsic structure of high dimensional data, whereas XGBoost consistently outperforms KNN in downstream classification, achieving an accuracy exceeding 99%. These findings indicate that the UMAP-K-Means-XGBoost pipeline provides a more robust, scalable, and interpretable framework for customer segmentation, thereby offering more reliable decision support for data-driven CRM strategies in e-commerce environments.

Copyrights © 2026






Journal Info

Abbrev

bits

Publisher

Subject

Computer Science & IT

Description

Building of Informatics, Technology and Science (BITS) is an open access media in publishing scientific articles that contain the results of research in information technology and computers. Paper that enters this journal will be checked for plagiarism and peer-rewiew first to maintain its quality. ...