Customer segmentation is widely used to analyze customer transaction patterns and support effective business strategies. However, previous studies have reported inconsistent findings regarding the impact of data normalization on clustering quality across different datasets and algorithms. This study investigates the effect of data normalization on RFM-based customer segmentation using K-Means and DBSCAN. Two transaction datasets, Online Retail II and TransJakarta, were analyzed under three preprocessing scenarios: no normalization, Min-Max normalization, and Z-Score normalization. Clustering performance was evaluated using the Silhouette Score and Davies–Bouldin Index (DBI). For the Online Retail II dataset, K-Means achieved the best performance without normalization (Silhouette Score = 0.9845), while DBSCAN produced valid clusters only after Z-Score normalization. For the TransJakarta dataset, both algorithms performed best without normalization, whereas DBSCAN identified up to 20 clusters and noise points. These findings highlight that the effectiveness of normalization depends on dataset characteristics and the clustering algorithm used.
Copyrights © 2026