E-commerce multi-warehouse distribution systems face the challenge of logistics data segmentation due to heterogeneous and unbalanced transaction characteristics, as well as high overlap between variables. This study aims to segment warehouses based on logistics costs and transaction patterns using the K-Means and Fuzzy C-Means algorithms. The study integrates ratio-based feature engineering through the Logistics Cost Ratio (LCR), Value Density (VD), and Transaction Intensity (TI), and applies robust scaling and 1% outlier trimming to improve clustering stability. The dataset consists of 13,550 transactions from four main warehouses, yielding 13,284 valid data points after preprocessing. Evaluation was conducted using the Silhouette Score, Davies-Bouldin Index (DBI), and Calinski-Harabasz Index (CHI). The results show that K-Means with k = 3 yields the best performance with a Silhouette Score of 0.421, a DBI of 0.584, and a CHI of 12,273.98. Ratio-based feature transformation was proven to produce a more balanced and interpretable cluster distribution compared to the use of raw data. The resulting segmentation consists of high-efficiency clusters, regular transaction clusters, and high-logistics-cost clusters that can be used as a basis for data-driven logistics distribution decision-making in multi-warehouse e-commerce systems.
Copyrights © 2026