Inaccurate or noisy data presents a significant challenge in machine learning, particularly in unsupervised clustering tasks. This study evaluates the robustness and performance of two popular clustering algorithms, K-Means and DBSCAN, against various levels of Gaussian noise (5%, 10%, 20%, and 30%) injected into a customer dataset. Evaluation was conducted using Silhouette Score and Davies-Bouldin Index (DBI). Initial results indicated that DBSCAN performed slightly better with a Silhouette Score of 0.4817 compared to K-Means at 0.4101. However, after noise injection, K-Means demonstrated superior stability by maintaining more consistent cluster memberships, whereas DBSCAN was more sensitive to distance variations, leading to significant fluctuations in cluster assignments. The study concludes that K-Means is more reliable for datasets where cluster integrity must be preserved despite minor data irregularities.
Copyrights © 2026