The increasing complexity of healthcare data presents significant challenges in identifying meaningful patient patterns, particularly in chronic disease management such as diabetes. Traditional analytical approaches often fail to capture hidden structures within multidimensional clinical data. This study aims to uncover latent patient profiles by applying unsupervised learning techniques to integrated demographic and physiological attributes, including age, body mass index, blood pressure, and gender. A comparative clustering framework was employed using K-Means, Hierarchical Clustering, and DBSCAN, with performance evaluated through Silhouette Score and Davies–Bouldin Index. The experimental results indicate that the optimal clustering configuration is achieved with K=3, where K-Means outperforms the other methods by producing more compact and well-separated clusters. Visual validation using Principal Component Analysis further confirms the structural coherence of the clusters. Specifically, Cluster 0 represents 60.4% of the dataset, with an average BMI of 24.31 and systolic blood pressure of 125.77 mmHg. Cluster 1, containing 29.3% of the patients, exhibits a higher average age of 61.16 years, with stable BMI and blood pressure values. Cluster 2, accounting for 10.2% of the patients, shows elevated physiological characteristics, including an average BMI of 26.49 and systolic pressure of 159.55 mmHg. The findings reveal three distinct patient groups: a majority group with moderate physiological conditions, an older-age group with relatively stable health indicators, and a smaller high-risk group characterized by elevated body mass index and significantly higher blood pressure levels. These results demonstrate the capability of unsupervised learning to identify clinically meaningful subpopulations without labeled data. In conclusion, this study provides a robust and scalable framework for patient segmentation using multidimensional healthcare data. The approach supports data-driven insights for healthcare analysis and risk stratification. Future work should explore the integration of additional clinical variables and advanced clustering techniques to enhance model interpretability and performance.
Copyrights © 2026