Routine disease-surveillance data combine heterogeneous clinical, demographic, spatial, and programmatic variables that cannot be represented adequately by a single clustering technique. This study developed a shared analytical framework and applied it to three datasets from the Yogyakarta Special Region Health Office: 20,660 HIV visits, 663 valid dengue cases aggregated into 14 district profiles, and 307 malaria records. Disease-specific preprocessing was followed by K-Prototypes clustering for mixed-type HIV data, Ward agglomerative hierarchical clustering for district-level dengue profiles, and K-Means clustering for malaria records. Candidate solutions were evaluated using elbow behavior, Silhouette Score, Davies-Bouldin Index, epidemiological interpretability, and operational usefulness. The analyses produced four HIV clusters with distinct demographic and clinical profiles, three dengue surveillance zones including four high-priority districts, and three malaria profiles differentiated by age, occupation, residency, imported-case status, and temperature. The selected configurations represented practical compromises between internal validity and actionable interpretation. Prototype black-box testing passed all documented HIV, dengue, and malaria scenarios. The framework enables consistent multi-disease surveillance while preserving disease-appropriate analytical choices, although external and prospective validation remains necessary.
Copyrights © 2026