Community-based noncommunicable disease (NCD) surveillance data are increasingly utilized as secondary sources in public health research. Variations in dataset structure, variable nomenclature, data types, and coding systems across surveillance periods frequently impede data integration and subsequent analyses. This study developed a harmonized community-based NCD surveillance dataset using secondary data collected at UPTD Puskesmas Johar Baru, Jakarta, Indonesia, from 2019 to 2021. The methodology included dataset inventory, variable mapping, standardization of variable names, data types, and coding systems, data cleaning, derivation of analytical variables, and dataset integration using Python with the Pandas and NumPy libraries. Three surveillance datasets, comprising 196,949 observations, were successfully integrated. The harmonization process identified 158 unique variables, which were standardized into 17 analytical variables while preserving all observations. The resulting dataset offers a consistent and interoperable structure suitable for epidemiological studies, health program evaluation, and other public health research. Furthermore, the harmonization framework established in this study may serve as a practical reference for standardizing multi-period health surveillance datasets and promoting the broader use of secondary health data in future public health research
Copyrights © 2026