Despite various mitigation efforts, Human Immunodeficiency Virus (HIV)/Acquired Immunodeficiency Syndrome (AIDS) remains a significant public health issue with widespread impacts in Indonesia. One of the challenges in HIV/AIDS classification using machine learning is data imbalance, where the number of HIV cases is smaller than Non-HIV cases. The aim of this study is to analyze the performance of the CatBoost algorithm in classification tasks and to evaluate the impact of the Synthetic Minority Oversampling Technique (SMOTE) on improving model performance in imbalanced datasets. The research method involves applying the CatBoost algorithm to the original dataset as well as to data that has been processed using SMOTE-based oversampling. Furthermore, model performance is evaluated using Precision, Recall, F1-Score, and Precision-Recall Area Under Curve (PR-AUC) metrics. The SMOTE + CatBoost model achieved an accuracy of 95%, precision of 93%, recall of 92%, F1-Score of 93%, and PR-AUC of 0.953, all of which are higher than those of the CatBoost Baseline model. In addition, the number of undetected HIV cases was reduced from 28 to 13 cases. The findings indicate that the integration of SMOTE with the CatBoost algorithm improves model performance, resulting in better classification outcomes on imbalanced datasets compared to the CatBoost Baseline, and potentially supports a more effective HIV/AIDS early detection system.
Copyrights © 2026