Tuberculosis remains a major global health issue due to its high transmission rate and delays in early detection. Early identification of suspected tuberculosis cases based on patient symptoms is important for timely screening and reducing disease transmission. This study aims to compare the performance of Logistic Regression, Support Vector Machine, K-Nearest Neighbor, and Extreme Gradient Boosting in classifying suspected tuberculosis cases using symptom-based data. The dataset, consisting of clinical indicators such as cough, fever, shortness of breath, and other tuberculosis-related symptoms, was preprocessed and divided into training and testing sets. Model performance was evaluated using accuracy, confusion matrix, sensitivity/recall, specificity, precision, F1-score, and balanced accuracy. The results show that K-Nearest Neighbor achieved the highest accuracy of 87%, compared with Support Vector Machine at 80%, Logistic Regression at 72%, and XGBoost at 71%. However, the confusion matrix showed that KNN produced 110 false-negative cases and only 19 true-positive TB cases, resulting in a TB recall of approximately 14.7%. These findings indicate that high accuracy does not necessarily reflect good screening performance, especially when sensitivity is low. Therefore, although KNN showed the highest accuracy, it cannot yet be considered adequate as a standalone tuberculosis screening model. Further improvement and validation using larger, balanced, and clinically confirmed datasets are required.
Copyrights © 2026