Muhammad Ivan Ardiansyah
Departement of Data Science, Universitas Muhammadiyah Semarang, Central Java, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Comparative analysis of machine learning algorithms for tuberculosis classification based on symptom data Ihsan Fathoni Amri; Muhammad Ivan Ardiansyah; Wikanastri Hersoelistyorini
Journal Focus Action of Research Mathematic (Factor M) Vol. 9 No. 1 (2026): June 2026
Publisher : Universitas Islam Negeri (UIN) Syekh Wasil Kediri

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30762/f_m.v9i1.8357

Abstract

Tuberculosis remains a major global health issue due to its high transmission rate and delays in early detection. Early identification of suspected tuberculosis cases based on patient symptoms is important for timely screening and reducing disease transmission. This study aims to compare the performance of Logistic Regression, Support Vector Machine, K-Nearest Neighbor, and Extreme Gradient Boosting in classifying suspected tuberculosis cases using symptom-based data. The dataset, consisting of clinical indicators such as cough, fever, shortness of breath, and other tuberculosis-related symptoms, was preprocessed and divided into training and testing sets. Model performance was evaluated using accuracy, confusion matrix, sensitivity/recall, specificity, precision, F1-score, and balanced accuracy. The results show that K-Nearest Neighbor achieved the highest accuracy of 87%, compared with Support Vector Machine at 80%, Logistic Regression at 72%, and XGBoost at 71%. However, the confusion matrix showed that KNN produced 110 false-negative cases and only 19 true-positive TB cases, resulting in a TB recall of approximately 14.7%. These findings indicate that high accuracy does not necessarily reflect good screening performance, especially when sensitivity is low. Therefore, although KNN showed the highest accuracy, it cannot yet be considered adequate as a standalone tuberculosis screening model. Further improvement and validation using larger, balanced, and clinically confirmed datasets are required.