Wayan Andre Pratama
Universitas Pendidikan Ganesha

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Random Forest and LightGBM Comparison for Acute Pain Diagnosis Using SMOTE on an Expert-Labeled Dataset Wayan Andre Pratama; I Made Gede Sunarya; Putu Hendra Suputra
Journal of Innovation and Technology Polbeng Series on Informatics (INOVTEK Polbeng - Seri Informatika) Vol. 11 No. 3 (2026): August (Inpress)
Publisher : P3M Politeknik Negeri Bengkalis

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35314/x2getn56

Abstract

Limited healthcare personnel may delay early pain assessment and encourage self-medication, increasing medication-error risk. However, evidence remains limited regarding whether bagging or boosting is more suitable for multiclass acute pain classification using imbalanced, expert-system-derived symptom data and whether SMOTE improves performance. This study compared Random Forest as a bagging approach and LightGBM as a boosting approach for classifying nine acute pain diagnostic classes without SMOTE and with SMOTE using k_neighbors=1 and 5. The dataset comprised 2,722 records and 36 discrete symptom features. Of 125 representative symptom combinations reviewed by a medical expert, 115 were considered appropriate; the remaining records were synthetically generated using the same expert-system knowledge base and inference mechanism. Data were divided using stratified 80:20 sampling, while model configuration was evaluated using five-fold cross-validation. SMOTE was applied only to training data within each fold. LightGBM without SMOTE achieved the best performance, with 83.49% accuracy, a macro F1-score of 0.81, and a weighted F1-score of 0.83, compared with 80.18%, 0.77, and 0.80 for Random Forest. With SMOTE, Random Forest achieved 78.35% and 77.61% accuracy, while LightGBM achieved 81.10% and 82.39%. Thus, LightGBM without SMOTE performed best for this dataset. Validation using real clinical data and multiple experts is required