Kevin Boy Sinaga
Department of Informatics Engineering, Universitas Prima Indonesia Medan, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Machine Learning-Based Obesity Prediction and Feature Importance Analysis Using Random Forest Anjelina Dwi Puspita; Nelly Astuti Hasibuan; Kevin Boy Sinaga; Krisna Apta Jaya Zendrato; Muhammad Alief Rahman Susanto
Journal of Information System Exploration and Research Vol. 4 No. 3: July 2026
Publisher : shmpublisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.52465/joiser.v4i3.82

Abstract

Obesity prevalence continues to rise both globally and in Indonesia, yet most predictive studies rely on datasets from Western populations or clinical settings, leaving a gap in machine-learning-based evidence from Southeast Asian university communities. This study applies Random Forest to survey data from 300 respondents in Medan, Indonesia, to identify dominant determinants of Body Mass Index (BMI) category and to evaluate predictive reliability under class imbalance and limited sample size. After cleaning, 297 valid records remained; twelve behavioral and demographic features were used as model inputs. Beyond a conventional 80:20 train-test split, this study applied SMOTE to address minority-class imbalance and 10-fold stratified cross-validation to assess stability. The 80:20 split yielded 83.33% accuracy (weighted F1 = 0.82); 10-fold cross-validation produced a more conservative mean accuracy of 77.05% (SD = 6.37%), confirming that single-split evaluation overstated performance. SMOTE improved minority-class (Obesity) recall from 0.67 to 0.78 without reducing overall accuracy. Age emerged as the dominant predictor (importance = 0.25), followed by meal frequency, physical activity, and vegetable consumption frequency. These findings support Random Forest as a viable obesity-risk screening tool in resource-limited, questionnaire-based settings, while highlighting the need for imbalance-aware evaluation on modest, real-world health-survey samples.