Marini
Institut Sains dan Bisnis Atma Luhur

Published : 2 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 2 Documents
Search

Ensemble Learning for Pediatric Stunting Detection: A Comparative Study of XGBoost, Random Forest, and LightGBM with Oversampling Techniques Tri Sugihartono; Djoko Soetarno; Rahmat Sulaiman; Sarwindah; Marini; Fitriyani
Journal of Information System and Informatics Vol 8 No 2 (2026): April
Publisher : Asosiasi Doktor Sistem Informasi Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.63158/journalisi.v8i2.1568

Abstract

Stunting, driven by chronic childhood malnutrition, remains a critical global public health concern. Early detection is persistently challenged by class imbalance in pediatric health datasets and the absence of systematic comparisons between oversampling strategies and ensemble classifiers. This study develops and evaluates an ensemble learning pipeline for stunting detection, benchmarking XGBoost, Random Forest, and LightGBM across five oversampling configurations — Original, SMOTE, ADASYN, Borderline-SMOTE, and SMOTE-ENN — using 10,000 pediatric health records from posyandu activities in Bangka Belitung Province, Indonesia. Seven anthropometric and demographic features were utilized, with stratified 80:20 train-test splitting and five-fold cross-validation. XGBoost with original imbalanced data achieved the highest Recall (0.9573) and a competitive F1-Score (0.9158), while LightGBM with SMOTE delivered the strongest balanced performance (F1-Score: 0.9160, ROC-AUC: 0.8431). SMOTE-ENN consistently underperformed across all classifiers. To our knowledge, this is the first study to simultaneously compare five oversampling strategies across three ensemble models within a unified framework, offering a foundation for high-sensitivity stunting surveillance in resource-constrained healthcare settings.
Two-Stage Tuning of Machine Learning Models for Heart Disease Classification on Synthetic Data Marini; Tri Sugihartono; Chandra Kirana; Benny Wijaya; Hamidah
Journal of Information System and Informatics Vol 8 No 3 (2026): June
Publisher : Asosiasi Doktor Sistem Informasi Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.63158/journalisi.v8i3.1599

Abstract

Heart disease remains a leading global cause of mortality, highlighting the need for accurate early risk classification. This study benchmarks Random Forest, XGBoost, and Logistic Regression for heart disease risk classification using a synthetic, perfectly balanced dataset, while addressing performance limitations caused by inadequate hyperparameter configuration. The dataset comprised 70,000 samples with a 50/50 class distribution and 18 clinical and demographic features. Although useful for controlled benchmarking, synthetic balanced data may yield optimistic estimates and may not fully represent real-world clinical variability. Each model was implemented in a scikit-learn Pipeline with median imputation and, where applicable, standard scaling. A two-stage tuning strategy was applied by combining RandomizedSearchCV with GridSearchCV refinement to optimize model configurations systematically. Under these benchmarking conditions, XGBoost achieved the best test performance, with an F1-score of 99.34%, AUC-ROC of 99.97%, and accuracy of 99.34%. Random Forest obtained an F1-score of 99.20% and AUC-ROC of 99.95%, while Logistic Regression achieved an F1-score of 99.12% and AUC-ROC of 99.95%. Age, pain in the arms/jaw/back, and cold sweats/nausea were the most influential predictors. The proposed framework is reproducible, computationally efficient, and suitable for validation on heterogeneous clinical datasets.