Claim Missing Document
Check
Articles

Found 15 Documents
Search

Evaluation of Tree-Based Models for Predicting Social Assistance Recipient Status Based on National Socio-Economic Survey (SUSENAS) 2024 Yani Prihantini Hiola; Zulhijrah; I Gusti Ngurah Sentana Putra; Syella Zignora Limba; Bagus Sartono; Aulia Rizki Firdawanti; Budi Susetyo; Gerry Alfa Dito
Journal of Mathematics, Computations and Statistics Vol. 9 No. 1 (2026): Volume 09 Issue 01 (March 2026)
Publisher : Jurusan Matematika FMIPA UNM

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35580/xyyv0f37

Abstract

Abstract. Poverty is a major socioeconomic challenge in Indonesia that affects the effectiveness of social protection programs. In response to this challenge, the government has created social assistance programs to improve the welfare of the people. However, the distribution of social assistance is often considered to be inaccurate, resulting in households that are deemed eligible for social assistance not being identified as recipients. One solution to improve the accuracy of distribution is the application of machine learning in the context of classification. Several tree-based models, such as LightGBM, Random Forest, and XGBoost, were selected because of their superior capabilities compared to classical models such as logistic regression, especially in handling complex data and fulfilling model assumptions. This study compares the performance of these three models in predicting social assistance recipient status using data from the 2024 West Java Provincial National Socioeconomic Survey (SUSENAS). Model evaluation was conducted on several data pre-processing scenarios involving outlier handling, class balancing, and feature engineering. The results show that LightGBM consistently outperforms the other models on six metrics, namely Accuracy, Balanced Accuracy, F1-Score, ROC-AUC, PR-AUC, and Brier Score, out of a total of eight evaluation metrics used. SHAP analysis identifies Social Assistance History and Asset Score as the most influential features for model prediction. Friedman and Nemenyi nonparametric tests confirmed significant performance differences between LightGBM and other models based on the F1-Score, PR-AUC, and Brier Score metrics. These findings indicate that tree-based models, particularly LightGBM, can support the development of a more targeted and data-driven social assistance targeting system. Keywords: Social Assistance; Tree-Based; SHAP; SUSENAS; Hybrid Bayesian Optimization
Modeling Monthly Rainfall Data Using the Alpha Power Transformed X-Lindley Distribution in the Toba Lake Region Mohamad Khoirun Najib; Sri Nurdiati; Elis Khatizah; Aulia Rizki Firdawanti; Hendri Irwandi; Mirza Farhan Azhari; David Vijanarco Martal; Nicholas Abisha
ZERO: Jurnal Sains, Matematika dan Terapan Vol 9, No 3 (2025): Zero: Jurnal Sains Matematika dan Terapan
Publisher : UIN Sumatera Utara

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30829/zero.v9i3.25692

Abstract

Modeling rainfall is crucial for hydrological studies and climate adaptation, especially in regions with complex topography such as the Toba Lake area, North Sumatra. Classical probability distributions often struggle to represent skewness, heavy tails, and variability observed in tropical rainfall. This study explores APTXL distribution as a flexible two-parameter model. Through the alpha power transformation, APTXL extends the X-Lindley distribution by introducing an additional shape parameter, allowing better accommodation of asymmetrical and extreme values while maintaining analytical tractability. Statistical properties are derived, and parameters are estimated using maximum likelihood. The model is applied to a long-term dataset from 13 meteorological stations, covering 408 monthly observations per station. Comparative analysis against Gamma, Lognormal, and Generalized Extreme Value distributions using multiple goodness-of-fit criteria indicates that APTXL provides consistently improved performance. These results suggest APTXL as a practical tool for rainfall modeling and water-resource applications in climate-sensitive regions.
Performance Analysis of Tree-Based Models for Classifying Complete Basic Childhood Immunization in West Java Windi Pangesti; Mega Maulina; Hazelita Dwi Rahmasari; Bagus Sartono; Budi Susetyo; Aulia Rizki Firdawanti; Gerry Alfa Dito
Inferensi Vol 9 No 1 (2026)
Publisher : Department of Statistics ITS

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.12962/j27213862.v9i1.9167

Abstract

Complete basic immunization is a key public health indicator, and disparities in coverage remain a major concern in West Java. The 2024 Universal Child Immunization (UCI) rate in West Java reached only 77.47%, declining from 2023 and reflecting persistent disparities in access to immunization services, particularly between urban and rural areas. This study aims to identify the determinants of complete basic immunization among children in West Java using SUSENAS 2024 survey data (n = 4,672). Three tree- based classification algorithms CART, Random Forest, and LightGBM were applied, with class imbalance addressed using SMOTE, Tomek Links, and the SMOTE–Tomek Links hybrid method. Model performance was evaluated using balanced accuracy. The Random Forest model combined with SMOTE achieved the highest performance, with a balanced accuracy of 60.3% and an overall accuracy of 70.3%. This model also demonstrated superior capability in identifying children with incomplete immunization. Global feature importance results indicate that household spending category, KIA book ownership, maternal age at first birth, and maternal education are the strongest predictors of complete basic immunization. SHAP analysis reveals contrasting patterns: knowledge- based factors dominate in urban areas, while structural and socioeconomic constraints are more influential in rural areas. These findings underscore the importance of geographically targeted immunization strategies to support equitable access across urban and rural communities in West Java.
Evaluasi Perbandingan Model XGBoost, Random Forest, LightGBM, dan Artificial Neural Network dalam Klasifikasi Kerawanan Pangan Mardatunnisa Isnaini; Dela Gustiara; Rizqi Annafi Muhadi; Shalshabilla Shafa; Bagus Sartono; Aulia Rizki Firdawanti; Budi Susetyo; Gerry Alfa Dito
Euler : Jurnal Ilmiah Matematika, Sains dan Teknologi Volume 14 Issue 1 April 2026
Publisher : Universitas Negeri Gorontalo

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.37905/euler.v14i1.36227

Abstract

Food insecurity remains a serious household-level issue, particularly in densely populated regions such as West Java, highlighting the need for analytical approaches capable of accurately identifying vulnerable groups. Machine learning algorithms offer the potential to improve the accuracy and precision of food insecurity classification based on survey data. This study aims to compare the predictive performance and variable importance identification of four machine learning algorithms—Random Forest, Light Gradient Boosting Machine (LightGBM), Extreme Gradient Boosting (XGBoost), and Artificial Neural Network (ANN)—in predicting household food insecurity status. The analysis employs SUSENAS 2023 data covering 26,012 households with 14 predictor variables, and food insecurity is classified using the Food Insecurity Experience Scale (FIES). Class imbalance is addressed using the Synthetic Minority Over-sampling Technique (SMOTE) within a 10-fold cross-validation framework. The results show that XGBoost achieves the highest accuracy of 71%, while Random Forest provides the best balanced accuracy under the SMOTE scenario. Moreover, all algorithms consistently identify the Wealth Index as the most influential predictor based on their respective Variable Importance measures, followed by variables related to water access and food assistance. Accordingly, XGBoost is recommended in terms of accuracy, whereas Random Forest demonstrates superior balanced accuracy and prediction stability.
Classification of Drinking Water Source Suitability in West Java Using XGBoost and Cluster Analysis Based on SHAP Values: Klasifikasi Kelayakan Sumber Air Minum di Jawa Barat Menggunakan XGBoost dan Analisis Klasterisasi Berdasarkan Nilai SHAP Annisa Permata Sari; Billy; Denanda Aufadlan Tsaqif; Bagus Sartono; Aulia Rizki Firdawanti
Indonesian Journal of Statistics and Applications Vol 8 No 2 (2024)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v8i2p202-214

Abstract

Water is essential for meeting the basic needs of living organisms. In Indonesia, ensuring safe and quality drinking water is crucial for public health. However, in some regions, particularly in West Java Province, people still rely on unsuitable water sources, which can negatively impact health. The classification of water source suitability can be achieved using machine learning, such as the Extreme Gradient Boosting (XGBoost) model. XGBoost with feature selection is effective in improving prediction accuracy and minimizing overfitting. This study evaluates the performance of the XGBoost model in classifying household drinking water sources in West Java and uses the K-Means algorithm for cluster SHAP values to identify key characteristics of households with safe drinking water. The results show that the XGBoost model, with an accuracy of 77.43% and an F1-Score of 80.17%, successfully classified 4187 households, with 2349 having safe drinking water and 1838 having unsuitable sources. SHAP value analysis identified location, water collection time, and monthly per capita expenditure as significant factors influencing water source suitability. Households with water sources inside the house's fence, a short water collection time, and high monthly per capita expenditure tend to have safe drinking water sources. There are 4 clusters formed, with cluster 1 and cluster 3 needing immediate quality of drinking water sources improvement with cluster 2 as an indicator of success. Cluster 4 consists of households with high expenditure, marking it as a potential household for the government to make water quality improvements.