Bagus Sartono
Study Program on Statistics and Data Science, IPB University, Indonesia

Published : 3 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 3 Documents
Search

Comparative Study of Multiclass Models for Identifying Household Drinking Water Sources in West Java Defri Ramadhan Ismana; Bagus Sartono; Erfiani; Yuri Nurdiantami; Mohammad Masjkur; Aam Alamudi; Rahma Anisa
Indonesian Journal of Statistics and Applications Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v10i1p119-128

Abstract

Data imbalance is a common challenge in classification modeling, typically driven by rare events such as fraud, credit default, student dropout, or infectious diseases. Multi-class imbalance is inherently more complex than binary classification because a class can be a minority relative to one class yet a majority to another. Access to safe drinking water is one of the key factors for public health. West Java, as the province with the largest population in Indonesia, has yet to achieve the target of 100% of households having access to safe drinking water. Identifying the types of drinking water sources can be approached using multi-class classification modeling. However, this case presents an issue of data imbalance. The CatBoost method is one of the recommended approaches for multi-class classification with imbalanced data. Additionally, the TabNet method is also considered to perform well in such cases. Therefore, this study aims to apply both methods and compare their performance in classifying household drinking water sources in West Java using 2023 SUSENAS data. The results show that CatBoost yielded an average MCC of 0.225, whereas TabNet exhibited an average of 0.209. In terms of computational efficiency, CatBoost recorded an average execution time of 25 seconds, compared to 134 seconds for TabNet. Therefore, the CatBoost method provides better multi-class classification performance and faster execution compared to TabNet. Although TabNet is superior in classifying minority classes, CatBoost is more accurate in predicting majority classes and identifying the most influential variables, such as building ownership status and regional classification
Digital Newsworthiness Scores Model Using a Combination of Unsupervised and Supervised Learning Approaches: Pemodelan Skor Kelayakan Berita Digital dengan Pendekatan Kombinasi Unsupervised dan Supervised Learning Reza Felix Citra; Aji Hamim Wigena; Bagus Sartono
Indonesian Journal of Statistics and Applications Vol 9 No 1 (2025)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v9i1p86-99

Abstract

The rapid evolution of digital technology has transformed the media landscape, making news more accessible while also introducing challenges related to content quality and accuracy. The rise of misinformation and fake news has diminished public trust in traditional media. A method for evaluating the quality and potential impact of news articles prior to publication. By adapting credit risk scoring principles, a model was used to predict the suitability of news content based on factors such as title length, number of images, news category, and publication timing. A variable target was firstly formed using three clustering methods: K-Means, K-Modes, and K-Medoids. The results indicated that K-Means outperformed the other methods, leading us to use its outcomes for determining publication suitability. Subsequently, stepwise logistic regression was applied to implement the credit risk scoring approach, allowing for variable selection and assessment of importance. Ultimately, ten variables were identified to generate a newsworthiness score, with minimum and maximum scores of 997 and 1407, respectively. The average scores for articles deemed publishable and not publishable were 1137 and 1110. A cutoff score of 1123 was established based on these averages, categorizing 6708 articles (57.9%) as suitable for publication. These findings aim to assist media organizations in refining their content curation processes, thereby enhancing the overall quality of news consumption.
Comparing Self-Paced Ensemble and RUSBoost for Imbalanced Poverty Classification in West Java Nur Andi Setiabudi; Bagus Sartono; Utami Dyah Syafitri; Komang Budi Aryasa
Indonesian Journal of Statistics and Applications Vol 9 No 2 (2025)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v9i2p218-229

Abstract

Class imbalance remains a major challenge in classification modelling that frequently leads to biased predictive models. This study aimed to compare two ensemble techniques based on an undersampling approach, namely Self-Paced Ensemble and RUSBoost, for handling imbalanced classification in poverty identification in West Java. The results suggested that RUSBoost consistently outperformed Self-Paced Ensemble across the most critical metrics. It showed better balance in classification outcomes. When the objective is to maximize the identification of poor households, the default threshold in the RUSBoost model was prefered. On the other hand, if precision is prioritized due to limited resources, the Youden Index threshold offers a better alternative. Given the overall evaluation metrics, RUSBoost with the default threshold was suggested as the most reliable and well-balanced option among the compared models for classifying poor households in West Java under imbalanced data condition