Rahma Anisa
Study Program on Statistics and Data Science, IPB University, Indonesia

Published : 2 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 2 Documents
Search

Comparative Study of Multiclass Models for Identifying Household Drinking Water Sources in West Java Defri Ramadhan Ismana; Bagus Sartono; Erfiani; Yuri Nurdiantami; Mohammad Masjkur; Aam Alamudi; Rahma Anisa
Indonesian Journal of Statistics and Applications Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v10i1p119-128

Abstract

Data imbalance is a common challenge in classification modeling, typically driven by rare events such as fraud, credit default, student dropout, or infectious diseases. Multi-class imbalance is inherently more complex than binary classification because a class can be a minority relative to one class yet a majority to another. Access to safe drinking water is one of the key factors for public health. West Java, as the province with the largest population in Indonesia, has yet to achieve the target of 100% of households having access to safe drinking water. Identifying the types of drinking water sources can be approached using multi-class classification modeling. However, this case presents an issue of data imbalance. The CatBoost method is one of the recommended approaches for multi-class classification with imbalanced data. Additionally, the TabNet method is also considered to perform well in such cases. Therefore, this study aims to apply both methods and compare their performance in classifying household drinking water sources in West Java using 2023 SUSENAS data. The results show that CatBoost yielded an average MCC of 0.225, whereas TabNet exhibited an average of 0.209. In terms of computational efficiency, CatBoost recorded an average execution time of 25 seconds, compared to 134 seconds for TabNet. Therefore, the CatBoost method provides better multi-class classification performance and faster execution compared to TabNet. Although TabNet is superior in classifying minority classes, CatBoost is more accurate in predicting majority classes and identifying the most influential variables, such as building ownership status and regional classification
Optimization of Fuzzy C-Means Clustering with Particle Swarm Optimization on Socioeconomic Indicators of ASEAN Countries Cindy Indriyani; Siti Arbaynah; Ananda Putra Wijaya; Lusi Oktaviani; Fadhilah Yumna; Norashida Othman; Sachnaz Desta Oktarina; Rahma Anisa
Indonesian Journal of Statistics and Applications Vol 9 No 2 (2025)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v9i2p274-288

Abstract

Grouping data based on similarity in characteristics is commonly applied in various exploratory analyses. The Fuzzy C-Means algorithm offers flexibility through the degree of membership of data points in each cluster, but it is vulnerable to poor cluster center initialization, which increases the risk of getting trapped in local optima. To enhance the performance of Fuzzy C-Means, this study integrates the Particle Swarm Optimization method for determining cluster centers. The evaluation is conducted by comparing Fuzzy C-Means and Fuzzy C-Means-Particle Swarm Optimization across several cluster counts using three internal validation metrics, namely the silhouette coefficient, partition coefficient, and Xie-Beni Index. The results show that Fuzzy C-Means-Particle Swarm Optimization consistently yields higher silhouette coefficient and partition coefficient values, along with lower Xie-Beni Index values, compared to standard Fuzzy C-Means. This indicates that the integration of Particle Swarm Optimization can improve clustering quality in terms of cluster compactness and separation. This hybrid approach demonstrates significant potential in complex data clustering scenarios.