Mohammad Masjkur
Study Program on Statistics and Data Science, IPB University, Indonesia

Published : 2 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 2 Documents
Search

Comparative Study of Multiclass Models for Identifying Household Drinking Water Sources in West Java Defri Ramadhan Ismana; Bagus Sartono; Erfiani; Yuri Nurdiantami; Mohammad Masjkur; Aam Alamudi; Rahma Anisa
Indonesian Journal of Statistics and Applications Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v10i1p119-128

Abstract

Data imbalance is a common challenge in classification modeling, typically driven by rare events such as fraud, credit default, student dropout, or infectious diseases. Multi-class imbalance is inherently more complex than binary classification because a class can be a minority relative to one class yet a majority to another. Access to safe drinking water is one of the key factors for public health. West Java, as the province with the largest population in Indonesia, has yet to achieve the target of 100% of households having access to safe drinking water. Identifying the types of drinking water sources can be approached using multi-class classification modeling. However, this case presents an issue of data imbalance. The CatBoost method is one of the recommended approaches for multi-class classification with imbalanced data. Additionally, the TabNet method is also considered to perform well in such cases. Therefore, this study aims to apply both methods and compare their performance in classifying household drinking water sources in West Java using 2023 SUSENAS data. The results show that CatBoost yielded an average MCC of 0.225, whereas TabNet exhibited an average of 0.209. In terms of computational efficiency, CatBoost recorded an average execution time of 25 seconds, compared to 134 seconds for TabNet. Therefore, the CatBoost method provides better multi-class classification performance and faster execution compared to TabNet. Although TabNet is superior in classifying minority classes, CatBoost is more accurate in predicting majority classes and identifying the most influential variables, such as building ownership status and regional classification
K-Prototypes Algorithm for School Indexing in Report Card-Based Student Admissions: Algoritma K-Prototypes untuk Indeks Sekolah pada Penerimaan Mahasiswa Baru Jalur Rapor Ervina Dwi Anggrahini; Mohammad Masjkur; Utami Dyah Syafitri
Indonesian Journal of Statistics and Applications Vol 9 No 1 (2025)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v9i1p117-135

Abstract

Institut Pertanian Bogor, also known as IPB University, is a state university that was ranked first as the best university in Indonesia by the Ministry of Research and Technology in 2020. It has three main channels in the new student admission selection system. The selection method is called “Seleksi Nasional Berdasarkan Prestasi”. “Seleksi Nasional Berdasarkan Prestasi” is one of the new student admission pathways at IPB University based on report cards without a test. The selection of new student admissions based on report cards requires creating a school index to assess the quality and commitment of each school by grouping schools among “Seleksi Nasional Berdasarkan Prestasi” applicants. One method that can be used is the K-Prototypes algorithm. K-Prototypes can be used to cluster large and mixed-type data (numeric and categorical) by combining distance measures from two non-hierarchical methods, namely the K-Means and K-Modes algorithms. Based on the analysis, the K-Prototypes algorithm yields three optimal clusters, each with distinct characteristics. Cluster 1 is the lowest cluster because it comprises schools with the lowest quality and commitment to new student admissions at IPB University, as indicated by the report card. Cluster 2 has a quality that is not superior to Cluster 3 but is higher than that of Cluster 1. Cluster 3 is the best cluster because it consists of schools that have high quality and commitment to new student admissions at IPB University through the report card route.