Claim Missing Document
Check
Articles

Found 24 Documents
Search

Handling Unbalanced Data with SMOTE Algorithm for Unemployment Classification in Lima Puluh Kota Regency Using CART Method Aldwi Riandhoko; Nonong Amalita; Dodi Vionanda; Admi Salma
Indonesian Journal of Statistics and Applications Vol 8 No 2 (2024)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v8i2p166-177

Abstract

Unemployment is a problem that occurs in the labor force, where high unemployment is caused by the low ability of the labor force. A region that is still experiencing unemployment problems in West Sumatera is Lima Puluh Kota Regency. Unemployment in Lima Puluh Kota Regency is caused by the low competence of human resources to fulfill employment market requirements. Based on the results of the Sakernas survey in August 2023, Lima Puluh Kota Regency has more employed labor force than unemployed labor force, so this results in unbalanced data. A method that can overcome unbalanced data is Synthetic Minority Oversampling Technique (SMOTE). SMOTE is a technique with addition of synthetic data in minority class so that the proportion is balanced. Data imbalance conditions need to be handled so as to improve the performance of the classification model. Classification and Regression Trees (CART) is a classification technique with a decision tree method that can obtain the characteristics of a classification. The purpose of this research is to compare the CART model before and after applying SMOTE which can be measured by comparing the highest Area Under Curve (AUC) value. The AUC value in the CART method before SMOTE applied has a value of 62.1% while the AUC value in the CART method after SMOTE applied has a value of 70.2%. Therefore, it can be concluded that the CART classification analysis after SMOTE applied is able to provide better performance compared to the CART classification analysis before SMOTE applied.
Implementation of Fuzzy C-Means Algorithm for Clustering Provinces in Indonesia Based on Micro and Small Industry Ratio in Village Areas Frandito Rahmanesta; Zamahsary Martha; Dodi Vionanda; Zilrahmi Zilrahmi
Indonesian Journal of Statistics and Applications Vol 8 No 2 (2024)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v8i2p178-190

Abstract

Post-economic crisis, the micro and small industries contribute the most labor compared to other industries. Regional development sourced from small micro industries is a strategic force in developing a country because the development of small micro industries leads to realizing equitable welfare to reduce income inequality. Development in village areas is an important factor for regional development, reducing inequality between regions, and alleviating poverty. However, based on the 2018 PODES survey, there are regional imbalances in Indonesia in the small micro industry which is centralized on Java Island. Therefore, clustering and characteristics of the province were carried out based on the PODES survey of the small micro industry sector. This research uses the Fuzzy C-Means algorithm to cluster 34 provinces in Indonesia based on the ratio of small micro industries in village areas in 2021, to see how the development of small micro industries in village areas in each province in Indonesia. Fuzzy C-Means is one of the data clustering techniques that uses a fuzzy clustering model, where cluster formation is based on a membership degree value that varies between 0 and 1. The Fuzzy C-Means algorithm generates 4 clusters, cluster 1 and 2 represents provinces with high and very high micro and small industry development in village areas and cluster 3 and 4 represents provinces with medium and low micro and small industry development in village areas. The Fuzzy C-Means algorithm produces a good cluster structure with a silhouette coefficient value of 0,6406.
Comparison of K-Means and K-Medoids in Clustering Regency/City in West Sumatra Province Based on Environmental Indicators Silfi Robiati; Dina Fitria; Dodi Vionanda; Dwi Sulistiowati
Indonesian Journal of Statistics and Applications Vol 8 No 2 (2024)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v8i2p191-201

Abstract

The Environmental Quality Index is an index that describes the condition of environmental management results nationally, and generalises from all regencies/cities and provinces in Indonesia. Although the Environmental Quality Index of West Sumatra Province has increased, there are still regencies/cities in West Sumatra Province have decreasing Environmental Quality Index. Therefore, it is necessary to conduct further analysis, one of which is to form a group of regencies/cities into a group according to their similarities or characteristics. This study aims to compare the K-Means and K-Medoids methods in grouping regencies/cities in West Sumatra Province based on environmental quality indicators in 2023. The data used in this research is secondary data, which is orginally the publication of Central Bureau of Statistics namely Sumatera Barat Dalam Angka in 2024. The research compares the K-Means cluster method and the K-Medoids cluster method. It concludes K-Means better than K-Medoids methods based on DB index with three clusters. First cluster has 12 regencies/cities with a high average air quality index, the second cluster has 6 regencies/cities that have small amounts of waste, and the third cluster has 1 city with a high average water quality index and land quality index, but a large amount of waste.   Keywords: Cluster, Comparison, Environmental, K-Means, K-Medoids
Do Prestigious Schools Still Exist in Padang? An Exploratory Study on State Junior High School Admission 2025 in Padang Dodi Vionanda; Raihan Attaya Wood; Amelia Susrifalah
Rangkiang Mathematics Journal Vol. 4 No. 2 (2025): Rangkiang Mathematics Journal
Publisher : Department of Mathematics, Universitas Negeri Padang (UNP)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/rmj.v4i2.106

Abstract

In this study, we perform an exploratory study of New Student Admission datasets for public Junior High School in Padang in 2025. We utilized tables, barplots, and boxplots to present information contained in datasets and we carried out cluster analysis using HDBSCAN algorithm. For this study we made use of admitted students’ datasets for each admission pathway of all state Junior High Schools in Padang in 2025. We carried out this study to investigate the emergence of prestigious schools among public Junior High School in Padang amid the implementation on zoning system. Our study reveals the presence of group of prestigious schools along with group of schools that admitted students mostly live nearby the schools. Hence, it is recommended for Padang Municipal government to improve the quality of schools that are not considered as prestigious schools since there are many schools that admitted students mostly live nearby the school.