Sahri Sahri
Universitas Nahdlatul Ulama Sunan Giri, Bojonegoro, Jawa Timur, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Comparison of Feature Selection Methods in Classifying Poverty Levels in Indonesia Using Comparative Machine Learning Methods Nuniska Dwi Kamayanti; Ifnu Wisma Dwi Prastya; Sahri Sahri
Journal of Applied Informatics and Computing Vol. 10 No. 3 (2026): June 2026
Publisher : Politeknik Negeri Batam

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30871/jaic.v10i3.12604

Abstract

Poverty classification requires models capable of handling multidimensional data and imbalanced class distributions. This study aims to develop and compare several machine learning algorithms for classifying poverty levels in Indonesia, as well as to analyze the impact of feature selection and reduction methods on model performance. The study employs a comparative approach using a secondary dataset consisting of 514 districts/cities with socio-economic indicators and a binary target variable. The methodology includes data preprocessing, the application of Chi-Square, Pearson Correlation, and Principal Component Analysis (PCA), and the handling of imbalanced data using the Synthetic Minority Oversampling Technique (SMOTE). Modelling is conducted using Random Forest, Support Vector Machine (SVM), Logistic Regression, and Artificial Neural Network (ANN), with evaluation performed using Stratified K-Fold Cross Validation and metrics including accuracy, precision, recall, and F1-score. The results indicate that Chi-Square and Pearson Correlation outperform PCA, with Random Forest achieving the best performance, attaining an accuracy of 0.9854 and an F1-score of 0.9507, while effectively detecting the minority class. Therefore, the combination of Chi-Square and Random Forest is identified as the most effective approach in this study, as it produces a model that is accurate, stable, and capable of handling imbalanced data.