Claim Missing Document
Check
Articles

Found 1 Documents
Search

Optimization of the SVM Algorithm with Chi-Square and SMOTE for High-Dimensional Stunting Data in Samarinda Taghfirul Azhima Yoga Siswa; Naufal Azmi Verdikha
Prisma Sains : Jurnal Pengkajian Ilmu dan Pembelajaran Matematika dan IPA IKIP Mataram Vol. 14 No. 3: July 2026
Publisher : Universitas Pendidikan Mandalika

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.33394/j-ps.v14i3.21848

Abstract

Stunting is still considered a serious problem and has become a focus of government attention in the national research priorities for 2020-2024, with a stunting prevalence of 21.6% in 2022. The city of Samarinda ranked second highest after the Kutai Kartanegara district in East Kalimantan Province in 2022, with a percentage of 25.3%. Based on previous data mining research trends, the use of classification methods such as KNN, Naïve Bayes, Random Forest, Neural Network, CART, and SVM has generally produced quite high accuracy but still focuses on low-dimensional data, which can lead to significant information loss, potential overfitting, and difficulty in interpretation. Whereas in research topics related to high-dimensional stunting data, the majority still yield low accuracy. This is further supported by the still prevalent class imbalance found in other studies, which can affect the accuracy and recall values of the built model's performance. The objective of this research is to apply the Support Vector Machine (SVM) algorithm with Chi-Square feature selection and the Synthetic Minority Over-sampling Technique (SMOTE) to address high-dimensional stunting data and handle class imbalance. The dataset for this research is sourced from the Samarinda City Health Office, consisting of 26 community health centers with 20 attributes and 102,534 records. The division of training and testing data uses the k-fold cross-validation technique with k=10. The research results show the model's performance is very good, with an accuracy of 96.6%, supported by precision, recall, and f1-score values of 97% each.