Filter By Year

1945 2024


Found 1 documents
Search 10.58477/cj.v4i2.497 , by doi

Analisis Sentimen Ulasan Pengguna inDrive Menggunakan IndoBERT dan Algoritma Genetika pada Klasifikasi K-Nearest Neighbor Muhammad Sigit Nurhafid; Rudiman Rudiman; Taghfirul Azhima Yoga
Computer Journal Vol. 4 No. 2 (2026): August
Publisher : Yayasan Pendidikan Mitra Mandiri Aceh (YPMMA)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.58477/cj.v4i2.497

Abstract

This study analyzes sentiment in InDrive user reviews from the Google Play Store using IndoBERT, Genetic Algorithm (GA), and K-Nearest Neighbor (KNN). A total of 2,000 reviews were assigned to three sentiment classes: 1,695 negative, 235 positive, and 70 neutral reviews. The pretrained indobenchmark/indobert-base-p1 model was used as a feature extractor by taking the [CLS] representation to produce 768-dimensional embeddings. The dataset was divided using a stratified 80:20 split into 1,600 training and 400 testing samples. The optimal K value was determined through stratified five-fold cross-validation on the training data. GA was applied only to the training set using a population of 30 individuals, 25 generations, a crossover rate of 0.8, a mutation rate of 0.005, and a feature penalty of 0.002. GA selected 250 features, reducing the dimensionality by 67.45%. IndoBERT + KNN correctly classified 365 of 400 test samples, achieving 91.25% accuracy (95% CI: 88.07–93.64%) and a macro F1-score of 66.88%. IndoBERT + GA + KNN correctly classified 362 samples, achieving 90.50% accuracy (95% CI: 87.23–93.00%) and a macro F1-score of 62.71%. Both models exceeded the 84.75% majority-class baseline. However, only three and two of the 14 neutral samples were correctly classified, respectively. GA substantially reduced feature dimensionality but did not improve predictive performance, indicating a trade-off between representation compactness and minority-class classification performance.

Page 1 of 1 | Total Record : 1