A'yun Fa'yuni
Program Studi Teknik Informatika, Fakultas Sains dan Teknologi, Universitas Nahdlatul Ulama Sunan Giri Bojonegoro, Jawa Timur, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

The Effect of Increasing the Dataset Size on the Performance of Sentiment Analysis Models for Tokopedia Reviews Using SVM, Naïve Bayes, and Perceptron Based on TF-IDF: Pengaruh Peningkatan Dataset terhadap Kinerja Model Analisis Sentimen Ulasan Tokopedia Menggunakan SVM, Naïve Bayes, dan Perceptron Berbasis TF-IDF A'yun Fa'yuni; Afril Efan Pajri; Ita Aristia Sa’ida
JRST (Jurnal Riset Sains dan Teknologi) Volume 10 No. 2, September 2026 :JRST
Publisher : Universitas Muhammadiyah Purwokerto

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30595/jrst.v10i2.29997

Abstract

The growth in user reviews on e-commerce platforms generates a large volume of data that has the potential to be utilised in sentiment analysis. This study aims to analyse the impact of increasing data volume on the performance of sentiment analysis models. The data was obtained through web scraping of 4,071 Tokopedia reviews on the Google Play Store. The data was then processed through a pre-processing stage comprising cleaning, normalisation, tokenisation, stopword removal, and stemming. Sentiment labelling was performed using a rating-based method supported by a lexicon, employing three identical machine learning algorithms: Support Vector Machine (SVM), Multinomial Naïve Bayes, and Perceptron, with TF-IDF feature representation. The results of the study show that Multinomial Naïve Bayes achieved the highest accuracy of 90.35%, followed by SVM at 88.93%, and Perceptron at 85.46%. All models showed an improvement in performance compared to previous research using a smaller dataset. This study demonstrates that the size of the dataset influences the accuracy and stability of sentiment analysis models ABSTRAK (Bahasa Indonesia) Pertumbuhan ulasan pengguna pada platform e-commerce menghasilkan volume data yang besar dan berpotensi dimanfaatkan dalam analisis sentimen. Penelitian ini bertujuan untuk menganalisis pengaruh penambahan jumlah data terhadap kinerja model analisis sentimen. Data diperoleh melalui web scraping ulasan Tokopedia di Google Play Store sebanyak 4.071 data. Data kemudian diproses melalui tahap pre-processing yang meliputi cleaning, normalisasi, tokenisasi, stopword removal, dan stemming. Pelabelan sentimen dilakukan menggunakan metode rating based dengan dukungan leksikon, dengan menggunakan tiga algoritma pembelajaran mesin yang sama, yaitu Support Vector Machine (SVM), Multinomial Naïve Bayes, dan Perceptron dengan representasi fitur TF-IDF. Hasil penelitian menunjukkan bahwa Multinomial Naïve Bayes memperoleh akurasi tertinggi sebesar 90.35%, diikuti SVM sebesar 88.93%, dan Perceptron sebesar 85.46%. Seluruh model mengalami peningkatan performa dibandingkan penelitian sebelumnya dengan dataset lebih kecil. Penelitian ini menunjukkan bahwa jumlah dataset berpengaruh terhadap akurasi dan stabilitas model analisis sentimen