JRST (Jurnal Riset Sains dan Teknologi)
Volume 10 No. 2, September 2026 :JRST

The Effect of Increasing the Dataset Size on the Performance of Sentiment Analysis Models for Tokopedia Reviews Using SVM, Naïve Bayes, and Perceptron Based on TF-IDF: Pengaruh Peningkatan Dataset terhadap Kinerja Model Analisis Sentimen Ulasan Tokopedia Menggunakan SVM, Naïve Bayes, dan Perceptron Berbasis TF-IDF

A'yun Fa'yuni (Program Studi Teknik Informatika, Fakultas Sains dan Teknologi, Universitas Nahdlatul Ulama Sunan Giri Bojonegoro, Jawa Timur, Indonesia)
Afril Efan Pajri (Program Studi Teknik Informatika, Fakultas Sains dan Teknologi, Universitas Nahdlatul Ulama Sunan Giri Bojonegoro, Jawa Timur, Indonesia)
Ita Aristia Sa’ida (Program Studi Teknik Informatika, Fakultas Sains dan Teknologi, Universitas Nahdlatul Ulama Sunan Giri Bojonegoro, Jawa Timur, Indonesia)



Article Info

Publish Date
01 Sep 2026

Abstract

The growth in user reviews on e-commerce platforms generates a large volume of data that has the potential to be utilised in sentiment analysis. This study aims to analyse the impact of increasing data volume on the performance of sentiment analysis models. The data was obtained through web scraping of 4,071 Tokopedia reviews on the Google Play Store. The data was then processed through a pre-processing stage comprising cleaning, normalisation, tokenisation, stopword removal, and stemming. Sentiment labelling was performed using a rating-based method supported by a lexicon, employing three identical machine learning algorithms: Support Vector Machine (SVM), Multinomial Naïve Bayes, and Perceptron, with TF-IDF feature representation. The results of the study show that Multinomial Naïve Bayes achieved the highest accuracy of 90.35%, followed by SVM at 88.93%, and Perceptron at 85.46%. All models showed an improvement in performance compared to previous research using a smaller dataset. This study demonstrates that the size of the dataset influences the accuracy and stability of sentiment analysis models ABSTRAK (Bahasa Indonesia) Pertumbuhan ulasan pengguna pada platform e-commerce menghasilkan volume data yang besar dan berpotensi dimanfaatkan dalam analisis sentimen. Penelitian ini bertujuan untuk menganalisis pengaruh penambahan jumlah data terhadap kinerja model analisis sentimen. Data diperoleh melalui web scraping ulasan Tokopedia di Google Play Store sebanyak 4.071 data. Data kemudian diproses melalui tahap pre-processing yang meliputi cleaning, normalisasi, tokenisasi, stopword removal, dan stemming. Pelabelan sentimen dilakukan menggunakan metode rating based dengan dukungan leksikon, dengan menggunakan tiga algoritma pembelajaran mesin yang sama, yaitu Support Vector Machine (SVM), Multinomial Naïve Bayes, dan Perceptron dengan representasi fitur TF-IDF. Hasil penelitian menunjukkan bahwa Multinomial Naïve Bayes memperoleh akurasi tertinggi sebesar 90.35%, diikuti SVM sebesar 88.93%, dan Perceptron sebesar 85.46%. Seluruh model mengalami peningkatan performa dibandingkan penelitian sebelumnya dengan dataset lebih kecil. Penelitian ini menunjukkan bahwa jumlah dataset berpengaruh terhadap akurasi dan stabilitas model analisis sentimen

Copyrights © 2026






Journal Info

Abbrev

JRST

Publisher

Subject

Chemical Engineering, Chemistry & Bioengineering Chemistry Civil Engineering, Building, Construction & Architecture Computer Science & IT Engineering

Description

JRST (Jurnal Riset Sains dan Teknologi) adalah jurnal peer reviewed dan Open-Acces. JRST merupakan jurnal yang diterbitkan oleh Lembaga Publikasi Ilmiah dan Penerbitan (LPIP) Universitas Muhammadiyah Purwokerto. JRST mengundang para peneliti, dosen, dan praktisi di seluruh dunia untuk bertukar dan ...