The growth in user reviews on e-commerce platforms generates a large volume of data that has the potential to be utilised in sentiment analysis. This study aims to analyse the impact of increasing data volume on the performance of sentiment analysis models. The data was obtained through web scraping of 4,071 Tokopedia reviews on the Google Play Store. The data was then processed through a pre-processing stage comprising cleaning, normalisation, tokenisation, stopword removal, and stemming. Sentiment labelling was performed using a rating-based method supported by a lexicon, employing three identical machine learning algorithms: Support Vector Machine (SVM), Multinomial Naïve Bayes, and Perceptron, with TF-IDF feature representation. The results of the study show that Multinomial Naïve Bayes achieved the highest accuracy of 90.35%, followed by SVM at 88.93%, and Perceptron at 85.46%. All models showed an improvement in performance compared to previous research using a smaller dataset. This study demonstrates that the size of the dataset influences the accuracy and stability of sentiment analysis models ABSTRAK (Bahasa Indonesia) Pertumbuhan ulasan pengguna pada platform e-commerce menghasilkan volume data yang besar dan berpotensi dimanfaatkan dalam analisis sentimen. Penelitian ini bertujuan untuk menganalisis pengaruh penambahan jumlah data terhadap kinerja model analisis sentimen. Data diperoleh melalui web scraping ulasan Tokopedia di Google Play Store sebanyak 4.071 data. Data kemudian diproses melalui tahap pre-processing yang meliputi cleaning, normalisasi, tokenisasi, stopword removal, dan stemming. Pelabelan sentimen dilakukan menggunakan metode rating based dengan dukungan leksikon, dengan menggunakan tiga algoritma pembelajaran mesin yang sama, yaitu Support Vector Machine (SVM), Multinomial Naïve Bayes, dan Perceptron dengan representasi fitur TF-IDF. Hasil penelitian menunjukkan bahwa Multinomial Naïve Bayes memperoleh akurasi tertinggi sebesar 90.35%, diikuti SVM sebesar 88.93%, dan Perceptron sebesar 85.46%. Seluruh model mengalami peningkatan performa dibandingkan penelitian sebelumnya dengan dataset lebih kecil. Penelitian ini menunjukkan bahwa jumlah dataset berpengaruh terhadap akurasi dan stabilitas model analisis sentimen
Copyrights © 2026