Claim Missing Document
Check
Articles

Found 2 Documents
Search

ANALISIS PENANGANAN DATA TIDAK SEIMBANG TERHADAP KINERJA KLASIFIKASI SENTIMEN MULTIKELAS PADA ULASAN MARKETPLACE TOKOPEDIA Nauval Alfarizi; Satria Sinurat; Adi Putra; Muhammad Amin; Prima Lydia
JOURNAL OF SCIENCE AND SOCIAL RESEARCH Vol. 9 No. 1 (2026): February 2026
Publisher : Smart Education

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.54314/jssr.v9i1.5804

Abstract

Abstract: The development of digital marketplaces has led to an increasing number of user reviews, which can be used to understand consumer perceptions of products and services. However, sentiment analysis in marketplace reviews faces a major challenge: class imbalance, where positive sentiment often dominates to an extreme. This study aims to analyze the effects of various imbalanced data-handling techniques on the performance of machine-learning-based multiclass sentiment classification in Tokopedia marketplace reviews. The dataset used consists of 56,981 reviews with three sentiment classes, with more than 97% of them being positive. Feature extraction was performed using the TF-IDF method, resulting in 17,765 features. The handling of data imbalance was tested through four scenarios: class weighting, Random Oversampling, SMOTE, and ADASYN, with the Naive Bayes, Logistic Regression, and Random Forest algorithms. The experimental results show that Random Forest with SMOTE achieves the highest accuracy of 0.9749 but has limitations in recognizing minority classes, with a recall of 0.3786. In contrast, Logistic Regression with Random Oversampling provides the most balanced performance with the highest F1-score (macro) value of 0.4992 and recall of 0.5866. Keywords: Analysis, Sentiment, Imbalanced Data, Multi-Class Classification F1-Score Abstrak: Perkembangan marketplace digital menyebabkan meningkatnya jumlah ulasan pengguna yang dapat dimanfaatkan untuk memahami persepsi konsumen terhadap produk dan layanan. Namun, analisis sentimen pada ulasan marketplace menghadapi tantangan utama berupa ketidakseimbangan distribusi kelas, di mana sentimen positif sering kali mendominasi secara ekstrem. Penelitian ini bertujuan untuk menganalisis pengaruh berbagai teknik penanganan data tidak seimbang terhadap kinerja klasifikasi sentimen multikelas pada ulasan marketplace Tokopedia berbasis machine learning. Dataset yang digunakan terdiri dari 56.981 ulasan dengan tiga kelas sentiment, di mana proporsi sentimen positif mencapai lebih dari 97%. Ekstraksi fitur dilakukan menggunakan metode TF-IDF yang menghasilkan 17.765 fitur. Penanganan ketidakseimbangan data diuji melalui empat skenario, yaitu class weighting, Random Oversampling, SMOTE, dan ADASYN, dengan algoritma Naive Bayes, Logistic Regression, dan Random Forest. Hasil eksperimen menunjukkan bahwa Random Forest dengan SMOTE menghasilkan akurasi tertinggi sebesar 0,9749, namun memiliki keterbatasan dalam mengenali kelas minoritas dengan nilai recall 0,3786. Sebaliknya, Logistic Regression dengan Random Oversampling memberikan performa paling seimbang dengan nilai F1-score (macro) tertinggi sebesar 0,4992 dan recall 0,5866. Kata kunci: Analisis, Sentimen, Data Tidak Seimbang, Klasifikasi Multi Kelas F1-Score
ANALISIS OPINI PUBLIK MENGENAI BATAS USIA PENGGUNAAN SOSMED ANAK INDONESIA MENGGUNAKAN PENDEKATAN MACHINE LEARNING Prima Lydia; Muhammad Irfan Sarif; Nauval Alfarizi; Sri Hidayati; Adi Putra
JOURNAL OF SCIENCE AND SOCIAL RESEARCH Vol. 9 No. 2 (2026): April 2026
Publisher : Smart Education

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.54314/jssr.v9i2.6231

Abstract

Abstract: This research looks at how people feel about government rules that limit children under 16 from using social media. This study uses data consisting of 4,651 social media posts obtained through scraping, as well as questionnaire responses from 338 participants at a public junior high school in the Deli Serdang region.In this study, two types of classification models were used: Support Vector Machine (SVM) and Random Forest. These models were compared to see how well they perform. This method serves as a means of achieving relatively similar levels of accuracy. The results indicate that both models perform at comparable accuracy levels. However, SVM demonstrates superior performance in maintaining class balance, as indicated by a higher macro F1-score of 0.85, compared to Random Forest, which tends to be more biased toward the neutral class. Further analysis shows that the neutral class is the most common in the sentiment distribution. Additionally, the models often have difficulty accurately recognizing positive sentiment, probably because they tend to be cautious and classify uncertain inputs as neutral. This study doesn't use special methods to balance the data, so it can check how well the model works when the data is naturally unbalanced. This method helps us understand how public opinion reacts to government policies, especially those affecting children, as the main focus of the analysis. Keywords: Access Restriction, Sentiment Analysis, Public Opinion, Children, Social Media Abstrak: Penelitian ini bertujuan menganalisis perasaan masyarakat umum terhadap kebijakan yang dibuat pemerintah yang membatasi akses anak di bawah usia 16 tahun ke media sosial. Data yang digunakan berasal dari hasil pengumpulan data melalui scraping media sosial sebanyak 4651 data dan data dari kuesioner yang diisi oleh 338 responden dari salah satu SMP Negeri Wilayah Deli Serdang. Model klasifikasi yang digunakan dalam penelitian ini adalah Support Vector Machine (SVM) dan Random Forest sebagai perbandingan dalam melihat performa. Hasil pengujian menunjukkan kedua model memiliki tingkat akurasi relativ sama. Namun, SVM menunjukkan hasil yang lebih baik dalam menjaga keseimbangan antar kelas dengan skor F1 makro sebesar 0,85, dibandingkan dengan Random Forest yang cenderung lebih condong ke kelas netral. Analisis lebih lanjut menunjukkan bahwa kelas netral mendominasi dalam penyebaran sentimen, Kemudian, model juga sering kesulitan dalam mengenali perasaan positif dikarenakan model cenderung konservatif dengan memilih netral. Dalam penelitian ini tidak mengklasifikasi secara khusus seperti model imbalance data untuk melihat performa tanpa perlu Teknik sampling guna melihat seberapa dan kenapa opini publik bereaksi terhadap suatu kebijakan pemerintahan khususnya terhadap anak anak sebagai output utama Kata Kunci: Pembatasan, Sentimen, Opini, Anak Anak, Media Sosial