Muhammad Derry Oktaviandi
Sekolah Tinggi Ilmu Komputer Cipta Karya Informatika

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Public Sentiment Analysis of the #KaburAjaDulu Hashtag Using a Combination of Support Vector Machine (SVM) and Random Forest Algorithms Muhammad Derry Oktaviandi; Yuma Akbar; Mesra Betty Yel
International Journal Software Engineering and Computer Science (IJSECS) Vol. 6 No. 3 (2026): DECEMBER 2026
Publisher : Lembaga Komunitas Informasi Teknologi Aceh (KITA), Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35870/ijsecs.v6i3.8068

Abstract

This study aims to analyze public sentiment toward the #KaburAjaDulu hashtag on platform X and compare the classification performance of Support Vector Machine (SVM), Random Forest, and a Voting Classifier ensemble. A total of 1,502 posts were collected via web scraping. The text preprocessing pipeline comprised case folding, text cleaning, tokenization, stopword removal, and stemming using the Sastrawi library. Processed texts were transformed into numerical feature vectors using Term Frequency-Inverse Document Frequency (TF-IDF), followed by an 80:20 train-test split. The empirical distribution revealed an extreme class imbalance: negative sentiment dominated at 96.54% (1,450 posts), followed by neutral at 3.33% (50 posts) and positive at 0.13% (2 posts), generating an imbalance ratio of 725:1. SVM and Random Forest achieved identical aggregate scores with 95.35% accuracy, 90.91% precision, 95.35% recall, and a 93.08% F1-score; however, both completely failed to detect neutral and positive classes. The Voting Classifier achieved the highest aggregate performance, reaching 96.01% accuracy, 95.84% precision, 96.01% recall, and a 94.55% F1-score by successfully identifying two neutral instances. Nevertheless, the ensemble model was unable to recognize positive sentiment, yielding zero sensitivity for the extreme minority class. These findings demonstrate that combining SVM and Random Forest offers marginal improvements in aggregate metrics but remains constrained by data distribution. Consequently, relying solely on aggregate accuracy produces misleading evaluations in severely skewed datasets, emphasizing the necessity of per-class metrics, confusion matrices, and data-balancing strategies for social media sentiment classification.