Sentiment analysis of IndiHome users on social media X faces a severe class imbalance, with negative tweets dominating 88.26% of the dataset. This study compares four machine learning algorithms, Support Vector Machine (SVM), Naive Bayes, Decision Tree, and Random Forest, for sentiment classification using SMOTE to address the imbalance. Initially, 20,001 Indonesian tweets were scraped using Tweet Harvest with the keyword "indihome". After duplicate removal and preprocessing, 7,199 tweets were retained. Each tweet was manually annotated into positive, negative, and neutral categories. TF-IDF was applied for feature extraction, and Stratified 5-Fold Cross Validation was used for evaluation. Algorithms were tested under two conditions: without and with SMOTE. Before SMOTE, SVM achieved the highest accuracy (94.55%) and F1-score (94.05%). After SMOTE, Random Forest outperformed others with 94.14% accuracy and 93.83% F1-score, as the only algorithm showing consistent improvement across all metrics, including balanced accuracy and MCC. Although Wilcoxon tests showed no statistically significant differences between algorithms, Random Forest demonstrated the most stable and consistent performance. These findings confirm that Random Forest with SMOTE is the most effective strategy for imbalanced sentiment classification in this context.
Copyrights © 2026