TIN: TERAPAN INFORMATIKA NUSANTARA
Vol 7 No 1 (2026): June 2026

Evaluasi Pseudo-Labeling IndoRoBERTa dan InSet Lexicon dengan SVM pada Komentar TikTok

Yehezkiel Juandro Metta (Universitas Pelita Bangsa, Bekasi)
Aswan Supriyadi Sunge (Universitas Pelita Bangsa, Bekasi)
Asep Suprianto (Universitas Pelita Bangsa, Bekasi)



Article Info

Publish Date
30 Jun 2026

Abstract

TikTok generates large volumes of public comments that can be used to identify trends in public opinion, including responses to the accident involving a vehicle used in the Free Nutritious Meal program. However, informal social media language and imbalanced class distributions may affect sentiment labeling and classification performance. This study evaluates sentiment pseudo-labels generated by IndoRoBERTa and InSet Lexicon through Support Vector Machine classification with the application of the Synthetic Minority Over-sampling Technique. A quantitative experimental approach was applied to 10,309 TikTok comments collected through Apify scraping. Since no human annotations were used as ground truth, the positive, negative, and neutral labels produced by both methods were treated as pseudo-labels. The research stages included filtering, preprocessing, pseudo-label generation, Term Frequency-Inverse Document Frequency feature extraction, an 80:20 data split, SMOTE application to the training data, SVM classification, and evaluation using accuracy, precision, recall, and F1-score. The results show that SVM reproduced the InSet Lexicon pseudo-labels most effectively, achieving an accuracy of 0.865 without SMOTE. After SMOTE was applied, precision increased to 0.870 while the F1-score remained at 0.865, and neutral-class recall increased from 0.767 to 0.807. These findings indicate that InSet Lexicon produced pseudo-labels that were more consistently learned by SVM on this dataset, while SMOTE primarily improved minority-class recognition.

Copyrights © 2026