Ucta Pradema Sanjaya
Ngudi Waluyo University, Kabupaten Semarang, Jawa Tengah 50512, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

COMPARISON OF LINEARSVC AND COMPLEMENT NAIVE BAYES ON CORETAX SENTIMENT ANALYSIS USING INDOBERT PSEUDO-LABELING AND ADASYN Riza Febyana Shollis; Ucta Pradema Sanjaya
JIKO (Jurnal Informatika dan Komputer) Vol 9 No 2 (2026)
Publisher : Program Studi Teknik Informatika Universitas Khairun

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.33387/jiko.v9i2.11615

Abstract

Coretax is an Indonesian digital tax service platform that has attracted extensive public discussion on social media, particularly on X/Twitter. Most discussions are related to system stability and user experience when accessing the platform. This study collected sentiment data from December 1, 2024, to December 20, 2024, resulting in 10,140 tweets, which were filtered into 7,900 valid data points. The dataset underwent standard text preprocessing and was represented using TF-IDF features. Sentiment labeling was performed automatically (pseudo-labeling) using the IndoBERT model, producing an imbalanced class distribution (Negative 63.20%; Neutral 30.24%; Positive 6.56%). To address this imbalance, ADASYN was applied to the training data, and two classification models were compared: Complement Naïve Bayes and SVM (LinearSVC). Evaluation using an 80:20 train–test split showed that LinearSVC + ADASYN achieved the best performance with an accuracy of 85.25%, F1-Macro of 0.7423, MCC of 0.7120, ROC-AUC of 0.9428, and Hamming Loss of 0.1475, outperforming ComplementNB + ADASYN with an accuracy of 78.67%. Furthermore, the McNemar test confirmed that the performance difference between the two models is statistically significant. These findings indicate that LinearSVC is more effective in distinguishing sentiments related to technical complaints and procedural inquiries in the context of digital tax services.