M Rizki Hardika
Universitas Internasional Semen Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Weakly Supervised Sentiment Analysis of Gold Price Discussions Using Conventional Machine Learning and IndoBERT M Rizki Hardika; Brina Miftahurrohmah
Indonesian Journal of Data and Science Vol. 7 No. 2 (2026): Indonesian Journal of Data and Science
Publisher : yocto brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijodas.v7i2.445

Abstract

Introduction: Gold price movements attract substantial public and investor attention because gold serves as both a safe-haven asset and a hedging instrument. This study investigates Indonesian public sentiment toward gold price discussions on Platform X using a weakly supervised sentiment-analysis framework. Method: A total of 7,283 Indonesian-language tweets containing the keyword “harga emas” were collected during 2023–2025, with 4,429 tweets retained after preprocessing. Sentiment labels were generated using a domain-specific lexicon and validated through manual annotation. Naïve Bayes, K-Nearest Neighbor, Support Vector Machine, and IndoBERT were evaluated using the same train–test partition. Results and Discussion: Manual validation achieved a Cohen’s Kappa of 0.8718, indicating almost perfect inter-annotator agreement, while the lexicon-based labels achieved 70.62% accuracy against the manually annotated reference. IndoBERT achieved the highest performance on weakly supervised labels with 98.31% accuracy and a 98.16% macro F1-score, outperforming SVM, Naïve Bayes, and KNN. However, its accuracy decreased to 69.49% when evaluated against manually annotated data, demonstrating that downstream performance remains strongly influenced by weak-label quality. Conclusion: Weak supervision provides an efficient and scalable approach for large-scale Indonesian financial sentiment annotation, while contextual models such as IndoBERT offer superior classification performance; however, reliable manual validation remains essential to mitigate label noise and improve generalizability.