Journal of Defense Technology and Engineering
Vol. 2 No. 1 (2026): July, Journal of Defense Technology and Engineering

Cyber threat detection on social media using indoBERT and sentiment analysis

Bagus Hendra Saputra (Universtias Pertahanan Republik Indonesia, Bogor, Indonesia)
Jonson Manurung (Universitas Pertahanan Republik Indonesia, Bogor, Indonesia)
Baringin Sianipar (Universitas HKPB Nommensen Medan, Medan, Indonesia)
R. Fanry Siahaan (Universitas HKBP Nommensen, Medan, Indonesia)



Article Info

Publish Date
31 Jul 2026

Abstract

The rapid growth of Indonesian social media has increased the spread of cyber threat-related content, creating significant challenges for digital security monitoring due to the informal language, code-switching, and sentiment-rich expressions commonly used in online communication. Existing detection approaches, particularly those based on multilingual or English-centric language models, often fail to capture the linguistic characteristics of Indonesian text effectively. This study aims to develop an accurate cyber threat detection model by fine-tuning IndoBERT, a transformer-based language model pretrained on a large-scale Indonesian corpus, for binary Threat and Non-Threat classification. The model was trained and evaluated using the Tweet ID Sentiment Dataset containing 10,800 annotated tweets, which were partitioned into training, validation, and test sets, and its performance was compared with four baseline methods: SVM with TF-IDF features, CNN with FastText embeddings, BiLSTM with Word2Vec representations, and multilingual BERT. Experimental results demonstrate that the proposed IndoBERT model achieved the best performance, obtaining an accuracy of 0.9389, a macro-F1 score of 0.9292, and a Threat-class recall of 0.9486, consistently outperforming all baseline models. The novelty of this study lies in demonstrating the effectiveness of a monolingual Indonesian pretrained transformer for cyber threat detection, highlighting the importance of language-specific contextual representations in improving classification performance. These findings indicate that the proposed approach provides a robust and practical solution for automated cyber threat detection, supporting early warning systems and digital security monitoring in Indonesian social media environments. Future work will investigate multiclass cyber threat categorization and cross-platform evaluation to improve model generalizability in real-world applications.

Copyrights © 2026