Iwan Setiawan Wibisono
Universitas Ngudi Waluyo, Semarang

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Comparative IndoBERT, Logistic Regression, SVM, Random Forest, and Naïve Bayes for TikTok Deepfake Misinformation Tri Pidianto; Iwan Setiawan Wibisono
Building of Informatics, Technology and Science (BITS) Vol 8 No 2 (2026): September 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i2.10768

Abstract

Generative AI has made deepfake content, especially face-swap videos and cloned voices, easier to produce and spread, and short-video platforms such as TikTok have become a major channel for this material. Most existing work still focuses on detecting the manipulated audio or video itself, leaving audience reactions in the comment section relatively unexplored. The problem this study addresses is that, without reading these comments, moderation systems have no way to know whether audiences are being misled by AI-generated content; the objective is to determine whether a fine-tuned IndoBERT model can flag such misinformation-leaning comments more reliably than conventional machine-learning classifiers. This paper takes a comparative approach to that gap, testing whether a fine-tuned IndoBERT model can flag potential misinformation in Indonesian-language TikTok comments more reliably than four TF-IDF-based classifiers: Logistic Regression, Random Forest, Multinomial Naive Bayes, and Linear SVM. Starting from 1,345 comments scraped across 15 TikTok videos, a multi-stage cleaning process left 1,071 usable comments, manually sorted into three labels: Misinformation, Skeptical, and Other. IndoBERT was fine-tuned with a weighted cross-entropy loss to offset class imbalance and evaluated through a train-validation-test split and Stratified 5-Fold Cross Validation. The fine-tuned model reached 82.41% accuracy and 81.58% F1-Macro, beating every machine-learning baseline by at least 13.84 points in F1-Macro, with five-fold results holding steady at an average F1-Macro of 80.17%. These results suggest IndoBERT is a stronger option than conventional machine learning for flagging misinformation-leaning comments on AI-generated content, offering a practical foundation for automated moderation on social platforms.