Minggar P. D. Ramadhan
Universitas Islam Negeri Maulana Malik Ibrahim, Malang

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Analisis Pengaruh Preprocessing Regex dan Cosine Similarity terhadap Performa IndoBERT dalam Klasifikasi Berita Hoaks Berbahasa Indonesia Minggar P. D. Ramadhan; Zainal Abidin; Mochamad Imamudin
JURIKOM (Jurnal Riset Komputer) Vol. 13 No. 3 (2026): Juni 2026
Publisher : Universitas Budi Darma

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30865/jurikom.v13i3.9730

Abstract

The rapid growth of online media has significantly improved access to information, but it has also accelerated the spread of misinformation and hoax news. In hoax detection research, datasets are commonly derived from fact-checking platforms, which typically contain structured components such as claims, narratives, and clarification statements explicitly indicating that certain information is false. The presence of such clarification sentences has the potential to cause bias, a condition in which the model learns text patterns that explicitly indicate the label, thereby reducing the model's ability to fully understand the content of the news. This study aims to analyze the impact of preprocessing techniques based on regular expression (regex) and cosine similarity on the performance of the IndoBERT model for Indonesian hoax news classification. Both approaches are employed to identify and handle clarification sentences, enabling the model to focus more on contextual and semantic understanding of the news content. Experimental results show that the cosine similarity-based preprocessing outperforms the regex-based approach, achieving accuracy, precision, recall, and F1-score of 92.8%. In comparison, the regex-based method obtains an accuracy of 90.7%, precision of 91.3%, recall of 90.7%, and F1-score of 90.6%. These findings indicate that the semantic-based approach is more effective in handling linguistic variability and reducing potential bias caused by explicit clarification patterns. Overall, this study highlights the importance of appropriate preprocessing strategies in improving classification performance and provides insights into the impact of clarification statements in fact-checking datasets on transformer-based hoax detection models