Riadhul Muttaqin
Universitas Islam Kalimantan Muhammad Arsyad Al Banjari Banjarmasin

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Deteksi Hoaks Berita Berbahasa Indonesia Menggunakan IndoBERT dengan Penanganan Ketidakseimbangan Kelas Berbasis Gabungan Class Weighting dan Focal Loss Riadhul Muttaqin; Muhammad Edya Rosadi; Muhammad Iqbal Firdaus; Dian Agustini
Infotek: Jurnal Informatika dan Teknologi Vol. 9 No. 2 (2026): Infotek : Jurnal Informatika dan Teknologi
Publisher : Fakultas Teknik Universitas Hamzanwadi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29408/jit.v9i2.35214

Abstract

The spread of hoaxes in Indonesian-language digital media has risen sharply in the past five years with wide societal harm. Most prior work on Indonesian hoax detection emphasizes model architecture, while the class-imbalance problem in field data receives less attention. This study presents a data-centric approach combining the IndoBERT pretrained language model with a hybrid class-imbalance objective of class weighting and focal loss. The dataset is the indonesiafalsenews corpus of 4231 labelled articles with an approximate 4.5 to 1 hoax-to-fact ratio. Evaluation uses five-seed runs with McNemar and paired-bootstrap significance testing. The proposed configuration achieves an F1-macro of 0.7240, accuracy of 0.8392, ROC-AUC of 0.8083, and a Matthews correlation coefficient of 0.4510, significantly outperforming every TF-IDF baseline at the 0.05 level. Text augmentation does not provide consistent gains, which implies that imbalance handling is the more effective lever for Indonesian hoax detection.