This study focuses on the implementation of the IndoBERT method for detecting hoax news in Indonesian digital media and examining the model’s performance and generalization ability. The dataset consists of primary data obtained from a Kaggle dataset and secondary data collected through web scraping from various sources, which are then combined and preprocessed. The model is trained using a fine-tuning approach with variations in parameters such as learning rate, batch size, and epoch to achieve optimal results. The experimental results indicate that the best configuration is achieved at epoch 4, learning rate 5e-5, and batch size 16, producing an accuracy of 0.9868 along with the lowest validation loss. Evaluation using a confusion matrix shows a relatively low error rate for both classes. Testing on new data reveals that the model correctly classifies 26 out of 30 samples, indicating good generalization capability, although some misclassifications still occur in factual news that share similar characteristics with hoaxes.
Copyrights © 2026