Jurnal Teknik Informatika (JUTIF)
Vol. 7 No. 4 (2026): JUTIF Volume 7, Number 4, August 2026

Indonesian Hate Speech Detection: A Cross-Validated Benchmark of Machine Learning and Pre-trained Transformer Models with Statistical Significance Analysis

Dodo Zaenal Abidin (Magister of Information System, Faculty of Computer Science, Universitas Dinamika Bangsa, Indonesia)
Agus Siswanto (Informatics Engineering, Faculty of Computer Science, Universitas Dinamika Bangsa, Indonesia)
Chindra Saputra (Informatics Engineering, Faculty of Computer Science, Universitas Dinamika Bangsa, Indonesia)
Imelda Yose (Informatics Engineering, Faculty of Computer Science, Universitas Dinamika Bangsa, Indonesia)
Kaslin Kaslin (Magister of Information System, Faculty of Computer Science, Universitas Dinamika Bangsa, Indonesia)



Article Info

Publish Date
18 Aug 2026

Abstract

Automated hate speech detection in Indonesian social media remains a persistent challenge due to dataset fragmentation, heterogeneous annotation schemes, and the lack of reproducible cross-model benchmarks with formal statistical validation. This study presents a cross-validated benchmark that systematically evaluates seven models — one majority baseline, three classical machine learning models (Logistic Regression, SVM, Random Forest), and three transformer-based pre-trained language models (DistilBERT, IndoRoBERTa, IndoBERT) — on a standardized multi-source Indonesian hate speech corpus comprising 14,043 samples from Twitter and Instagram. All models were trained and evaluated under identical conditions using stratified three-way splits (70/15/15) replicated across three random seeds, reporting mean ± standard deviation for F1-macro, precision-macro, recall-macro, and ROC-AUC. Paired t-tests and Cohen's d effect size analysis were applied to formally assess statistical significance and practical magnitude of performance differences. Results show that all transformer-based models significantly outperformed all classical ML models, with IndoBERT achieving the highest mean F1-macro of 0.8834 (±0.0026) and the lowest cross-seed variance among all models. Notably, IndoBERT and IndoRoBERTa were found to be statistically equivalent (p=0.7273, d=0.283), indicating that neither model is definitively superior for this task. Among classical models, SVM attained the best F1-macro of 0.8509. These findings confirm that domain-specific pre-training on Indonesian corpora contributes to both higher performance and superior cross-seed stability. The proposed benchmarking framework, standardized corpus, and statistical evaluation protocol provide a reproducible reference for future Indonesian hate speech detection research, thereby advancing the methodological standards of automated text classification in computer science and informatics, particularly for under-resourced language NLP research.

Copyrights © 2026






Journal Info

Abbrev

jurnal

Publisher

Subject

Computer Science & IT

Description

Jurnal Teknik Informatika (JUTIF) is an Indonesian national journal, publishes high-quality research papers in the broad field of Informatics, Information Systems and Computer Science, which encompasses software engineering, information system development, computer systems, computer network, ...