Imelda Yose
Informatics Engineering, Faculty of Computer Science, Universitas Dinamika Bangsa, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Indonesian Hate Speech Detection: A Cross-Validated Benchmark of Machine Learning and Pre-trained Transformer Models with Statistical Significance Analysis Dodo Zaenal Abidin; Agus Siswanto; Chindra Saputra; Imelda Yose; Kaslin Kaslin
Jurnal Teknik Informatika (Jutif) Vol. 7 No. 4 (2026): JUTIF Volume 7, Number 4, August 2026
Publisher : Informatika, Universitas Jenderal Soedirman

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.52436/1.jutif.2026.7.4.5176

Abstract

Automated hate speech detection in Indonesian social media remains a persistent challenge due to dataset fragmentation, heterogeneous annotation schemes, and the lack of reproducible cross-model benchmarks with formal statistical validation. This study presents a cross-validated benchmark that systematically evaluates seven models — one majority baseline, three classical machine learning models (Logistic Regression, SVM, Random Forest), and three transformer-based pre-trained language models (DistilBERT, IndoRoBERTa, IndoBERT) — on a standardized multi-source Indonesian hate speech corpus comprising 14,043 samples from Twitter and Instagram. All models were trained and evaluated under identical conditions using stratified three-way splits (70/15/15) replicated across three random seeds, reporting mean ± standard deviation for F1-macro, precision-macro, recall-macro, and ROC-AUC. Paired t-tests and Cohen's d effect size analysis were applied to formally assess statistical significance and practical magnitude of performance differences. Results show that all transformer-based models significantly outperformed all classical ML models, with IndoBERT achieving the highest mean F1-macro of 0.8834 (±0.0026) and the lowest cross-seed variance among all models. Notably, IndoBERT and IndoRoBERTa were found to be statistically equivalent (p=0.7273, d=0.283), indicating that neither model is definitively superior for this task. Among classical models, SVM attained the best F1-macro of 0.8509. These findings confirm that domain-specific pre-training on Indonesian corpora contributes to both higher performance and superior cross-seed stability. The proposed benchmarking framework, standardized corpus, and statistical evaluation protocol provide a reproducible reference for future Indonesian hate speech detection research, thereby advancing the methodological standards of automated text classification in computer science and informatics, particularly for under-resourced language NLP research.