Claim Missing Document
Check
Articles

Found 1 Documents
Search
Journal : kinetik game technology information system computer network computing electronics and control

Hate Speech Analysis of YouTube Comments on the 2024 Indonesian Presidential Debate Using IndoBERT Agus Sasmito Aribowo; Yuli Fauziah; Yusna Bantulu; Shoffan Saifullah; Azfa Mutiara Ahmad Fubalo
Kinetik: Game Technology, Information System, Computer Network, Computing, Electronics, and Control Vol. 11, No. 3, August 2026
Publisher : Universitas Muhammadiyah Malang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.22219/kinetik.v11i3.2604

Abstract

The rapid digitization of political campaigns has intensified the spread of hate speech on social media, threatening democratic discourse and social cohesion. In Indonesia, YouTube comments on the 2024 presidential election debates have emerged as a critical yet underexplored source of polarizing content. However, existing detection systems struggle with Indonesian-specific linguistic features, code-mixing, and implicit political sarcasm, while high annotation costs and class imbalance further limit the scalability of supervised approaches. To address these challenges, this study introduces a large-scale dataset of 38,742 YouTube comments collected from the five official debate stages and labeled using a cost-effective semi-supervised framework (20% expert-annotated, 80% pseudo-labeled). We systematically evaluate four classification models —IndoBERT, mBERT, SVM, and Random Forest—under identical experimental conditions using evaluation metrics optimized for imbalanced data. Experimental results demonstrate that IndoBERT consistently outperforms all baseline models, achieving an average accuracy of 89.7% and a macro F1-score of 0.89 across all debate stages. Notably, IndoBERT maintains high recall for the minority hate speech class (0.81–0.90), confirming its superior ability to capture localized political rhetoric and contextual nuances that multilingual and classical models frequently miss. This study contributes a publicly available Indonesian political hate speech dataset, validates a scalable semi-supervised annotation pipeline, and provides empirical evidence that domain-specific Transformer models are essential for reliable content moderation in politically charged, low-resource environments.