The rise of political hate speech on social media calls for detection systems that are not only accurate but also transparent to avoid moderation bias. Although Transformer models achieve high performance, their “black box” nature creates the risk of a “false sense of security” in content moderation, where high accuracy can mask systemic bias. This study aims to transparently audit the decision-making mechanisms of a lexical-feature-based Support Vector Machine (SVM) model and a contextual-representation-based IndoBERT model using an Explainable AI (XAI) approach via the LIME method. Experimental results show that IndoBERT significantly outperforms SVM with a Macro-F1 score of 90.8% versus 84.0%. However, the XAI audit revealed the presence of data-driven bias in both models toward specific political entities such as “Jokowi,” “Prabowo,” “cebong,” and “kampret,” which often trigger negative labels automatically without a comprehensive contextual review. These findings underscore that transparency audits through XAI serve as a crucial bridge for building a content moderation system that is fair, accountable, and capable of protecting freedom of expression within the digital democratic ecosystem.
Copyrights © 2026