Witta Listiya Ningrum
Gunadarma University

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

WVisionBERT-VL: a multimodal model architecture for toxicity classification on social media platforms using large language models Witta Listiya Ningrum; Achmad Benny Mutiara; Diana Ikasari
IAES International Journal of Artificial Intelligence (IJ-AI) Vol 15, No 4: August 2026
Publisher : Institute of Advanced Engineering and Science

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.11591/ijai.v15.i4.pp3365-3375

Abstract

The increasing prevalence of toxic content on social media, conveyed through text, images, and videos, poses significant challenges for automated content moderation systems. Although prior studies have reported promising results in unimodal and bimodal settings, they often fail to capture implicit and contextual toxicity emerging from interactions across multiple modalities, particularly in non-English environments. This paper proposes WVisionBERT-VL, an end-to-end multimodal framework for toxicity detection that integrates text, image, and video modalities within a unified architecture. The proposed model incorporates modality-specific encoders, bidirectional multi-head cross attention (BMHCA) for cross-modal synchronization, and an adaptive fusion gate to dynamically balance modality contributions. A balanced multimodal dataset is constructed from social media platforms, including X, Instagram, and TikTok, and refined using a model-based labeling strategy with limited human-in-the-loop validation. Experimental results on a custom Indonesian dataset demonstrate strong in-domain performance, achieving an accuracy of 94.12%, a macro-F1 of 0.9407, and a receiver operating characteristic - area under the curve (ROC-AUC) of 0.9721, with robustness further validated through five-fold cross-validation. Cross-dataset evaluation highlights challenges related to domain shift, underscoring the need for future research on robust and domain-adaptive multimodal toxicity detection.