Jemmy Edwin Bororing
Janabadra University

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search
Journal : journal of deep learning computer vision and digital image processing

Automatic Detection of Toxic Content in Short Videos Using Deep Learning-Based Text and Audio Feature Integration Heri Agus Supriyanto; Ryan Ari Setyawan; Jemmy Edwin Bororing
Journal of Deep Learning, Computer Vision, and Digital Image Processing Volume 4 Issue 2 June 2026
Publisher : CV. Sakura Digital Nusantara

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.61255/decoding.v4i2.1540

Abstract

Purpose – This study develops and evaluates a multimodal classification model combining audio and text features via configurable weighted late fusion to detect toxic content in keyword-retrieved Indonesian short videos containing profanity-related contexts, addressing text-only detection limitations.Methods – The pipeline utilized FFmpeg for audio extraction, MFCC for audio features, and Google Speech Recognition for text. Three fusion configurations (Audio-Text: 40%–60%, 50%–50%, and 60%–40%) combining a DNN and BiLSTM were evaluated on 1,484 manually labeled Indonesian short videos.Findings – The Audio 60%–Text 40% configuration achieved the numerically highest test accuracy of 93.94% with a 95% confidence interval of 90.62%–96.13%, using a decision threshold of 0.60 selected from the validation set. The model obtained an F1-score of 0.95 for the toxic class. Compared with unimodal baselines, all fusion models achieved higher accuracy, indicating the benefit of integrating audio and text features.Research implications – The findings suggest that multimodal audio-text fusion can improve Indonesian short-form video toxicity detection compared with audio-only or text-only models. However, the differences among the three fusion weighting schemes were not statistically significant based on McNemar’s test, so the audio-dominant configuration should be interpreted as the numerically best configuration in this dataset rather than as a universally superior setting.Originality – This study systematically compares unimodal baselines and configurable audio-text late-fusion weighting strategies for Indonesian short-form video toxicity detection. The study also applies validation-based threshold selection, confidence intervals, and pairwise McNemar testing to provide a more reliable evaluation of multimodal model performance.