Cyberbullying involving children has become a major concern on social media, generating diverse public responses. This study aims to analyze public sentiment toward child cyberbullying cases on social media and to evaluate the effectiveness of data-balancing techniques in improving the performance of a Naïve Bayes (NB) classifier. The motivation for using NB lies in its computational efficiency, strong performance in text classification, and suitability for high-dimensional Term Frequency–Inverse Document Frequency (TF-IDF) features. A dataset consisting of 3,137 comments was collected from X/Twitter and processed through text preparation, lexicon-based sentiment labeling using Indonesian Sentiment Lexicon (InSet), and TF-IDF feature representation. Three balancing scenarios, namely Random Oversampling, Random Undersampling, and Non-Balancing, were compared by varying the proportion of training and testing samples across four split ratios: 90:10, 80:20, 70:30, and 60:40. The experimental results show that balancing techniques contribute to classification effectiveness, with Random Oversampling producing the best outcome. The highest accuracy of 86.36% was achieved with Random Oversampling and a 90:10 training-test split, outperforming Random Undersampling and Non-Balancing. Furthermore, the sentiment distribution indicates that negative sentiment dominates public responses to child cyberbullying incidents.
Copyrights © 2026