The rapid growth of TikTok has increased interest in identifying factors that influence content virality, particularly audio elements that play a central role in trend formation and user engagement. This study aims to develop a model for predicting TikTok content virality based on audio characteristics using a Convolutional Neural Network (CNN). The proposed approach focuses exclusively on audio information to evaluate its independent contribution to virality prediction without incorporating visual, textual, or engagement-based features. The research employed a quantitative experimental design using audio extracted from publicly available TikTok videos categorized into viral and non-viral classes. Audio signals were preprocessed through normalization and duration standardization before being transformed into Mel-spectrogram representations. These spectrogram images were then used as input to a CNN model for automatic feature extraction and classification. Model performance was evaluated using precision, recall, and F1-score metrics. The experimental results demonstrate that the proposed CNN model effectively distinguished viral and non-viral TikTok content. Evaluation on the testing dataset produced a precision of 0.833, recall of 1.000, and F1-score of 0.909. In addition, Mel-spectrogram visualizations revealed distinct frequency-energy patterns between viral and non-viral audio samples, indicating that acoustic characteristics contain meaningful information associated with content virality. In conclusion, audio features can serve as reliable predictors of TikTok content virality, and the CNN-based framework successfully extracted discriminative acoustic patterns from Mel-spectrogram representations. This study contributes to the fields of audio analytics and social media intelligence by providing an audio-centered approach for early-stage virality prediction prior to content publication.
Copyrights © 2026