Claim Missing Document
Check
Articles

Found 1 Documents
Search

Multi-Class Humor Level Classification of Indonesian Stand-Up Comedy Transcripts Using Fine-Tuned IndoBERT Najib, Jihan; Supriyono, Supriyono; Aziz, Okta Qomaruddin; Andika, Daffa; Davissyah, Asfa
ILKOMNIKA Vol 8 No 2 (2026): Volume 8, Number 2, August 2026
Publisher : Lembaga Penelitian dan Pengabdian Masyarakat

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.28926/ilkomnika.v8i2.935

Abstract

Stand-up comedy has become a popular form of entertainment in Indonesia, but humor assessment remains subjective and is generally performed manually. This study develops an automatic classification model for Indonesian stand-up comedy humor levels using transcripts and the IndoBERT language model. The dataset was collected from stand-up comedy videos on the Kompas TV YouTube channel and used audience laughter counts as a pragmatic indicator of humor response. After removing records without transcripts and duplicate transcripts, 2,774 independent records were obtained and categorized into four humor levels: Not Funny, Slightly Funny, Funny, and Very Funny. The dataset was divided into training, validation, and test sets using an 80:10:10 stratified split. IndoBERT was fine-tuned and compared with a majority-class classifier and TF-IDF-based conventional baselines. On the test set, IndoBERT achieved 67.99% accuracy, 67.83% weighted precision, 67.99% weighted recall, and 67.19% weighted F1-score, outperforming the strongest conventional baseline by 11.87 percentage points. Cohen’s Kappa values were 0.5647 unweighted and 0.8309 with quadratic weighting. Moreover, 91.0% of misclassifications occurred in adjacent categories, indicating that the model captures the ordinal structure of humor levels. These results demonstrate that IndoBERT provides a viable baseline for Indonesian humor-level classification, although distinguishing between adjacent humor categories remains challenging.