Ahmad Fauzi
Badan Riset dan Inovasi Nasional

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

KLASIFIKASI EMOSI LIRIK LAGU BERBAHASA INDONESIA MENGGUNAKAN MODEL INDOBERT: EMOTIONAL CLASSIFICATION OF INDONESIAN SONG LYRICS USING THE INDOBERT MODEL Wildan Suharso; Briansyah Setio Wiyono; Bahrul Ulum; Ruslan Yusuf; Ahmad Fauzi
Rabit : Jurnal Teknologi dan Sistem Informasi Univrab Vol 11 No 2 (2026): Juli
Publisher : LPPM Universitas Abdurrab

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.36341/rabit.v11i2.8263

Abstract

Song lyrics reflect human feelings, but manually identifying emotions in a dataset is a significant challenge. The urgency of this research lies in utilizing the latest model, IndoBERT-base-p1, in analyzing lyrics, so that it can be used by the music industry to understand audiences and be useful in psychology and language research. The dataset used consists of 360 Indonesian song lyrics obtained from Kaggle. This data is then broken down into chunks per line to allow for more detailed emotion analysis. The IndoBERT-base-p1 model will be trained to classify each line of lyrics into 10 predetermined emotion categories. The target output of this research is a trained model capable of classifying emotions with high accuracy. Based on the research results, it can be seen that the augmentation has proven effective. Song lyrics reflect human feelings, but manually identifying emotions in a dataset is a significant challenge. The urgency of this research lies in utilizing the latest model, IndoBERT-base-p1, in analyzing lyrics, so that it can be used by the music industry to understand audiences and be useful in psychology and language research. The dataset used consists of 360 Indonesian lyrics and 10 emotion categories. This data is then broken down into line-by-line chunks to allow for more detailed emotion analysis. The IndoBERT-base-p1 model will be trained to classify each line of lyrics into 10 predetermined emotion categories. The target output of this research is a trained model capable of classifying emotions with high accuracy. Based on the research results, it can be seen that the augmentation has proven to be quantitatively effective. Before augmentation, the number of data for some emotions was in the range of 250–350 samples, while dominant emotions such as love and sadness were above 1,400 samples. After augmentation, all emotion labels increased significantly and were in a more balanced range, namely around 1,100 to 2,600 samples per emotion. The precision values ​​for each emotion include sad 0.87, happy 0.85, angry 0.89, afraid 0.94, longing 0.89, love 0.89, disappointed 0.86, nostalgia 0.91, heartbroken 0.93.