RABIT: Jurnal Teknologi dan Sistem Informasi Univrab
Vol 11 No 2 (2026): Juli

KLASIFIKASI EMOSI LIRIK LAGU BERBAHASA INDONESIA MENGGUNAKAN MODEL INDOBERT: EMOTIONAL CLASSIFICATION OF INDONESIAN SONG LYRICS USING THE INDOBERT MODEL

Wildan Suharso (Universitas Muhammadiyah Malang)
Briansyah Setio Wiyono (Universitas Muhammadiyah Malang)
Bahrul Ulum (Universitas Muhammadiyah Malang)
Ruslan Yusuf (Universitas Mataram)
Ahmad Fauzi (Badan Riset dan Inovasi Nasional)



Article Info

Publish Date
30 Jul 2026

Abstract

Song lyrics reflect human feelings, but manually identifying emotions in a dataset is a significant challenge. The urgency of this research lies in utilizing the latest model, IndoBERT-base-p1, in analyzing lyrics, so that it can be used by the music industry to understand audiences and be useful in psychology and language research. The dataset used consists of 360 Indonesian song lyrics obtained from Kaggle. This data is then broken down into chunks per line to allow for more detailed emotion analysis. The IndoBERT-base-p1 model will be trained to classify each line of lyrics into 10 predetermined emotion categories. The target output of this research is a trained model capable of classifying emotions with high accuracy. Based on the research results, it can be seen that the augmentation has proven effective. Song lyrics reflect human feelings, but manually identifying emotions in a dataset is a significant challenge. The urgency of this research lies in utilizing the latest model, IndoBERT-base-p1, in analyzing lyrics, so that it can be used by the music industry to understand audiences and be useful in psychology and language research. The dataset used consists of 360 Indonesian lyrics and 10 emotion categories. This data is then broken down into line-by-line chunks to allow for more detailed emotion analysis. The IndoBERT-base-p1 model will be trained to classify each line of lyrics into 10 predetermined emotion categories. The target output of this research is a trained model capable of classifying emotions with high accuracy. Based on the research results, it can be seen that the augmentation has proven to be quantitatively effective. Before augmentation, the number of data for some emotions was in the range of 250–350 samples, while dominant emotions such as love and sadness were above 1,400 samples. After augmentation, all emotion labels increased significantly and were in a more balanced range, namely around 1,100 to 2,600 samples per emotion. The precision values ​​for each emotion include sad 0.87, happy 0.85, angry 0.89, afraid 0.94, longing 0.89, love 0.89, disappointed 0.86, nostalgia 0.91, heartbroken 0.93.  

Copyrights © 2026






Journal Info

Abbrev

rabit

Publisher

Subject

Computer Science & IT Engineering

Description

This journal is called RABIT, where the name comes from two words namely, RAB which means Abdurrab University and IT which means information technology, it can be interpreted as a journal of this journal Journal of Informatics Engineering Study Program Pekanbaru Abdurrab University. This RABIT ...