EXPLORER
Vol 6 No 2 (2026): July 2026

Klasifikasi Sentimen Teks Code-Mixed Indonesia–Inggris Non-Formal Pada X Menggunakan Model Fine-Tuned DistilBERT

Syifa Arifah Nurbayani (Program Studi Teknik Informatika, Fakultas Sains dan Teknologi, Universitas Islam Negeri Sunan Gunung Djati Bandung)
Dian Sa'adillah Maylawati (Program Studi Teknik Informatika, Fakultas Sains dan Teknologi, Universitas Islam Negeri Sunan Gunung Djati Bandung, Indonesia)
Aldy Rialdy Atmadja (Program Studi Teknik Informatika, Fakultas Sains dan Teknologi, Universitas Islam Negeri Sunan Gunung Djati Bandung, Indonesia)



Article Info

Publish Date
25 Jul 2026

Abstract

The rapid growth of social media has increased the use of Indonesian–English code-mixed language in digital communication, particularly on social media X. The non-formal characteristics of social media text, such as slang, abbreviations, emojis, and language switching within a single sentence, make sentiment analysis more challenging than monolingual text. This study aims to perform sentiment classification on code-mixed text by evaluating the performance of the lightweight Transformer model DistilBERT and comparing it with IndoBERTweet. The study adopts the CRISP-DM methodology, which consists of Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation, and Deployment stages. The dataset comprises 1,108 primary data collected from social media X between 2022 and 2026 and 5,048 secondary data obtained from previous research. Three experimental scenarios were applied: translation into Indonesian, translation into English, and raw data without translation. The results show that DistilBERT achieved its best performance on the English translation scenario, with accuracies of 73.66% on the primary dataset and 72.00% on the secondary dataset. Meanwhile, IndoBERTweet obtained the highest performance on the Indonesian translation scenario, achieving an accuracy of 79.01%. These findings indicate that the alignment between the language of the input data and the pre-training characteristics of the model significantly affects sentiment classification performance on non-formal code-mixed text.

Copyrights © 2026






Journal Info

Abbrev

Explorer

Publisher

Subject

Computer Science & IT

Description

EXPLORER Journal of Computer Science and Information Technology is a scientific journal published by the FKPT (Forum Kerjasama Pendidikan Tinggi). This journal contains scientific papers from Academics, Researchers, and Practitioners about research on Computer Science and Information Technology. ...