This Author published in this journals
All Journal EXPLORER
Syifa Arifah Nurbayani
Program Studi Teknik Informatika, Fakultas Sains dan Teknologi, Universitas Islam Negeri Sunan Gunung Djati Bandung

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Klasifikasi Sentimen Teks Code-Mixed Indonesia–Inggris Non-Formal Pada X Menggunakan Model Fine-Tuned DistilBERT Syifa Arifah Nurbayani; Dian Sa'adillah Maylawati; Aldy Rialdy Atmadja
Explorer Vol 6 No 2 (2026): July 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/explorer.v6i2.2856

Abstract

The rapid growth of social media has increased the use of Indonesian–English code-mixed language in digital communication, particularly on social media X. The non-formal characteristics of social media text, such as slang, abbreviations, emojis, and language switching within a single sentence, make sentiment analysis more challenging than monolingual text. This study aims to perform sentiment classification on code-mixed text by evaluating the performance of the lightweight Transformer model DistilBERT and comparing it with IndoBERTweet. The study adopts the CRISP-DM methodology, which consists of Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation, and Deployment stages. The dataset comprises 1,108 primary data collected from social media X between 2022 and 2026 and 5,048 secondary data obtained from previous research. Three experimental scenarios were applied: translation into Indonesian, translation into English, and raw data without translation. The results show that DistilBERT achieved its best performance on the English translation scenario, with accuracies of 73.66% on the primary dataset and 72.00% on the secondary dataset. Meanwhile, IndoBERTweet obtained the highest performance on the Indonesian translation scenario, achieving an accuracy of 79.01%. These findings indicate that the alignment between the language of the input data and the pre-training characteristics of the model significantly affects sentiment classification performance on non-formal code-mixed text.