Irwan Budiman
Faculty of Mathematics and Natutal Science, Department of Computer Science, Lambung Mangkurat University, Kalimantan, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Quantifying Cross-Lingual Sentiment Prediction Difference in Bilingual Pop-Culture Reviews via Dual Monolingual Transformers and Jensen-Shannon Divergence Azfani Naurotul Jannah; Triando Hamonangan Saragih; Irwan Budiman; Muliadi Muliadi; Radityo Adi Nugroho
International Journal of Advances in Data and Information Systems Vol. 7 No. 2 (2026): August 2026 - International Journal of Advances in Data and Information Systems
Publisher : Indonesian Scientific Journal

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59395/ijadis.v7i2.1678

Abstract

Machine translation has been widely used to support cross-lingual sentiment analysis when labeled data in the target language are limited. However, differences in linguistic representation between the source text and its translation may affect the probabilistic outputs of sentiment classification models. This study investigated differences in sentiment predictions between original Mandarin reviews and their Indonesian translations using two independently fine-tuned monolingual transformer models. Chinese-RoBERTa-wwm-ext was applied to the original Mandarin reviews, whereas IndoBERT was used to classify the Indonesian translations. Jensen–Shannon Divergence was employed to compare the sentiment probability distributions generated by the two models. The results showed that Chinese-RoBERTa achieved an accuracy of 74%, whereas IndoBERT achieved 69% on the pseudo-labeled evaluation dataset. Furthermore, 73.81% of the review pairs retained consistent sentiment predictions, while 26.19% exhibited prediction shifts, with Polarity Amplification being the most frequently observed category and most transitions occurring between adjacent sentiment classes. The probability-distribution analysis also revealed substantial differences in prediction confidence for some review pairs, even when the predicted sentiment labels remained identical. These findings demonstrated that comparing probability distributions provided complementary information beyond label-based evaluation for analyzing prediction differences between independently trained monolingual sentiment models on bilingual review pairs.Â