Muhammad Ridlo Nu'man Hakim
Universitas Kristen Satya Wacana

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Comparing the Interpretability of Classical Machine Learning Models and IndoBERT in Sentiment Analysis of Indonesian Health Service Applications using Multi-Review Generalization Testing Muhammad Ridlo Nu'man Hakim; Dwi Hosanna Bangkalang
SISTEMASI Vol 15, No 9 (2026): Sistemasi: Jurnal Sistem Informasi
Publisher : Universitas Islam Indragiri

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32520/stmsi.v15i9.6899

Abstract

The increasing use of healthcare service applications in Indonesia has led to a growing volume of user reviews on the Google Play Store. This development makes manual sentiment analysis increasingly inefficient, particularly for reviews containing diverse linguistic expressions and contrastive sentences that involve sentiment shifts across clauses and are prone to causing classification errors. However, the ability of sentiment analysis models to handle such contrastive sentences has received limited attention. Moreover, previous sentiment analysis studies have tended to report performance metrics without systematically evaluating model generalization and interpretability. This study evaluates classical machine learning models, namely Naïve Bayes, Logistic Regression, Support Vector Machine (SVM), Random Forest, and XGBoost, alongside the IndoBERT transformer model for sentiment classification of Indonesian-language reviews of healthcare service applications, with particular emphasis on contrastive sentences. The evaluation covers performance comparison, out-of-distribution generalization testing, and interpretability analysis using Integrated Gradients. The research methodology includes collecting 20,000 reviews from the Google Play Store, labeling the reviews with validation against manually assigned labels, developing models using classical machine learning algorithms and fine-tuning IndoBERT, and conducting out-of-distribution generalization testing. The results show that IndoBERT achieved the highest F1-score of 98.10%, outperforming the classical machine learning models, which achieved F1-scores ranging from 96.28% to 97.57%. On contrastive reviews, IndoBERT demonstrated more consistent generalization performance, achieving an accuracy of 92.50% in the in-domain scenario and 77.50% in the cross-domain scenario, compared with ranges of 65.00%–72.50% and 67.50%–75.00%, respectively, for the classical models. Further analysis using Integrated Gradients revealed that IndoBERT's predictions were influenced more strongly by sentiment-bearing tokens, particularly adjectives and negations, than by contrastive conjunctions themselves. These findings provide a reference for developing Indonesian-language sentiment analysis models and support healthcare service application developers in evaluating service quality based on user perceptions.