Muhammad Saifuddin Eka Nugraha
Universitas Sebelas Maret

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Benchmarking Indonesian Transformer Models and Explainable AI for Disaster-Related Sentiment Analysis Muhammad Saifuddin Eka Nugraha; Afrizal Doewes; Herdito Ibnu Dewangkoro
Journal of Embedded Systems, Security and Intelligent Systems Vol 7 No 3 (2026): September 2026
Publisher : Program Studi Teknik Komputer

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59562/jessi.v7i3.13295

Abstract

Purpose – This study evaluates the comparative performance of Indonesian transformer models for disaster-related sentiment analysis and examines the faithfulness of explainable artificial intelligence methods applied to the best-performing model. Methods – YouTube comments related to the 2025 Sumatra flood were processed using a hybrid labeling approach combining automatic classification and expert annotation. After preprocessing and class balancing, IndoBERT, IndoBERTweet, and IndoRoBERTa were fine-tuned using Optuna-based hyperparameter optimization. Model performance was assessed using accuracy, precision, recall, and F1-score. Integrated Gradients (IG), Local Interpretable Model-Agnostic Explanations (LIME), and SHapley Additive exPlanations (SHAP) were subsequently evaluated using the Area Under the Threshold-Performance Curve (AUC-TP) to quantify explanation faithfulness. Findings – IndoRoBERTa achieved the strongest overall classification performance among the evaluated models. Faithfulness analysis showed that IG provided the strongest overall explanation performance and performed particularly well for negative and positive sentiment, whereas LIME showed better performance for neutral sentiment. SHAP produced comparatively weaker faithfulness under the applied evaluation protocol. Research Implications – The findings demonstrate the potential of Indonesian transformer models and quantitative XAI evaluation for analyzing disaster-related social media discourse. However, the results should be interpreted cautiously because automatic labeling, confidence-based undersampling, and the absence of inferential significance testing may affect generalizability. Originality – This study integrates comparative benchmarking of Indonesian transformer models with systematic deletion-based faithfulness evaluation of multiple XAI methods, extending explainable sentiment analysis beyond predominantly qualitative interpretation.