Purpose – This study evaluates user-perceived experience in Indonesia's M-Pajak mobile tax application using Google Play reviews and identifies temporal, sentiment, and lexical patterns in users' expressed evaluations. Design/methodology/approach – The study analyzed a relevance-ranked, retrievable corpus scraped on 24 May 2026. Of 7,009 retrieved reviews, 6,980 valid reviews from 2021–2026 were retained. IndoRoBERTa sentiment labels were evaluated against 1,500 reviews selected through proportionate year-stratified sampling and manually adjudicated by two annotators. The analysis combined temporal ratings, temporal sentiment, rating–sentiment alignment, and transparent bigram-to-dimension coding. Findings/Results – Human annotators achieved 97.20% agreement (Cohen's kappa = 0.9132), while IndoRoBERTa achieved 94.20% agreement with the human gold standard (kappa = 0.8243). Within the retrieved corpus, 78.81% of reviews were classified as negative, and later-year review shares were increasingly concentrated in low ratings and negative sentiment. These patterns describe expressed review-based evaluation and do not establish that overall application quality or population-level satisfaction necessarily declined. Authentication, OTP, login, NPWP, EFIN, payment, and system-error barriers dominated negative expressions. Originality/Value – The study offers a human-validated, multi-layer review-analysis framework for evaluating mobile public services, makes the lexical grouping procedure auditable, and translates recurrent complaints into operational service-improvement priorities.
Copyrights © 2026