Transformer-based language models are increasingly used in economic text analysis, yet their internal reasoning remains insufficiently interpretable, particularly in multilingual policy contexts. This study examines the mechanistic interpretability of attention heads in analyzing Indonesian fiscal policy discourse using a bilingual dataset comprising 47 YouTube transcripts and 23,500 news articles collected between 2018 and 2025. A fine-tuned IndoBERT model is implemented for fiscal stance classification, achieving strong performance with accuracy of 0.91, precision of 0.89, recall of 0.88, and F1-score of 0.885, ensuring a reliable basis for interpretability analysis. To uncover internal model behavior, an integrated framework combining attention entropy, head importance scoring, and Bias Sensitivity Index (BSI) is applied. Results show a consistent decline in attention entropy across layers from 2.85 to 1.62, indicating a shift from distributed contextual processing to focus semantic representation. Head-level analysis reveals functional specialization, with key attention heads reaching importance scores up to 0.25. Ablation experiments further demonstrate measurable prediction shifts, reflected in elevated BSI values, confirming sensitivity to specific heads associated with fiscal risk terms. These findings indicate that transformer models encode structured economic reasoning while exhibiting latent bias, highlighting the need for interpretability-driven approaches in AI-based policy analysis.
Copyrights © 2026