Background: In Indonesia’s digitally dense and linguistically diverse environment, LLM safety cannot be evaluated only through English or standardized Indonesian because harm is often expressed through local registers, indirectness, stigma, and culturally situated forms of advice-seeking. Objective: This study examines whether large language models produce safe, culturally sensitive, and verifiably helpful responses to safety-relevant prompts in Indonesian, Javanese, Sundanese, and Buginese. Method: Using a public-source corpus of 24 verified documents and 160 coded model responses, this study combines safety response auditing, cultural-pragmatic analysis, and grounded helpfulness assessment across domains including bullying, online gender-based violence, mental health, misinformation, and digital harm. Results: The analysis shows that safe refusal and mitigated completion dominate the response distribution, but unsafe compliance, over-refusal, evasion, and culturally inadequate redirection remain visible. Cultural-pragmatic adequacy declines in local-language prompts, where register mismatch, missed indirect distress, stigma reproduction, and unsupported cultural assumptions become more prominent. Implication: Grounded helpfulness is also uneven, as local-language responses more often contain incomplete referral pathways or unverifiable local claims in public-risk contexts. Novelty: This study contributes a multilingual Indonesian safety-evaluation framework that treats local languages not as translated inputs but as pragmatic environments where AI safety may shift, weaken, or become socially inadequate
Copyrights © 2026