Achraf El Bouazzaoui
Embedded Electronics and Intelligent Systems, Ibn Tofail University

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Beyond translated benchmarks: a culturally grounded evaluation of large language models for Indonesian language understanding Achraf El Bouazzaoui
Indonesian Journal of Computational Language Studies Vol. 1 No. 2 (2026): Large language models, linguistic diversity, and responsible NLP in Indonesia
Publisher : CV Narasi Khatulistiwa Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.67490/ijcl.v1i2.911

Abstract

Background: Translated benchmarks have expanded multilingual large language model evaluation, yet Indonesian remains a sociolinguistically dense setting where meaning is shaped by public institutions, regional languages, religious discourse, and culturally situated inference. Objective: This study aims to examine how Indonesian language understanding can be evaluated through source-grounded, culturally embedded corpus items rather than through translated English-centric tasks. Method: Using a qualitative-diagnostic benchmark design, this study constructs and codes an auditable corpus of public Indonesian documents, benchmark references, and contextual materials across semantic, pragmatic, cultural, regional, institutional, safety, and multimodal dimensions. Results: The analysis shows that semantic comprehension appears as a baseline requirement but does not sufficiently capture Indonesian understanding. Institutional register and cultural commonsense emerge as dominant dimensions, indicating that Indonesian prompts frequently require recognition of public authority, local values, collective memory, and socially constrained meaning. Implication: Task sensitivity is highest when documents combine policy language, cultural inference, safety cues, regional references, and pragmatic restraint, revealing why ordinary accuracy metrics may obscure fragile understanding. Novelty: This study contributes a culturally grounded evaluation perspective that reframes Indonesian LLM assessment as source-verifiable, context-sensitive interpretation rather than transferable language performance, and offers methodological resources for future multilingual benchmarking in low-resource and culturally plural settings within Southeast Asian digital ecologies