Allan Desi Alexander
Universitas Bhayangkara Jakarta Raya, Jakarta, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Evolution and Adaptation of Large Language Models for Bahasa Indonesia Allan Desi Alexander; Siti Setiawati
Dinasti Information and Technology Vol. 4 No. 1 (2026): Dinasti Information and Technology (July - September 2026)
Publisher : Dinasti Research & Yayasan Dharma Indonesia Tercinta (DINASTI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.38035/dit.v4i1.3623

Abstract

This study evaluates the systematic evolution and computational adaptation of pre-trained language models and Large Language Models (LLMs) for Bahasa Indonesia and its low-resource regional dialects. Initially centered on bidirectional encoder-based representations like IndoBERT, the regional natural language processing (NLP) field has transitioned toward generative sequence-to-sequence structures and massive decoder-only architectures. This paper investigates the engineering methodologies of cross-lingual vocabulary adaptation, parameter initialization heuristics, and language-adaptive pre-training strategies designed to address text overfragmentation, representational misalignment, and tokenization cost inefficiencies. Through extensive structural benchmarks, this analysis compares discriminative and generative performances across tasks including sentiment classification, extractive question answering, text style normalization, domain-specific retrieval-augmented pipelines, and entity linking. While localized generative models such as Komodo, Sailor, and the SEA-LION suite improve contextual reasoning, colloquial style transfers, and regional dialect preservation, they remain susceptible to architectural anomalies like template leakage and entity hallucination. This study provides foundational benchmarks and methodological frameworks for adapting massive language models to morphologically rich, culturally diverse, and low-resource linguistic environments.