Background: The rapid growth of Indonesian digital news content increases the demand for accurate Named Entity Recognition (NER), yet linguistic complexity and limited annotated data remain key challenges. Objective: This study evaluates the effectiveness of a BiLSTM-CRF model enhanced with domain-specific word embeddings for Indonesian NER across person, location, and organization entities. Method: A supervised sequence-labeling approach was applied using annotated news data, with embeddings trained on large-scale political and economic corpora and evaluated via precision, recall, and F1-score. Results: The model shows stable performance for person and location entities, while domain-specific embeddings improve all categories, especially organization entities with high variation; remaining errors relate to boundary detection and semantic ambiguity. Implication: These findings highlight the importance of domain-adaptive representations for improving NER systems in low-resource languages and support more reliable information extraction in Indonesian digital media. Novelty: This study demonstrates that embedding-level domain adaptation significantly enhances Indonesian NER without increasing model complexity, while clarifying distinctions between corpus resources and annotated data for better reproducibility.
Copyrights © 2026