cover
Contact Name
Achmad Fawaid
Contact Email
achmad_fawaid.linguistik@upnjatim.ac.id
Phone
+6282318007953
Journal Mail Official
khatulistiwanarasi@gmail.com
Editorial Address
Dusun Krajan, RT 015 / RW 007 Desa Karanganyar, Kecamatan Paiton, Kabupaten Probolinggo, Provinsi Jawa Timur, Kodepos 67291
Location
Kab. probolinggo,
Jawa timur
INDONESIA
Indonesian Journal of Computational Language Studies
ISSN : -     EISSN : 31637884     DOI : -
Core Subject :
Indonesian Journal of Computational Language Studies is a double blind peer-reviewed scholarly journal that publishes original research articles and critical studies at the intersection of language, computation, and data-driven methodologies. This journal is published quarterly as a platform for the dissemination of theoretical, empirical, and interdisciplinary findings that explore how computational approaches contribute to the analysis, modeling, and understanding of linguistic phenomena. It addresses a broad range of topics, including but not limited to computational linguistics, natural language processing, corpus linguistics, language modeling, machine learning for language analysis, discourse and text mining, digital humanities, language technologies, and computational approaches to language use across social, cultural, and digital contexts.
Arjuna Subject : -
Articles 12 Documents
Indonesian named entity recognition using BiLSTM-CRF with domain-specific word embeddings Joko Santoso
Indonesian Journal of Computational Language Studies Vol. 1 No. 1 (2026): Data-driven approaches to Indonesian language processing
Publisher : CV Narasi Khatulistiwa Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.67490/ijcl.v1i1.548

Abstract

Background: The rapid growth of Indonesian digital news content increases the demand for accurate Named Entity Recognition (NER), yet linguistic complexity and limited annotated data remain key challenges. Objective: This study evaluates the effectiveness of a BiLSTM-CRF model enhanced with domain-specific word embeddings for Indonesian NER across person, location, and organization entities. Method: A supervised sequence-labeling approach was applied using annotated news data, with embeddings trained on large-scale political and economic corpora and evaluated via precision, recall, and F1-score. Results: The model shows stable performance for person and location entities, while domain-specific embeddings improve all categories, especially organization entities with high variation; remaining errors relate to boundary detection and semantic ambiguity. Implication: These findings highlight the importance of domain-adaptive representations for improving NER systems in low-resource languages and support more reliable information extraction in Indonesian digital media. Novelty: This study demonstrates that embedding-level domain adaptation significantly enhances Indonesian NER without increasing model complexity, while clarifying distinctions between corpus resources and annotated data for better reproducibility.
Sentiment analysis on Indonesian political discourse using transformer-based models Wayan Novitasari
Indonesian Journal of Computational Language Studies Vol. 1 No. 1 (2026): Data-driven approaches to Indonesian language processing
Publisher : CV Narasi Khatulistiwa Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.67490/ijcl.v1i1.549

Abstract

Background: The expansion of digital political communication in Indonesia has intensified the circulation of evaluative language, making accurate sentiment analysis essential for understanding public opinion and ideological polarization. Objective: This study examines the effectiveness of transformer-based models in classifying sentence-level sentiment in Indonesian political discourse and the impact of domain-specific fine-tuning. Method: A supervised experiment was conducted using annotated political texts from social media and online news, comparing transformer models with baseline classifiers. Results: Transformer-based models outperform traditional and recurrent approaches in accuracy and F1-scores, while domain-specific fine-tuning improves the detection of negative and neutral sentiments, especially in implicit and framed expressions; challenges remain in sarcasm, metaphor, and context-dependent meaning. Implication: These results underscore the need for context-aware and domain-adapted models to enhance sentiment analysis reliability in politically nuanced and low-resource settings. Novelty: This study proposes a linguistically informed, domain-adapted transformer framework that demonstrates how contextual modeling and fine-tuning improve sentiment interpretation in Indonesian political discourse.
Code-mixing detection in Indonesian–English tweets using machine learning and linguistic features Arif Nugroho
Indonesian Journal of Computational Language Studies Vol. 1 No. 1 (2026): Data-driven approaches to Indonesian language processing
Publisher : CV Narasi Khatulistiwa Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.67490/ijcl.v1i1.550

Abstract

Background: The increasing prevalence of Indonesian–English code-mixing in social media reflects sociolinguistic shifts driven by globalization, while challenging language processing systems that assume monolingual input. Objective: This study evaluates the effectiveness of machine learning models in detecting Indonesian–English code-mixing in Twitter data and the role of linguistic features in improving accuracy. Method: A supervised approach was applied using SVM and Random Forest classifiers on annotated tweets, enriched with features such as part-of-speech patterns, token-level language identification, and morphological markers. Results: SVM models outperform baselines with high accuracy and balanced precision–recall, while linguistic features significantly enhance detection, especially for intra-word mixing; errors mainly arise from lexical borrowing, short contexts, and morphologically integrated forms. Implication: These findings emphasize the importance of integrating linguistic knowledge into computational models to improve robustness in multilingual and low-resource settings. Novelty: This study demonstrates that linguistically informed machine learning frameworks enhance both performance and interpretability in detecting Indonesian–English code-mixing.
Lexical patterns in Indonesian academic writing: a corpus-based analysis for writing assistance Abdul Wahid
Indonesian Journal of Computational Language Studies Vol. 1 No. 1 (2026): Data-driven approaches to Indonesian language processing
Publisher : CV Narasi Khatulistiwa Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.67490/ijcl.v1i1.551

Abstract

Background: Despite growing demands for academic publication in Indonesian higher education, many students and early-career researchers struggle with appropriate lexical usage. Objective: This study identifies lexical features that characterize proficient Indonesian academic writing through a corpus-based analysis. Method: A large corpus of journal articles and theses was analyzed using frequency, lexical bundle, and keyword analysis across sections and registers. Results: Academic texts are dominated by abstract nouns and procedural verbs, while lexical bundles show strong sectional specialization; keyword analysis highlights contrasts with non-academic texts in abstraction, impersonality, and epistemic stance. Implication: These findings support the development of corpus-informed writing instruction and tools to enhance academic literacy in Indonesian contexts. Novelty: This study integrates section-sensitive lexical analysis with register-based comparison to provide a comprehensive empirical model of academic lexical proficiency.
Automatic readability assessment of Indonesian educational texts using hybrid NLP approaches Tifani Yuliyana
Indonesian Journal of Computational Language Studies Vol. 1 No. 1 (2026): Data-driven approaches to Indonesian language processing
Publisher : CV Narasi Khatulistiwa Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.67490/ijcl.v1i1.552

Abstract

Background: The increasing complexity of Indonesian educational texts across print and digital platforms raises concerns about mismatches with students’ reading capacities, while existing readability formulas remain largely English-centric. Objective: This study develops and evaluates an automatic readability assessment framework for Indonesian texts integrating surface metrics, linguistic features, and machine learning. Method: A stratified corpus of textbooks and digital materials was analyzed using readability indices, lexical–morphological–syntactic features, and supervised models with cross-validation. Results: Surface complexity rises across levels but shows overlap, indicating limits of traditional metrics; linguistic features such as lexical density, nominalization, and morphological complexity strongly predict readability, while hybrid ensemble models achieve the highest accuracy and lowest misclassification. Implication: These findings support the need for language-specific, multidimensional readability tools to improve text design and educational alignment in Indonesian contexts. Novelty: This study proposes a linguistically grounded hybrid NLP framework that reconceptualizes readability and enables scalable, automated evaluation of Indonesian educational texts.
Morphological segmentation of low-resource Indonesian dialects using unsupervised neural models Muhammad Iqbal Ibrahim
Indonesian Journal of Computational Language Studies Vol. 1 No. 1 (2026): Data-driven approaches to Indonesian language processing
Publisher : CV Narasi Khatulistiwa Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.67490/ijcl.v1i1.553

Abstract

Background: Indonesia’s linguistic diversity, with many underrepresented dialects, poses challenges for morphological analysis under low-resource conditions. Objective: This study examines whether unsupervised neural models can learn morphological structure in Indonesian dialects without annotated data. Method: A comparative unsupervised approach was applied using probabilistic segmentation, character-level BiLSTM, and Bayesian neural models on raw dialectal corpora. Results: Probabilistic methods capture frequent roots and affixes but struggle with reduplication and clitics; character-level models better handle phonological variation, while Bayesian models achieve the most balanced performance with stronger coherence and cross-dialect generalization. Implication: These findings highlight the potential of unsupervised, probabilistic approaches to support inclusive language technology for low-resource dialects. Novelty: This study reconceptualizes dialectal morphology as a latent probabilistic system and demonstrates the effectiveness of Bayesian unsupervised neural segmentation.
Beyond translated benchmarks: a culturally grounded evaluation of large language models for Indonesian language understanding Achraf El Bouazzaoui
Indonesian Journal of Computational Language Studies Vol. 1 No. 2 (2026): Large language models, linguistic diversity, and responsible NLP in Indonesia
Publisher : CV Narasi Khatulistiwa Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.67490/ijcl.v1i2.911

Abstract

Background: Translated benchmarks have expanded multilingual large language model evaluation, yet Indonesian remains a sociolinguistically dense setting where meaning is shaped by public institutions, regional languages, religious discourse, and culturally situated inference. Objective: This study aims to examine how Indonesian language understanding can be evaluated through source-grounded, culturally embedded corpus items rather than through translated English-centric tasks. Method: Using a qualitative-diagnostic benchmark design, this study constructs and codes an auditable corpus of public Indonesian documents, benchmark references, and contextual materials across semantic, pragmatic, cultural, regional, institutional, safety, and multimodal dimensions. Results: The analysis shows that semantic comprehension appears as a baseline requirement but does not sufficiently capture Indonesian understanding. Institutional register and cultural commonsense emerge as dominant dimensions, indicating that Indonesian prompts frequently require recognition of public authority, local values, collective memory, and socially constrained meaning. Implication: Task sensitivity is highest when documents combine policy language, cultural inference, safety cues, regional references, and pragmatic restraint, revealing why ordinary accuracy metrics may obscure fragile understanding. Novelty: This study contributes a culturally grounded evaluation perspective that reframes Indonesian LLM assessment as source-verifiable, context-sensitive interpretation rather than transferable language performance, and offers methodological resources for future multilingual benchmarking in low-resource and culturally plural settings within Southeast Asian digital ecologies
Do large language models reason consistently across Indonesian languages? a multilingual evaluation of Indonesian, Javanese, Sundanese, and Buginese Fitriya Dessi Wulandari
Indonesian Journal of Computational Language Studies Vol. 1 No. 2 (2026): Large language models, linguistic diversity, and responsible NLP in Indonesia
Publisher : CV Narasi Khatulistiwa Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.67490/ijcl.v1i2.912

Abstract

Background: Multilingual large language model evaluation increasingly recognises Indonesian as a significant benchmark language, yet Indonesia’s wider linguistic ecology requires assessment across local languages whose digital representation remains uneven. Objective: This study examines whether large language models reason consistently across Indonesian, Javanese, Sundanese, and Buginese when semantically aligned reasoning tasks are presented in each language. Method: Using a controlled multilingual evaluation design, this study compares model outputs through three analytic dimensions: final-answer consistency, explanation coherence and faithfulness, and language-linked error patterns. Results: Findings show that answer stability is not evenly preserved across language pairs, with stronger alignment in Indonesian–Javanese and Indonesian–Sundanese comparisons than in Indonesian–Buginese comparison. Explanation analysis further indicates that apparently correct answers may be accompanied by compressed, partially faithful, or weakly grounded reasoning. Implication: Error analysis reveals that inconsistency emerges through multiple pathways, including lexical-semantic drift, register mismatch, translation-related distortion, cultural inferential misreading, and language fallback. Novelty: The novelty of this study lies in shifting Indonesian multilingual LLM evaluation from isolated benchmark accuracy to cross-linguistic reasoning consistency, positioning local languages as central analytic sites for assessing reliability, equity, and epistemic accountability in multilingual artificial intelligence
When safety fails in local languages: evaluating culturally sensitive responses of large language models in multilingual Indonesia Nadiyah
Indonesian Journal of Computational Language Studies Vol. 1 No. 2 (2026): Large language models, linguistic diversity, and responsible NLP in Indonesia
Publisher : CV Narasi Khatulistiwa Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.67490/ijcl.v1i2.914

Abstract

Background: In Indonesia’s digitally dense and linguistically diverse environment, LLM safety cannot be evaluated only through English or standardized Indonesian because harm is often expressed through local registers, indirectness, stigma, and culturally situated forms of advice-seeking. Objective: This study examines whether large language models produce safe, culturally sensitive, and verifiably helpful responses to safety-relevant prompts in Indonesian, Javanese, Sundanese, and Buginese. Method: Using a public-source corpus of 24 verified documents and 160 coded model responses, this study combines safety response auditing, cultural-pragmatic analysis, and grounded helpfulness assessment across domains including bullying, online gender-based violence, mental health, misinformation, and digital harm. Results: The analysis shows that safe refusal and mitigated completion dominate the response distribution, but unsafe compliance, over-refusal, evasion, and culturally inadequate redirection remain visible. Cultural-pragmatic adequacy declines in local-language prompts, where register mismatch, missed indirect distress, stigma reproduction, and unsupported cultural assumptions become more prominent. Implication: Grounded helpfulness is also uneven, as local-language responses more often contain incomplete referral pathways or unverifiable local claims in public-risk contexts. Novelty: This study contributes a multilingual Indonesian safety-evaluation framework that treats local languages not as translated inputs but as pragmatic environments where AI safety may shift, weaken, or become socially inadequate
Hallucinating the archipelago: detecting factual and cultural errors in retrieval-augmented generation for Indonesian knowledge Radna Tulus Wibisono
Indonesian Journal of Computational Language Studies Vol. 1 No. 2 (2026): Large language models, linguistic diversity, and responsible NLP in Indonesia
Publisher : CV Narasi Khatulistiwa Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.67490/ijcl.v1i2.915

Abstract

Background: Retrieval-augmented generation has been promoted as a practical response to hallucination in large language models, yet Indonesian knowledge presents a harder evidential problem because factual claims are often inseparable from regional attribution, cultural terminology, legal classification, and administrative scope. Objective: This study examines how factual and cultural errors appear in RAG-generated answers about Indonesian knowledge and how such errors can be traced across claim alignment, retrieval provenance, and cultural specificity. Method: Using a qualitative-computational design, this study evaluates 186 claim units derived from a public-source corpus of Indonesian cultural-heritage records, legal documents, statistical portals, lexical references, government explainers, and local institutional sources. Results: The findings indicate that RAG performs most reliably when generated claims reproduce explicit institutional facts, such as official entities, legal instruments, statistical categories, and cultural-heritage designations. However, unsupported and partially supported claims emerge when retrieved evidence is general, incomplete, or transformed into causal, nationalising, or culturally overextended explanations. Implication: Cultural errors are especially visible in regional scope reduction, terminological loss, ritual simplification, and temporal or administrative decontextualisation. Novelty: The novelty of this study lies in integrating claim–evidence alignment, retrieval-provenance diagnosis, and cultural-specificity coding into one framework for evaluating Indonesian RAG hallucination

Page 1 of 2 | Total Record : 12