Mohammad Asaduzzaman Rasel
Monash University

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Construction of a Dialect-Sensitive Javanese Semantic Lexicon to Support Machine Translation Systems Musthofa Galih Pradana; Ridwan Raafi'udin; Nurul Afifah Arifuddin; Mohammad Asaduzzaman Rasel
Journal of Applied Informatics and Computing Vol. 10 No. 4 (2026): August 2026
Publisher : Politeknik Negeri Batam

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30871/jaic.v10i4.13379

Abstract

The development of linguistic resources for natural language processing (NLP) in Javanese remains limited, especially regarding the representation of semantic relationships between different speech levels. This study aims to construct a Javanese semantic lexicon that integrates Indonesian lexical equivalents with three Javanese speech levels: ngoko, krama alus, and krama inggil. A research design based on lexical resource construction was employed, using a Javanese digital dictionary as the primary data source. The methodology included data extraction, preprocessing, semantic lexicon construction, analysis of speech level variation, and a preliminary exploration of polysemous lexical entries using automatic identification, followed by validation by native speakers. The resulting semantic lexicon successfully represents lexical relationships between levels in a structured manner. Analysis of speech-level variation revealed that partially distinct lexical patterns were the most dominant, with 733 entries, followed by fully distinct patterns (193 entries) and identical patterns (21 entries). These findings indicate that speech-level differences in Javanese are selectively realized and should be explicitly considered in the development of linguistic resources. Furthermore, preliminary exploration of polysemous candidates demonstrated that dictionary-based automatic identification can overestimate polysemy without linguistic validation. Only a limited number of lexical entries exhibited features consistent with genuine polysemous relationships. This study provides an initial basis for the development of Javanese semantic resources that are sensitive to speech-level variation and semantic complexity. The constructed semantic lexicon has the potential to support future research in NLP applications in Javanese, including politeness identification, lexical normalization, word sense disambiguation, and machine translation.