This Author published in this journals
All Journal Jurnal Teknoinfo
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Building a Digital Lexical Resource for Banyumasan Javanese: A Low-Resource Language Approach Nisrina Hanifa Setiono; Angga Kurniawan; Viga Laksa Hardjanto
Jurnal Teknoinfo Vol. 20 No. 2 (2026): Period July 2026
Publisher : Universitas Teknokrat Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.33365/teknoinfo.v20i2.1705

Abstract

Banyumasan Javanese, widely recognized through the Ngapak dialect, remains culturally significant but is still underrepresented in reusable computational resources. This study develops a digital Banyumasan-Indonesian lexical corpus and frames it as a reusable research artifact rather than a static appendix. The corpus was constructed from a Banyumasan-Indonesian dictionary, normalized into a structured bilingual dataset, and packaged as an installable Python resource so that it can be used directly in computational experiments. The implemented resource supports dataset loading, Banyumasan lookup, Indonesian lookup, simple translation, structured translation analysis, batch translation, and corpus statistics. The resulting corpus contains 2,000 lexical pairs, 1,996 unique Banyumasan forms, 1,444 unique Indonesian equivalents, and 4 duplicated Banyumasan headwords that preserve lexical ambiguity from the source material. To demonstrate practical utility, the study includes a 100-sentence implementation example in which Banyumasan text is translated with the published banyumasan-corpus package and evaluated against Indonesian ground truth using the Indonesian-focused embedding model LazarusNLP/all-indo-e5-small-v4. The average semantic similarity rises from 0.4833 for direct Banyumasan-versus-ground-truth comparison to 0.6427 after translation, producing an absolute gain of 0.1594 and a relative improvement of approximately 33.0% over the baseline. These findings indicate that a structured lexical corpus, when distributed in a directly reusable computational form, can strengthen both resource accessibility and small-scale downstream experimentation for a low-resource regional language.