In the digital era, folktales as part of Indonesia's cultural heritage are increasingly being digitized and widely disseminated through various media platforms. This study aims to apply the Latent Dirichlet Allocation (LDA) method to analyze 300 digitized Indonesian folktales. After the data collection and preprocessing stages, the dataset was divided into two parts: 70% for training and 30% for testing. The training process produced 9 topics that represent various themes found in the folktales. Evaluation was carried out using four main metrics: Coherence Score, Precision, Recall, and F1-Score. The Coherence Score measures the quality of the semantic relationship between words in each topic, while Precision, Recall, and F1-Score are used to assess the system’s accuracy, completeness, and balance in presenting relevant folktales based on the topic entered by the user. The results show that the LDA model achieved the highest coherence score of 0.6250 with 9 topics. Meanwhile, evaluation showed a precision of 76%, recall of 71.2%, and an F1-score of 73.5%, indicating that the system can effectively present relevant folktales. This study contributes to the development of a topic-based search system implemented in the Nusantara Panji Kediri application, enabling users to easily find stories according to their preferred topics. In addition, the study fills a research gap by applying LDA to a corpus of Indonesian folktales, which remains underexplored in the context of thematic search systems. The system also demonstrates great potential in the development of information technology based on local cultural content. However, the limited amount of data used—only 300 folktales—presents a challenge and opens opportunities for future improvements with a broader dataset, as well as testing the model on more diverse types of texts to enhance the system’s generalizability.
Copyrights © 2025