cover
Contact Name
Aji Prasetya Wibawa
Contact Email
keds.journal@um.ac.id
Phone
+62818539333
Journal Mail Official
keds.journal@um.ac.id
Editorial Address
Semarang St. No. 5, Malang, East Java, 65145, Indonesia
Location
Kota malang,
Jawa timur
INDONESIA
Knowledge Engineering and Data Science
ISSN : -     EISSN : 25974637     DOI : http://dx.doi.org/2597-4637
The journal welcomes experimental and theoretical findings on data science and knowledge engineering along with their applications to real-life situations.
Articles 117 Documents
Social Media Mining with Fuzzy Text Matching: A Knowledge Extraction on Tourism After COVID-19 Pandemic Manuaba, Ida Bagus Putra; Sentana, I Wayan Budi; Astawa, I Nyoman Gede Arya; Suasnawa, I Wayan; Pradnyana, I Putu Bagus Arya
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Social media mining is an emerging technique for analyzing data to extract valuable knowledge related to various domains. However, traditional text matching techniques, such as exact matching, are not always suitable for social media data, which can contain spelling mistakes, abbreviations, and variations in the use of words. Fuzzy matching is a text matching technique that can handle such variations and identify similarities between two texts, even if there are differences in spelling or phrasing. The gap in existing research is the limited use of fuzzy matching in social media mining for tourism recovery analysis. By applying fuzzy matching to social media data related to COVID-19 and tourism recovery, this research seeks to bridge this gap and extract valuable insights related to the impact of the pandemic on tourism recovery. We manually retrieved 19,462 Twitter records and differentiated the data sources using four diver parameters to indicate data related to the impact of COVID-19 on the tourism industry, such as the economy, restrictions, government policies, and vaccination. We conducted text mining analysis on the collected 7,352 words and identified 25 highly recommended words that indicated COVID-19 recovery from a tourism perspective. We separated the four words representing the tourism perspective to perform fuzzy matching as a dataset. We then used the inbound dataset on the fuzzy matching process, with the 7,352-word data collected from the text mining process. The matching process resulted in 18 words representing COVID-19 recovery from a tourism perspective.
Can Multinomial Logistic Regression Predicts Research Group using Text Input? Rosyid, Harits Ar; Putra, Aulia Yahya Harindra; Akbar, Muhammad Iqbal; Dwiyanto, Felix Andika
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

While submitting proposals in SISINTA, students often confuse or falsely submit their proposals to the less relevant or incorrect research group. There are 13 research groups for the students to choose from. We proposed a text classification method to help students find the best research group based on the title and/or abstract. The stages in this study include data collection, preprocessing data, classification using Logistic Regression, and evaluation of the results. Three scenarios in research group classification are based on 1) title only, 2) abstract only, and 3) title and abstract. Based on the experiments, research group classification using title-only input is the best overall. This scenario gets the most optimal results with accuracy, precision, recall, and f1-score successively at 63.68%, 64.91%, 63.68%, and 63.46%. This result is sufficient to help students find the best research group based on the text titles. In addition, lecturers can comment more elaborately since the proposals are relevant to the research group’s scope.
Traffic Density Prediction using IoT-based Double Exponential Smoothing Asmara, Rosa Andrie; Noprianto, Noprianto; Ilmy, Muhammad Ainur, State Polytechnic of Malang; Arai, Kohei
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

The number of vehicles and currents that tend to increase causes traffic density. A system is proposed to calculate the number of vehicles and predict real-time traffic density. This research uses Haar Cascade to detect the number of cars and motorcycles and the Double Exponential Smoothing (DES) for forecasting the number of vehicles on the road. MAPE describes forecasting accuracy as a base for selecting the best smoothing constant (Alpha). The best test results from June 13 to 20, 2020, are cars on June 14, 2020 (alpha 0.5, MAPE 0%) and Motorcylecycles on June 18, 2020 (alpha 0.5, MAPE 0.1134% ). The most significant MAPE results of the car were on June 15, 2020, with alpha 0.5 and MAPE 2.1073%. The 3 minutes haar cascade detects 72.58% of cars and 81.90% of motorcycles.
Predicting Heart Disease using Logistic Regression Anshori, Mochammad; Haris, M. Syauqi
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

A common risk of death is caused by heart disease. It is critical in the field of medicine to be able to diagnose cardiac disease in order to adequately prevent and treat patients. The most accurate method of prediction has the potential to both extend the patient's life and reduce the severity of their cardiac disease. The use of machine learning is one approach that may be taken to generate predictions. In this study, patient medical record information was used in conjunction with an algorithm for logistic regression in order to make heart disease diagnoses. The outcomes of the logistic regression have been utilized to achieve a high level of accuracy in the prediction of heart disease. To get the model coefficients needed for the equation, the experiment uses an iterative form of the logistic regression test. Iteration 14 produced the best results, with an accuracy of 81.3495% and an average calculation time of 0.020 seconds. The best iteration was reached at that point. The percentage of space that lies beneath the ROC curve is 89.36%. The findings of this study have significant implications for the field of heart disease prediction and can contribute to improved patient care and outcomes. Accurate predictions obtained through logistic regression can guide healthcare professionals in identifying individuals at risk and implementing preventive measures or tailored treatment plans. The computational efficiency of the model further enhances its applicability in real-time decision support systems.
Stable Numerical Solution of an Elliptic PDE Inverse Problem Subject to Incomplete Boundary Conditions Tayyeh, Qasim Abd Ali
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

This study addresses the challenging problem of solving inverse elliptic Partial Differential Equations (PDE) with incomplete boundary data, data available only on a part of the domain boundary. The aim is to develop a robust, effective numerical framework that consistently recovers parameters and/or sources from incomplete, ill-posed data. In the case of a variational problem discretized by the Finite Element Method (FEM) and solved by an adjoint-based optimization strategy, the framework uses Tikhonov regularization. Morozov's Discrepancy Principle is used to determine regularization parameters that achieve the best balance between accuracy and stability. Even with 5% noise in the measurement data, the unknown model parameters can be successfully reconstructed with an L2 error of less than 5%, according to numerical results. The method can be used to recover parameters for a variety of domain geometries, including those with complex shapes and re-entrant edges. This research demonstrates that the framework is robust and stable, delivering reliable solutions to this challenging problem class. The solution approach has direct application in medical imaging, geophysics, and engineering diagnostics, where boundary data cannot be obtained.
A Bounded Custom GPT Structures Operational Knowledge for Small-Industry Decision Support Arief, Ikhwan; Hasan, Alizar; Putri, Nilda Tri; Rahmann, Hafiz
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Small manufacturing and craft-based firms increasingly use Generative Artificial Intelligence (GenAI) through public chat interfaces, low-cost tools, and informal experimentation. However, these firms often make operational decisions with incomplete records, tacit owner knowledge, fragmented spreadsheets, and limited managerial capacity. Under such conditions, open-ended chatbots may generate fluent but unsafe recommendations by overlooking missing information, contradictory evidence, feasibility constraints, and implementation constraints. This study presents Asisten Cerdas Industri Kecil as a bounded Custom GPT artifact for operational diagnosis, priority selection, and short-horizon action planning in small industries. Using a Design Science Research approach, the study develops a documented artifact corpus comprising a master instruction contract, 15 curated knowledge modules, structured output rules, a 60-scenario evaluation dataset, and automated runner and judge scripts. The artifact formalizes bounded operational reasoning through minimum-field gating, evidence separation, Knowledge-Resource Advantage synthesis, feasibility-weighted priority ranking, stopping rules, and conjunctive batch evaluation thresholds. A technical verification layer was conducted on a three-scenario local-smoke subset using smollm2:135m, qwen2.5:0.5b, and qwen2.5:1.5b with a heuristic judge. Results show that the evaluation scaffold is executable and that model capacity affects bounded instruction following, although hard-guardrail and priority-compliance failures indicate that the results should be interpreted as technical verification rather than final validation. The study contributes an audit-ready bounded Custom GPT architecture, a formal decision model, and a reproducible evaluation protocol for future comparative assessment of bounded and unbounded GenAI decision-support artifacts.
Assessing Deep Learning Models and Hyperparameter Optimization for Stable Time-Series Electricity Load Forecasting Patrya, Sukma; Wibawa, Aji Prasetya; Aripriharta, Aripriharta
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Long-term electricity load forecasting plays an important role in ensuring system reliability, optimizing energy management, and making operational plans in face of continuously rising electricity demands. This study suggests a complete deep learning method for univariate forecasting of future electricity loads based on climatology and electricity consumption data for the period between 2019 and 2023. The initial dataset was cleaned, normalized, and partitioned chronologically into train/test datasets. Four train/test split cases (20/80, 40/60, 60/40, 80/20) were considered to explore the impact of different levels of historical data availability on the performance of the suggested framework from data-poor to data-rich situations. Then five deep learning structures (CNN, RNN, GRU, LSTM, and BiLSTM) were trained and tuned by using three different hyperparameter optimization methods. Grid Search provided an extensive exploration of parameter space to obtain solid baseline configurations for the considered neural networks, Random Search allowed efficient sampling of the search space to find high-quality deep learning models with lower computational expenses, and Particle Swarm Optimization (PSO) enabled adaptive optimization of near-optimal solutions via population-based optimization technique. Forecasting models were assessed by means of Mean Absolute Percentage Error (MAPE) and Root Mean Square Error (RMSE) metrics. Technical novelty of the study is related to the consideration of the impact of various hyperparameter optimization approaches on the quality and robustness of predictions provided by several deep learning architectures depending on the degree of historical data availability. Experiment results showed that the CNN tuned with the help of Random Search gave the best forecasting results in case of abundant training data (80/20 split) with RMSE=22.13. In case of scarcity of training samples (20/80 split) CNN tuned by Grid Search and PSO provided stable forecasts with MAPE≈0.077 and RMSE=37.56 with efficient reduction of prediction deviation.

Page 12 of 12 | Total Record : 117


Filter by Year

2018 2025


Filter By Issues
All Issue