cover
Contact Name
Husni Teja Sukmana
Contact Email
husni@bright-journal.org
Phone
+62895422720524
Journal Mail Official
jads@bright-journal.org
Editorial Address
Gedung FST UIN Jakarta, Jl. Lkr. Kampus UIN, Cemp. Putih, Kec. Ciputat Tim., Kota Tangerang Selatan, Banten 15412
Location
Kota adm. jakarta pusat,
Dki jakarta
INDONESIA
Journal of Applied Data Sciences
Published by Bright Publisher
ISSN : -     EISSN : 27236471     DOI : doi.org/10.47738/jads
One of the current hot topics in science is data: how can datasets be used in scientific and scholarly research in a more reliable, citable and accountable way? Data is of paramount importance to scientific progress, yet most research data remains private. Enhancing the transparency of the processes applied to collect, treat and analyze data will help to render scientific research results reproducible and thus more accountable. The datasets itself should also be accessible to other researchers, so that research publications, dataset descriptions, and the actual datasets can be linked. The journal Data provides a forum to publish methodical papers on processes applied to data collection, treatment and analysis, as well as for data descriptors publishing descriptions of a linked dataset.
Articles 628 Documents
Course-Disjoint Evaluation and Capacity-Aware Triage for Student Dropout Risk Prediction Wijiyanto Wijiyanto; Aris Marjuni; Ahmad Zainul Fanani; Ruri Suko Basuki
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1419

Abstract

Early-warning systems for student dropout prevention require evaluation protocols and outputs that remain reliable when applied across heterogeneous academic contexts. This study quantifies how conventional random splits can overestimate performance when models are expected to generalize across different courses and proposes a decision-support layer that translates predicted risk into capacity-aware intervention policies. Using a benchmark higher-education dataset (N=4,424; 34 predictors; three classes: Dropout, Enrolled, Graduate) with 17 Course groups, phased prediction is implemented to reflect increasing evidence availability: S0 (pre-enrollment), S1 (plus semester-1 academic evidence), and S2 (plus semester-2 academic evidence). Baseline results are replicated with leakage-safe preprocessing (imputation, one-hot encoding, scaling) and Synthetic Minority Over-sampling Technique (SMOTE) applied strictly within training folds, comparing multinomial logistic regression, random forest, and tree-based boosting models. Deployment-oriented performance is assessed using StratifiedGroupKFold by Course to enforce course-disjoint testing. Discrimination is reported with Macro-F1 and Balanced Accuracy, while probability quality is evaluated using LogLoss, Brier score, expected calibration error, maximum calibration error, and reliability diagrams. Calibrated probabilities are translated into capacity-aware risk bands (Top-k% triage), selective prediction is evaluated via abstention to defer low-confidence cases, and split conformal prediction sets are optionally reported for multiclass uncertainty communication. Results show consistent performance drops under course-disjoint validation, confirming a non-trivial generalization gap. Error decomposition indicates that Enrolled is the most ambiguous class and exhibits phase-dependent confusion toward both terminal outcomes. Calibration shows phase-specific trade-offs between likelihood-based and worst-case calibration metrics, while risk bands yield high-precision triage under limited capacity, and abstention improves decision quality at reduced coverage. Overall, the study provides a deployment-oriented evaluation and decision-support workflow for translating dropout risk models into actionable capacity planning.
Modeling Thermomechanical Effects in Rods with Variable Cross-Section and Local Heat Sources Zhuldyz Tashenova; Shirin Amanzholova; Zhanat Abdugulova; Elmira Nurlybaeva
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1340

Abstract

This study is a numerical study of the transient thermomechanical behavior of rods with variable cross-section exposed to partial thermal insulation, localized heat flows, and convective heat exchange. The main goal is to increase the accuracy and reliability of modeling related thermomechanical reactions in heterogeneous structural elements widely used in engineering applications. The main idea is based on the development of a mathematically consistent model based on the principle of conservation of energy, which allows simultaneous assessment of temperature fields and stress-strain state under inhomogeneous boundary conditions. The contribution of this work is to develop a computational algorithm capable of detecting spatial and temporal temperature changes and deformations with increased stability and efficiency. Numerical calculations show that the temperature along the rod varies non-linearly, reaching a maximum increase of about 18-25% near the zones of concentrated heat flow, while the thermal elongation differs by up to 12% compared with homogeneous models. The calculated stress values indicate an increase in peak thermal stresses by 20-30% in areas with a reduced cross-sectional area, which confirms the strong influence of geometric variability. Convergence analysis shows that reducing the sampling step by 50% increases the accuracy of the solution by about 8-10% while maintaining computational costs within acceptable limits (an increase of 15%). The results confirm that the proposed method provides stable solutions with an error of less than 5% compared to reference analytical solutions. The novelty of the research lies in the integration of variable geometry, mixed boundary conditions, and transient effects into a single numerical model that provides more realistic predictions than traditional simplified models. The results obtained can be effectively applied in the design and optimization of the main components operating under coupled thermomechanical loads.
Enhancing Low-Resource Lampung Speech Recognition through Cross-Lingual XLSR-Wav2Vec 2.0 Pretraining Hendra Kurniawan; Akmal Junaidi; Favorisen Rosyking Lumbanraja; Wamiliana Wamiliana
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1388

Abstract

This study investigates the application of Wav2Vec 2.0 (W2V2) and Cross-Lingual Speech Representation (XLSR) models to Lampung language speech recognition. LampungNyow v1.0 is introduced, a speech corpus designed to provide a baseline for training and evaluating Automatic Speech Recognition (ASR) for this low-resource regional language of Indonesia. The dataset enables supervised fine-tuning and standardized evaluation, addressing the lack of publicly available linguistic resources for Lampung. Several pre-trained W2V2 models on Lampung speech recognition using Word Error Rate (WER) as the evaluation metric. The evaluated models include W2V2-Base, W2V2-Large, W2V2-Large-XLSR-Indonesian, W2V2-Large-XLSR-Sundanese, W2V2-Large-XLSR-53, and the multilingual W2V2-Large-XLSR-Indonesia-Javanese-Sundanese model. Monolingual models have higher WER values, according to experimental results: W2V2-Base achieved 36,23%, while W2V2-Large achieved 36,30%. XLSR models, such as XLSR-53 (33,88%), Sundanese (33,99%), and Indonesian (33,70%), demonstrated modest improvements. The W2V2-Large-XLSR-Indonesian-Javanese-Sundanese model, which was the foundation for the Lampung automatic speech recognition system in this study, achieved lower WER of 17,39%. These findings suggest that, in contrast to more comprehensive multilingual or monolingual pretraining models, multilingual pretraining utilizing a number of Indonesian regional languages can produce acoustic and contextual speech representations that are better suited for the resource-constrained Lampung automatic speech recognition task. When compared to the baseline W2V2-Large model, the obtained WER of 17,39% indicates a relative improvement of more than 50%.
FCI-ANTREE: A GUI-Centric Method for Schema Recovery and Conceptual Database Model Reconstruction in Legacy Form-Based Systems Juanda Hakim Lubis; Elviawaty Muisa Zamzami; Mahyuddin K. M Nasution; Mohammad Andri Budiman
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1383

Abstract

Legacy systems that lack technical documentation present significant challenges for database schema recovery, particularly when access to source code and SQL queries is unavailable. Existing reverse engineering approaches predominantly rely on backend artifacts such as database logs, schema definitions, or program code, limiting their applicability in undocumented environments. Although GUI-based approaches offer an alternative by utilizing interface-level information, many existing methods still rely on shallow visual parsing and lack systematic mechanisms to capture structural and interaction semantics. To address this limitation, this study proposes FCI–ANTREE, a GUI-centric method for reconstructing conceptual database schemas from legacy form-based systems. The method integrates Form-Centric Interaction (FCI) to extract candidate entities, attributes, relationships, constraints, and data types from user interactions and validation logic, and Admin Interface Tree (ANTREE) to model hierarchical relationships and transform them into logical and relational schemas. The objective of this research is to provide a systematic, interpretable, and semi-automated approach for database reverse engineering without relying on backend access. The proposed method was evaluated using three case studies: an online store application, a library system, and an inventory application. The evaluation employed structural consistency analysis, confusion-matrix-based metrics, and quantitative error measurements. The results show that the method achieved a mean MAE of 1.77, RMSE of 3.16, and R² of 97.09%, along with an average F1-score of 0.8768, indicating a high level of agreement between reconstructed and reference schemas. These findings demonstrate that FCI–ANTREE provides an effective and practical solution for database schema reconstruction in legacy systems with limited or no backend accessibility. The method contributes by introducing an interaction-aware and rule-based framework that enhances the accuracy, interpretability, and applicability of GUI-driven reverse engineering.
Diffusion2D and Anchored Inference for Asymptotic Stabilization of Diffusion-Convolutional Neural Networks in Multidomain Medical Image Classification Hanna Willa Dhany; Sutarman Sutarman; Poltak Sihombing; Mohammad Andri Budiman
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1215

Abstract

Medical image classification across heterogeneous domains remains challenging due to domain shift, spatial variability, and unstable inference behavior. This study proposes a diffusion-stabilized Diffusion-Convolutional Neural Network (DCNN) framework that integrates Diffusion2D and post-hoc Anchored Diffusion to improve inference stability, probabilistic consistency, and robustness in multidomain medical image classification. The main contribution of this work is the introduction of a two-stage stabilization mechanism in which Diffusion2D performs controlled intra-image diffusion on feature representations before graph construction, while Anchored Diffusion refines uncertain predictions in the logit space through a k-nearest neighbors graph without retraining. The framework was evaluated on heterogeneous medical imaging datasets consisting of brain MRI, leukemia microscopy, and COVID-19 chest radiographs. Experimental results show that the proposed approach maintained baseline classification performance with an accuracy of 64.70% while improving the Macro-F1 score from 0.7045 to 0.7061. The diffusion mechanism reduced the average Laplacian value from 0.864355 to 0.187525, corresponding to a 78.23% reduction in spatial gradient variability. Internal analysis further demonstrated stable diffusion coefficients with a mean value of 0.141734 and a standard deviation of 0.003757, indicating controlled diffusion behavior. Anchored Diffusion selectively refined uncertain predictions, affecting only 0.6% of evaluated samples while preserving overall decision consistency. Repeated inference experiments across 40 iterations also revealed highly stable confidence trajectories with no observable variance after diffusion stabilization. The novelty of this research lies in combining feature-level diffusion stabilization, post-hoc anchored inference, and asymptotic regularization within a unified DCNN framework, providing a theoretically grounded and uncertainty-aware approach for robust multidomain medical image classification.
Performance Comparison of K-Means and Hybrid Hierarchical–Partitioning Methods for Clustering Efficiency Bowo Winarno; Budi Warsito; Bayu Surarso
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1140

Abstract

Clustering is a fundamental technique in data analysis, particularly for exploring patterns in large-scale datasets. While K-Means is widely used for its simplicity and efficiency, its performance is highly sensitive to centroid initialization, which can affect both clustering quality and convergence speed. Hierarchical clustering methods, such as agglomerative and divisive approaches, provide more structured and deterministic initialization but incur higher computational cost. This study evaluates two hybrid models—Agglomerative K-Means and Divisive K-Means—where hierarchical clustering is used to initialize centroids, followed by K-Means refinement. This approach aims to reduce the limitations of random initialization while improving clustering stability and efficiency in large-scale data environments. Experiments on poverty data from Central Java Province show that hybrid methods accelerate K-Means convergence: Agglomerative K-Means reduced iterations to 2 (from 3 in standard K-Means), while Divisive K-Means converged in 1 iteration. Silhouette, Davies–Bouldin, and Calinski–Harabasz indices indicate that Agglomerative K-Means achieves the most compact and well-separated clusters, whereas Divisive K-Means performed worse than standard K-Means. Execution time measured only during the K-Means refinement phase shows that hybrids converge faster (Agglomerative: 2.04 ms; Divisive: 1.91 ms; K-Means: 116.68 ms), though this does not account for the hierarchical initialization cost. These findings provide practical insights into the trade-offs between clustering quality and computational efficiency when applying hybrid clustering methods. Overall, these results demonstrate that hybrid approaches can improve clustering stability and convergence efficiency, with Agglomerative K-Means providing the best balance between cluster quality and computational performance.
Enhancing SMOTE-ENN Efficacy on Imbalanced Datasets Using Decision Tree Leaf Feature Extraction: A Case Study on Student Employability Data Rizkysari Meimaharani; Widowati Widowati; Ahmad Abdul Chamid
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1396

Abstract

This study looks at the challenge of classifying tabular data that is highly imbalanced and overlapping, where standard predictive models often lose performance and tend to focus too much on the majority class. Another problem is that many advanced ensemble models are highly complex and lack transparency. These models are often viewed as black boxes, making it difficult for users to clearly and explain how each feature contributes to the final prediction result.This study offers a hybrid classification approach to address the problem, by combining rule extraction from decision tree leaves, SMOTE-ENN resampling technique, and XGBoost algorithm to improve prediction performance more accurately and reliably.The leaf extraction process helps reorganize the data by separating overlapping class regions into clearer and more structured groups before synthetic samples are generated. The test results show that the proposed approach is able to exceed the performance of the baseline model, by obtaining an F1-score of 0.8554 which indicates increased effectiveness and balance in prediction. In addition to improving performance, this method also keeps the model interpretable. Instead of relying only on abstract engineered features, the model allows us to trace important features back to the original decision tree rules. This approach helps explain the prediction formation process more transparently, so that each model decision can be understood clearly, logically, and easily interpreted. Overall, the combination of Decision Tree, SMOTE-ENN, and XGBoost is effective in handling extreme class imbalance, while producing a clear, stable, and easy-to-understand model, making it more reliable and trustworthy in various real-world applications.
Factors Influencing Customer Repurchase Intention and Word-Of-Mouth Behavior in Food and Beverage Chain Stores Dat Tuan Nguyen; Trang Thi Kieu Ta
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1316

Abstract

This study investigates how customer experience attributes shape perceived value, satisfaction, repurchase intention, and word-of-mouth behavior in Vietnamese food and beverage chain stores. The study addresses the managerial problem of retaining customers in a competitive food and beverage market and contributes to the literature by clarifying a sequential value-satisfaction-behavior mechanism rather than examining isolated direct effects. Drawing on expectation conformation theory and perceived value theory, the research proposes a framework in which store atmosphere, price, perceived quality, product quality, and service quality influence perceived value; perceived value influences satisfaction; and satisfaction affects repurchase intention and word-of-mouth. A mix-methods design was used. The qualitative stage involved expert review and customer pretesting to refine measurement items and improve contextual relevance, while the quantitative stage used a structured survey of customers who had recently experienced food and beverage chain-store services in Vietnam. The research framework and seven empirical tables report the model, respondent profile, measurement items, reliability and validity indices, discriminant validity, hypothesis testing, mediation effects, and explanatory and predictive capability. After data screening, 432 valid responses were analyzed through reliability assessment, validity testing, mediation analysis, and partial least squares structural equation modeling, results show that all five  experiential factors positively influence perceived value. Store atmosphere has the strongest effect among the antecedents (β = 0.313), while price has a smaller but significant effect (β = 0.177), suggesting that customers evaluate price through the overall experience rather than cost alone. Perceived value strongly increases satisfaction (β = 0.552), and satisfaction significantly enhances both repurchase intention (β = 0.400) and word-of-mouth (β = 0.369). The findings imply that chain-store managers should coordinate atmosphere, product consistency, service responsiveness, and value-based pricing as an integrated experience system to strengthen satisfaction, repeat purchase, and customer advocacy.
Hybrid MINLP-FNS Framework for Solving Large-Scale LIRP in E-Retail Logistics Kristian Telaumbanua; Syahril Efendi; Poltak Sihombing; Maya Silvi Lydia
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1114

Abstract

This research proposes a hybrid optimization framework to address the large-scale, multi-echelon Location-Inventory-Routing Problem (LIRP) in e-retail logistics. The proposed method combines a Mixed-Integer Nonlinear Programming (MINLP) model with a tailored Feasible Neighborhood Search (FNS) algorithm to solve complex decision-making problems involving facility location, inventory control, and vehicle routing simultaneously. Distinct from conventional models, the framework is integrated with a Business Intelligence (BI) environment to enable real-time decision support and dynamic data processing. Experimental evaluations were conducted using realistic logistics scenarios involving four distribution echelons. The results show that the hybrid MINLP-FNS approach achieves a total logistics cost reduction of 13.6% and a runtime improvement of 48% compared to the baseline MINLP-only model. It also significantly outperforms traditional GRG-based methods in scalability and computational stability across large datasets. These findings demonstrate that the proposed hybrid framework offers a more effective and scalable solution for complex logistics optimization, while its BI integration ensures practical applicability in real-world operations. This study contributes a novel data-driven framework that advances current research in intelligent supply chain and optimization systems.
Determinants of Stock Volatility of Indonesian State-Owned Enterprises on the Indonesia Stock Exchange: A Panel-Data Modelling Approach with Firm Size as a Moderating Variable I Dewa Made Tirta Meirsha; I Gusti Ketut Agung Ulupui; Adler Haymans Manurung
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1500

Abstract

This study examines the determinants of stock-price volatility among Indonesian State-Owned Enterprises (SOEs) listed on the Indonesia Stock Exchange (IDX) and tests whether Firm Size (FS) moderates these relationships. Five groups of explanatory features are integrated into a single model: macro fundamentals (MF), micro fundamentals (MiF), capital structure measured by Government Funding (GF), Stock-Market Volatility (SMV), and Market Liquidity (ML). Using a balanced panel of 24 listed SOEs over 2019–2024 (144 firm-year observations), model selection across common-, fixed-, and random-effects specifications identifies the Common Effects Model (CEM) as the most appropriate. The model is jointly significant (Prob. F = 0.0001) and explains 40.37% of the variation in volatility (R2 = 0.4037). The findings reveal that inflation (β = -29.412, p = 0.022), the debt-to-equity ratio (β = -0.9996, p = 0.035), return on assets (β = -1.6373, p = 0.011), and government funding (β = -0.0197, p = 0.002) each significantly and negatively shape stock volatility. Broad market volatility and trading volume are statistically insignificant. Furthermore, firm size significantly moderates the exchange rate effect (β = 5.1933, p = 0.003) and the return on assets effect (β = -0.4727, p = 0.023). This paper contributes to literature by offering a unified, data-driven framework and extending efficient-market, behavioural-finance, and trade-off theories to state ownership, demonstrating that state-backed funding acts as a stabilizing capital component rather than a risk driver. Practical recommendations for the Ministry of SOEs include establishing performance-linked funding benchmarks and public-service account separation to mitigate market distortions and stock volatility.