cover
Contact Name
Husni Teja Sukmana
Contact Email
husni@bright-journal.org
Phone
+62895422720524
Journal Mail Official
jads@bright-journal.org
Editorial Address
Gedung FST UIN Jakarta, Jl. Lkr. Kampus UIN, Cemp. Putih, Kec. Ciputat Tim., Kota Tangerang Selatan, Banten 15412
Location
Kota adm. jakarta pusat,
Dki jakarta
INDONESIA
Journal of Applied Data Sciences
Published by Bright Publisher
ISSN : -     EISSN : 27236471     DOI : doi.org/10.47738/jads
One of the current hot topics in science is data: how can datasets be used in scientific and scholarly research in a more reliable, citable and accountable way? Data is of paramount importance to scientific progress, yet most research data remains private. Enhancing the transparency of the processes applied to collect, treat and analyze data will help to render scientific research results reproducible and thus more accountable. The datasets itself should also be accessible to other researchers, so that research publications, dataset descriptions, and the actual datasets can be linked. The journal Data provides a forum to publish methodical papers on processes applied to data collection, treatment and analysis, as well as for data descriptors publishing descriptions of a linked dataset.
Articles 628 Documents
An Integrated Linguistic and Metaheuristic-Optimized Elman Neural Network Framework for Cyberbullying Detection Siti Aisyah; Arnes Sembiring; Faadhil Faadhil; Hartono Hartono; Rahmad Syah; M. Khahfi Zuhanda
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1337

Abstract

The rapid growth of social media platforms has intensified the need for accurate cyberbullying detection systems capable of understanding contextual and linguistically complex expressions. Existing machine learning and deep learning approaches often suffer from limited interpretability, insufficient contextual understanding, and suboptimal parameter optimization, reducing their effectiveness in identifying harmful online content. This study proposes a novel cyberbullying detection framework that integrates Linguistic Rule-Based Feature Extraction, an Elman Neural Network (ENN), and the Local Search-Based Improved Bat Algorithm (LSBIA). The main contribution of this research lies in the synergistic combination of interpretable linguistic knowledge, contextual sequence modeling, and metaheuristic optimization within a unified classification framework. Linguistic rules are employed to capture negation patterns, intensifiers, and adjective–noun relationships, while ENN models contextual dependencies through recurrent memory structures. LSBIA is utilized to optimize network parameters and improve convergence stability. Experiments were conducted using textual data collected from Instagram, Twitter, and Facebook and evaluated using stratified 10-fold cross-validation. The proposed method achieved an accuracy of 99.12%, precision of 94.73%, recall of 97.45%, and F1-score of 93.91%, outperforming Support Vector Machine (91.20% accuracy), Naïve Bayes (89.75%), and Decision Tree (90.10%). Ablation experiments further demonstrated the importance of each component, where removing linguistic rules reduced accuracy to 94.90%, removing sentiment scoring reduced accuracy to 96.30%, and replacing ENN with LSTM, GRU, or Transformer architectures resulted in lower accuracies of 92.50%, 91.90%, and 93.20%, respectively. These findings confirm that integrating linguistic feature engineering, contextual neural modeling, and metaheuristic optimization significantly enhances cyberbullying detection performance while maintaining interpretability. The novelty of this study resides in the integration of linguistic rule-based representation with LSBIA-optimized ENN for context-aware cyberbullying classification.
Determinants of Consumers’ Behavioural Intention and E-Wallet Usage Behavior: Evidence from an Extended UTAUT Model via Structural Equation Modeling Bui Hong Diep; Dam Tri Cuong
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1436

Abstract

The rapid proliferation of digital financial technology is speeding up the acceptance of e-wallets as a favoured payment method owing to their simplicity, quickness, and accessibility. Despite the increasing prevalence of e-wallet services, prior research has yielded inconsistent results regarding the factors influencing consumers’ behavioural intentions and actual usage, especially in relation to performance expectancy, effort expectancy, facilitating conditions, trust, and perceived risk. This research enhances the Unified Theory of Acceptance and Use of Technology (UTAUT) model by integrating trust and perceived risk as supplementary antecedents of e-wallet adoption to explain these inconsistencies. Data were gathered using a structured questionnaire sent to 280 e-wallet users in Ho Chi Minh City, Vietnam. The suggested study model was evaluated by Structural Equation Modelling (SEM) to investigate the interrelations among performance expectancy, effort expectancy, social influence, facilitating conditions, trust, perceived risk, behavioural intention, and actual e-wallet use behaviour. The results demonstrate that performance expectancy, social influence, favourable conditions, and trust considerably and favourably affect customers’ propensity to use e-wallets. Moreover, facilitating conditions and behavioural intention were identified as having substantial beneficial impacts on actual e-wallet-using behaviour. In the case of an e-wallet, however, effort expectancy and perceived risk do not have a big effect on behaviour intentions. This research enhances the technology adoption literature by enhancing the explanatory ability of the UTAUT framework within the realm of digital payment systems. The results also offer practical implications for e-wallet providers and digital financial service firms by highlighting the necessity to improve service performance, bolster trust mechanisms, enhance supporting infrastructure, and utilise social influence to promote e-wallet adoption and utilisation.
Feature Selection Methods as Attribute Weighting Schemes for Clustering Corporate Financial Statements Tora Fahrudin; Dedy Rahman Wijaya
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1410

Abstract

Financial clustering plays an important role in uncovering hidden patterns and supporting decision-making in high-dimensional financial statement analysis. However, clustering performance is often constrained by noisy, redundant, and heterogeneous attributes that reduce cluster quality, interpretability, and robustness. Although previous studies in financial analytics have extensively explored clustering algorithms, limited attention has been given to systematically evaluating feature-selection-based attribute weighting strategies for improving clustering effectiveness in complex financial datasets. This study investigates the effectiveness of 6 feature-selection-based weighting strategies for improving financial statement clustering using annual financial data from 604 publicly listed companies on the Indonesia Stock Exchange during 2020-2023. Rather than relying solely on feature filtering, the evaluated methods were utilized to prioritize informative financial attributes and improve clustering structure. Clustering performance was assessed using internal validation metrics, including the Silhouette Coefficient, Dunn Index, Calinski Harabasz Score, and Davies Bouldin Index to evaluate cluster compactness, separation, and overall quality. The results demonstrate that feature-selection-based weighting substantially improves clustering quality compared with the baseline. Fisher Score achieved the strongest overall performance with a Silhouette score of 0.9899, Dunn Index of 1.7238, and Davies Bouldin Index of 0.0034, outperforming the baseline values of 0.8678, 0.3509, and 0.8716, respectively. Autoencoder-based ranking also produced highly competitive results, achieving a Silhouette score of 0.9757 and a Calinski Harabasz Index of 37,105.45. These findings indicate that prioritizing informative financial attributes significantly enhances cluster compactness, separation, and interpretability. The study contributes a comprehensive comparative evaluation of feature-selection-based weighting schemes and provides practical insights for financial segmentation, risk profiling, decision support, and data-driven financial analytics. This study further offers a foundation for developing more robust clustering frameworks for complex financial datasets.
Data Driven Evaluation of the New Learning Paradigm Using Machine Learning for Optimizing Graduate Outcomes Mesra Betty Yel; Relita Buaton; Yuma Akbar; Novriyenni Novriyenni
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1152

Abstract

In preparing students to face digital literacy and critical thinking transformations rapid, universities are therefore required to design and implement learning processes that are innovative, adaptive, and differentiated, in accordance with the new learning paradigm. However, the unemployment rate in Indonesia remains high approximately 5.98% vocational high school graduates and 4.8% diploma and university graduates. To develop highly skilled human resources, higher education must strengthen the competencies of students as future agents of change entering the workforce. The persistent unemployment rate among diploma and university graduates presents a national challenge that may hinder the progress of human capital development. Therefore, this study aims to develop a classification and association model linking new learning paradigm programs to student learning outcomes, in order to generate new knowledge and identify correlations among grade point average, employment waiting period, occupational field, and graduate income. The research employs a machine learning approach using association rule mining and the K-Nearest Neighbors algorithm to analyze correlations and predict graduate outcomes. Based on data processing of 450 graduate data who participated in the new paradigm learning program, the findings indicate that graduates under the new learning paradigm with grade point average ≥ 3.50 are significantly more likely to secure employment within ≤ 2 months, support = 20%, confidence = 100% based on processing a data set of 450 data. Participants in the teaching assistance program tend to experience longer waiting periods ≥ 6 months and lower initial earnings compared to those from other new learning paradigm pathways. Conversely, graduates involved in certified internships or independent study programs demonstrate higher earnings potential and stronger academic performance. The results confirm that the new learning paradigm exerts a positive influence on graduate employability, income level, and academic achievement, especially through experiential and industry-oriented learning mechanisms.
An Interpretable Machine Learning Framework for Imbalanced Audit Opinion Prediction Ha Thanh Nguyen
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1373

Abstract

Audit opinion prediction is important for capital-market monitoring because non-unqualified audit opinions may signal financial distress, reporting uncertainty, and elevated information risk. This study develops an interpretable machine learning framework for predicting non-unqualified audit opinions under severe class imbalance using firm-year observations of Vietnamese listed companies from 2015 to 2024. The proposed framework is based on a Modified Random Forest approach that integrates imbalance-aware sub-sampling, performance-based tree selection, ensemble probability aggregation, and decision-rule extraction. The framework is designed to improve minority-class detection while preserving transparent and economically interpretable risk signals. The empirical analysis compares the proposed framework with several benchmark models, including logistic regression, Support Vector Machine, k-nearest neighbours, Random Forest, XGBoost, Random Forest with synthetic minority oversampling, cost-sensitive Random Forest, and Balanced Random Forest. Predictive performance is evaluated using three validation strategies: a random 70/30 train-test split, an out-of-time split in which observations from 2015-2021 are used for training and observations from 2022-2024 are reserved for testing, and k-fold cross-validation. The proposed framework achieves the strongest overall performance across these settings, with area under the receiver operating characteristic curve values of 0.829, 0.811, and 0.793, respectively, while also improving minority-class recall and F1-score relative to benchmark models. The extracted decision rules indicate that audit opinion modifications are associated with persistent and interacting financial vulnerabilities, particularly prior modified opinions, weak profitability, leverage pressure, limited debt-servicing capacity, liquidity constraints, and asset-structure risk. Methodologically, the study contributes by integrating imbalance-aware ensemble learning with interpretable rule-based analysis. Practically, the framework may serve as a complementary early-warning tool for auditors, regulators, and investors in financial reporting risk assessment within emerging markets.
Extending Hybrid GRG-NS With LSTM-Based Demand Forecasting for Dynamic Multi-Depot Routing in Disaster Logistics Dedy Hartama; Poningsih Poningsih; Lili Tanti
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1444

Abstract

Disaster logistics management requires accurate demand forecasting and efficient routing optimization to ensure timely distribution of emergency supplies under dynamic and uncertain conditions. Conventional routing approaches often experience limitations in handling fluctuating disaster demand, resulting in inefficient distribution performance and increased operational costs. This study proposes an integrated LSTM–Hybrid Generalized Reduced Gradient and Neighborhood Search (LSTM–Hybrid GRG–NS) framework for disaster-demand forecasting and routing optimization. The proposed approach combines Long Short-Term Memory (LSTM) for sequential demand prediction with a hybrid GRG–NS optimization mechanism to improve routing efficiency and solution convergence. Experimental evaluation was conducted using disaster-demand scenarios and routing datasets to assess forecasting and optimization performance. The forecasting results demonstrated strong predictive capability with low MAE, RMSE, and MAPE values, indicating that the LSTM model effectively captured temporal demand patterns. Furthermore, the routing optimization results showed that the proposed framework successfully generated stable and near-optimal routing solutions while maintaining full demand fulfillment and efficient vehicle utilization. The convergence analysis also confirmed that the optimization process converged consistently within a limited number of iterations. Overall, the proposed LSTM–Hybrid GRG–NS framework provides an effective and reliable decision-support approach for proactive humanitarian logistics and disaster-routing management.
A Text Summarization-Based Similarity Algorithm for Inter-Subchapter Coherence Analysis in Indonesian Academic Documents Rogayah Rogayah; Achmad Benny Mutiara; Dina Anggraini
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1425

Abstract

In academic writing, the logical relationship between subchapters is essential for maintaining document coherence; however, manual evaluation of inter-subchapter consistency is time-consuming and subjective. This study proposes a hybrid text summarization and similarity-based algorithm integrating feature-based extractive summarization with BERT-based neural summarization and semantic similarity measurement using TF-IDF, cosine similarity, and Latent Semantic Analysis (LSA) to analyze inter-subchapter coherence in Indonesian dissertation qualification proposals. The principal novelty lies in operating at the subchapter level rather than the sentence or document level, enabling structural relationship analysis not addressed by existing approaches. Experiments were conducted on 30 Indonesian dissertation qualification proposals split into 21 documents (70%) for training, 3 (10%) for validation, and 6 (20%) for testing, annotated by domain expert evaluators using six quality criteria. Similarity analysis results show that the proposed method identifies strong semantic alignment between logically connected section pairs, with cosine similarity scores reaching 1.00 for the problem formulation - objectives pair and the background -methodology pair on the test set; these perfect scores reflect the structural consistency of academic proposals rather than normalization artifacts. In quality assessment, the model achieves an average exact-match accuracy of 50% against expert evaluations, with per-proposal accuracy ranging from 17% to 83%. The lower overall accuracy is attributed to BERT's tendency to over-predict quality in poorly structured documents, and these findings are reported as exploratory results given the limited test set size (n=6). The proposed framework makes a meaningful contribution toward automated academic writing assessment tools for Indonesian higher education, providing a structured, data-driven approach to evaluating proposal coherence that can serve as a foundation for future large-scale deployment.
Beyond Positive and Negative: A Directional SHAP Framework for Sequential Multi-Class Classification Ahmet Yalcin; Selim Cetin; Bekir Cetintav
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1375

Abstract

In this study, we address the interpretative limitations of the standard Shapley Additive Explanations (SHAP) method in sequential multi-class classification problems within the scope of Explainable Artificial Intelligence (XAI). Stemming from the observation that classical SHAP is restricted to revealing only positive and negative contributions for a single class, we propose a novel directional framework that categorizes feature effects as 'Lower' (driving towards a lower class), 'Upper' (driving towards a higher class), and 'Ambiguous' (representing inconsistent effects). To validate this approach, a Random Forest model predicting obesity levels across seven hierarchical classes was trained on an open-source dataset, achieving a classification accuracy of 95.5%. Furthermore, a stability analysis comprising 10,000 sampling iterations demonstrated the robustness of the proposed framework, with dominant features retaining their directional categorizations consistently in over 99.9% of the trials. The findings indicate that unlike standard SHAP, our method successfully isolates the specific variables that prevent an instance from ascending to a higher class or descending to a lower one, particularly clarifying the role of ambiguous boundary features. In conclusion, this modification significantly enhances model transparency for complex hierarchical scenarios, and the framework has been released as an open-source Python library to provide researchers with a practical tool for automated directional feature analysis.
A Diagnostic Framework for Staged AI Adoption in Batik Motif Recognition: Integrating CNN Evidence and Implementation Readiness Irwan Sembriring; Paminto Agung Christianto; Eko Budi Susanto; Suharyadi Suharyadi; Cheryl Louisa Loedwyca
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1427

Abstract

This study proposes a diagnostic dual-layer decision-support framework for staged artificial intelligence adoption in batik motif recognition. The objective is to examine whether technical evidence from convolutional neural network classification and perceived implementation-side readiness can be jointly interpreted to prevent premature deployment in cultural-heritage recognition. The contribution of the study is not a new classifier architecture, but an operational diagnostic logic that treats model performance, class-level instability, readiness perception, governance, security, and feedback mechanisms as complementary but non-substitutable evidence. Methodologically, the technical layer evaluated three transfer-learning baselines, namely VGGNet-16, ResNet50, and MobileNetV2, using 983 batik images across 20 motif classes. The implementation layer assessed perceived readiness among 173 information technology practitioners using the Technology-Organization-Environment-Human plus Feedback dimensions. The integration layer then mapped technical-readiness evidence and readiness perception into explicit staged-adoption decisions rather than averaging them as interchangeable indicators.  The analysis used performance summaries, readiness profiles, decision matrices, security checklists, learning curves, and confusion-matrix diagnostics to connect empirical observations with staged adoption recommendations. The best-performing baseline was ResNet50, with 45% accuracy and a macro F1-score of 0.40, showing low technical readiness and substantial motif-specific instability. In contrast, the readiness survey indicated high perceived implementation-side readiness, with an average agreement score of 78.7%. This mismatch reveals a readiness asymmetry: implementation support may exist even when the recognition model remains technically immature. The findings imply that batik-recognition systems should prioritize dataset expansion, expert label validation, model refinement, moderated feedback, security governance, and controlled pilot testing before operational deployment. The framework provides a transparent basis for risk-aware, staged adoption decisions in artificial-intelligence-assisted cultural heritage preservation.
Rapid Ecosystem-Driven Deep Learning: On-Device Grain Type Classification and Authentication using iOS Swift and Core ML Trianggoro Wiradinata
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1414

Abstract

The 2025 Indonesian rice scandal highlighted major shortfalls in food security and the pressing need for robust, data-based verification of authenticity. The goal of this work is to design a fast, lightweight, and fully offline rice grain classification and verification system that can run directly on consumer mobile hardware. The basic idea to overcome the technical bottleneck of deploying complex computer vision models on edge devices is to use macOS and Apple’s unified ecosystem as a rapid prototyping and deployment platform. The deliberate avoidance of fragmented and high-latency workflows such as external environments (e.g. TensorFlow or PyTorch) and intermediate formats (e.g. ONNX) is mentioned. The study contributes a streamlined pipeline that incorporates an Image Feature Print V1 feature extractor, natively trained with Create ML on a publicly available dataset of 75,000 balanced images of five rice varieties (Arborio, Basmati, Ipsala, Jasmine, and Karacadag), directly into a native iOS application built with SwiftUI. The novelty of this approach is the usage of native tools like VisionKit and Core ML, which enables the complete elimination of third-party bridging code that normally bloats the binary overhead. The results show excellent edge efficiency on an iPhone 15 with a median prediction rate of 2.75 ms, an initial load time of 0.54 ms and a compilation latency of 3.66 ms. Moreover, the results reveal that by employing aggressive data augmentations, including the addition of visual noise, blur, exposure adjustment, flipping and rotation to ensure robustness, and by deliberately not cropping to preserve absolute grain dimensions, the model achieved a remarkable overall accuracy of 97% with an ultra-compact deployment footprint of only 66 KB. These metrics demonstrate that a fast, fully offline and privacy-preserving verification system is well within reach with modern consumer hardware.