cover
Contact Name
Husni Teja Sukmana
Contact Email
husni@bright-journal.org
Phone
+62895422720524
Journal Mail Official
jads@bright-journal.org
Editorial Address
Gedung FST UIN Jakarta, Jl. Lkr. Kampus UIN, Cemp. Putih, Kec. Ciputat Tim., Kota Tangerang Selatan, Banten 15412
Location
Kota adm. jakarta pusat,
Dki jakarta
INDONESIA
Journal of Applied Data Sciences
Published by Bright Publisher
ISSN : -     EISSN : 27236471     DOI : doi.org/10.47738/jads
One of the current hot topics in science is data: how can datasets be used in scientific and scholarly research in a more reliable, citable and accountable way? Data is of paramount importance to scientific progress, yet most research data remains private. Enhancing the transparency of the processes applied to collect, treat and analyze data will help to render scientific research results reproducible and thus more accountable. The datasets itself should also be accessible to other researchers, so that research publications, dataset descriptions, and the actual datasets can be linked. The journal Data provides a forum to publish methodical papers on processes applied to data collection, treatment and analysis, as well as for data descriptors publishing descriptions of a linked dataset.
Articles 628 Documents
ACO-Optimized DSTATCOM for Reactive Power Compensation in PV-Integrated Weak Grids Rajasree R; Lakshmi D; Stalin K; Karthick Manoj R; P Jeyarani
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1467

Abstract

The integration of photovoltaic generation into weak distribution grids introduces voltage instability, reactive power imbalance, harmonic distortion, and prolonged transient response due to low short-circuit capacity and high grid impedance. This study proposes an Ant Colony Optimization-based Distribution Static Synchronous Compensator for improving voltage regulation and power quality in a PV-integrated weak grid. The optimization algorithm is used to tune the proportional–integral controller parameters by minimizing a composite objective involving PCC voltage deviation, reactive power error, total harmonic distortion, feeder power loss, and settling time. The proposed system is modeled in MATLAB/Simulink and evaluated against two benchmark configurations: an uncompensated weak-grid system and a conventional PI-controlled DSTATCOM. The results show that the proposed controller reduces PCC voltage deviation from 0.060 pu to 0.012 pu and reactive power demand from 150 kVAR to 5 kVAR. Total harmonic distortion decreases from 6.5% to 1.6%, while the power factor improves from 0.78 lagging to 0.995. In addition, feeder real-power losses are reduced from 45 kW to 28 kW, and the settling time decreases from 1.8 s to 0.5 s. These findings demonstrate that ACO-based controller tuning provides consistent improvements in voltage regulation, reactive power compensation, harmonic mitigation, feeder efficiency, and dynamic recovery under weak-grid operating conditions.
The Application of the Analytical Network Process for Select Leaders in Higher Education Institutions Phan Hong Hai; Nguyen Anh Tuan; Vo Van Tuyen
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1251

Abstract

This study aims to examine the application of the Analytic Network Process (ANP) in selecting leaders for public universities in Vietnam. The research seeks to develop a transparent, scientific, and decision-making framework for evaluating leadership candidates based on multiple interrelated criteria in higher education governance. Empirical data were collected from experts, lecturers, and students at a public university in Vietnam, with a total sample size of 257 respondents. The research employed pairwise comparison matrices and ANP procedures, including criteria identification, consistency testing, weight calculation, and candidate ranking. Five criteria were identified for evaluating leadership candidates: professional qualifications and scientific research capacity, ethics and integrity, experience in management and organization, ability to inspire colleagues and students, and ability to mobilize financial resources. The results indicate that ethics and integrity (C2) with weight (0.2550), ability to mobilize financial resources (C5) with weight (0.2350), and experience in management and organization (C3) with weight (0.2309) were the most influential factors in leadership selection. Based on these weighted criteria, three leadership candidates were evaluated and ranked, demonstrating the suitability and effectiveness of the ANP model in supporting decisions for selecting university leaders. The study provides a scientifically based, transparent assessment tool while suggesting policy implications in senior human resource management, in line with the requirements of innovation and sustainable development of Vietnamese higher education.
A Multi-Model Framework for Autonomous Schema Discovery and Hybrid Natural Language Generation-Driven Augmented Business Intelligence Randy Permana; Sarjon Defit; Gunadi Widi Nurcahyo
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1479

Abstract

Traditional Business Intelligence (BI) frameworks rely heavily on data professionals to manually map relational metadata into structured star schemas during the ETL process, creating a significant operational bottleneck in data preparation and downstream interpretation of insights. To address these limitations, this study introduces an augmented analytics framework for autonomous schema discovery and hybrid Natural Language Generation (NLG) driven Augmented BI. In the data representation layer, a weighted hybrid feature fusion mechanism combines structural database metadata with contextual text embeddings produced by a pre-trained Sentence Transformer (all-MiniLM-L6-v2). In the multi-model machine learning layer, a multi-paradigm execution engine combines unsupervised geometric clustering models (K-Means, K-Medoids, DBSCAN) and supervised classifiers (SVM, Random Forest), and performance is evaluated using Leave-One-Out Cross-Validation (LOOCV). The resulting schemas are then dynamically projected into an in-memory OLAP cube. At the downstream insight interpretation layer, a Hybrid NLG engine combines a context-aware, rule-based router with an autoregressive generative decoding mechanism to autonomously produce adaptive, actionable business commentaries triggered by user-driven OLAP exploration states. Experimental results demonstrate that applying linear semantic scaling optimization (α) substantially mitigates statistical semantic blindness and protects the framework from structural schema misclassification. On the E-Commerce dataset, the proposed K-Medoids+Semantic configuration demonstrated topological superiority, achieving a peak Silhouette Score of 0.611 and a compressed Davies-Bouldin Index of 0.514. Meanwhile, on the high-dimensional Superstore dataset, the pipeline maintained high functional flexibility, stabilizing overall classification accuracy up to 94.74%. Furthermore, the downstream Hybrid NLG engine (Template+Generative) demonstrated high factual integrity and linguistic flexibility, achieving a ROUGE-1 score of 0.85, a ROUGE-2 score of 0.82, and a BLEU score of 0.15. This research provides ABI frameworks that enable accelerated executive decision-making through seamless, data-to-insight automation.
Multi Domain Feature Fusion and Boosting Based Learning for Robust Gallbladder Ultrasound Image Classification Gede Angga Pradipta; Pharan Chawaphan; Sutikno Sutikno; Putu Desiana Desiana Ayu; Dandy Pramana Hostiadi; Made Liandana
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1268

Abstract

Accurate identification of gallbladder conditions is critically important in the medical sector, as early detection of diseases such as gallstones, cholecystitis, carcinoma, and polyps can substantially improve treatment outcomes, reduce complications, and guide timely surgical or therapeutic interventions. Existing literature on gallbladder disease classification still presents notable gaps. Most prior works rely solely on single-domain feature extraction either deep CNN-based spatial descriptors or handcrafted statistical/texture features without exploiting the complementary strengths of multi-domain feature fusion. This study addresses these gaps by proposing a hybrid framework that combines advanced preprocessing, multi-domain feature fusion, feature selection, and ensemble classification. The preprocessing pipeline applies Non-Local Means (NLM) denoising, Contrast Limited Adaptive Histogram Equalization (CLAHE), frequency-domain low-pass filtering, and Gabor filtering to enhance image quality and highlight diagnostically relevant structures. Features are extracted by fusing deep spatial descriptors from a pretrained Inception V3 network with handcrafted statistical and texture-based features, including Gray Level Dependency Matrix (GLDM) measures. Dimensionality reduction is performed using ANOVA k-best selection to retain the most discriminative attributes. The refined features are classified using LightGBM, XGBoost, Histogram Gradient Boosting, and AdaBoost, enabling a comprehensive performance comparison. Experiments conducted on the balanced UIdataGB dataset (10,692 annotated images across nine diagnostic categories) demonstrate that LightGBM, XGBoost, and Histogram Gradient Boosting achieve near-perfect performance, with accuracies exceeding 98.6% and AUC values of 0.98 across all classes, while AdaBoost shows markedly lower discriminative capability.The results suggest that gradient boosting approaches are a promising option for multi-class gallbladder disease detection, particularly when combined with multi-domain feature fusion and appropriate preprocessing and feature selection techniques.
Analysis of Machine Learning Models Based on Predictive Performance, Energy Consumption, and Carbon Emissions Willy Permana Putra; Eko Marpanaji; Septafiansyah Dwi Putra
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1482

Abstract

This study aims to evaluate intrusion detection models by jointly considering predictive performance and computational sustainability. The main problem addressed is that many intrusion detection studies emphasize classification accuracy while providing limited evidence about execution time, energy use, and carbon dioxide-equivalent emissions, even though these factors affect repeated training and practical deployment. The contribution of this work is a comparative assessment of four supervised learning models, Random Forest, Histogram-based Gradient Boosting, Support Vector Machine, and Extreme Gradient Boosting, under a unified experimental workflow. The methodology uses the Wednesday working-hours subset of the Canadian Institute for Cybersecurity intrusion detection dataset released in 2017, which contains benign traffic and several denial-of-service attack classes. The procedure includes dataset selection, data cleaning, preprocessing, stratified training and testing, model fitting, predictive evaluation, sustainability measurement, and comparative interpretation. The evaluation is supported by one workflow figure, tables describing predictive results and sustainability measurements, and comparative visualizations of classification and resource-efficiency outcomes. The results show that Extreme Gradient Boosting achieved the strongest overall classification performance, with an accuracy of 0.9994, macro recall of 0.9966, macro F1-score of 0.9956, macro precision of 0.9946, and one-versus-rest area under the curve of 0.9999, while requiring 0.002559 kilowatt-hours of energy and 11.83 seconds of execution time. Random Forest produced highly comparable predictive results with similarly low resource consumption. Histogram-based Gradient Boosting was the most efficient model in terms of time, energy use, and emissions, but its macro-level performance was substantially lower. Support Vector Machine produced acceptable predictive results but required substantially higher computational resources. These findings imply that sustainable intrusion detection should select models through a balanced evaluation of detection capability and computational cost rather than accuracy alone.
An IoT-Driven Hybrid Stacking Ensemble with Deep Meta-Learning for Vending Machine Sales Forecasting Yulisman Yulisman; Zupri Henra Hartomi; Rian Ordila; Uci Rahmalisa; Arie Linarta; Yuda Irawan
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1395

Abstract

Accurate sales prediction is essential for optimizing inventory management and supporting dynamic pricing strategies in the retail industry, particularly for vending machines (VMs) integrated with IoT technologies. The availability of real-time transactional and environmental data from IoT sensors provides opportunities to improve forecasting accuracy by capturing complex temporal patterns and external influences on consumer behavior. However, traditional time series models and single machine learning approaches often struggle to model nonlinear relationships and long-term dependencies in such data. This study proposes a hybrid stacking ensemble model that integrates machine learning and deep learning techniques to enhance the prediction of daily sales volume per Stock Keeping Unit (SKU) in IoT-enabled vending machines. The proposed framework employs Random Forest Regressor (RF), Support Vector Regression (SVR), and XGBoost Regressor (XGB) as Level-0 base learners. Their predictions, along with corresponding residuals, are utilized as meta-features for a Long Short-Term Memory (LSTM)-based meta-learner, enabling effective modeling of both nonlinear and temporal characteristics. The model incorporates diverse features derived from IoT data, including lagged sales, rolling statistics, temporal attributes (day of week and weekend indicators), and environmental variables such as temperature and humidity collected from IoT sensors. Hyperparameter optimization of the LSTM meta-model is performed using Optuna to improve model stability and generalization. The proposed approach is evaluated using 10-Fold Time Series Cross-Validation to preserve temporal data structure. Experimental results show that the proposed model achieves an R² of 0.9967 and an RMSE of 0.0899, outperforming the best individual base model, XGBoost (R² = 0.9946, RMSE = 0.1121). Although the improvement is marginal, it consistently demonstrates the advantage of combining machine learning and deep learning through a stacking ensemble strategy. These findings indicate that integrating meta-features, residual learning, and IoT-based feature engineering can improve predictive performance and support adaptive decision-making in real-time vending machine operations.
Explainable EEG-Based Framework for Student Attention Classification with LLM Instructional Recommendations Samaher Dawood; Wafaa Alsaggaf; Miada Almasre
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1379

Abstract

Student attention is a central factor in learning effectiveness, yet most existing electroencephalography-based attention monitoring systems focus exclusively on classification accuracy without addressing interpretability or practical instructional utility. This paper presents a unified framework that integrates attention classification, explainable artificial intelligence, and large language model-driven instructional recommendations to bridge the gap between physiological monitoring and actionable teaching decisions. The framework derives binary attention labels from electroencephalography spectral features using the Theta-Beta Ratio combined with age-adjusted Monastra thresholds, applied here as population-level heuristic guidelines rather than individual diagnostic criteria. Five machine learning classifiers were trained and evaluated under stratified five-fold cross-validation on a publicly available dataset of 983 samples with precomputed electroencephalography spectral power features across five frequency bands. Gradient Boosting achieved the strongest performance, reaching 98.3% accuracy, 97.8% balanced accuracy, and a macro F1-score of 97.6%, with particularly reliable detection of the low-attention class. These figures represent an upper bound, given the shared spectral origin of the labeling mechanism and model inputs, and the absence of subject-level identifiers in the dataset. To ensure transparency, Shapley Additive Explanations analysis was applied at both global and local levels, identifying Beta and Theta band activity as the dominant predictors of attention state. A large language model component then maps each classified attention state to a concrete, knowledge-grounded instructional strategy drawn from peer-reviewed educational literature, providing teachers with an actionable pedagogical recommendation rather than a classification label alone. The proposed framework advances the field by combining classification accuracy, model interpretability, and practical instructional guidance within a single deployable system, providing a foundation for intelligent, transparent, and educationally responsive attention monitoring in real classroom settings.
Lion-Optimized ANFIS Control for a Hybrid PV–Wind–Battery System with a Reduced-Switch 31-Level Inverter for Induction Motor Drives R K Padmashini; D Lakshmi; J N Rajesh Kumar; C N Ravi; P Jeyarani
Journal of Applied Data Sciences Vol 7, No 3: September 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i3.1466

Abstract

This study proposes and experimentally validates a hybrid photovoltaic–wind–battery energy conversion system for driving a three-phase induction motor under variable renewable-source conditions. The system integrates a high-gain SEPIC converter, a PWM rectifier, a bidirectional battery converter, a reduced-switch 31-level cascaded H-bridge inverter, and a Lion Optimization Algorithm–tuned adaptive neuro-fuzzy inference system controller. The proposed controller performs maximum power point tracking and regulates the common DC-link voltage by adjusting the converter and rectifier switching commands in response to changes in solar irradiance, photovoltaic temperature, and wind-side voltage. The system was evaluated using MATLAB/Simulink and an FPGA-based experimental prototype. The photovoltaic array produced approximately 134 V and 75 A under the tested operating conditions, while the SEPIC converter and PWM rectifier maintained the DC-link voltage near 400 V. The battery state of charge remained around 60%, indicating balanced charging and discharging operation. The 31-level inverter generated a stable multilevel output with an RMS voltage of approximately 499 V and real power of approximately 4990 W. The simulated total harmonic distortion was 1.39%, while the experimental prototype achieved 2.85%. The experimental results were consistent with the simulation outcomes and confirmed the effectiveness of the proposed controller in improving transient response, reducing steady-state fluctuations, and maintaining stable power delivery to the induction motor. The proposed configuration provides a feasible solution for renewable-powered motor-drive and marine propulsion applications requiring high voltage gain, adaptive control, energy-storage coordination, and low harmonic distortion.