cover
Contact Name
Edi Sutoyo
Contact Email
journalijadis@gmail.com
Phone
+62895410194922
Journal Mail Official
info@ijadis.org
Editorial Address
Indonesian Scientific Journal (Jurnal Ilmiah Indonesia) Jl. Pasar Atas No 3, Kompleks Setramas Kota Cimahi, Bandung
Location
Unknown,
Unknown
INDONESIA
International Journal of Advances in Data and Information Systems
ISSN : -     EISSN : 27213056     DOI : https://doi.org/10.25008/ijadis
International Journal of Advances in Data and Information Systems (IJADIS) (e-ISSN: 2721-3056) is a peer-reviewed journal in the field of data science and information system that is published twice a year; scheduled in April and October. The journal is published for those who wish to share information about their research and innovations and for those who want to know the latest results in the field of Data Science and Information System. The Journal is published by the Indonesian Scientific Journal. Accepted paper will be available online (free access), and there will be no publication fee. The author will get their own personal copy of the paperwork. IJADIS welcomes all topics that are relevant to data science, and information system. The listed topics of interest are as follows: Data clustering and classifications Statistical model in data science Artificial intelligence and machine learning in data science Data visualization Data mining Data intelligence Business intelligence and data warehousing Cloud computing for Big Data Data processing and analytics in IoT Tools and applications in data science Vision and future directions of data science Computational Linguistics Text Classification Language resources Information retrieval Information extraction Information security Machine translation Sentiment analysis Semantics Summarization Speech processing Mathematical linguistics NLP applications Information Science Cryptography and steganography Digital Forensic Social media and social network Crowdsourcing Computational intelligence Collective intelligence Graph theory and computation Network science Modeling and simulation Parallel and distributed computing High-performance computing Information architecture
Articles 195 Documents
AugLog-LightGBM: A Log-Based Feature AugmentationFramework for Class Imbalance in Credit RiskClassification Hana Azizah; Eni Sumarminingsih; Adji Achmad Rinaldo Fernandes
International Journal of Advances in Data and Information Systems Vol. 7 No. 2 (2026): August 2026 - International Journal of Advances in Data and Information Systems
Publisher : Indonesian Scientific Journal

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59395/ijadis.v7i2.1623

Abstract

Non-performing loan (NPL) detection is inherently a class-imbalance problem because defaulting borrowers represent a persistent minority. Standard gradient boosting often favors the majority class. This paper proposes AugLog-LightGBM, an extension of LightGBM that improves initialization through Log-Based Feature Augmentation (LBFA). Instead of using an uninformative constant, boosting starts from an informed prior combining a logistic-regression logit score and a kernel-density-estimation log-density ratio (LDR), which capture complementary global and local information. These representations are incorporated as augmented features and as the init_score, reformulating boosting as residual correction over an informed Bayesian prior. The proposed framework is evaluated on a dataset of 2,700 home-mortgage borrowers collected from partner banks in Malang, Indonesia (NPL rate = 16.11%), using repeated stratified cross-validation and comparison against four imbalance-aware baselines. AugLog-LightGBM achieves the highest ROC-AUC (0.815 ± 0.018), PR-AUC (0.679), and F1-score (0.631). DeLong tests show statistically significant ROC-AUC improvements over class-weighted Logistic Regression and Random Forest, while gains over XGBoost and SMOTE + LightGBM are positive but not statistically significant. Robustness analyses and SHAP interpretation further support the consistency and practical applicability of the proposed framework.
An Optimization Framework for Retrieval Augmented Generation in Indonesian Educational Question Answering I Ketut Resika Arthana; Nyoman Gunantara; Made Sudarma; I Made Sukarsa
International Journal of Advances in Data and Information Systems Vol. 7 No. 2 (2026): August 2026 - International Journal of Advances in Data and Information Systems
Publisher : Indonesian Scientific Journal

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59395/ijadis.v7i2.1633

Abstract

The Retrieval-Augmented Generation (RAG) approach has been widely adopted to produce responses that are more closely aligned with a predefined knowledge context. However, many RAG implementations have not undergone systematic optimization of their retrieval and generation components, resulting in outputs that do not always correspond accurately to the reference context. This study developed a RAG optimization framework for Indonesian-language educational question answering using a Human-Computer Interaction learning corpus as a case study. In the retrieval stage, the study evaluated chunking strategies, multilingual embedding models, and the use of a reranker. Evaluation was conducted using Mean Reciprocal Rank (MRR), Normalized Discounted Cumulative Gain (nDCG@K), and Hit@K. In the generation stage, candidate Large Language Models (LLMs) were assessed using RAGAS metrics, namely Context Precision (CP), Context Recall (CR), Faithfulness (F), Answer Relevancy (AR), and Answer Correctness (AC). Experimental results showed that the GTE configuration with fixed-size chunking and a reranker yielded the best retrieval performance, achieving an MRR of 0.9082, nDCG@5 of 0.9215, and Hit@5 of 0.9655. In the generation stage, Gemma 4 E4B exhibited the most balanced answer quality. The resulting framework provides a procedure for selecting retrieval and generation settings for a given corpus.
An Optimized Heterogeneous Stacking Ensemble with Hyperparameter Optimization for Multi Class Hypertension Risk Classification Novi Yona Sidratul Munti; Yuda Irawan; Muhammad Habib Yuhandri
International Journal of Advances in Data and Information Systems Vol. 7 No. 2 (2026): August 2026 - International Journal of Advances in Data and Information Systems
Publisher : Indonesian Scientific Journal

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59395/ijadis.v7i2.1637

Abstract

Hypertension is a major global health concern and a leading risk factor for cardiovascular disease, stroke, kidney failure, and premature mortality. Accurate and interpretable prediction of hypertension risk is essential for supporting early intervention and preventive healthcare. This study proposes an optimized and explainable stacking ensemble framework for multi class hypertension classification by integrating LASSO feature selection, SMOTEENN data balancing, heterogeneous ensemble learning, CatBoost meta learning, GridSearchCV optimization, and SHAP based explainability. The proposed architecture combines five base learners, namely XGBoost, LightGBM, Random Forest, Extra Trees, and Support Vector Machine, whose probability outputs are transformed into meta features and processed by an optimized CatBoost meta learner. Experiments conducted on a dataset containing 3,000 hypertension related records demonstrated superior classification performance, achieving 97.50% accuracy, 97.46% precision, 97.50% recall, and 97.48% F1 score. 10 fold cross validation further confirmed the robustness of the framework with a mean accuracy of 97.50% ± 0.0027. ROC analysis produced AUC values above 0.97 for all classes, indicating excellent discriminative capability. The results demonstrate that the proposed framework provides both high predictive accuracy and strong interpretability, making it a promising solution for intelligent hypertension risk assessment and clinical decision support.
Towards Explainable Multidimensional Digital Divide Classification in Online Learning Using Artificial Neural Networks and SHAP Imalatul Hidayah; Ririen Kusumawati; Mochamad Imamudin
International Journal of Advances in Data and Information Systems Vol. 7 No. 2 (2026): August 2026 - International Journal of Advances in Data and Information Systems
Publisher : Indonesian Scientific Journal

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59395/ijadis.v7i2.1666

Abstract

The transformation of online learning increases access to education, but also increases the risk of digital divide influenced by various technological and educational factors. This study aims to develop a framework Towards Explainable Multidimensional Digital Divide Classification in Online Learning Using Artificial Neural Networks and SHAP to classify cluster-derived digital divide categories transparently. Data were obtained from 324 students through a questionnaire covering Primary Device, Internet Stability, Equity Score, and Accessibility Score. The target Digital Divide categories were analytically derived using K-Means clustering based on respondents' multidimensional characteristics and subsequently used as labels for supervised classification. The data were processed through one-hot encoding, Min-Max normalization, and splitting the training and test data before training the model. The performance of ANN was compared with Logistic Regression, Decision Tree, Random Forest, and Support Vector Machine (SVM). The results showed that ANN achieved 90.77% accuracy, with competitive performance against the comparison models. Furthermore, SHapley Additive exPlanations (SHAP) was used to interpret the model's decisions at both global and local levels. The analysis shows that Primary Device is the most influential factor, followed by Accessibility Score, Equity Score, and Internet Stability. The main contribution of this research is the presentation of an Explainable AI framework that combines the predictive capabilities of ANN with the interpretability of SHAP to support a more transparent multidimensional digital divide classification and can serve as a basis for developing data-driven digital inclusion policies.
Quantifying Cross-Lingual Sentiment Prediction Difference in Bilingual Pop-Culture Reviews via Dual Monolingual Transformers and Jensen-Shannon Divergence Azfani Naurotul Jannah; Triando Hamonangan Saragih; Irwan Budiman; Muliadi Muliadi; Radityo Adi Nugroho
International Journal of Advances in Data and Information Systems Vol. 7 No. 2 (2026): August 2026 - International Journal of Advances in Data and Information Systems
Publisher : Indonesian Scientific Journal

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59395/ijadis.v7i2.1678

Abstract

Machine translation has been widely used to support cross-lingual sentiment analysis when labeled data in the target language are limited. However, differences in linguistic representation between the source text and its translation may affect the probabilistic outputs of sentiment classification models. This study investigated differences in sentiment predictions between original Mandarin reviews and their Indonesian translations using two independently fine-tuned monolingual transformer models. Chinese-RoBERTa-wwm-ext was applied to the original Mandarin reviews, whereas IndoBERT was used to classify the Indonesian translations. Jensen–Shannon Divergence was employed to compare the sentiment probability distributions generated by the two models. The results showed that Chinese-RoBERTa achieved an accuracy of 74%, whereas IndoBERT achieved 69% on the pseudo-labeled evaluation dataset. Furthermore, 73.81% of the review pairs retained consistent sentiment predictions, while 26.19% exhibited prediction shifts, with Polarity Amplification being the most frequently observed category and most transitions occurring between adjacent sentiment classes. The probability-distribution analysis also revealed substantial differences in prediction confidence for some review pairs, even when the predicted sentiment labels remained identical. These findings demonstrated that comparing probability distributions provided complementary information beyond label-based evaluation for analyzing prediction differences between independently trained monolingual sentiment models on bilingual review pairs. 

Filter by Year

2020 2026


Filter By Issues
All Issue Vol. 7 No. 2 (2026): August 2026 - International Journal of Advances in Data and Information Systems Vol. 7 No. 1 (2026): April 2026 - International Journal of Advances in Data and Information Systems Vol. 6 No. 3 (2025): December 2025 - International Journal of Advances in Data and Information Syste Vol. 6 No. 2 (2025): August 2025 - International Journal of Advances in Data and Information Systems Vol. 6 No. 1 (2025): April 2025 - International Journal of Advances in Data and Information Systems Vol. 5 No. 2 (2024): October 2024 - International Journal of Advances in Data and Information System Vol. 5 No. 1 (2024): April 2024 - International Journal of Advances in Data and Information Systems Vol. 4 No. 2 (2023): October 2023 - International Journal of Advances in Data and Information System Vol. 4 No. 1 (2023): April 2023 - International Journal of Advances in Data and Information Systems Vol. 3 No. 2 (2022): October 2022 - International Journal of Advances in Data and Information System Vol. 3 No. 1 (2022): April 2022 - International Journal of Advances in Data and Information Systems Vol. 2 No. 2 (2021): October 2021 - International Journal of Advances in Data and Information System Vol. 2 No. 1 (2021): April 2021 - International Journal of Advances in Data and Information Systems Vol. 1 No. 2 (2020): October 2020 - International Journal of Advances in Data and Information System Vol. 1 No. 1 (2020): April 2020 - International Journal of Advances in Data and Information Systems More Issue