cover
Contact Name
Jumanto
Contact Email
jumanto@mail.unnes.ac.id
Phone
+628164243462
Journal Mail Official
sji@mail.unnes.ac.id
Editorial Address
Ruang 114 Gedung D2 Lamtai 1, Jurusan Ilmu Komputer Universitas Negeri Semarang, Indonesia
Location
Kota semarang,
Jawa tengah
INDONESIA
Scientific Journal of Informatics
ISSN : 24077658     EISSN : 24600040     DOI : https://doi.org/10.15294/sji.vxxix.xxxx
Scientific Journal of Informatics (p-ISSN 2407-7658 | e-ISSN 2460-0040) published by the Department of Computer Science, Universitas Negeri Semarang, a scientific journal of Information Systems and Information Technology which includes scholarly writings on pure research and applied research in the field of information systems and information technology as well as a review-general review of the development of the theory, methods, and related applied sciences. The SJI publishes 4 issues in a calendar year (February, May, August, November).
Articles 190 Documents
A Comparative Evaluation of Model-Based and SHAP-Based Feature Analysis for Robust Health Data Classification Indra Waspada; Satriawan Rasyid Purnama; Alwey Hakim; Alfonso Clement Sutantio
Scientific Journal of Informatics Vol. 13 No. 2: May 2026
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v13i2.42406

Abstract

Purpose: Feature selection is one critical element of healthcare data classification, which directly affects predictive performance, model robustness, and interpretability. Nevertheless, traditional model-based feature importance methods are unstable in robustness and provide random or misleading results on high-dimensional and heterogeneous healthcare data. The purpose of this paper was to assess model-based and SHapley Additive exPlanations (SHAP) - based feature analysis on multimodal healthcare data classification in a comparative manner. Methods: This study employed a quantitative comparative experimental design using real Electronic Medical Records (EMR) data captured from primary care clinics. The dataset comprises 2,158 patient records, with numerical and textual features in a multimodal feature space. The analytical pipeline included data acquisition, preprocessing, multimodal feature integration, and model development using six supervised learning algorithms from ensemble-based and margin-based categories. The model-based feature importance was used with the SHAP-based feature importance. We examined the robustness of this method across different scenarios through systematic feature ablation using baseline, strong, and weak features. Result: Experimental results demonstrate that model-based feature importance exhibits unpredictable behavior when features are removed and is sometimes counterintuitive. On the other hand, feature importance based on SHAP is consistent and linearly proportional across the models under consideration. The highest macro F1-score was achieved by Extra Trees with SHAP-based strong features (0.824), exceeding the baseline (0.811). In robustness testing, SHAP-based weak-feature removal reduced the LightGBM F1-score from 0.764 to 0.575, suggesting a clearer distinction between informative and non-informative features. Overall, SHAP-based feature selection provided a more reliable and interpretable framework for multimodal healthcare classification. Novelty: Our experiment results demonstrate empirically that the SHAP-based feature importance is more robust and reliable than conventional model-based approaches when it comes to extracting features from multimodal medical records. This work shows that SHAP is more than a post hoc explanation by presenting it as an interpretable feature selection criterion guiding feature relevance analysis in healthcare machine learning.
A Multi-Model Forecasting Framework for New Student Admissions: SMA, ARIMA, and Random Forest Approaches Falentino Sembiring; Rieska Rahayu Ayuningsih; Adhitia Erfina; Risky
Scientific Journal of Informatics Vol. 13 No. 2: May 2026
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v13i2.42746

Abstract

Purpose: This study evaluates and compares the forecasting performance of traditional statistical methods (Simple Moving Average and ARIMA) and a machine learning approach (Random Forest) in predicting student applicant numbers across multiple study programs at XYZ University for 2024–2026. The objective is to identify the most accurate model to support data-driven strategic planning and enrollment management. Methods: A quantitative comparative forecasting design was applied using historical admission data from 2018–2023. Three models SMA, ARIMA, and Random Forest were implemented and assessed using Mean Absolute Percentage Error (MAPE), Mean Absolute Deviation (MAD), and Mean Squared Error (MSE). Model robustness was evaluated across several study programs with different growth patterns. Result: The findings reveal an overall upward enrollment trend, particularly in Management and Informatics Engineering. Random Forest achieved the highest predictive accuracy, with MAPE values ranging from 6.17% to 17.95%, outperforming ARIMA (17.95%–33.13%) and SMA (14.5%–28.25%). The results indicate that Random Forest more effectively captures complex and non-linear enrollment dynamics. Novelty: This study provides a systematic multi-program comparison between classical time-series models and a machine learning approach within a single institutional context. It demonstrates the superior robustness of Random Forest and supports integrating machine learning–based forecasting into higher education information systems for improved strategic decision-making.
Analysis of the Dominance of Natural Factors and Human Activities on Recurrent Floods on Sumatra Island Using SHAP-Based Random Forest Sudin saepudin; April Lia Hananto; Ibnu Mu'ti Lidinillah Abidin; Muhamad Rivqi Nurridwan; Muhamad Zidane Mustofa; Hidayat Abdul Wahab; Difa Nafis Conoras Algifari
Scientific Journal of Informatics Vol. 13 No. 2: May 2026
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v13i2.44661

Abstract

Purpose: The objectives of this work are to conduct a comprehensive analysis of the primary influences of both natural and anthropogenic causes on flooding in Kalimantan, utilising a transparent machine learning approach, and to provide policymakers with valuable insights for formulating strategies to mitigate flooding. Methods: The research employs the Random Forest model, integrating SHAP (Shapley Additive Explanations), to examine non-linear multivariate correlations and ascertain the significance of each variable. The data consists of both natural elements (precipitation, elevation, slope gradient, and proximity to the river) and anthropogenic activity (land fire hotspots, planting, mining, and building new infrastructure). The data also includes public sentiment data in text form. We got these data points per year from 2021 to 2025. We used R-squared and SHAP scores to figure out how accurate the model was. Result: The model has a high R² score of 0.81, which shows that it can make accurate predictions. Using SHAP (SHapley Additive exPlanations), we can see that natural factors, such as how far away the river is and how much it rains, are what make the area vulnerable in the first place. Human actions, on the other hand, are what cause the floods to happen again and again. The indicator for hotspots of land burning is the most important factor. Plantation and mining operations make up more than 90% of the overall contributions, followed by other predictors, in the case of flood-induced deforestation. Novelty: The current study proposes a framework for explicable AI that integrates the use of random forests and SHAP to assess the significance of various flood risk indicators via quantitative analysis of geographical and public opinion data.
Hybrid CNN–LSTM–Transformer for Electrical Energy Consumption Forecasting: A Multi-Scenario Evaluation Alief Flostyo ZUlfiqor Roshif; Oky Dwi Nurhayati; Dinar Mutiara Kusumo Nugraeni
Scientific Journal of Informatics Vol. 13 No. 2: May 2026
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v13i2.44973

Abstract

Purpose: The purpose of this study is to develop hybrid deep learning model combining Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), and Transformer mechanisms for short-term electricity consumption forecasting. The motivation arises from the limitations of existing studies, which are mostly based on single or dual hybrid models and have not fully captured complex temporal dependencies in electricity demand. To address this gap, the proposed CNN–LSTM–Transformer architecture is introduced to jointly model local patterns, sequential dependencies, and global temporal relationships, which have not been simultaneously explored in prior electricity forecasting studies. In addition, the study investigates the impact of multivariate inputs and feature engineering on forecasting performance. Methods: The proposed model is evaluated using the Tétouan City Electricity Consumption dataset, which includes load data from three zones and meteorological variables. Data preprocessing involves cleaning, Min–Max normalization, and sequence windowing for supervised learning. Model performance is assessed using MAE, RMSE, and MAPE. To ensure comprehensive evaluation, six experimental scenarios are designed, including univariate and multivariate settings, per-zone and combined-zone configurations, as well as feature selection-based scenarios, to analyze accuracy, robustness, and generalization capability. Result: The experimental results demonstrate that the proposed model achieves consistent forecasting performance across all scenarios. The best performance is obtained in the multivariate combined scenario (S4), with MAE values of 232–257 kWh, RMSE values of 385–433 kWh, and MAPE values ranging from 0.82% to 1.43% across all zones. The feature selection scenarios (S5 and S6) also show competitive performance, with MAE ranging from 234 to 324 kWh, RMSE from 393 to 590 kWh, and MAPE between 0.82% and 1.83%, indicating that engineered features can maintain prediction accuracy while reducing input complexity. In contrast, the univariate per-zone (S1) and multivariate per-zone (S3) scenarios produce higher errors, while the combined univariate scenario (S2) yields moderate improvements in certain zones. Overall, these findings confirm that integrating multivariate features with cross-zone data leads to the most accurate forecasting performance. Novelty: The novelty of this study lies in the proposed integration of CNN, LSTM, and Transformer architectures into a unified hybrid framework for electricity consumption forecasting, which has not been widely explored in the energy demand domain. The model effectively combines local feature extraction, sequential dependency learning, and global attention mechanisms, and demonstrates strong capability in handling both univariate and multivariate electricity consumption data, including additional meteorological variables. Practical Implications: This study shows that hybrid CNN-LSTM-Transformer can applied in electrical domain with multiple meteorological variables.
An Evaluation of AES-256 Based PDF Document Security Using Cryptographic Metrics and Blockchain Based Hash Storage Ellen Chandra; Nur Rochmah Dyah Puji Astuti
Scientific Journal of Informatics Vol. 13 No. 2: May 2026
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v13i2.45421

Abstract

Purpose: This research aims to design and implement a digital document security by integrating the Advanced Encryption Standard (AES) 256-bit algorithm with blockchain technology to protect documents from manipulation, forgery, and information leakage. Methods: The study involved implementation using Python. AES-256 was applied to encrypt digital documents, while the hash value of the ciphertext was stored on the blockchain to ensure integrity and authenticity. System performance was evaluated through avalanche effect testing, entropy analysis, and process time measurement. Result: The average efficiency is 49.989% for small documents, 49.988% for documents 50-200 KB and 49.998% for documents larger than 200 KB, using good bits. An average entropy value of 7.99 indicates a high performance. The average sending time is 1.62 ms, and the average blockchain sending time is 1061.862 ms making it possible to send in sending digital documents. Novelty: This study integrates AES-256 encryption with blockchain-based hash storage to simultaneously ensure confidentiality, integrity, and authenticity of digital documents with high efficiency.
Performance Evaluation of Distributed Lag, Autoencoder, and LSTM Autoencoder Methods in Detecting Anomalies in Simulated Data Based on Air Quality Index in Jakarta Yenni Angraini; Adelia Putri Pangestika; I Made Sumertajaya
Scientific Journal of Informatics Vol. 13 No. 2: May 2026
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v13i2.45581

Abstract

Purpose: This study evaluates the performance of distributed lag, autoencoder, and LSTM autoencoder methods in detecting point anomalies in simulated data generated from Jakarta's Air Quality Index (AQI). The evaluation is conducted across several simulation scenarios that represent factors that may influence anomaly detection performance. Methods: Simulation scenarios were constructed by varying two anomaly characteristics: anomaly percentage (0.3%, 0.5%, and 1.0%) and anomaly depth (4.4σ, 4.7σ, and 5.0σ), yielding 90 datasets generated via repeated experiments. Anomaly detection was performed using a forecasting-based approach with a 4σ threshold on prediction errors. Model performance was evaluated using mean absolute percentage error (MAPE) for forecasting accuracy and balanced accuracy for anomaly detection. Result: Increasing anomaly percentage significantly degrades both forecasting and anomaly detection performance across all methods. In contrast, anomaly depth has no significant effect on forecasting accuracy but strongly influences detection performance. Among the evaluated methods, the distributed lag model consistently shows the most robust anomaly detection performance across scenarios, outperforming the autoencoder and LSTM autoencoder, particularly at higher anomaly percentages and depths. Novelty: This study introduces a more rigorous and structured evaluation framework for anomaly detection by integrating three key contributions. First, it provides a unified comparison of distributed lag, autoencoder, and LSTM autoencoder methods within a single experimental setting, a limitation in prior studies. Second, it employs a controlled simulation design that systematically varies the anomaly percentage and depth while preserving key characteristics of empirical AQI data, enabling a more objective assessment of model performance across diverse anomaly conditions. Third, it uses ANOVA and interaction analysis to formally examine the effects of anomaly characteristics on both forecasting and detection performance, moving beyond purely descriptive comparisons commonly used in previous research.
Factors Influencing the Adoption of Digital Waste Banks: A Systematic Literature Review Nur Isni Nirwan; Dana Indra Sensuse
Scientific Journal of Informatics Vol. 13 No. 2: May 2026
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v13i2.46754

Abstract

Purpose: This study aims to identify and analyze the factors influencing the adoption of digital waste banks, considering the gap between their potential to improve waste management and their limited implementation in many regions. Methods/Study design/approach: This research use a Systematic Literature Review (SLR) approach by analyzing 22 selected studies published between 2020 and 2025 from reputable databases. The study follows structured stages, including planning, literature selection based on inclusion and exclusion criteria, quality assessment, data extraction, and data synthesis using the PICOC framework. Result/Findings: The results show that there are 26 factors influencing the adoption of digital waste banks, which are grouped into four main categories: technology, organization, environment, and people. Technological factors include ease of use, transparency, system quality, and security. Organizational factors involve infrastructure, innovation, and human resources. Environmental factors include regulations, social influence, environmental awareness, and economic incentives. Meanwhile, people-related factors include age, education, income, and training. These factors are interconnected and play an important role in determining adoption success. Novelty/Originality/Value: This study provides a comprehensive classification of factors influencing digital waste bank adoption based on recent literature. The findings offer practical insights for policymakers and practitioners in designing strategies to increase community participation and improve sustainable waste management through digital innovation.
Towards Achieving Smart Healthcare Through the Adoption of Information Technology Mahar Nooruldeen Saaed; Osamah Fadhil Taher; Adnan Salih Mahmood; Suhair Abd Dawwod; Ramadan Mahmood Ramo; Muhammad Mustafa Hussien
Scientific Journal of Informatics Vol. 13 No. 2: May 2026
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v13i2.46790

Abstract

Purpose: The study focuses on the impact of Virtual Reality as a therapeutic technology in the context of Smart Healthcare on the rehabilitation of hand amputees and the ability of out-patients to gain computer access in a cost-effective manner. Methods: A qualitative, descriptive bibliometric survey of the literature was undertaken, utilizing peer-reviewed literature relevant to the integration of virtual medical care (V-Med) with the IoT, EHR, and CDSS. The literature examined was categorized into two analytical domains - physical and clinical opportunities and technical challenges. Results: The evidence suggests that cloud focused Virtual Reality (VR) helps people with Phantom limb pain and reduces the time required for amputees to achieve the adaptation to a prosthesis. The disadvantages of VR technology will only be realized when VR technology, EHR and the Internet of Everything (IoE) systems coalesce. This is where the central issue of technical feasibility lies. Novelty: The presentation here outlines a shift in attitude towards Virtual Reality. Rather than isolating it from a continuum encompassing all virtualization in the cloud with amputee rehabilitation modalities, we embrace it. We propose a framework that integrates cloud-based IoT wearables, EHRs, and CDSS into a Smart Healthcare ecosystem, thus outlining a feasible, affordable scalable virtual rehabilitation path.
Performance Analysis of VNC-Based Remote Monitoring for Flight Information Display Systems Using ANOVA and Regression Models Ceorido Ghalib Wibowo; Much Aziz Muslim
Scientific Journal of Informatics Vol. 13 No. 2: May 2026
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v13i2.47378

Abstract

Purpose: This study evaluates the performance of a Virtual Network Computing (VNC)-based remote monitoring system within a Flight Information Display System (FIDS) environment. The research aims to identify infrastructure factors affecting monitoring performance and to develop a data-driven framework for evaluating monitoring reliability in distributed airport systems. Methods: Correlation analysis, Analysis of Variance (ANOVA), and multiple linear regression were applied to analyze the relationship between system resource utilization, network characteristics, and monitoring performance. The dataset consisted of 1,000 observations collected under various simulated monitoring conditions representing variations in latency, throughput, CPU utilization, and memory usage. Residual analysis and model evaluation were also performed to validate the statistical model. Result: The results showed that most infrastructure variables had very weak correlations (−0.02 to 0.05), indicating minimal multicollinearity. ANOVA testing revealed no statistically significant latency differences across low, medium, and high CPU load categories (F = 0.1625, p = 0.8500), with average latency remaining stable at approximately 52.82 ms. Regression evaluation demonstrated stable residual distribution and acceptable model consistency. The findings indicate that monitoring performance is influenced more by network conditions, particularly latency and throughput variability, than by computational load. Novelty: This study proposes an integrated analytical framework combining correlation analysis, ANOVA, and regression modeling to evaluate VNC-based monitoring performance in distributed systems. The framework provides a practical and reproducible approach for monitoring performance evaluation and infrastructure optimization in airport monitoring environments.
A Cross-Platform Assessment of Personal Data Protection Compliance among Electronic System Providers Under Indonesia’s PDP Law Rindy; Wahyu Catur Wibowo
Scientific Journal of Informatics Vol. 13 No. 2: May 2026
Publisher : Universitas Negeri Semarang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.15294/sji.v13i2.47516

Abstract

Purpose: This study evaluates the level of Personal Data Protection (PDP) compliance among Electronic System Providers (ESPs) in Indonesia based on the Personal Data Protection Law (PDP Law) and the Government Regulation concerning the Implementation of Electronic Systems and Transactions. Methods: A descriptive-evaluative approach was conducted through observations of website and mobile application interfaces from 20 ESPs registered with the Ministry of Communication and Digital Affairs. Compliance was assessed using a binary scoring system based on six PDP indicators: consent mechanisms, privacy notices, TLS/SSL implementation, data disclosure, malicious libraries, and device data access. Descriptive statistical analysis was used to evaluate compliance levels. Instrument validity was established through content validity and expert judgment. Result: Most ESPs were classified within the moderate compliance category, covering 90% of websites and 80% of mobile applications. Governance-related indicators showed the lowest compliance levels, particularly website consent mechanisms (15%) and website privacy notices (40%) and mobile consent and privacy notice compliance (20%). In contrast, all ESPs complied with technical indicators, including TLS/SSL, malicious library, and device data access requirements. Novelty: Unlike previous studies that focused on single sectors or platforms, this study provides a cross-platform PDP compliance assessment integrating both technical and governance indicators within a single framework. The findings indicate that governance practices remain the primary challenge in PDP implementation, providing practical recommendations for regulators and ESPs in strengthening personal data protection implementation in Indonesia.