cover
Contact Name
Mesran
Contact Email
mesran.skom.mkom@gmail.com
Phone
-
Journal Mail Official
jurnal.bits@gmail.com
Editorial Address
-
Location
Kota medan,
Sumatera utara
INDONESIA
Building of Informatics, Technology and Science
ISSN : 26848910     EISSN : 26853310     DOI : -
Core Subject : Science,
Building of Informatics, Technology and Science (BITS) is an open access media in publishing scientific articles that contain the results of research in information technology and computers. Paper that enters this journal will be checked for plagiarism and peer-rewiew first to maintain its quality. This journal is managed by Forum Kerjasama Pendidikan Tinggi (FKPT) published 2 times a year in Juni and Desember. The existence of this journal is expected to develop research and make a real contribution in improving research resources in the field of information technology and computers.
Arjuna Subject : -
Articles 1,045 Documents
Comparison of XGBoost and Random Forest for Prediction of Male Fertility Status Based on Semen Analysis Parameters Kecitaan Harefa; Joko Priambodo
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.10264

Abstract

Male infertility is a major reproductive health problem that contributes to approximately half of infertility cases among couples of reproductive age. Accurate evaluation of male fertility status commonly relies on semen analysis, including parameters such as semen volume, sperm concentration, motility, morphology, and vitality. However, manual interpretation of these parameters remains time-consuming and is highly dependent on clinical expertise. This study aims to compare the performance of the XGBoost and Random Forest algorithms in predicting male fertility status based on semen analysis parameters. The study employed a secondary dataset consisting of 1,000 semen analysis records with 11 predictor variables and one target variable representing fertility status. Data preprocessing included categorical encoding, data cleaning, and an 80:20 train–test split before model development and evaluation using accuracy, precision, recall, F1-score, and ROC-AUC. Experimental results showed that XGBoost outperformed Random Forest, achieving an accuracy of 99.00%, precision of 100.00%, recall of 96.88%, F1-score of 98.41%, and ROC-AUC of 99.86%, while Random Forest achieved an accuracy of 97.50%. Feature importance analysis identified Total Motility, Vitality, and Progressive Motility as the most influential predictors of male fertility status. The main contribution of this study is the direct empirical comparison of Random Forest and XGBoost under identical experimental settings using comprehensive semen analysis parameters, providing evidence on the relative effectiveness of ensemble learning algorithms for male fertility status prediction. Although the proposed model demonstrates excellent predictive performance, it was developed using secondary data and is intended to support, rather than replace, clinical decision-making. Future studies should validate the model using larger multicenter clinical datasets to improve its generalizability and practical applicability.
Comparative Study of Agglomerative Hierarchical Clustering and K-Means for Student Academic Stress Grouping Irfan Arifin; Iwan Iskandar; Elvia Budianita; Novi Yanti; Fitri Insani
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.10265

Abstract

Academic stress is a common problem experienced by college students due to high academic demands, parental expectations, and social pressures during their college years. The high levels of academic stress experienced by students underscore the need for a data-driven approach to more accurately identify and map students’ stress levels. This research aims to compare the performance of the Agglomerative Hierarchical Clustering (AHC) and K-Means methods in clustering students’ academic stress levels and to determine which method produces the best clustering quality. Data were obtained from the distribution of the Perception of Academic Stress Scale (PAS) questionnaire, consisting of 18 statement items, with 361 valid respondents from the Informatics Engineering Program at UIN SUSKA Riau, class of 2022–2025. The selection of the best linkage method in AHC was performed using the Cophentic Correlation Coefficient (CCC), where Ward Linkage was selected with the highest CCC value of 0.8180. Comparative evaluation was conducted using the Silhouette Coefficient, Davies-Bouldin Index, and Calinski-Harabasz Index for variations in the number of clusters from K=2 to K=7. The test results showed that AHC Ward Linkage with K=2 was the best configuration with a Silhouette Coefficient of 0.4407 and a Davies-Bouldin Index of 0.8373, outperforming K-Means, which only excelled in the Calinski-Harabasz Index with a value of 419.7405 The clustering resulted in two clusters: High Stress with 244 students (67.6%) and Low Stress with 117 students (32.4%). The 2023 and 2024 cohorts had the highest proportions of high stress at 90.4% and 90.6%, respectively. This research contributes empirical evidence comparing hierarchy-based and partition-based clustering methods for academic stress data, while also demonstrating the use of the Cophenetic Correlation Coefficient as an objective basis for linkage method selection in AHC. It is hoped that the results of this study can serve as a basis for the institution in designing targeted mental health intervention programs for students.
ShelfMind: Demand Forecasting and Stock Control for Retail SMEs using Global XGBoost and Context-Injected AI Chatbot Aulia Syafitri; Aryanti Aryanti; Sholihin Sholihin
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.10280

Abstract

SME retail stores in Indonesia face serious challenges in inventory management: stockouts of popular products, waste from perishable items that expire before being sold, and restocking decisions relying on owner intuition. This research develops ShelfMind, a web-based inventory management system integrating three main components: a demand prediction model using a global XGBoost algorithm, an AI chatbot based on real-time data context injection, and an active notification system for low-stock and expiry risk alerts. Dataset from Toko Tika Baru covered 47 products in 11 categories with 8,225 daily sales records over 205 days. The XGBoost model was trained with 29 time-engineered features and achieved an MAE of 0.4107 units/day and RMSE of 0.5081 units/day on a 40-day test set. The global XGBoost method and context injection approach were selected due to their computational efficiency, avoiding the high costs associated with LLM fine-tuning, thus making it an ideal solution for SMEs with limited budget and infrastructure data. The AI chatbot, using context injection from inventory, sales, and XGBoost forecast data, achieved an average relevance and data accuracy of 4.94/5 across 20 test scenarios with an average response time of 2.94 seconds. Black-box functional testing of 31 scenarios passed completely. Usability evaluation scored 3.77/5, with notes on improving the communication of prediction features to non-technical users. This system proves that integrating XGBoost and LLM in an SME inventory management platform is technically feasible and provides tangible value for retail operations.
Peramalan Volume Transaksi Uang Elektronik Menggunakan Hybrid SARIMA-LSTM dengan Integrasi Klasifikasi Berita IndoBERT Nabila Sya’bani Wardana; Wiwik Anggraeni
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.10344

Abstract

Amid the accelerating transformation of digital payments, the volume of electronic money transactions in Indonesia has grown rapidly, exhibiting volatile patterns and frequent structural breaks. These characteristics make forecasting particularly challenging, even though accurate projections are needed to support monitoring and planning. Existing approaches generally rely on univariate models or macroeconomic variables that suffer from publication lag, rendering them less responsive to rapidly evolving economic shocks. To address these limitations, this study proposes a hybrid SARIMA-LSTM framework that incorporates online news classification, generated using IndoBERT, as an exogenous variable providing more timely signals than macroeconomic indicators. Approximately 13,000 news headlines were collected through web scraping and classified into three impact categories pendorong (driving), penghambat (inhibiting), and informatif (informative) using a fine-tuned IndoBERT model. The classification results were aggregated monthly into exogenous features, with SARIMA capturing the linear trend and seasonal patterns while LSTM modeled the nonlinear residuals. The best-performing model, a hybrid SARIMA-LSTM with the informative news variable, achieved a MAPE of 14.42%, outperforming the hybrid model without exogenous variables (15.24%), the standalone SARIMA (17.77%), and the standalone LSTM (23.37%). The main contribution is to evaluate the hybrid SARIMA-LSTM architecture and asses whether integrating online news classification can serve as a complementary indicator, reducing reliance on lag-affected data. Although the improvement from the news variable was marginal and more pronounced during volatile periods, its consistent superiority indicates that informative news carries nonlinear signals beneficial for the forecasting model.
Optimasi Model Evaluasi Kinerja Karyawan Berbasis Rekam Jejak Digital Menggunakan PCA dan Algoritma Machine Learning Yan Yang Thanri; Juli Iriani; Angel Gowasa; Luthfi Zaidi
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.8402

Abstract

In the era of digital transformation, organizations face challenges in evaluating employee performance objectively and based on data. Traditional performance appraisal systems often contain subjectivity and limitations in data integration, making them less effective in dynamic work environments. This study aims to develop a performance evaluation model based on digital footprints using machine learning and multivariate analysis. Digital footprints include work activity data (daily working hours, screen time, meetings, and emails), wearable data (physical steps, sleep duration, and stress levels), satisfaction (work-life balance, organizational support), capability (tech skills score, job level, and training), and organizational data (salary, incentives, and overtime). Principal Component Analysis (PCA) is used to reduce data dimensions and identify key performance indicators. Three machine learning algorithms—Decision Tree, Random Forest, and Gradient Boosting—are applied to classify employee performance into Low, Average, Good, and Excellent categories. Model evaluation is performed using accuracy, precision, recall, and F1-score metrics. The results show that the Gradient Boosting model combined with PCA delivers the best performance with an accuracy of 0.887 and an F1-score of 0.884. The application of PCA significantly improved classification model performance by reducing noise and multicollinearity in high-dimensional data. These findings highlight the great potential of leveraging employees' digital behavioral data to build a transparent and adaptive performance evaluation system. This study contributes to the development of intelligent HR management and supports data-driven decision-making in modern organizations.
Analisis Pengaruh Chi-Square Feature Selection terhadap Kinerja Random Forest dan XGBoost dalam Prediksi Konversi Pengunjung Website Muhamad Rosdiana; Teti Desyani; Perani Rosyani
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.9662

Abstract

The increasing number of visitors to e-commerce websites is not always accompanied by a corresponding increase in purchase transactions, making it difficult for companies to identify visitors with high conversion potential. In addition, using all available attributes may increase model complexity without necessarily improving predictive performance. This study analyzes the impact of Chi-Square Feature Selection on the performance of Random Forest and Extreme Gradient Boosting (XGBoost) in predicting website visitor conversion. The study uses the Online Shoppers Purchasing Intention dataset consisting of 12,330 instances with 17 predictor attributes and one target attribute. The research process includes exploratory data analysis, preprocessing, Chi-Square-based feature selection, classification model development using Random Forest and XGBoost, and evaluation using Accuracy, Precision, Recall, F1-Score, Matthews Correlation Coefficient (MCC), and Area Under the Curve (AUC-ROC). Four experimental scenarios were evaluated: all features (baseline), Top-15, Top-10, and Top-5 selected features. The results show that the baseline model using all features achieved the best overall performance. The Random Forest baseline model obtained an Accuracy of 90.05%, Precision of 73.94%, F1-Score of 63.13%, and MCC of 0.5835, while the XGBoost baseline model achieved the highest AUC-ROC of 0.9271. Furthermore, PageValues, BounceRates, ExitRates, ProductRelated_Duration, and ProductRelated were identified as the most influential features affecting visitor conversion. The main contribution of this study is providing empirical evidence that Chi-Square Feature Selection is more effective in reducing feature complexity and identifying relevant attributes than improving classification performance on the Online Shoppers Purchasing Intention dataset, offering practical guidance for feature selection strategies in machine learning-based website conversion prediction
Symptom-Based Classification of Migraine Severity Using Random Forest and LassoNet Syehan Fariz Gustomo; Kemas Muslim Lhaksmana
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.9900

Abstract

This study aims to develop and compare classification models for predicting symptom-based migraine intensity using a public dataset from Kaggle. The research was conducted on multiclass data with an imbalanced class distribution. Therefore, target creation and result interpretation must be carefully designed so that the resulting evaluation remains consistent with the characteristics of the data used. In this study, the Intensity variable was recoded into three operational classes: low, moderate, and high. The features used included symptoms, characteristics of migraine episodes, and `symptom_count`, which represents the number of symptoms in each sample. The two models compared were Random Forest and LassoNet, both of which were tested using Stratified 5-Fold Cross-Validation. Model performance was assessed using the Macro F1-score as the primary metric, Balanced Accuracy as the main supplementary metric, and Accuracy as a complementary metric. The test results showed that Random Forest performed better, with a Macro F1-score of 0.6366, Balanced Accuracy of 0.6181, and Accuracy of 0.6225. Meanwhile, LassoNet achieved a Macro F1-score of 0.2474, a Balanced Accuracy of 0.3333, and an Accuracy of 0.5900. These results indicate that symptom patterns in the dataset can still be utilized to distinguish migraine intensity within a computational classification framework, although the separation between closely related classes is not yet fully robust.
Hybrid Two-Stage CatBoost and Multilayer Perceptron Model for Sleep Disorder Classification Visal Ady Yanuar; Aripin Aripin
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.9958

Abstract

Sleep disorders are a health problem that significantly impacts quality of life and potentially increases the risk of various chronic diseases. Conventional sleep disorder diagnosis generally requires expensive and complex examinations, so an alternative, non-invasive data-driven approach is needed. This study proposes a sleep disorder classification approach based on tabular medical data using a Hybrid Two-Stage architecture. The proposed approach integrates the CatBoost algorithm as a binary screening stage to distinguish healthy individuals from individuals with sleep disorders, and a Multilayer Perceptron (MLP) as a latent feature extractor, combined with CatBoost to classify sleep disorder subtypes, namely insomnia and sleep apnea. The datasets used were obtained from two public data sources and evaluated using a stratified k-fold cross-validation scheme. Class imbalance was addressed using the SMOTE-ENN technique, while hyperparameter optimization was applied as part of the model training pipeline. Performance evaluation was conducted using accuracy, Macro-F1, and Matthews Correlation Coefficient (MCC) metrics. Experimental results show that the Hybrid Two-Stage architecture achieves an accuracy of 94.1%, a Macro-F1 of 0.90, and an MCC of 0.88, and exhibits stable performance across a wide range of fold variations. These results demonstrate that the hybrid two-stage approach is effective in improving the performance of sleep disorder classification based on medical tabular data. The main contribution of this study is the development of a two-stage hybrid classification framework that explicitly separates healthy-disorder screening and sleep disorder subtype classification, while integrating SMOTE-ENN, Optuna-based hyperparameter optimization, and MLP-derived latent feature representation to improve classification stability on imbalanced medical tabular data.
Perbandingan Kinerja Model ARIMA dan LSTM pada Multi-Horizon Forecasting Harga Emas dengan Evaluasi Mean Directional Accuracy Muhamad Prasetyo Bayu Aji; Aris Rakhmadi
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.10010

Abstract

Precious metals, particularly gold, represent one of the most sought-after value-preserving investment instruments, yet their dynamic price fluctuations present significant challenges, making gold a difficult-to-predict yet crucial asset for investment decision-making. This study aims to forecast gold prices by comparing the performance of Autoregressive Integrated Moving Average (ARIMA) and Long Short-Term Memory (LSTM) forecasting models through three testing scenarios: single-step, multi-step forecasting, and rolling forecasting. This study utilizes 40 years of historical gold price data obtained from the public source Kaggle. The ARIMA model was implemented on stationary data, while LSTM was optimized with additional lag, volatility, and momentum features. Experimental results indicate that in the single-step scenario, both models produced equivalent accuracy with a MAPE below 1%. In the multi-step scenario, LSTM significantly outperformed ARIMA with a MAPE of 1.89% compared to 3.54%. In the rolling scenario, LSTM again performed better with a MAPE of 1.83% versus 3.52% for ARIMA. Conversely, ARIMA consistently recorded higher Mean Directional Accuracy (MDA) values across all scenarios, reaching 57.30% in the rolling forecast compared to LSTM's 46.07%, indicating ARIMA's advantage in identifying trend direction. This study concludes that the LSTM approach is more optimal for achieving numerical prediction precision over medium-term horizons, while the statistical ARIMA method is more reliable for accurately projecting market movement direction.
Perbandingan Kinerja Naive Bayes dan SVM dalam Analisis Sentimen Program Makanan Bergizi Gratis (MBG) sebagai Pendukung Pengambilan Keputusan Vebi Adeka Putra; Nirwana Hendrastuty
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.10032

Abstract

The Free Nutritious Meal Program is one of the Indonesian government programs aimed at improving community nutritional quality and reducing stunting rates. The program has generated various publik responses expressed through media social platforms, particularly in YouTube comment sections. This study was conducted to analyze publik sentiment toward the Free Nutritious Meal Program (MBG) and to compare the performance of the Naive Bayes and Support Vector Machine (SVM) algorithms in classifying sentiment from YouTube user comments. The research data were obtained through a YouTube comment scraping process and then processed through several preprocessing stages, including cleaning, case folding, normalization, tokenization, stopword removal, and stemming. Furthermore, feature weighting was performed using the TF-IDF method, and data labeling was carried out using a lexicon-based approach. The sentiment classification process employed the Naive Bayes and Support Vector Machine (SVM) algorithms, while model evaluation was conducted using confusion matrix, accuracy, precision, precision, and f1-score metrics. The results showed that the Support Vector Machine (SVM) algorithm achieved better performance than Naive Bayes. The SVM algorithm obtained an accuracy of 77.4%, precision of 78.4%, precision of 77.4%, and f1-score of 77.6%, whereas the Naive Bayes algorithm achieved an accuracy of 70.5%, precision of 74.4%, precision of 70.5%, and f1-score of 67.7%. The main contribution of this study is the comparative evaluation of Naive Bayes and Support Vector Machine (SVM) for classifying public sentiment from YouTube comments related to the MBG program, providing empirical evidence on the most effective classification approach for supporting social media–based public opinion analysis of government policies.