cover
Contact Name
Teuku Rizky Noviandy
Contact Email
trizkynoviandy@gmail.com
Phone
+6282275731976
Journal Mail Official
editorial-office@heca-analitika.com
Editorial Address
Jl. Makam T. Nyak Arief Kompleks BUPERTA Blok L7B, Lamgapang, Aceh Besar, Provinsi Aceh
Location
Kab. aceh besar,
Aceh
INDONESIA
Infolitika Journal of Data Science
ISSN : -     EISSN : 30258618     DOI : https://doi.org/10.60084/ijds
Infolitika Journal of Data Science is a distinguished international scientific journal that showcases high caliber original research articles and comprehensive review papers in the field of data science. The journals core mission is to stimulate interdisciplinary research collaboration, facilitate the exchange of knowledge, and drive the advancement and application of innovative strategies within the data science domain. Topics of this journal includes, but not limited to Data Mining and Analysis, Machine Learning and Artificial Intelligence, Big Data and Data Engineering, Predictive Modeling and Forecasting, Natural Language Processing, Computer Vision, Data Visualization and Interpretation, Ethics and Privacy in Data Science, Applications of Data Science, Interdisciplinary Approaches
Articles 35 Documents
Assessing LightGBM Performance in Automated Leukemia Cell Classification Rara Syifa Qaisa; Hayatun Maghfirah; Suryadi Suryadi; Noviana Husdayanti; Rivansyah Suhendra
Infolitika Journal of Data Science Vol. 4 No. 1 (2026): May 2026
Publisher : Heca Sentra Analitika

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.60084/ijds.v4i1.351

Abstract

Leukemia is a type of blood cancer that requires fast and accurate diagnosis for effective treatment. Manual identification of leukemia blood cell subtypes is often challenging, time-consuming, and prone to observer variability, making automated image-based classification essential. This study evaluates the performance of the Light Gradient-Boosting Machine (LightGBM) as a computationally efficient and interpretable alternative to deep learning models for classifying leukemia subtypes. The dataset includes 3,000 microscopic images representing five classes: acute lymphocytic, acute myelogenous, chronic lymphocytic, chronic myelogenous, and healthy blood cells. Images were preprocessed using bilinear interpolation to balance quality and efficiency, and 90 statistical features were extracted across 13 distinct color spaces. The model was trained on an 80% subset and validated on a 20% hold-out set after hyperparameter optimization. LightGBM achieved robust performance with an accuracy of 93.3%, precision of 99.1%, recall of 94.9%, and an F-measure of 96.8%. Feature importance analysis revealed that texture variance in the YIQ color space (STD_YIQ_I) was the most critical predictor, highlighting the biological relevance of chromatin texture in classification. These results indicate that LightGBM is an effective, lightweight, and reliable approach for leukemia subtype classification, holding strong potential for implementation in resource-constrained automated diagnostic systems.
Time Series Analysis of UV Radiation and Temperature Using Seasonal ARIMA Rahmatul Fauzi; Tasyaul Husna; Izzul Akrami; Novi Reandy Sasmita
Infolitika Journal of Data Science Vol. 4 No. 1 (2026): May 2026
Publisher : Heca Sentra Analitika

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.60084/ijds.v4i1.401

Abstract

High exposure to ultraviolet (UV) radiation in Banda Aceh poses significant risks to public health and the environment. Daily forecasts of UVA, UVB, and maximum temperature are important for climate planning. But current models often overlook daily changes in tropical regions or fail to incorporate them into their forecasts. This study develops a climatologically informed SARIMA framework incorporating a semiannual seasonal structure (s = 180) to model ultraviolet radiation and temperature dynamics in an equatorial tropical region. The variables used were UVA (W/m²), UVB (W/m²), and maximum temperature (°C) in Banda Aceh during the period December 2018-July 2024. The SARIMA method was applied after data pre-processing, such as Box-Cox transformation that stabilizes the variance (λ ≈ 1) and seasonal differencing (s = 180 days) to overcome non-stationarity. Model identification using ACF/PACF plots, with diagnostic tests (Ljung-Box white noise test, Shapiro-Wilk normality test) and accuracy metrics (MAPE, MASE, BIC, and AIC) for optimization. SARIMA(1,0,2)(1,1,0)¹⁸⁰ was selected as the optimal model for all variables. The selected SARIMA models yielded MAPE values of 18.93% (UVA), 0.48% (UVB), and 13.03% (temperature), indicating that the selected SARIMA specifications were able to capture the dominant temporal patterns observed in the analyzed dataset. The peak values for March-April 2025 were predicted to be 17.69 W/m² (UVA), 0.69 W/m² (UVB), and 31.85°C.
Ensemble Variable Importance: Combining Random Forest, Neural Network, and Support Vector Machine via Genetic Algorithm (Case Study: Student Productivity) Asep Rusyana; Marzuki Marzuki; Siti Rusdiana; Fitriana AR; Nurhasanah Nurhasanah; Nany Salwa; Mahmudi Mahmudi
Infolitika Journal of Data Science Vol. 4 No. 1 (2026): May 2026
Publisher : Heca Sentra Analitika

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.60084/ijds.v4i1.424

Abstract

This study proposes and evaluates an ensemble variable‑importance framework that integrates permutation‑based importance scores from three distinct supervised learning algorithms: Random Forest, Neural Network, and Support Vector Machine, using a genetic‑algorithm optimizer. The approach addresses the well‑known problem that algorithm‑specific importance diagnostics can yield divergent feature rankings, complicating substantive interpretation and downstream decision‑making. Using a large publicly available student‑productivity dataset (N = 20,000), predictors describing study behavior, digital‑media use, lifestyle, and academic indicators were normalized with Min–Max scaling, and permutation variable importance (PVI) was estimated repeatedly within each model to obtain stable mean PVI values and standard errors. A genetic algorithm was then employed to search the space of ensemble weightings (rank‑aggregation solutions) that maximize a chosen fitness criterion—Spearman rank concordance with out‑of‑sample predictive relevance—thereby producing a consensus ranking of predictors. Empirical results indicate rapid GA convergence (fitness ≈ 0.82 within 20–30 generations) and strong cross‑model agreement for a small core of predictors: study hours (X3) and focus score (X15) consistently emerged as the most salient features across individual models and in the ensemble ranking. A secondary set of variables (e.g., sleep hours, phone usage, attendance, and stress level) displayed moderate importance, while several features exhibited model‑dependent variability in ranks. The ensemble procedure thereby yields stable, model‑agnostic importance estimates that enhance interpretability and reduce dependence on any single algorithm’s idiosyncrasies. We discuss implications for educational analytics and recommend external validation, targeted feature engineering, and sensitivity analyses (alternate scalings and GA settings) to assess robustness and to support reliable, actionable inferences from machine‑learning models in applied settings.
Short-Term Reliability and Long-Term Limits of Stable Log-ARIMA for Population Forecasting Across Southeast Asia Ghifari Maulana Idroes; Fatih Avicenna Hizir; Muhammad Zhafran Abiyyu
Infolitika Journal of Data Science Vol. 4 No. 1 (2026): May 2026
Publisher : Heca Sentra Analitika

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.60084/ijds.v4i1.425

Abstract

Population forecasting is essential for long-term planning in labor markets, healthcare, education, infrastructure, and social protection. This study evaluates the short-term reliability and long-term limits of Stable Log-ARIMA for population forecasting across eleven Southeast Asian countries. Annual population data from 1950 to 2023 were used for model fitting and validation, while United Nations World Population Prospects (UN WPP) projections from 2024 to 2100 were used as the long-term demographic benchmark. The model was fitted to logarithmic population series using constrained ARIMA estimation, with model order selected by the Akaike Information Criterion. Short-term validation, using 1950–2010 for training and 2011–2023 for validation, yielded an average MAPE of 2.01%, indicating strong short-term forecasting performance. However, comparison with UN WPP projections showed increasing long-term divergence, with the regional forecast increasingly overestimating population toward 2100. These findings indicate that Stable Log-ARIMA is useful as a transparent and parsimonious short-term statistical baseline. Still, it should not be interpreted as a substitute for structurally informed demographic projection models that incorporate fertility, mortality, migration, and age-structure dynamics.
Internet Bandwidth Forecasting by Using Fuzzy Time Series in Zainal Abidin General Hospital, Indonesia Khalid Rianda; Muhd Iqbal; Muslim Amiren; Maulyanda Maulyanda; Afdhaluzzikri Afdhaluzzikri; Intan Syahrini; Siti Rusdiana; Abdul Fikri; Irvanizam Irvanizam
Infolitika Journal of Data Science Vol. 4 No. 1 (2026): May 2026
Publisher : Heca Sentra Analitika

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.60084/ijds.v4i1.428

Abstract

Data bandwidth capacity is a critical component of internet infrastructure management, directly impacting network efficiency and operational costs. Accurate measurement and forecasting of bandwidth requirements are essential to optimize resource allocation. This study utilizes a Fuzzy Time Series (FTS) approach for bandwidth forecasting, leveraging its ability to capture complex patterns from historical data without requiring the rigid statistical assumptions of classical forecasting methods. A forecasting model was developed and implemented to predict data bandwidth requirements at the Zainal Abidin General Hospital (RSUZA). Utilizing historical data collected from February 1, 2019, to April 29, 2019, the model's performance was evaluated using the Mean Absolute Percentage Error (MAPE). The proposed method achieved a MAPE of 6.45%, demonstrating high accuracy and falling into the "highly accurate" category.

Page 4 of 4 | Total Record : 35