cover
Contact Name
Sachnaz Desta Oktarina
Contact Email
sachnazdes@apps.ipb.ac.id
Phone
-
Journal Mail Official
ijsa@apps.ipb.ac.id
Editorial Address
sachnazdes@apps.ipb.ac.id
Location
Kota bogor,
Jawa barat
INDONESIA
Indonesian Journal of Statistics and Its Applications
ISSN : 25990802     EISSN : 25990802     DOI : -
Core Subject : Science, Education,
Indonesian Journal of Statistics and Its Applications (eISSN:2599-0802) (formerly named Forum Statistika dan Komputasi), established since 2017, publishes scientific papers in the area of statistical science and the applications. The published papers should be research papers with, but not limited to, the following topics: experimental design and analysis, survey methods and analysis, operation research, data mining, statistical modeling, computational statistics, time series and econometrics, and statistics education. All papers were reviewed by peer reviewers consisting of experts and academicians across universities and agencies
Articles 210 Documents
Deep Learning–Based Semantic Segmentation for Evaluating Urban Environmental Quality and Walkability in Dongdaemun Su Myat Thwin
Indonesian Journal of Statistics and Applications Vol 9 No 2 (2025)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v9i2p194-217

Abstract

This study evaluates environmental quality and urban walkability in the Dongdaemun district through geospatial semantic segmentation of street-view imagery. A DeepLab ResNet101 model, pre-trained on the ADE20K dataset and implemented using the GluonCV framework, was applied to Google Street View images collected at 40-meter intervals in four cardinal directions. Pixel-level segmentation was used to quantify key environmental features such as greenery, sky visibility, pavement, and road surfaces. Based on these visual attributes, composite indicators representing comfort, convenience, and safety were derived, leading to the calculation of an Integrated Visual Walkability index. The results reveal clear spatial variations in walkability across the study area, highlighting areas with favorable pedestrian environments and zones requiring improvement. Although the analysis is constrained by image quality and spatial coverage, the findings demonstrate the effectiveness of deep learning–based semantic segmentation for large-scale environmental assessment. This approach provides a scalable and data-driven framework to support evidence-based urban planning and sustainable city development.
Bayesian Vector Autoregressive Modeling on Macroeconomic Variables in Indonesia Indra Mahib Zuhair Riyanto; Muhammad Firlan Maulana; Nur Anggraini Fadhilah; Laras Suprapti; Salsabila Fayiza; Eliza Rahmadania; Bulan Cahyani Suhaeri; Anang Kurnia; Laily Nissa Atul Mualifah; Aulia Akhrian Syahidi
Indonesian Journal of Statistics and Applications Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v10i1p105-118

Abstract

This research studies a Bayesian Vector Autoregressive (BVAR) model to analyze the dynamic interactions among the rupiah exchange rate, exports, imports, gold futures prices, and inflation in Indonesia during the 2015-2024 period. The BVAR method was chosen to overcome the limitations of conventional VAR models on overparameterization problem by utilizing hierarchical Minnesota priors and Markov-Chain Monte Carlo (MCMC) estimation. Data were stationary through first order differencing and normalized using z-score. Lag selection based on the Akaike Information Criterion (AIC) showed that lag 6 is optimal. Model evaluation using Mean Absolute Percentage Error (MAPE) shows good overall model performance on training data, especially on the gold price variable (MAPE 10,09%) and inflation (MAPE 3,74%). On test data, the model struggles to perform well on prediction due to the high uncertainty of the test data period. Impulse Response Function (IRF) analysis is used to reveal short-term responses between variables, such as the effect of exchange rate depreciation on inflation and the impact of export value on a temporary decline in import value. The result highlights the BVAR model’s ability to capture general macroeconomic relationships, especially when many parameters need to be estimated and the available data is limited.
Comparative Study of Multiclass Models for Identifying Household Drinking Water Sources in West Java Defri Ramadhan Ismana; Bagus Sartono; Erfiani; Yuri Nurdiantami; Mohammad Masjkur; Aam Alamudi; Rahma Anisa
Indonesian Journal of Statistics and Applications Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v10i1p119-128

Abstract

Data imbalance is a common challenge in classification modeling, typically driven by rare events such as fraud, credit default, student dropout, or infectious diseases. Multi-class imbalance is inherently more complex than binary classification because a class can be a minority relative to one class yet a majority to another. Access to safe drinking water is one of the key factors for public health. West Java, as the province with the largest population in Indonesia, has yet to achieve the target of 100% of households having access to safe drinking water. Identifying the types of drinking water sources can be approached using multi-class classification modeling. However, this case presents an issue of data imbalance. The CatBoost method is one of the recommended approaches for multi-class classification with imbalanced data. Additionally, the TabNet method is also considered to perform well in such cases. Therefore, this study aims to apply both methods and compare their performance in classifying household drinking water sources in West Java using 2023 SUSENAS data. The results show that CatBoost yielded an average MCC of 0.225, whereas TabNet exhibited an average of 0.209. In terms of computational efficiency, CatBoost recorded an average execution time of 25 seconds, compared to 134 seconds for TabNet. Therefore, the CatBoost method provides better multi-class classification performance and faster execution compared to TabNet. Although TabNet is superior in classifying minority classes, CatBoost is more accurate in predicting majority classes and identifying the most influential variables, such as building ownership status and regional classification
Spatio-Temporal Clustering of PM2.5 Estimation in Jakarta Lutfiah Nursabiliyanti; Imas Sukaesih Sitanggang; Hendra Rahmawan; Muhammad Asyhar Agmalaro; Nor Azura Husin
Indonesian Journal of Statistics and Applications Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v10i1p79-93

Abstract

PM2.5 has negative impacts on human health because it can penetrate the alveoli of the lungs. This study aims to develop a PM2.5 estimation model for Jakarta based on Himawari AOD (Aerosol Optical Depth) and weather data from 2022 to 2024, utilizing a Random Forest Regressor and clustering with ST-DBSCAN. The data used in this study were PM2.5 measurements from eight air quality monitoring stations, recorded on the Jakarta Low Emissions website, and AOD Level 2 data from the Himawari-8 and Himawari-9 satellites, with a spatial resolution of 0.05 degrees, obtained from the Japan Aerospace Exploration Agency (JAXA) website. The results showed that the best PM2.5 estimation model was obtained with R2 = 0.63 and MAE = 8.035. Smoothing techniques for PM2.5 data have also been shown to improve model performance. Furthermore, based on the feature importance of the best PM2.5 estimation model, the Himawari AOD data are considered to have a less significant contribution to the PM2.5 estimation model. Clustering using ST-DBSCAN was successfully implemented on PM2.5 data for the 2022-2024 period, divided into 12 sub-datasets based on the seasons per year: the rainy season (DJF), the transition from rainy to dry (MAM), the dry season (JJA), and the transition from dry to rainy (SON). The best clustering result yielded a Silhouette coefficient of 0.75 on the 2023 rainy season (December–February) dataset. Three clusters were formed, consisting of 2094, 1458, and 878 data points, along with 45 noise points. The average PM2.5 levels in each cluster were 70.12 µg/m³, 62.71 µg/m³, and 52.18 µg/m³, respectively. The results of this study are expected to benefit other stakeholders involved in air pollution control in Jakarta.
Kernel Analog Forecasting with Koopman Operator Theory for Indo-Pacific Climate Variability: A Nonlinear Spatio-Temporal Statistical Framework Dhilshath Shajahan; Ali Akg¨ul
Indonesian Journal of Statistics and Applications Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v10i1p47-59

Abstract

Skillful long-range forecasting of the El~Niño-Southern Oscillation (ENSO) andthe Indian Ocean Dipole (IOD) remains, despite decades of sustained research effort,an open problem of substantial societal weight.The difficulty is attributable to a basic tension involving the structuralinappropriateness of classical linear statistical frameworks, despite theiranalytical tractability, to regimes whose commencement and termination ofextreme events are nonlinear, with lead times above five months.We describe here a forecasting architecture that fuses \emph{Kernel AnalogForecasting} (KAF) with the operator-theoretic machinery of \emph{Koopmanspectral analysis}, thus re-formulating a nonlinear, high-dimensional predictionproblem as linear regression in a reproducing kernel Hilbert space (RKHS).The predictor representation takes the form of an anisotropic, Markov-normalisedGaussian kernel constructed on delay-coordinate embeddings of Indo-Pacific seasurface temperature (SST) fields, with the anisotropy parameter optimised topreferentially weight pairs of states evolving along coherent dynamical directions.Assuming ergodicity and mild regularity, we prove convergence of the KAF estimatorto the Koopman-propagated conditional expectation of the target observable the $L^2$-optimal predictor (Theorem~\ref{thm:convergence}).Against calibrated multifrequency synthetic signals, KAF sustains patterncorrelation above 0.5 to roughly 14~months ahead, versus 7~months for alinear inverse model baseline; root-mean-square error falls by 18-26\%across the full verification horizon.The notorious spring predictability barrier weakens materially, and probabilisticevent forecasts based on Brier Skill Score and the Kullback-Leibler relativeentropy exhibit sharper, less-biased distributions than either the linearor \LSTM{} comparators.
Spatial Analysis of Primary-Sector GRDP in Pemalang Regency Using Remote Sensing Vegetation Indices Khusnudin Tri Subhi; Vincentia Anggita Puspitasari
Indonesian Journal of Statistics and Applications Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v10i1p23-36

Abstract

Agriculture has contributed with the highest percentage of Pemalang Grodd Regional Product in 2022 and still continue to become regency’s primary source of income. However, the GRDP figures published at regency level do not really show how agricultural productivity varies from one place to another, something that matters a lot for making good policy decisions. In this paper, we combine official economic records with satellite-derived vegetation indicators to map where primary-sector GRDP is concentrated. We used a Geographically Weighted Regression (GWR) approach for this analysis. We downscaled GRDP data from BPS (Statistics Indonesia) onto a 500 × 500 m grid as the response variable. We used vegetation indices like NDVI, EVI, MSAVI, LAI, and MNDWI, which we calculated from Landsat and MODIS imagery through Google Earth Engine, as predictors that reflect crop condition, land productivity, and water availability. The GWR model showed a much better fit than ordinary global regression, with an Adjusted R² of 0.77. This indicates that spatial variation is significant here. The highest GRDP coefficients appeared in northern and northeastern coastal subdistricts such as Ulujami, Comal, and Petarukan, as well as in areas of the fertile hinterland, where irrigated farming and aquaculture thrive. Lower values emerged in central and southern areas that are more urbanized. A clear dual landscape exists in Pemalang: strong agricultural output in rural regions, alongside a gradual shift toward manufacturing and services in urban centers. Our results highlight how remote sensing, spatial modeling, and economic statistics can be brought together to produce detailed agricultural GRDP maps, and we think these maps can be genuinely useful for protecting farmland, targeting agricultural support, and planning more balanced regional development.
Asymmetric Laplace Stochastic Volatility Model and its Applications Rahul Thekkedath; Shiji Kavungal
Indonesian Journal of Statistics and Applications Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v10i1p60-78

Abstract

This paper proposes a stochastic volatility model driven by a first-order autoregressive process with an asymmetric Laplace marginal distribution. The autoregressive structure with asymmetric Laplace marginal is incorporated into the variance equation to better capture asymmetry and heavy tails in financial return series. The model parameters are estimated using the generalized method of moments. A simulation study is conducted to evaluate the performance of the estimators. Finally, a real-data application is presented to illustrate the practical utility of the proposed model and to demonstrate that it captures the stylized features of financial return series.
A Two-Stage Framework for Unsupervised Sentiment Analysis with Model Selection and Semantic Similarity Evaluation Cici Suhaeni; Fani Fahira; Hari Wijayanto; La Ode Abdul Rahman; Hwan-Seung Yong
Indonesian Journal of Statistics and Applications Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v10i1p01-22

Abstract

Sentiment analysis is widely used to extract user opinions from large-scale textual data. However, in practical settings, sentiment labels are often unavailable, making it difficult not only to perform sentiment classification but also to evaluate whether the predicted labels are reliable. This study proposes a two-stage framework for unsupervised sentiment analysis with model selection and semantic similarity evaluation. The dataset consists of Gemini app reviews, in which a labeled subset was used in the first stage to investigate the behavioral characteristics and predictive patterns of three sentiment analysis approaches: lexicon-based, transformer-based, and large language model (LLM)-based methods. The model with the most suitable performance was then selected and applied to predict sentiment labels for the remaining unlabeled data in the second stage. The predicted labels were further evaluated using embedding-based cosine similarity to assess semantic consistency within sentiment classes and separability between classes. The results show that the LLM-based method using Gemini 2.0 Flash achieved the best performance, with accuracy, balanced accuracy, and F1-score values above 0.91, followed by the transformer-based IndoBERT model, while the InSet lexicon-based method showed the weakest performance. In the second stage, semantic similarity evaluation revealed a high average intra-class similarity of 0.6699 and a low inter-class similarity of 0.1917, resulting in a similarity gap of 0.4783. These findings indicate that the proposed framework can support reliable sentiment prediction in largely unlabeled datasets by combining sample-based model selection with semantic validation.
A Comparative Study of Sequential Biclustering and Fuzzy C-Means, K-Nearest Neighbors, and Mean Imputation for Missing Value Estimation in Gene Expression Data Yolanda Azzahra; Titin Siswantining; Setia Pramana; Mogana Darshini Ganggayah; Alhadi Bustamam
Indonesian Journal of Statistics and Applications Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v10i1p37-46

Abstract

Missing values are a common issue in gene expression data and can significantly affect downstream analysis. This study aims to compare the performance of a hybrid sequential biclustering and centroid-based clustering method with conventional imputation approaches for handling missing values. The proposed method integrates sequential biclustering based on mean squared residue to identify coherent submatrices, followed by centroid-based clustering to estimate missing entries. The dataset used in this study consists of gene expression data of patients with type 2 diabetes mellitus, with missing values introduced under various proportions ranging from 5% to 55%. The performance of the proposed method is evaluated and compared with mean imputation and nearest neighbor imputation using mean squared error, root mean squared error, and mean absolute error. The experimental results show that the proposed method consistently produces lower errors across all missing rates than the baseline methods. This indicates that incorporating local pattern structures through biclustering improves the accuracy of missing value estimation. The findings suggest that the proposed hybrid framework is more effective in preserving the underlying structure of gene expression data and provides a reliable approach for handling missing data in high-dimensional biological datasets.
Segmentation of Health Status in Coastal Areas Using the Finite Mixture Partial Least Squares (FIMIX-PLS) Method: An Analysis Based on Socioeconomic Factors Riwi Dyah Pangesti; Idhia Sriliana; Susi Wijuniamurti; Alus Ahmad Suhaimi; Athaya Fairuzindah; Anne Mudya Yolanda
Indonesian Journal of Statistics and Applications Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v10i1p94-104

Abstract

The purpose of this study is to model and segment the health status at the regency/city level in the Southern Sumatra region by considering the complexity of relationships among the variables of Poverty, Economic Welfare, Environment, Utilization of Health Services, and Educational Attainment. The method used to form latent segments is Finite Mixture Partial Least Square Structural Equation Modeling (FIMIX-PLS SEM), which classifies latent heterogeneity based on finite mixture distributions into membership probabilities for each segment. The results obtained from the FIMIX-PLS SEM analysis indicate the formation of 2 optimal segments based on nearly all segment selection criteria. Segment 1 consists of 49 regencies/cities that are more strongly influenced by poverty and economic welfare. Segment 2 consists of 11 regencies/cities that are more strongly influenced by environmental factors, utilization of health services, and educational attainment. The formation of segments also increases the R² value, thereby increasing the variation in Health Degree that can be explained by the exogenous latent variables (Poverty, Economic Welfare, Environment, Utilization of Health Services, and Educational Attainment). Therefore, FIMIX-PLS SEM is capable of improving the accuracy of the analysis results.