Articles
PENGGEROMBOLAN SUBSEKTOR INDUSTRI BERDASARKAN PERKEMBANGAN INDEKS PRODUKSI MENGGUNAKAN PREDICTION-BASED CLUSTERING
Agustin Faradila;
Utami Dyah Syafitri;
I Made Sumertajaya
Indonesian Journal of Statistics and Applications Vol 4 No 3 (2020)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)
Show Abstract
|
Download Original
|
Original Source
|
Check in Google Scholar
|
DOI: 10.29244/ijsa.v4i3.585
Statistics Indonesia (BPS) noted that there has been a decrease in the contribution of the industrial sector to the national GDP even though it had provided a significant multiplier effect on national economic growth. Therefore, it is necessary to cluster the industrial subsector based on its growth patterns so that the optimization of development results can be achieved. Prediction-based clustering is part of time series clustering (TSclust) which aims to form clusters based on prediction characteristics so that it can be used to choose a cluster that will become a mainstay industry in the future. This study focused on applying prediction-based clustering in the large and medium industrial sub-sector for a prediction period of 1 month, 1 quarter, and 1 semester. The data used in this study was the production index data from January 2010 to December 2018. The results showed that the best cluster for 1 month consisted of 5 groups, for 1 quarter consisted of 4 groups and for 1 semester consisted of 2 groups. Thus, it was concluded that the food industry; leather industry, leather goods, and footwear; and the pharmaceutical industry, chemical drug products, and traditional medicine could be chosen to be the mainstay industry in the future.
PENGGEROMBOLAN DERET WAKTU DENGAN PENDEKATAN UKURAN KEMIRIPAN PICCOLO UNTUK PERAMALAN CURAH HUJAN PROVINSI BANTEN
Sarah Fadhlia;
I Made Sumertajaya;
Anik Djuraidah
Indonesian Journal of Statistics and Applications Vol 4 No 2 (2020)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)
Show Abstract
|
Download Original
|
Original Source
|
Check in Google Scholar
|
DOI: 10.29244/ijsa.v4i2.607
Time series data modeling can be done by modeling each object one by one. Monthly rainfall data is an example of time series data. The purpose of time series analysis is to find patterns of past data and then forecast the future characteristics of data. The data used in this study is the Banten Province rainfall data which contained 19 rainfall stations. So it will require 19 models to forecast the rainfall data. The pattern of time series data in Banten Province monthly rainfall data in several locations has similarities. So that the similarity of this pattern can be considered in the clusters. In time series clustering, the idea is to investigate the similarity of time series in a cluster. The accuracy of distance similarity size measurements is performed on the generation data generated from 3 models, namely AR (1), AR (2), and AR (3). The piccolo method has an average accuracy of 0.62. While the maharaj method has an average accuracy of 0.41. This means that the Ward hierarchical clustering method using the Piccolo distance approach has a greater accuracy value than the Maharaj distance approach. Furthermore, the Piccolo method can be used as an alternative to the excellent distance method for grouping time series data in case data. The Banten Province rainfall station has 3 optimal clusters. Modeling individual level and cluster level has accuracy values that are not much different.
EVALUASI KINERJA METODE CLUSTER ENSEMBLE DAN LATENT CLASS CLUSTERING PADA PEUBAH CAMPURAN
Debora Chrisinta;
I Made Sumertajaya;
Indahwati Indahwati
Indonesian Journal of Statistics and Applications Vol 4 No 3 (2020)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)
Show Abstract
|
Download Original
|
Original Source
|
Check in Google Scholar
|
DOI: 10.29244/ijsa.v4i3.630
Most of the traditional clustering algorithms are designed to focus either on numeric data or on categorical data. The collected data in the real-world often contain both numeric and categorical attributes. It is difficult for applying traditional clustering algorithms directly to these kinds of data. So, the paper aims to show the best method based on the cluster ensemble and latent class clustering approach for mixed data. Cluster ensemble is a method to combine different clustering results from two sub-datasets: the categorical and numerical variables. Then, clustering algorithms are designed for numerical and categorical datasets that are employed to produce corresponding clusters. On the other side, latent class clustering is a model-based clustering used for any type of data. The numbers of clusters base on the estimation of the probability model used. The best clustering method recommends LCC, which provides higher accuracy and the smallest standard deviation ratio. However, both LCC and cluster ensemble methods produce evaluation values that are not much different as the application method used potential village data in Bengkulu Province for clustering.
KAJIAN VARIANCE MEAN RATIO PADA SIMULASI SEBARAN DATA BINOMIAL NEGATIF
Choirun Nisa;
Muhammad Nur Aidi;
I Made Sumertajaya
Indonesian Journal of Statistics and Applications Vol 4 No 4 (2020)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)
Show Abstract
|
Download Original
|
Original Source
|
Check in Google Scholar
|
DOI: 10.29244/ijsa.v4i4.689
The negative binomial distribution is one of the data collection counts that focuses on success and failure events. This study conducted a study of the distribution of negative binomial data to determine the characterization of the distribution based on the value of Variance Mean Ratio (VMR). Simulation data are generated based on negative binomial distribution with a combination of p and n parameters. The results of the VMR study on negative binomial distribution simulation data show that the VMR value will be smaller when the p-value is large and the VMR value is more stable as the sample size increases. Simulation data of negative binomial distribution when p≥0.5 begins to change data distribution to the distribution of Poisson and binomial. The calculation VMR value can be used as a reference for detecting patterns of data count distribution.
ANALISIS INFLASI MENGGUNAKAN DATA GOOGLE TRENDS DENGAN MODEL ARIMAX DI DKI JAKARTA
Newton Newton;
Anang Kurnia;
I Made Sumertajaya
Indonesian Journal of Statistics and Applications Vol 4 No 3 (2020)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)
Show Abstract
|
Download Original
|
Original Source
|
Check in Google Scholar
|
DOI: 10.29244/ijsa.v4i3.694
Inflation is an important economic indicator in showing the economic symptoms of a region's price level. DKI Jakarta is the capital of Indonesia chosen as the center of the economic barometer because it can provide the greatest contribution and influence on the Indonesian economy. The ARIMAX model was used for forecasting by adding independent variables in the Google trends data. Google trends data were explored based on seven expenditure groups published by IHK. The purpose of this study was to determine the effect of forecast Google trends using BPS inflation data in DKI Jakarta. The result of the exploration of Google Trends data was forecasted to get the best forecast model results. The result of data analysis indicates that the forecast results approached the original BPS data with the best forecast model is ARIMAX (2.0.3) all variables X. Google Trends data can be used as forecasting but cannot be used as a reference policy decision.
Geographically Weighted Regression with Kernel Weighted Function on Poverty Cases in West Java Province: Regresi Terboboti Geografis dengan Fungsi Pembobot Kernel pada Data Kemiskinan di Provinsi Jawa Barat
Winda Nurpadilah;
I Made Sumertajaya;
Muhamad Nur Aidi
Indonesian Journal of Statistics and Applications Vol 5 No 1 (2021)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)
Show Abstract
|
Download Original
|
Original Source
|
Check in Google Scholar
|
DOI: 10.29244/ijsa.v5i1p173-181
Spatial regression analysis is a form of regression model that considers spatial effects. Geographically weighted regression (GWR) is the spatial regression methods that can be used to deal with the problem of spatial diversity. This method generates local model parameter estimates for each observation location. The application of spatial statistics can be done in all areas such as the problem of poverty. Poverty can be influenced by factors of proximity between regions, so that in determining the poverty factor, the proximity factor of the region cannot be ignored. West Java Province is a province with the largest population, so this study aims to model the poverty data in West Java Province by incorporating spatial effects. The weighting function used for the GWR model is the function of the fixed and adaptive kernels. The analysis results show that the fixed exponential kernel function has the smallest cross validation (CV) value, so the weighting matrix used in the model is determined by the exponential kernel function. The largest  value and the smallest AIC value are owned by the GWR model with an exponential kernel function. Based on the results obtained by the the ANOVA table to test GWR's global goodness, the GWR model is more effective than global regression. Therefore, the GWR model is the best model when it used in West Java’s poverty cases. The effect of each explanatory variable on the percentage of poverty varies in each district/city in West Java Province.
Study of Clustering Time Series Forecasting Model for Provincial Grouping in Indonesia Based on Rice Price: Kajian Model Peramalan Clustering Time Series untuk Penggerombolan Provinsi Indonesia berdasarkan Harga Beras
Muhammad Ulinnuha;
Farit M Afendi;
I Made Sumertajaya
Indonesian Journal of Statistics and Applications Vol 6 No 1 (2022)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)
Show Abstract
|
Download Original
|
Original Source
|
Check in Google Scholar
|
DOI: 10.29244/ijsa.v6i1p50-62
Most indonesians consume rice as the main staple. The high low price of rice has an impact on farmers and communities, especially those who cannot afford it. Rice price forecasting is one of the important information to be considered for future rice prices. The data used is secondary data sourced from bps publication, Rural Consumer Price Statistics: Food Group, from January 2008 to December 2019 for 32 provinces in Indonesia. Time series modeling and forecasting is usually done on a single variable using ARIMA. however, modeling becomes inefficient if there are many variables, so clustering time series analysis is performed using correlation distance with the clustering method of average linkage hierarchy. Cluster level ARIMA modeling with 4 clusters provides high efficiency because only by doing 4 times modeling results in accuracy values not much different from individual level modeling. the results obtained by individual-level ARIMA Modeling resulted in an average MAPE of 3.36%, while cluster-level ARIMA modeling with 4 clusters resulted in an average MAPE value of 4.27%, with a second MAPE difference of -0.91%. Formally conducted z test, the results obtained there is no difference between individual-level MAPE and cluster-level MAPE. This means that cluster-level modeling is relatively good and representative.
Application of Fuzzy C-Means and Weighted Scoring Methods for Mapping Blankspot Villages in Pemalang Regency
Imam Adiyana;
I Made Sumertajaya;
Farit M Afendi
Indonesian Journal of Statistics and Applications Vol 6 No 1 (2022)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)
Show Abstract
|
Download Original
|
Original Source
|
Check in Google Scholar
|
DOI: 10.29244/ijsa.v6i1p77-89
Covid-19 pandemic affects habits people around the world. The education sector in Indonesia is also undergoing policy changes, namely policy of transitioning face-to-face teaching and learning process to distance learning process (PJJ/online learning). Several studies have been conducted to examine the constraints PJJ process, resulting in finding that quality of internet network is majority obstacle in PJJ process. Conditions where there is no internet network in an area is commonly called a blankspot. In order to minimize the problem of blankspots, President and Ministry of Communication and Informatics of Indonesia realized the program "Indonesia is free signals to the corners of the country". This program involves all districts in Indonesia to conduct network quality surveys in the smallest areas of the village. Basically, network quality survey activities require relatively no small resources and costs. So as to conduct the efficiency of field survey activities, early detection of village blankspot status is required based on the characteristics blankspot village in general. While the commonly used method of grouping village based on village characteristics is the fuzzy c-means and weighted scoring method. These two methods were chosen because they have good cluster convergence rate and easily interpreted display results of the group by user in the form diagrams and scores. This study aims to prove that fuzzy c-means and weighted scoring method are good for grouping cases of blankspot villages according to previous studies with different cases. The result comparison goodness value of clustering, it is known that fuzzy c-means method more suitable for clustering characteristics blankspot village than the k-means method. Meanwhile, weighted scoring method cannot be said better method for village classification than the decision tree method.
Energy Sector Stock Price Forecasting with Time Series Clustering Approach: Peramalan Harga Saham Sektor Energi dengan Pendekatan Penggerombolan Data Deret Waktu
Linda Sakinah;
Rahma Anisa;
I Made Sumertajaya
Indonesian Journal of Statistics and Applications Vol 8 No 2 (2024)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)
Show Abstract
|
Download Original
|
Original Source
|
Check in Google Scholar
|
DOI: 10.29244/ijsa.v8i2p132-142
Stock investment promises higher returns but carries high risks because unpredictable price fluctuations. Energy sector shows potential due to its highest sectoral index growth in 2022. However, this doesn’t indicate that stock price increases occur evenly among all issuers. Therefore, it’s necessary to analyze clustering of issuers based on similarity of their stock price movements and used for forecasting stock prices at cluster level. This study aims to evaluate performance of clustering energy sector issuers using autocorrelation-based distance and dynamic time warping(DTW), and to forecast stock prices at cluster level. The data used consists weekly closing stock prices. The clustering used hierarchical average linkage method. Stock price forecast for each cluster used ARIMA model and its performance was evaluated using rolling-cross validation. The results showed that DTW distance had the best clustering performance. Energy sector issuers were grouped into four clusters with strong cluster category, indicated by silhouette coefficient >0.71. ARIMA models for each cluster produced MAPE values between 10-20%, categorizing them as good forecasting models. Clusters A and D were recommended for investors because have highest potential for capital gain based on forecasted stock prices. That clusters also consisted of companies with strong fundamentals and dividend policies.
Stochastic Residual Selection in Simulated Annealing for Clusterwise Panel Optimization
Luh Putu Widya Adnyani;
Bagus Sartono;
Asep Saefuddin;
I Made Sumertajaya;
Gerry Alfa Dito
Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi) Vol 10 No 4 (2026): August 2026
Publisher : Ikatan Ahli Informatika Indonesia (IAII)
Show Abstract
|
Download Original
|
Original Source
|
Check in Google Scholar
|
DOI: 10.29207/resti.v10i4.7608
Modeling heterogeneity in panel data requires solving a complex combinatorial partition problem under structural constraints. Although clusterwise regression captures latent group structures with distinct parameters, determining the optimal partition remains computationally challenging due to the vast solution space and susceptibility to local minima. This study proposes a modified simulated annealing (SA) algorithm incorporating a Stochastic Residual Selection (SRS) mechanism, in which candidate units are selected from a high-residual subset rather than deterministically relocating only the unit with the largest residual. The stochastic candidate-size parameter was evaluated using m=1 and m=5, where m=1 represents deterministic selection of the largest residual unit, while m=5 randomly selects one unit from the five largest-residual units for clusterreassignment. The stochastic perturbation enhances global exploration and improves convergence stability in non-convex optimization landscape. Simulation experiments involving 200 individuals observed over three time periods demonstrate that the proposed SRS-SA outperforms standard SA, achieving an Adjusted Rand Index of approximately 0.95 at 1,000 iterations while producing lower Mean Absolute Bias and Mean Squared Error. An empirical application to improved sanitation data across districts and municipalities in Java, Indonesia, further confirms its effectiveness in identifying latent structural heterogeneity. These findings highlight the robustness and computational efficiency gained through stochastic diversification in metaheuristic optimization for constrained clusterwise panel modeling.