Claim Missing Document
Check
Articles

Found 30 Documents
Search

Comparison of ARIMA, LSTM, and Ensemble Averaging Models for Short-Term and Long- Term Forecasting of Non-Stationary Time Series Data Pratiwi, Windy Ayu; Sumertajaya, I Made; Notodiputro, Khairil Anwar
Inferensi Vol 8, No 3 (2025)
Publisher : Department of Statistics ITS

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.12962/j27213862.v8i3.22643

Abstract

This study aims to forecast the highest weekly selling rate of the Indonesian Rupiah (IDR) against the US Dollar (USD) and identify the most accurate model among ARIMA, LSTM, and Ensemble Averaging. The evaluation results indicate that ARIMA achieves an accuracy of 97.75%, demonstrating strong performance in short-term forecasting, while LSTM achieves 99.98% accuracy, excelling in capturing complex and dynamic patterns in long-term predictions. The Ensemble Averaging approach attains the highest accuracy of 99.99%, proving to be the optimal solution by combining ARIMA’s stability with LSTM’s adaptability, resulting in more precise and stable predictions. The findings of this study highlight that the ensemble approach is more effective than individual models, as it balances accuracy and prediction stability across various forecasting scenarios. This method serves as a reliable tool for addressing market volatility and contributes significantly to the advancement of financial and economic forecasting techniques that are more adaptive and accurate.
Comparison of ARIMA, LSTM, and Ensemble Averaging Models for Short-Term and Long- Term Forecasting of Non-Stationary Time Series Data Windy Ayu Pratiwi; I Made Sumertajaya; Khairil Anwar Notodiputro
Inferensi Vol 8 No 3 (2025)
Publisher : Department of Statistics ITS

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.12962/j27213862.v8i3.22643

Abstract

This study aims to forecast the highest weekly selling rate of the Indonesian Rupiah (IDR) against the US Dollar (USD) and identify the most accurate model among ARIMA, LSTM, and Ensemble Averaging. The evaluation results indicate that ARIMA achieves an accuracy of 99.37%, demonstrating strong performance in short-term forecasting, while LSTM achieves an accuracy of 99.99%, excelling in capturing complex and dynamic patterns for long-term predictions. The Ensemble Averaging approach achieves an accuracy of 99.87%, proving to be the optimal solution by combining ARIMA’s stability with LSTM’s adaptability, resulting in relatively accurate and stable predictions. Although the Ensemble Averaging model has higher RMSE and MSE values compared to the individual models (ARIMA and LSTM), this approach remains quite effective in forecasting both short-term and long-term time series data. This shows that, despite larger prediction errors, Ensemble Averaging provides more stable and accurate results in the long term. The findings highlight that the ensemble approach is more effective than individual models, as it balances accuracy and prediction stability across various forecasting scenarios. This method serves as a reliable tool for addressing market volatility and contributes significantly to the advancement of more adaptive and accurate financial and economic forecasting techniques.
PENERAPAN ANALISIS LASSO DAN GROUP LASSO DALAM MENGIDENTIFIKASI FAKTOR-FAKTOR YANG BERHUBUNGAN DENGAN TUBERKULOSIS DI JAWA BARAT Stephan Chen; Khairil Anwar Notodiputro; Septian Rahardiantoro
Indonesian Journal of Statistics and Applications Vol 4 No 1 (2020)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v4i1.510

Abstract

Tuberculosis is the deadliest infectious disease in Indonesia, and West Java is a province with the largest number of tuberculosis cases in Indonesia. This research was conducted to identify variables and groups of variables that could explain the number of tuberculosis cases in West Java. The data used has many explanatory variables, and these variables form groups. LASSO and group LASSO analysis can be used for variables selection and handle data that has many explanatory variables, and group LASSO analysis can be used on data with grouped variables. The results of the LASSO analysis, variables that can explain the number of tuberculosis cases in West Java are the number of people with disabilities, the number of pharmacy staff, the number of malnourished people, the number of people working and the number of cities. According to the group LASSO analysis, the variables that can explain the number of tuberculosis cases in West Java are variables in the health and environmental groups. The government can focus on these factors if they want to reduce the number of tuberculosis cases in West Java.
A REPEATED CROSS-SECTIONAL MODEL FOR ANALYZING UNEMPLOYMENT DATA IN BOGOR Ulfah Sulistyowati; Khairil Anwar Notodiputro; I Made Sumertajaya
Indonesian Journal of Statistics and Applications Vol 4 No 2 (2020)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v4i2.513

Abstract

In general, the form of data encountered in statistical problems is panel data and cross-sectional data. There are times in certain conditions, the data formed in the form of a combination of panel data with cross-sectional data, which is commonly referred to as repeated cross-sectional data. Repeated cross-sectional data is often done in research with individual observations. In this study, a repeated cross-sectional analysis was carried out using a fixed influence model with observations in the form of an area (village) in Bogor, West Java to analyze unemployment factors. The results obtained are that ongoing village development affects the unemployment rate in Bogor
METODE ANALISIS DISKRIMINAN KUADRAT TERKECIL PARSIAL UNTUK KLASIFIKASI SEGMEN LOYALITAS KONSUMEN SUSU PERTUMBUHAN Herdina Kuswari; Farit Mochamad Afendi; Khairil Anwar Notodiputro
Indonesian Journal of Statistics and Applications Vol 4 No 2 (2020)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v4i2.586

Abstract

Consumer segmentation is the process of dividing consumers into different segments based on consumer characteristics, making it easier for companies to develop marketing strategies. The segmentation is carried out based on consumer loyalty using the RFM (Recency, Frequency, Monetary) approach a number of 7753 members of a nutritional product loyalty program is considered in the analysis. Partial least square discriminant analysis classification modeling is built using the results of consumer segmentation being the a response variable. The model is not good enough based on the AUC (Area Under Curve) value of the ROC (Relative Operating Characteristic) curve that quite low for each segment. The explanatory variables that have high contribution to the model is X5, X9, and X2 with VIP (Variable Importance in the Projection) values more than 1.
COMPARISON OF K-MEANS CLUSTERING METHOD AND K-MEDOIDS ON TWITTER DATA Cahyani Oktarina; Khairil Anwar Notodiputro; Indahwati Indahwati
Indonesian Journal of Statistics and Applications Vol 4 No 1 (2020)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v4i1.599

Abstract

The presidential election is one of the political events that occur in Indonesia once in five years. Public satisfaction and dissatisfaction with political issues have led to an increase in the number of political opinion tweets. The purpose of this study is to examine the performance of the k-means and k-medoids method in the Twitter data and to tweet about the presidential election in 2019. The data used in this study are primary data taken from Muhyi's research, then mining the text against data obtained. Because this data has been processed by Muhyi to analyze the electability of the 2019 presidential candidate pairs, for this journal needs a preprocessing was carried out to analyze the tendency of tweets to side with the candidate pairs of one or two. The difference in the pre-processing of this research with previous research is that there is a cleaning of duplicate data and normalizing. The results of this study indicate that the optimal number of clusters resulting from the k-means method and the k-medoid method are different.
Improving Classification Model Performances using an Active Learning Method to Detect Hate Speech in Twitter: Peningkatan Kinerja Model Klasifikasi dengan Pembelajaran Aktif dalam Mendeteksi Ujaran Kebencian di Twitter Muhammad Ilham Abidin; Khairil Anwar Notodiputro; Bagus Sartono
Indonesian Journal of Statistics and Applications Vol 5 No 1 (2021)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v5i1p26-38

Abstract

Efforts from the police to address hate speech on social media such as Twitter will not be sufficient to rely solely on manual checks. Therefore, it is necessary to use statistical modelling like the classification model to detect hate speech automatically. Classification is a type of predictive modelling to produce accurate predictions based on labelled data. Generally, the available data are usually unlabelled implying that the labelling process needs to be done beforehand. Data labelling is time consuming, high cost, and often fails to produce correct labels. This research aims to improve the performances of classification models by adding a small amount of data through the so called active learning method. The results showed that there was no significant difference in the performances of logistic regression and naïve bayes classification models in detecting hate speech. However, the results also showed that adding data through the active learning method substantially improved the logistics regression performance in detecting hate speech when compared to data addition based on a simple random sampling method. Therefore, the performances of classification models in detecting hate speech on Twitter could be improved by using an active learning method.
Determinant Factors of Working Children based on Conditional Logistics Regression for Matched Pairs Data: Determinan Anak Bekerja Berdasarkan Model Regresi Logistik Bersyarat untuk Data Berpasangan Rizky Zulkarnain; Tri Listianingrum; Khairil Anwar Notodiputro
Indonesian Journal of Statistics and Applications Vol 5 No 1 (2021)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v5i1p161-172

Abstract

Working children may create problem since it relates to human right as well as to the development of children especially in getting sufficient education. This paper discusses determinant factors of working children by using conditional logistics regression for matched pairs data. Matching is employed to adjust confounding factors and to avoid bias. In this paper there are three confounding factors that have been considered, i.e. residential area, gender, and income of household head. The results showed that the conditional regression model outperformed the standard regression model. The number of household members, whether the head of household was married or single, age of the head of household, educational attainment of the head of household, as well as the work status of the head of household were the determinant factors of the working children.
A Conditional Logistic Regression Model for Analyzing Unemployment Rates in West Java: Model Regresi Logistik Bersyarat untuk Analisis Tingkat Pengangguran di Provinsi Jawa Barat Dwi Jayanti; Septian P Palupi; Khairil Anwar Notodiputro
Indonesian Journal of Statistics and Applications Vol 5 No 1 (2021)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v5i1p195-204

Abstract

Unemployment is a critical problem faced by developing countries.  It is a complex problem which creates other social and economic problems such as poverty, economic gaps, and crimes. This paper discusses the determinant factors of unemployment rates based on empirical data using the conditional logistic regression model.  The model was used to analyze matched pair data using gender, age and residence as matching factors.  The result showed that household status, marriage status, as well as levels of education were the determinant factors of a person being unemployed in West Java.  It is also shown that the conditional logistic regression outperformed the standard logistic regression for analyzing the cause of unemployment.
Analyzing Low Birthweight in Java Based on Logistic Regression Model for Matched Pair Data: Analisis Berat Badan Lahir Rendah di Pulau Jawa Berdasarkan Model Regresi Logistik untuk Data Berpadanan Christiana Anggraeni Putri; Rini Irfani; Khairil Anwar Notodiputro
Indonesian Journal of Statistics and Applications Vol 7 No 2 (2023)
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v7i2p75-85

Abstract

Low birthweight is one of the leading causes of neonatal death. Generally, the study of low birth weight is done by modeling logistic regression without considering the influence of confounding variables that can deviate the actual relationship between the explanatory variables and the response. This paper aims to identify low birth weight determinants in Java based on the logistic regression model for conditional study design, in which the analysis is based on matching the education level of the mother with one control. The results of the analysis showed that matched logistic regression can be used to correct bias due to the influence of a confounding variable. It reveals that based on the results of modeling, the frequency of pregnancy examinations and the parity of children are significantly affect the risk of low birth weight in Java Island.