cover
Contact Name
Dania Siregar
Contact Email
jsamtk.unj@gmail.com
Phone
+6281316044605
Journal Mail Official
jsa@unj.ac.id
Editorial Address
Kampus A Universitas Negeri Jakarta, Lt.6 Gd. Dewi Sartika Jalan Rawamangun Muka, Jakarta Timur.
Location
Kota adm. jakarta timur,
Dki jakarta
INDONESIA
Jurnal Statistika dan Aplikasinya
ISSN : -     EISSN : 26208369     DOI : https://doi.org/10.21009/JSA.041
Jurnal Statistika dan Aplikasinya JSA is dedicated to all statisticians who wants to publishing their articles about statistics and its application. The coverage of JSA includes every subject that using or related to statistics.
Articles 191 Documents
ANALYSIS OF FACTORS EXPLAINING SENIOR HIGH SCHOOL DROPOUT RATE USING GEOGRAPHICALLY WEIGHTED REGRESSION Yekti Widyaningsih; Hana Adzania Nufaisah
Jurnal Statistika dan Aplikasinya Vol. 10 No. 1 (2026): Jurnal Statistika dan Aplikasinya
Publisher : LPPM Universitas Negeri Jakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/JSA.10106

Abstract

East Nusa Tenggara (NTT) Province is one of the areas experiencing the problem of dropping out of high school education. Even though NTT Province has adequate educational facilities and teaching staff, the high school dropout rate in NTT Province always ranks top 9 in Indonesia for the 2019/2020 academic year to 2021/2022. Dropping out of school can be influenced by region (spatial) and does not occur at one time, so research is needed using panel-structured spatial data that accommodates spatial effects over time. Geographically Weighted Panel Regression (GWPR) is a local regression analysis method that considers the effect of spatial heterogeneity on panel-structured spatial data. This study aims to analyse the factors that explain the high school dropout rate in NTT Province in 2019-2021 using GWPR. The results showed that the GWPR model with the Fixed Exponential weighting function was the best model compared to other weighting functions based on  and AIC. Population density, student-teacher ratio, regional minimum wage, open unemployment rate, student-to-school ratio, average length of schooling, and Smart Indonesia Program budget have a significant effect on explaining high school dropout rates in at least 21 regencies/cities in NTT Province. Grouping districts/cities based on variable significance using k-modes clustering produces 4 groups.
APPLYING TEXT MINING FOR EVIDENCE-BASED POLICY: SENTIMENT AND TOPIC ANALYSIS OF INDONESIA’S RESEARCH FUNDING REFORM Rolando Gultom; Raden Arthur Ario Lelono; Juhartono; Lianna Kusumawati; Dody Styawan; Rino; Setyo Purnomo; Agus Prihartono; Maiforlion; Maman Firmansyah; Akhmad Faishal; Danang Rita Handoko
Jurnal Statistika dan Aplikasinya Vol. 10 No. 1 (2026): Jurnal Statistika dan Aplikasinya
Publisher : LPPM Universitas Negeri Jakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/JSA.10104

Abstract

A substantial portion of governance-relevant information in Indonesia is contained in unstructured textual artefacts such as meeting minutes and focus group discussion (FGD) notes, yet these materials remain underutilized in policy evaluation. Existing studies have not systematically examined how computational text analysis can extract policy insights from such documents. The purpose of this research is to evaluate the implementation dynamics of the RIIM funding scheme during its transition to SBK, and to generate evidence-based recommendations that support the development of a more adaptive and context-sensitive research funding framework. This study employs a text-mining approach combining exploratory word cloud visualization, lexicon-based sentiment analysis using the Bing lexicon, emotion analysis using the NRC Emotion Lexicon, and Latent Dirichlet Allocation (LDA) topic modeling. The analysis is conducted on a corpus derived from Focus Group Discussion (FGD) verbatim transcripts involving RIIM stakeholders, consisting of 13 pages and 4,188 words. The findings reveal four major issue clusters: administrative and contractual inconsistencies, ambiguity in output classification across research fields, uncertainty in financial and temporal transitions, and challenges in aligning performance evaluation with diverse research outputs. These findings inform concrete recommendations for enhancing procedural clarity, establishing explicit transitional provisions, strengthening output classification, and institutionalizing stakeholder feedback mechanisms. This study acknowledges limitations related to the use of English-based sentiment lexicons for Indonesian data and the reliance on summarized minutes rather than verbatim transcripts. The study’s originality lies in demonstrating how computational text analysis can systematically extract policy insights from unstructured governance documents, offering a novel evidence-based approach for refining Indonesia’s research funding instruments.
MODELING THE AGGREGATE LOSS DISTRIBUTION IN MOTOR THIRD-PARTY LIABILITY INSURANCE USING MONTE CARLO SIMULATION Tajmahal Ghaza Antoni; Nur Rahmadani Ahmad; Viery Salsaputra Triana; Gemala Azzahra Ocan; Ruhiyat
Jurnal Statistika dan Aplikasinya Vol. 10 No. 1 (2026): Jurnal Statistika dan Aplikasinya
Publisher : LPPM Universitas Negeri Jakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/JSA.10101

Abstract

This study models the aggregate loss distribution for motor third-party liability insurance using Monte Carlo simulation. Aggregate loss estimation is essential because it depends on claim frequency and severity, which often exhibit overdispersion and heavy tails, making analytical solutions intractable and motivating simulation-based approaches for accurate tail-risk assessment. The objective of this study is to identify appropriate distributions for frequency and severity using French Motor Third-Party Liability (MTPL) insurance data and to construct the aggregate loss distribution through Monte Carlo simulation. The modeling procedure involves distribution selection, goodness-of-fit assessment using Chi-Square and Kolmogorov-Smirnov tests, graphical comparison, and model evaluation using the Akaike Information Criterion (AIC). The selected distributions are then combined to generate simulated aggregate losses, from which Value at Risk (VaR) and Tail Value at Risk (TVaR) are computed. The results show that the Zero-One-Two-Three Modified Negative Binomial (Z123M-NB) distribution provides the best fit for claim frequency, while the Burr XII distribution effectively represents claim severity. Monte Carlo simulation with 10 million iterations produces stable estimates of the aggregate loss mean and variance, and the estimated VaR at the 95%, 97.5%, and 99% confidence levels are 105.85, 1,506.61, and 3,629.14, with corresponding TVaR values of 4,122.93, 7,418.70, and 15,075.21, indicating substantial tail heaviness. The study is limited by the sensitivity of variance estimation under extreme severity values and the assumption of a continuous severity model. The novelty of this study lies in integrating the Z123M-NB frequency model with Burr XII severity within a Monte Carlo framework for real MTPL data, offering enhanced flexibility in modeling extreme aggregate losses.
SMALL-SAMPLE AND IMBALANCED DATA MODELING OF FAMILY QUALITY INDEX: A FIRTH LOGISTIC REGRESSION STUDY Dania Siregar; Rini Warti; Amita Rahmat; Anang Kurnia; Kusman Sadik
Jurnal Statistika dan Aplikasinya Vol. 10 No. 1 (2026): Jurnal Statistika dan Aplikasinya
Publisher : LPPM Universitas Negeri Jakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/JSA.10109

Abstract

The Family Quality Index (FQI) in Indonesia is not published annually, limiting the availability of timely information for monitoring and evaluating family development policies. This study addresses this issue by identifying determinants of provincial FQI categories and developing a classification model that can be applied when official FQI measurements are unavailable. Secondary data from 34 provinces in Indonesia for 2023 were analyzed using Firth logistic regression, a method designed to reduce bias in small samples and imbalanced datasets.  The response variable was the FQI category, classified into moderately responsive and responsive groups. Explanatory variables included mean years of schooling, open unemployment rate, poverty rate, and population density. The results show that poverty rate is the only predictor that remains statistically significant after bias correction, with higher poverty levels associated with a lower probability of a province being classified as responsive. The other variables were not statistically significant.  Model performance was evaluated using Leave-One-Out Cross-Validation (LOOCV). Compared with conventional logistic regression estimated by maximum likelihood, the Firth model achieved higher accuracy (88.24% vs. 85.29%), Kappa (0.531 vs. 0.459), and sensitivity (0.931 vs. 0.897), while maintaining the same specificity. Additional sensitivity analyses using a reduced model produced similar results, indicating that the effect of poverty was stable across model specifications.  These findings suggest that annually available socio-economic indicators may be used to provide provisional estimates of provincial FQI categories when official FQI data are unavailable, thereby supporting evidence-based family development policy evaluation.
COMPARISON OF PSO AND ABC IN CHENG FUZZY TIME SERIES FOR RICE PRICE FORECASTING Machdina Indira Laupa; Isran K. Hasan; Nisky Imansyah Yahya
Jurnal Statistika dan Aplikasinya Vol. 10 No. 1 (2026): Jurnal Statistika dan Aplikasinya
Publisher : LPPM Universitas Negeri Jakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/JSA.10102

Abstract

Rice prices as a primary food commodity in Indonesia play an important role in maintaining economic stability and public welfare, but tend to fluctuate, thus requiring accurate forecasting methods to support decision-making. Research on optimization in the Cheng Fuzzy Time Series (FTS Cheng) method remains limited, particularly in comparing the performance of Particle Swarm Optimization (PSO) and Artificial Bee Colony (ABC) in rice price forecasting. This study aims to compare the performance of PSO and ABC optimization in the FTS Cheng method using monthly data from January 2018 to October 2025, with accuracy evaluated using MAE, RMSE, and MAPE. The forecasting process is carried out through interval formation on training and testing data to obtain an optimal model. The results show that FTS Cheng-ABC performs better, with an MAE of 97.947, RMSE of 142.855, and MAPE of 0.633%, compared to FTS Cheng-PSO with an MAE of 118.579, RMSE of 153.354, and MAPE of 0.767%. However, this study is limited to the use of the Fuzzy Time Series Cheng method with two optimization algorithms, namely PSO and ABC, and does not incorporate adaptive parameter mechanisms or comparisons with more advanced methods. Therefore, the FTS Cheng-ABC method is more effective and can be used to support policy decision-making related to rice price stability. This study contributes by providing a comparative analysis of PSO and ABC optimization in improving the performance of the FTS Cheng method for rice price forecasting in Indonesia.
MODELING GEOGRAPHICALLY WEIGHTED BETA REGRESSION ON POVERTY DATA IN INDONESIA Mayang Puspitasari; Idhia Sriliana; Sigit Nugroho; Jose Rizal; Ramya Rachmawati
Jurnal Statistika dan Aplikasinya Vol. 10 No. 1 (2026): Jurnal Statistika dan Aplikasinya
Publisher : LPPM Universitas Negeri Jakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/JSA.10107

Abstract

Poverty is a major development issue influenced by socioeconomic factors and regional disparities. Previous studies on poverty in Indonesia have largely employed global regression approaches that assume homogeneous relationships between explanatory variables and poverty across regions. This assumption may overlook spatial heterogeneity, resulting in less accurate estimates and limited relevance for local policymaking. Therefore, this study applies Geographically Weighted Beta Regression (GWBR) to model the proportion of the poor population in Indonesia and compares its performance with global beta regression. The independent variables are the Human Development Index, Open Unemployment Rate, Mean Years of Schooling, Life Expectancy, Literacy Rate, and Access to Safe Drinking Water. Parameter estimation was performed using Maximum Likelihood Estimation (MLE), while statistical inference was conducted through simultaneous and partial significance tests. Spatial dependence was assessed using Moran’s I statistic. The main contribution of this study is the identification of spatially varying determinants of poverty using the GWBR approach, providing localized insights for more targeted poverty alleviation policies. The results indicate significant spatial autocorrelation in the residuals of the global beta regression model, suggesting that spatial effects should be considered in poverty analysis. Compared with the global beta regression model, GWBR demonstrated superior performance, yielding lower values of AIC (-909.462), AICc (-905.439), and MAAPE (0.218). Local parameter estimates revealed substantial spatial heterogeneity in the determinants of poverty. Life Expectancy was the most consistent factor, exhibiting a significant negative effect in 34 provinces, while Mean Years of Schooling showed a significant negative effect in 15 provinces. The effects of the remaining variables varied across regions, indicating that the determinants of poverty differ spatially across Indonesia. These findings suggest that GWBR more effectively captures local variations in the determinants of poverty and supports the development of region-specific poverty alleviation policies in Indonesia.
DATA IMPUTATION FOR BIVARIATE GAMMA-GENERATED DATA USING PREDICTIVE MEAN MATCHING AND RANDOM FOREST METHODS Muhammad Arib Alwansyah Arib; Jose Rizal; Ramya Rachmawati Ramya
Jurnal Statistika dan Aplikasinya Vol. 10 No. 1 (2026): Jurnal Statistika dan Aplikasinya
Publisher : LPPM Universitas Negeri Jakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/JSA.10103

Abstract

Missing data is a common problem in data analysis and can reduce the quality and accuracy of research results if not handled properly. This study aims to compare the Predictive Mean Matching (PMM) and Random Forest (RF) imputation methods in handling missing data with missing levels of 5%, 10%, 15%, and 20% using correlation indicators, p-values, and observing the smallest Mean Absolute Percentage Error (MAPE) and Root Mean Square Error (RMSE) values. The results show that both methods differ at each level of missing data. At 5% missing data, both methods show significant differences to the original data with a p-value smaller than α = 0.05, but the RF method produces smaller MAPE and RMSE values ​​than PMM. At 10% missing data, the PMM method still shows significant differences to the original data, while the RF method does not. At 15% missing data, the PMM method showed results that were not significantly different from the original data and had smaller MAPE and RMSE values ​​than RF. Meanwhile, at 20% missing data, the RF method produced the highest correlation value of 0.7788 compared to PMM at 0.7638. In general, the results of the study indicate that the greater the proportion of missing data, the imputation error rate also tends to increase. Therefore, the selection of imputation methods needs to be adjusted to the characteristics and proportion of missing data to obtain optimal imputation results.
COMPARISON OF XGBOOST AND SVM FOR SENTIMENT ANALYSIS MERAH PUTIH COOPERATIVE POLICY Wahyu Pratama Lasaleng; Fahrezal Zubedi; Siti Nurmardia Abdussamad
Jurnal Statistika dan Aplikasinya Vol. 10 No. 1 (2026): Jurnal Statistika dan Aplikasinya
Publisher : LPPM Universitas Negeri Jakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/JSA.10105

Abstract

Large volumes of textual data have been generated by the rapid growth of social media, making sentiment analysis an effective approach for understanding public perceptions of government policies. However, text classification still faces challenges such as high feature dimensionality, class imbalance, and non-linear relationships within the data. This study used data from the X social media platform to evaluate the performance of the Support Vector Machine (SVM) and XGBoost algorithms in classifying public sentiment toward the Koperasi Desa Merah Putih policy. The dataset consisted of 1,074 tweets collected through a scraping technique between July 21 and November 4, 2025, comprising 800 positive tweets and 274 negative tweets. The research process included data preprocessing, feature extraction using TF-IDF, data splitting with an 80:20 ratio, and hyperparameter tuning using GridSearchCV with 5-fold cross-validation. The models were evaluated using accuracy, precision, recall, and F1-score. Hyperparameter tuning successfully enhanced the performance of both models, with SVM benefiting the most from the optimisation process. The findings demonstrated that both models achieved strong classification performance; however, SVM outperformed XGBoost. The SVM model achieved an accuracy of 95%, with more balanced precision, recall, and F1-score values across both sentiment classes, whereas XGBoost achieved an accuracy of 91% and showed limitations in detecting negative sentiment as the minority class. The data exploration results also indicated that most users expressed positive sentiment toward the Koperasi Desa Merah Putih policy. Nevertheless, this study has several limitations, particularly the use of TF-IDF-based feature representation, which does not capture semantic relationships or sarcasm in textual data. The novelty of this study lies in the comparison of SVM and XGBoost with hyperparameter tuning using GridSearchCV in the context of sentiment analysis of the Koperasi Desa Merah Putih policy, a topic that has received limited attention in previous studies.
OPTIMIZING UNIVARIATE TIME SERIES IMPUTATION USING RANDOM FOREST REGRESSION AND LSTM FOR ACCURATE FORECASTING Maulana Baihaqi Ramadhan; Emli Rahmi; Isran Hasan
Jurnal Statistika dan Aplikasinya Vol. 10 No. 1 (2026): Jurnal Statistika dan Aplikasinya
Publisher : LPPM Universitas Negeri Jakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/JSA.10108

Abstract

Indonesia possesses high solar radiation potential, making solar energy a strategic pillar for the national clean energy transition. However, its utilization is hindered by incomplete data due to instrument failure, which significantly reduces prediction accuracy. Starting from this problem, this study aims to evaluate the performance of the Machine Learning Based Univariate Time Series Imputation-Random Forest Regression (MLBUI-RFR) method by comparing it with the Mean Imputation method and evaluating it through Long Short-Term Memory (LSTM) forecasting. The methodology begins with data preprocessing using the MLBUI-RFR scheme to handle missing values, which are then used as input for the LSTM architecture to forecast solar radiation in Indonesia. The findings demonstrate that the use of the MLBUI-RFR method contributes significantly to improving data quality, where the LSTM model trained with MLBUI-RFR imputed data achieves higher accuracy compared to Mean Imputation. The evaluation results show a lower error rate, with an NRMSE of (15.68%) and a MAPE of (18.98%), whereas the Mean Imputation method yields an NRMSE of (16.06%) and a MAPE of (19.15%) proving that the proposed method is more effective in capturing non-linear patterns in the data. However, this study is based exclusively on data obtained from the Gorontalo Climatology Station in Gorontalo Province, Indonesia. The contribution of this study lies in evaluating the integration of MLBUI-RFR and LSTM for solar radiation forecasting, demonstrating how machine learning based univariate time series imputation can improve data quality and subsequently enhance forecasting performance on solar radiation data.
Front Matter Jurnal Statistika dan Aplikasinya Vol. 10 No. 1, June 2026 Journal Editor JSA
Jurnal Statistika dan Aplikasinya Vol. 10 No. 1 (2026): Jurnal Statistika dan Aplikasinya
Publisher : LPPM Universitas Negeri Jakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/JSA.10100

Abstract